saved
The Most Important Market in AI is the Middle
by Tomasz Tunguz · X · published
Hraness republishes this public post from a saved copy. The post is the author’s own words.
Yesterday, Anthropic released a new model & cut its price. Ninety minutes later, OpenAI did the same.
Most business AI use is the messy middle. It is where the competition is fiercest.
The Most Important Market in AI is the Middle
Yesterday, Anthropic released a new model & cut its price. Ninety minutes later, OpenAI did the same.
Most business AI use is the messy middle: multi-step workflows that need a smart enough model at a price a company can afford. It is the most important part of the market today, & it is where the competition is fiercest. The price cuts are the evidence.
In June, Anthropic set the frontier price at $10 & $50 per million tokens with Fable 5. In July, OpenAI answered with GPT-5.6 Sol at $5 & $30, matching that capability at a third of the cost per task.
The Opus line had never moved. Opus 4.5, 4, 4.8 & 5 all listed at $5 & $25 per million tokens. Yesterday’s cut was the first.
It is even more extreme at the low end. OpenAI cut Luna by 80% in July, then cut it another 50% yesterday.
More than just closed source rivalry, open models deflate prices too. The generics on the AI grocery aisle run a majority of token volume on the gateways that publish data, at an 86% discount to the blended price of closed models.
Large customers pursue even greater savings with fine tuning. Cursor’s Composer 2 fine tuned Kimi K2.5, an open-weight base, cutting its overall cost 86% against its previous in-house model. Harvey did the same, cutting cost per cell 55% against Sonnet 5 while scoring higher than Fable 5.
But the right tail of the market is thinner than almost anyone forecast. Anthropic’s Fable 5.1, its most capable & most expensive model, commanded only 3.7% of gateway spending in its first twelve days. Its predecessor peaked at 13.2% when access was restored in July, then fell to 4.9% a month later when Opus 5 shipped at half the price. Among large corporate accounts, frontier models fell from 53% of token consumption in early August to 45% by September.
Demand for intelligence is not a pyramid with a small, wealthy peak paying for everything beneath it. It is a normal distribution with a fat middle. The middle buys intelligence per dollar.
Intelligence costs keep plummeting. What enterprises demand from AI does not change nearly as fast. So the tier that satisfies a fixed requirement keeps getting cheaper.
As intelligence per dollar explodes, the distribution of tokens may shift to commodity. Whether that happens will determine the economics of the AI market.
Image 1 transcription
Can : Price collapse at the frontier Blended API cost per million tokens across OpenAl, Anthropic, & Google flagship models, 2023 to 2026, logarithmic scale, $100.00 © Anthrople $5000 ora © Openat . puss ee . GPT-6 Astra $20.00 2 ee Fable 5:1 o ‘Opus 4.5 3 s1000 af opus 5.5 5 . _ivohs g $500 ——S—S eee ‘ if 2 e @ $200 8 = DU $100 z B $050 $0.20 $010 Yen 23 Jul23 Jan 24 Jul 24 Jon 25, Jul 25 Jan 26 Jul26 Jen 27 THEORY VENTURES
Image 2 transcription
A ”° My 2¢? Oh, it's actually a penny for your thought. Blended API cost per million tokens across OpenAl, Anthropic, & Google budget models, 2023 to 2026, logarithmic scale. $1000 © Anthrople $5.00 @ Openal Fash 3s © Google °. § s200 a °,
2 ‘.
5 -
= $100 ———
= Ow, S
8 ua
a
@ $050 2"
Z
= ma
3 .
5 50.20 e
° ma Fash 1s Jon 23 Jul23 Jan 24 Jul24 Jan 25, Jul25 Jan 26 Jul26 Jen 27 THEORY VENTURES
Image 3 transcription
°° The middle is where the money is: the mid tier claims 40% of spend & 30% of tokens Comparison of raw token volume share & dollar spend share by model tier across commercial routing platforms. A. Raw Token Volume Share B. Gateway Spend Share Forcariage i avril tokens foutng thsugh procutlon geteeys orca fifa cantor redo pant ot sock a Commodity Tier (Luna Hoi Llama 388) Commodity Tier (Luna. Hou Loma $88) Po iat _ m= Mid-Tier Workhorse (So) Liem 5.9 708) Mid-Tier Workhorse (So\ Lams 3.9 708) | a Fs “on Frontier SOTA. (Fable 5.1, GPT-6 Astra) Frontier SOTA (Fable 51, GPT-6 Astra) || m= Po ie The Barbell Paradox: SOTA drives 45% of customer spend but handles only 16% of raw token traffic. mrivenae
Image 4 transcription
The price of artificial thought may have fallen faster than for any other transformative technology in history Relative cost (cheaper) Starting price 10x a 1092-1079 1000x Lithium batteries - 1991-2024 = Al
Imillion* 599496 \
DNA sequencing
2001-22 1 billionx
Compute 1 trillionx 1940-2001 a ao an a —— 0 10 20 30 40 50 60 70 80 90
Years since start of price decline All prices adjusted for inflation. Compute series stops merely for lack of newer data. Sources: Electricity: Hersh (1999), Figure 2.1, and EEI (1995), Table 63. Compute: Nordhaus (2007), Appendix Table 2. Batteries: Fouquet (2026) via Our World in Data. DNA sequencing: National Human Genome Research Institute. Al line is 47%/quarter (12.7x/year), per estimates reported here for 2023-26; the extrapolation of that rate back to 2021 is broadly validated by Appenzeller (2024)’s finding that per-token prices for models clearing a certain capability threshold fell 10x/year in 2021-24. Z EPOCH AI | Cc-BY epoch.ai



