saved
tokens too cheap to meter
Hraness cites a source capture. The source author remains the source.
gist
jyn argues AI token and task costs are falling by orders of magnitude through GPU efficiency, cheaper-per-task model frontiers, inference-engine gains, MoE and Mamba architectures, and specialized classifiers like Jev. Combining those stacks yields roughly 2.5 orders of magnitude cheaper tokens in a year, with specialized yes/no models another 1-2 orders cheaper still. Quality and access, not raw token count, become the binding constraints; once tokens undercut tool calls, models embed into ordinary computing infrastructure amid supply- and demand-side Jevons effects.
ideas
- Cost-per-task is collapsing even when frontier price-per-token is not. Smaller models may burn more tokens; the pareto frontier still moves left and up across 2025 into 2026.
- Hardware, engines, and architectures stack. GPU efficiency roughly doubles every two years; serving stacks gain tens of percent yearly; MoE and Mamba cut compute or RAM for comparable quality.
- Specialized classifiers unlock another two orders. Jev-style yes/no models price input near $42 per billion tokens with free outputs, cheap enough for pipe tools like jgrep.
- When tokens undercut tools, models move into infrastructure. Adaptive schedulers and other once-expert systems become embeddable with general models.
- Cheap volume does not erase quality rents. Frontier labs can still monetize hard tasks while open weights and discounts absorb ordinary work under Jevons-driven demand.
quotes
“Extraordinary claims require extraordinary evidence, so I collected a whole bunch of evidence.”
“If we combine all this, we see about 2.5 orders of magnitude decrease in token cost in the last year.”
“Once models are cheaper than a tool, it becomes attractive to put models in tools.”
“Just because volume is getting cheap doesn't mean that quality is.”