saved
Small Models Have Arrived
gist
Calvin French-Owen argues that small, fast models have crossed a cost-quality threshold that unlocks consumer AI. His gpt-5.6-luna runs sit near 100 tokens per second and keep research-thread API bills in the tens of cents; a personalized news eval drops from a dollar on Sonnet-class models to ten cents. Inference cost, not product taste, is why consumer AI companies have been scarce. He still uses frontier models for hard coding, and expects that demand to keep compounding, but most company work is "token spewer" responsiveness—the hiring archetype cheap models can now serve.
ideas
- Cheap models crossed a consumer threshold. Luna keeps research-thread bills in the tens of cents and a personalized news eval near $0.10, versus about $1 on Sonnet-class models.
- Token cost, not taste, blocked consumer AI. The old consumer playbook assumed cheap-to-run websites; per-request inference blew up the capital needed before ads.
- Keep frontier models for IQ-180 work. French-Owen still reaches for Fable 5 and 5.6 Sol for coding, and expects frontier demand to keep compounding in engineering and discovery.
- Most company work is token-spewer work. Peter Reinhardt's claim that about 95% of operator work is responsiveness, not genius, is the demand curve cheap models can fill.
- Harnesses and permissions still missing. Fast/cheap/good-enough models need new harnesses, prompt-injection safety, roles, and permissions before they can run a business.
quotes
“There's a straightforward answer: token costs.”
“But looking at luna, the results are pretty decent, and the average cost is ~$0.10.”
“Peter mentioned that ~95% of the work he does falls into bucket 2.”
“Most of the "human tokens" at companies today are spent this way — hiring skews heavily toward the fast/cheap/good-enough archetype.”