saved
Are Open Models Catching Up?
gist
SemiAnalysis measures the open-versus-closed capability gap across three LLM eras (early scaling, reasoning, and agentic) with era-specific benchmarks instead of one continuous score. Their composites show a cycle: a frontier lab jumps ahead, open models close the gap, and catch-up time halves each generation, down to 4.8 to 6 months for agentic models. The authors still prefer Anthropic’s productized stack for daily work and treat public benchmarks as hill-climbable, incomplete proxies for real use.
ideas
- Evaluate each era on its own benchmarks. Saturated exams from the last era hide the new capability jump, so one historical scoreboard misreads the gap.
- Catch-up time is halving. Open models matched GPT-4-class systems only in late 2024, closed the o1-era gap in 8.5 months, then passed Opus 4.5 and GPT-5.2 in 4.8 to 6 months.
- The gap is cyclic, not secular. Each era starts with a closed-lab research lead that others reverse-engineer, including through distillation.
- The product, not the leaderboard, still decides daily use. GPT-5.2 scored higher than Opus 4.5 on their suite, but Anthropic’s model-plus-harness won the agentic market; the authors still prefer Fable 5 over higher-scoring Kimi K3.
- Public benchmarks invite hill-climbing. Labs can train RL environments that mimic the evals, so composite scores overstate how completely the gap has closed for real work.
quotes
“the open vs. closed gap moves in cycles.”
“with each generation, open-source models take half as long to catch up to the first closed-source model of the era.”
“the gap closed faster in Era 3 than either era before it.”
“public benchmarks, which model makers can easily hill climb by simply creating a bunch of RL environments that closely mimic the benchmark tasks.”