hraness

saved

Are Open Models Catching Up?

by Evan Cloutier, Max Kan, Jordan Nanos and Dylan PatelSemiAnalysispublished

gist

SemiAnalysis measures the open-versus-closed capability gap across three LLM eras (early scaling, reasoning, and agentic) with era-specific benchmarks instead of one continuous score. Their composites show a cycle: a frontier lab jumps ahead, open models close the gap, and catch-up time halves each generation, down to 4.8 to 6 months for agentic models. The authors still prefer Anthropic’s productized stack for daily work and treat public benchmarks as hill-climbable, incomplete proxies for real use.

ideas

  • Evaluate each era on its own benchmarks. Saturated exams from the last era hide the new capability jump, so one historical scoreboard misreads the gap.
  • Catch-up time is halving. Open models matched GPT-4-class systems only in late 2024, closed the o1-era gap in 8.5 months, then passed Opus 4.5 and GPT-5.2 in 4.8 to 6 months.
  • The gap is cyclic, not secular. Each era starts with a closed-lab research lead that others reverse-engineer, including through distillation.
  • The product, not the leaderboard, still decides daily use. GPT-5.2 scored higher than Opus 4.5 on their suite, but Anthropic’s model-plus-harness won the agentic market; the authors still prefer Fable 5 over higher-scoring Kimi K3.
  • Public benchmarks invite hill-climbing. Labs can train RL environments that mimic the evals, so composite scores overstate how completely the gap has closed for real work.

quotes

the open vs. closed gap moves in cycles.

Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel, stating the historical pattern.

with each generation, open-source models take half as long to catch up to the first closed-source model of the era.

Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel, stating the measured trend.

the gap closed faster in Era 3 than either era before it.

Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel, describing the agentic-era result.

public benchmarks, which model makers can easily hill climb by simply creating a bunch of RL environments that closely mimic the benchmark tasks.

Evan Cloutier, Max Kan, Jordan Nanos, and Dylan Patel, limiting what the composites prove.