saved
Terminal-Bench-Science 0.1
gist
Terminal-Bench-Science 0.1 is a Stanford-led, Harbor-built benchmark of AI agents on 70 expert-curated scientific research workflows, so scientists—not model vendors—set the bar. Of 920 proposals, only 70 survived domain, technical, and bar-raiser review. Claude Opus 5 with Claude Code leads at 30% resolution; GPT-5.6 Sol is second at 22.4% for under a third of Claude Fable 5's cost. The suite is deliberately harder than Terminal-Bench 3.0, reports Pareto frontiers on cost and tokens, and will keep adding and retiring tasks as the frontier moves.
ideas
- Scientists set the evaluation bar. Tasks come from researchers' own workflows, not textbooks; 70 of 920 proposals survived domain, technical, and bar-raiser review.
- Frontier agents still fail most of the work. Claude Opus 5 resolves 30%; most other systems sit under 11%, and GLM 5.3 is the strongest open model at 8.1%.
- The hardness is calibrated. Resolution rates sit more than 10 points below Terminal-Bench 3.0 because reviewers rejected tasks frontier agents already solve.
- Cost and tokens split the frontier. GPT-5.6 Sol matches Claude Fable 5 at under a third the cost; only Kimi K3 and Opus 5 sit on both Pareto fronts.
- It is a living benchmark. Releases will add, retire, and recalibrate tasks; 0.2 PRs are due October 5, 2026, and Harbor can re-run versioned trials.
quotes
“Scientists, not model developers or data vendors, set the bar for scientific capability in AI.”
“The strongest model evaluated, Claude Opus 5, achieves a 30% resolution rate on Terminal-Bench-Science 0.1.”
“Of 920 proposals, 464 were approved for implementation and 386 pull requests were opened, but only 70 tasks made it into Terminal-Bench-Science 0.1.”
“GPT-5.6 Sol matches Claude Fable 5's performance at less than a third of the cost ($4.2k vs $14.2k).”