hraness

saved

Terminal-Bench-Science 0.1

by Steven DillmannTerminal-Bench-Science

gist

Terminal-Bench-Science 0.1 is a Stanford-led, Harbor-built benchmark of AI agents on 70 expert-curated scientific research workflows, so scientists—not model vendors—set the bar. Of 920 proposals, only 70 survived domain, technical, and bar-raiser review. Claude Opus 5 with Claude Code leads at 30% resolution; GPT-5.6 Sol is second at 22.4% for under a third of Claude Fable 5's cost. The suite is deliberately harder than Terminal-Bench 3.0, reports Pareto frontiers on cost and tokens, and will keep adding and retiring tasks as the frontier moves.

ideas

  • Scientists set the evaluation bar. Tasks come from researchers' own workflows, not textbooks; 70 of 920 proposals survived domain, technical, and bar-raiser review.
  • Frontier agents still fail most of the work. Claude Opus 5 resolves 30%; most other systems sit under 11%, and GLM 5.3 is the strongest open model at 8.1%.
  • The hardness is calibrated. Resolution rates sit more than 10 points below Terminal-Bench 3.0 because reviewers rejected tasks frontier agents already solve.
  • Cost and tokens split the frontier. GPT-5.6 Sol matches Claude Fable 5 at under a third the cost; only Kimi K3 and Opus 5 sit on both Pareto fronts.
  • It is a living benchmark. Releases will add, retire, and recalibrate tasks; 0.2 PRs are due October 5, 2026, and Harbor can re-run versioned trials.

quotes

Scientists, not model developers or data vendors, set the bar for scientific capability in AI.

Steven Dillmann, stating who should set the evaluation bar.

The strongest model evaluated, Claude Opus 5, achieves a 30% resolution rate on Terminal-Bench-Science 0.1.

Steven Dillmann, reporting the headline result.

Of 920 proposals, 464 were approved for implementation and 386 pull requests were opened, but only 70 tasks made it into Terminal-Bench-Science 0.1.

Steven Dillmann, describing how selective the first task set was.

GPT-5.6 Sol matches Claude Fable 5's performance at less than a third of the cost ($4.2k vs $14.2k).

Steven Dillmann, separating cost from peak resolution.