hraness
Theme
Appearance

saved

Human Baselines for Benchmarks: AI Now Outperforms Junior Accountants

by Aden BartonMercorpublished

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Mercor hired twelve licensed CPAs, averaging about five and a half years of experience, to finish simplified month-end close tasks drawn from its APEX-Accounting benchmark. Frontier models now beat even the best participant on speed and accuracy, and they cost far less per rubric criterion. The tasks reward file search and exact instruction following rather than client communication, so the result is a warning about both accounting work and benchmarks that drift past what one person can complete.

ideas

  • Frontier models now beat the best CPA in the study. On these medium-length month-end tasks, recent models score at or near 100 percent in under ten minutes, while accountant accuracy ranged from 0 percent to about 90 percent and most attempts took 30 to 180 minutes.
  • The lead appeared in about eighteen months. The best models recently sat below the accountants' roughly 37 percent average; o3 passed that line in spring 2025, and current frontier models ace the same tasks.
  • The tasks test model strengths, not the whole job. They stack hard-to-spot but realistic requirements where one missed number tanks the score, and they omit coworkers, accumulated context, and client communication.
  • Unassisted humans cost far more per correct criterion. The write-up puts Claude Opus 5 at $0.21 per rubric criterion met versus $10.35 for accountants at the US median wage.
  • Frontier benchmarks can leave human work behind. Tasks built to stump models now take many experts tens of hours, so a high score may describe work one person cannot finish unaided.

quotes

“frontier AI models are now faster and more accurate than junior accountants, even the best one in our study”

Aden Barton, stating the study's main result.

“Just eighteen months ago, the best AI models fell short of the average accountant’s ~37% score.”

Aden Barton, describing how quickly models passed the human average.

“These results do not mean accountants are replaceable, since our tasks ended up testing what AI is best at”

Aden Barton, limiting the claim to the skills the tasks actually measure.

“We don’t ask humans to run 60 miles an hour, yet transporting people at that speed is exactly how cars boost productivity”

Aden Barton, explaining why benchmarks can move past work one human can do.