Study finds frontier AI models outperform junior accountants on specific tasks
A new study involving 12 licensed junior accountants reveals that frontier AI models now outperform humans on medium-length, well-defined accounting tasks. The researchers hired CPAs with an average of five and a half years of experience to complete four realistic month-end close scenarios. In these tests, which required digging through working files to find numbers and perform calculations, AI…
Key points
- Frontier AI models outperformed 12 junior accountants on specific month-end close tasks.
- AI models were faster, more accurate, and more than an order of magnitude cheaper than humans.
- The study notes AI excels at detail-oriented tasks but does not replace full accounting roles.
The study highlights a rapid shift in capability. Just eighteen months ago, the best AI models scored below the average accountant’s 37% mark, but today they ace these same tasks. The authors note that this does not mean accountants are replaceable, as the tasks tested only detail-oriented instruction following and file navigation, excluding broader job duties like client communication and context building. The models were also significantly cheaper, costing more than an order of magnitude less per task criterion than human labor.
The findings suggest substantial productivity gains in accounting are likely in the coming years, even if model progress stalls. The study also critiques current benchmarking practices, arguing that as AI improves, benchmarks are evolving to test tasks that are too complex for any single human to complete, shifting the focus from comparing AI to humans to measuring previously impossible or prohibitively costly work.
Human Baselines for Benchmarks: AI Now Outperforms Junior Accountants
mercor.com · 2 October 2026
Loading the full article…
This text was published by mercor.com and written by Aden Barton. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Reddit discussionreddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model · 1 src
- Korean legal study finds KLUE-BERT outperforms GPT models in sexual offense text classification · 1 src
- Researchers question human-derived bias measures for LLM evaluation · 1 src
- arXiv study finds reading LLM judges from first token overstates position bias · 1 src
- Study compares On-Device NER models for speed, cost and accuracy · 1 src
Comments
via GitHub Discussions