Study maps 14,767 LLM benchmark papers, noting shift toward action and professional tasks
Researchers systematically examined 14,767 arXiv submissions that introduced or updated evaluation resources for large language models (LLMs) between January 2022 and August 2026. Using staged screening and automated full‑text coding, they categorized changes in target systems, domains, evaluation materials, conditions, and scoring mechanisms.
Key points
- Researchers coded 14,767 arXiv benchmark papers from Jan 2022–Aug 2026.
- Benchmarks now prioritize action, interaction, and professional‑use scenarios.
- LLM‑based scoring grows, yet model‑generated test materials lack sustained rise.
The analysis shows a growing emphasis on benchmarks that require action, interaction, and professional‑application tasks, while older and newer design elements often coexist. Participation by LLMs in scoring has risen across both agent‑focused and non‑agent groups, but the production of model‑generated test materials has not shown a comparable sustained increase in recent cohorts. The authors highlight that as AI systems increasingly help construct tests, perform tasks, and judge responses, expanding evaluation may either provide more independent evidence of capability or risk reinforcing the preferences and blind spots of the models themselves.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- New framework optimizes LLM inference costs via adaptive model activation · 4 src
- Qwen3.5-4B outperforms larger LLMs on new user-side conflict benchmark · 1 src
- Neo-Classic benchmark evaluates linguistic-aesthetic reasoning in Classical Chinese poetry · 1 src
- Study finds trust and friction issues in major generative AI app reviews · 1 src
- Study finds PCA can detect stylistic axes in LLM activations without training · 1 src
Comments
via GitHub Discussions