AISI and EvalEval release evaluation cards for five benchmarks and six frontier models
The UK AI Security Institute (AISI) and the EvalEval coalition have published a new set of Evaluation Cards that bundle verified results, context, and configuration details for five major benchmarks: HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, and Terminal‑Bench 2.0. The release also includes data from two cyber‑focused evaluations, Cyber CTFs and The Last Ones.
Key points
- Evaluation Cards now contain verified results for five benchmarks and six frontier LLMs.
- Benchmarks include HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, and Terminal‑Bench 2.0.
- Release aims to improve reproducibility by providing full context, configuration, and run data via the Every Eval Ever schema.
The cards cover six frontier large‑language models – Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT‑5, GPT‑5.2 and GPT‑5.4 – and are built on the Every Eval Ever (EEE) schema that EvalEval developed to standardise reporting. AISI’s prior work on OptStop, HiBayES and transcript‑level transparency underpins the effort, aiming to make benchmark replication less costly and more statistically rigorous. By providing full run metadata, the initiative lets researchers diagnose how inference‑time compute and evaluation protocols influence reported performance, and offers a reference point for future meta‑research.
The collaboration hopes that wider adoption of the EEE schema and Evaluation Cards will improve reproducibility across the AI evaluation ecosystem, helping model developers, evaluators and policy analysts understand the conditions behind published scores.
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
Hugging Face · 22 September 2026
Loading the full article…
This text was published by Hugging Face and written by Avijit Ghosh. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- OpenAI's GPT-6 Astra attempts 97 unsafe instructions in robot safety test · 2 src
- Google DeepMind uses 7‑hour role‑playing game to explore AI’s impact on science · 1 src
- Opinion: Reproducibility challenges rise with large language models · 1 src
- Anthropic reports Claude leads 26% of its AI research and development tasks · 4 src
- Opinion: big tech’s AI cancer‑cure promises outpace proven progress · 1 src
Comments
via GitHub Discussions