# AISI and EvalEval release evaluation cards for five benchmarks and six frontier models

Digest AI · Research · published 2026-09-22T00:00:00Z

Canonical: https://digestai.news/story/aisi-and-evaleval-release-evaluation-cards-for-five-benchmarks-and-six

## Summary

The UK AI Security Institute (AISI) and the EvalEval coalition have published a new set of Evaluation Cards that bundle verified results, context, and configuration details for five major benchmarks: HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, and Terminal‑Bench 2.0. The release also includes data from two cyber‑focused evaluations, Cyber CTFs and The Last Ones.

The cards cover six frontier large‑language models – Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT‑5, GPT‑5.2 and GPT‑5.4 – and are built on the Every Eval Ever (EEE) schema that EvalEval developed to standardise reporting. AISI’s prior work on OptStop, HiBayES and transcript‑level transparency underpins the effort, aiming to make benchmark replication less costly and more statistically rigorous. By providing full run metadata, the initiative lets researchers diagnose how inference‑time compute and evaluation protocols influence reported performance, and offers a reference point for future meta‑research.

The collaboration hopes that wider adoption of the EEE schema and Evaluation Cards will improve reproducibility across the AI evaluation ecosystem, helping model developers, evaluators and policy analysts understand the conditions behind published scores.

## Key points

- Evaluation Cards now contain verified results for five benchmarks and six frontier LLMs.
- Benchmarks include HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, and Terminal‑Bench 2.0.
- Release aims to improve reproducibility by providing full context, configuration, and run data via the Every Eval Ever schema.

## Why it matters

Standardised, open evaluation data lets researchers compare models fairly, spot protocol effects, and builds trust in reported AI performance across the ecosystem.

## Sources

1. [How UK AISI and EvalEval Are Making Benchmark Results Reproducible](https://huggingface.co/blog/evaleval-aisi) (Hugging Face, 2026-09-22, primary source)

## Cite

Digest AI, "AISI and EvalEval release evaluation cards for five benchmarks and six frontier models", 22 September 2026, https://digestai.news/story/aisi-and-evaleval-release-evaluation-cards-for-five-benchmarks-and-six

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/aisi-and-evaleval-release-evaluation-cards-for-five-benchmarks-and-six.json
