DigestAI news desk

Cut through the AI noise.

Research3 min read

AISI and EvalEval release evaluation cards for five benchmarks and six frontier models

The UK AI Security Institute (AISI) and the EvalEval coalition have published a new set of Evaluation Cards that bundle verified results, context, and configuration details for five major benchmarks: HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, and Terminal‑Bench 2.0. The release also includes data from two cyber‑focused evaluations, Cyber CTFs and The Last Ones.

1 source primary source

Key points

  • Evaluation Cards now contain verified results for five benchmarks and six frontier LLMs.
  • Benchmarks include HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, and Terminal‑Bench 2.0.
  • Release aims to improve reproducibility by providing full context, configuration, and run data via the Every Eval Ever schema.

The cards cover six frontier large‑language models – Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT‑5, GPT‑5.2 and GPT‑5.4 – and are built on the Every Eval Ever (EEE) schema that EvalEval developed to standardise reporting. AISI’s prior work on OptStop, HiBayES and transcript‑level transparency underpins the effort, aiming to make benchmark replication less costly and more statistically rigorous. By providing full run metadata, the initiative lets researchers diagnose how inference‑time compute and evaluation protocols influence reported performance, and offers a reference point for future meta‑research.

The collaboration hopes that wider adoption of the EEE schema and Evaluation Cards will improve reproducibility across the AI evaluation ecosystem, helping model developers, evaluators and policy analysts understand the conditions behind published scores.

Full story from Hugging Face · by Avijit Ghosh primary sourceOpen source ↗

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

Hugging Face · 22 September 2026

Loading the full article…

This text was published by Hugging Face and written by Avijit Ghosh. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories