{"version":1,"type":"story","url":"https://digestai.news/story/aisi-and-evaleval-release-evaluation-cards-for-five-benchmarks-and-six","json":"https://digestai.news/story/aisi-and-evaleval-release-evaluation-cards-for-five-benchmarks-and-six.json","markdown":"https://digestai.news/story/aisi-and-evaleval-release-evaluation-cards-for-five-benchmarks-and-six.md","slug":"aisi-and-evaleval-release-evaluation-cards-for-five-benchmarks-and-six","headline":"AISI and EvalEval release evaluation cards for five benchmarks and six frontier models","summary":"The UK AI Security Institute (AISI) and the EvalEval coalition have published a new set of Evaluation Cards that bundle verified results, context, and configuration details for five major benchmarks: HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, and Terminal‑Bench 2.0. The release also includes data from two cyber‑focused evaluations, Cyber CTFs and The Last Ones.\n\nThe cards cover six frontier large‑language models – Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT‑5, GPT‑5.2 and GPT‑5.4 – and are built on the Every Eval Ever (EEE) schema that EvalEval developed to standardise reporting. AISI’s prior work on OptStop, HiBayES and transcript‑level transparency underpins the effort, aiming to make benchmark replication less costly and more statistically rigorous. By providing full run metadata, the initiative lets researchers diagnose how inference‑time compute and evaluation protocols influence reported performance, and offers a reference point for future meta‑research.\n\nThe collaboration hopes that wider adoption of the EEE schema and Evaluation Cards will improve reproducibility across the AI evaluation ecosystem, helping model developers, evaluators and policy analysts understand the conditions behind published scores.","keyPoints":["Evaluation Cards now contain verified results for five benchmarks and six frontier LLMs.","Benchmarks include HealthBench, FrontierMath, Humanity’s Last Exam, SWE‑Bench Pro, and Terminal‑Bench 2.0.","Release aims to improve reproducibility by providing full context, configuration, and run data via the Every Eval Ever schema."],"whyItMatters":"Standardised, open evaluation data lets researchers compare models fairly, spot protocol effects, and builds trust in reported AI performance across the ecosystem.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["Anthropic","OpenAI","EvalEval","UK AI Security Institute"],"models":["Claude Opus 4","Claude Opus 4.5","Claude Opus 4.6","GPT-5","GPT-5.2","GPT-5.4"],"people":[]},"firstPublishedAt":"2026-09-22T00:00:00Z","updatedAt":"2026-09-22T00:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"Hugging Face","title":"How UK AISI and EvalEval Are Making Benchmark Results Reproducible","url":"https://huggingface.co/blog/evaleval-aisi","publishedAt":"2026-09-22T00:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"AISI and EvalEval release evaluation cards for five benchmarks and six frontier models\", 22 September 2026, https://digestai.news/story/aisi-and-evaleval-release-evaluation-cards-for-five-benchmarks-and-six","publisher":"Digest AI","title":"AISI and EvalEval release evaluation cards for five benchmarks and six frontier models","datePublished":"2026-09-22T00:00:00Z","url":"https://digestai.news/story/aisi-and-evaleval-release-evaluation-cards-for-five-benchmarks-and-six"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}