Researchers introduce Benchy, a standardized language for AI task benchmarks
Researchers have proposed Benchy, a new semantic language and execution engine designed to standardize task-oriented AI benchmarks. The system defines benchmarks as a structured triplet of program, scoring function, and dataset (B=(P,S,D)), separate from the AI system being tested. Runs bind these components (R=(B,AI)), ensuring consistency and reproducibility.
Key points
- Benchmarks defined as structured triplets: program, scoring function, and dataset (B=(P,S,D))
- Authors use YAML with a shared ontology for tasks, domains, and languages, compiled to JSON
- Engine enforces a universal runtime contract to ensure consistent AI system integration
Benchmarks are authored in canonical YAML, where each semantic concept has a single valid syntax. This syntax is classified using a shared ontology for tasks, domains, and languages. The YAML is then deterministically compiled into a canonical JSON intermediate representation executed by the engine. The design ensures that compilation preserves meaning without altering definitions or injecting defaults. Programs adhere to fixed schemas with named input/output fields, and the engine enforces a universal runtime contract—a standardized input/output object interface—to which external AI systems adapt. This approach prevents integration mechanics from affecting benchmark semantics.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers release benchmark for AI in systematic review screening · 1 src
- SlideLab framework generates scientific presentations from research papers · 1 src
- Cartograph reduces AI agent tool discovery from O(n) to O(k) · 1 src
- Survey reviews 211 fake review detection studies from 2018 to 2026 · 1 src
- Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement · 1 src
Comments
via GitHub Discussions