DigestAI news desk

Cut through the AI noise.

Research

Researchers introduce Benchy, a standardized language for AI task benchmarks

Researchers have proposed Benchy, a new semantic language and execution engine designed to standardize task-oriented AI benchmarks. The system defines benchmarks as a structured triplet of program, scoring function, and dataset (B=(P,S,D)), separate from the AI system being tested. Runs bind these components (R=(B,AI)), ensuring consistency and reproducibility.

1 source primary source

Key points

  • Benchmarks defined as structured triplets: program, scoring function, and dataset (B=(P,S,D))
  • Authors use YAML with a shared ontology for tasks, domains, and languages, compiled to JSON
  • Engine enforces a universal runtime contract to ensure consistent AI system integration

Benchmarks are authored in canonical YAML, where each semantic concept has a single valid syntax. This syntax is classified using a shared ontology for tasks, domains, and languages. The YAML is then deterministically compiled into a canonical JSON intermediate representation executed by the engine. The design ensures that compilation preserves meaning without altering definitions or injecting defaults. Programs adhere to fixed schemas with named input/output fields, and the engine enforces a universal runtime contract—a standardized input/output object interface—to which external AI systems adapt. This approach prevents integration mechanics from affecting benchmark semantics.

Read the original at arXiv cs.AI · by Francis F Daniel, Mauro Iba\~nez, Francis Perelman, Marian Basti primary sourceOpen source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories