HakemBench releases 2,346-item Turkish benchmark for typed decisions
Researchers have released HakemBench, a new benchmark designed to evaluate AI models on typed decision-making tasks in Turkish. Version 1.0 is fully open under a CC BY 4.0 license and contains 2,346 items across seven tracks, including fact-checking, legal routing, and customer support. The benchmark uses a specific harness to score decision quality, calibration, and selective automation,…
Key points
- HakemBench v1.0 is released under CC BY 4.0 with 2,346 items and 4,275 questions.
- Gold labels are generated by AI model comparison, not human verification.
- The top model scores 0.888; the lab's model scores 0.660 overall and 0.678 on unbiased tracks.
A key methodological detail is that most gold labels are derived from blind passes of one AI model family compared against votes from other large language model families, rather than human verification. The authors note that these labels are not human-verified.
Initial results show a board of 16 models where the leading model achieved a composite score of 0.888. The lab's own model ranked 7th with a score of 0.660, though this figure is flagged because earlier test results influenced its training data. When scored only on the four tracks unaffected by this potential bias, the lab's model ranked 6th with a composite score of 0.678.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Google VP Yossi Matias says AI’s biggest impact may come from intersecting fields · 1 src
- Researchers adapt speech language model for simultaneous translation using prefix supervision · 1 src
- Study finds inductive prompting most consistent for LLM generalization in temporal extraction · 1 src
- AraBERT-based framework reaches 96.88% accuracy on Arabic DP ambiguity · 1 src
- Researchers introduce APDMem hierarchical memory for long-context LLM assistants · 1 src
Comments
via GitHub Discussions