# HakemBench releases 2,346-item Turkish benchmark for typed decisions

Digest AI · Research · published 2026-10-05T04:00:00Z

Canonical: https://digestai.news/story/hakembench-releases-2-346-item-turkish-benchmark-for-typed-decisions

## Summary

Researchers have released HakemBench, a new benchmark designed to evaluate AI models on typed decision-making tasks in Turkish. Version 1.0 is fully open under a CC BY 4.0 license and contains 2,346 items across seven tracks, including fact-checking, legal routing, and customer support. The benchmark uses a specific harness to score decision quality, calibration, and selective automation, combining these metrics via a geometric mean.

A key methodological detail is that most gold labels are derived from blind passes of one AI model family compared against votes from other large language model families, rather than human verification. The authors note that these labels are not human-verified.

Initial results show a board of 16 models where the leading model achieved a composite score of 0.888. The lab's own model ranked 7th with a score of 0.660, though this figure is flagged because earlier test results influenced its training data. When scored only on the four tracks unaffected by this potential bias, the lab's model ranked 6th with a composite score of 0.678.

## Key points

- HakemBench v1.0 is released under CC BY 4.0 with 2,346 items and 4,275 questions.
- Gold labels are generated by AI model comparison, not human verification.
- The top model scores 0.888; the lab's model scores 0.660 overall and 0.678 on unbiased tracks.

## Why it matters

It provides a specialized, open-source tool for evaluating AI performance in Turkish, addressing a gap in non-English benchmarks. The transparent methodology regarding AI-generated labels offers a realistic, if imperfect, view of current model capabilities in decision-making tasks.

## Sources

1. [HakemBench: A Turkish Benchmark of Typed Decisions](https://arxiv.org/abs/2610.02293) (arXiv cs.CL, 2026-10-05, primary source)

## Cite

Digest AI, "HakemBench releases 2,346-item Turkish benchmark for typed decisions", 5 October 2026, https://digestai.news/story/hakembench-releases-2-346-item-turkish-benchmark-for-typed-decisions

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/hakembench-releases-2-346-item-turkish-benchmark-for-typed-decisions.json
