DigestAI news desk

Cut through the AI noise.

Research

HakemBench releases 2,346-item Turkish benchmark for typed decisions

Researchers have released HakemBench, a new benchmark designed to evaluate AI models on typed decision-making tasks in Turkish. Version 1.0 is fully open under a CC BY 4.0 license and contains 2,346 items across seven tracks, including fact-checking, legal routing, and customer support. The benchmark uses a specific harness to score decision quality, calibration, and selective automation,…

1 source primary source

Key points

  • HakemBench v1.0 is released under CC BY 4.0 with 2,346 items and 4,275 questions.
  • Gold labels are generated by AI model comparison, not human verification.
  • The top model scores 0.888; the lab's model scores 0.660 overall and 0.678 on unbiased tracks.

A key methodological detail is that most gold labels are derived from blind passes of one AI model family compared against votes from other large language model families, rather than human verification. The authors note that these labels are not human-verified.

Initial results show a board of 16 models where the leading model achieved a composite score of 0.888. The lab's own model ranked 7th with a score of 0.660, though this figure is flagged because earlier test results influenced its training data. When scored only on the four tracks unaffected by this potential bias, the lab's model ranked 6th with a composite score of 0.678.

Read the original at arXiv cs.CL · by Sait Furkan Teke (ufak AI) primary sourceOpen source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories