Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement
Researchers examined the agreement between an LLM-as-judge and human annotators in a production text-to-SQL pipeline. The deployed gpt-4o-mini judge achieved a Cohen's kappa of 0.04 on a disagreement‑enriched set and 0.42 on a uniform‑random spot‑check, over‑flagging 77.1% of human‑faithful cases. The over‑flags were largely due to a mechanism called GRADE‑HALLUCINATION.
Key points
- gpt-4o-mini judge kappa 0.04 on enriched set, 0.42 on spot‑check, over‑flagging 77.1%
- Qwen3.6‑27B replacement kappa 0.72, costs 1/300 of gpt‑4o‑mini
- Three strong judges unanimity routing reach kappa 0.79 at 89.7% coverage
A self‑hosted Qwen3.6‑27B replacement scored a kappa of 0.72, comparable to Claude Opus 4.7’s 0.71, while costing roughly 1/300 of the gpt‑4o‑mini per call. Ensembling did not improve results; pairing weak and strong judges reduced agreement. However, three strong judges under unanimity routing reached a kappa of 0.79 with 89.7% auto‑coverage. Applied out‑of‑domain, the audit flagged 25.5% of BIRD‑financial’s expert‑authored gold SQLs.
The study’s code and pre‑registration are available on GitHub (https://github.com/JamesL404/synca-audit).
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers release benchmark for AI in systematic review screening · 1 src
- SlideLab framework generates scientific presentations from research papers · 1 src
- Cartograph reduces AI agent tool discovery from O(n) to O(k) · 1 src
- Survey reviews 211 fake review detection studies from 2018 to 2026 · 1 src
- Researchers introduce hierarchical memory system for LLM agents · 1 src
Comments
via GitHub Discussions