DigestAI news desk

Cut through the AI noise.

Research

Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement

Researchers examined the agreement between an LLM-as-judge and human annotators in a production text-to-SQL pipeline. The deployed gpt-4o-mini judge achieved a Cohen's kappa of 0.04 on a disagreement‑enriched set and 0.42 on a uniform‑random spot‑check, over‑flagging 77.1% of human‑faithful cases. The over‑flags were largely due to a mechanism called GRADE‑HALLUCINATION.

1 source primary source

Key points

  • gpt-4o-mini judge kappa 0.04 on enriched set, 0.42 on spot‑check, over‑flagging 77.1%
  • Qwen3.6‑27B replacement kappa 0.72, costs 1/300 of gpt‑4o‑mini
  • Three strong judges unanimity routing reach kappa 0.79 at 89.7% coverage

A self‑hosted Qwen3.6‑27B replacement scored a kappa of 0.72, comparable to Claude Opus 4.7’s 0.71, while costing roughly 1/300 of the gpt‑4o‑mini per call. Ensembling did not improve results; pairing weak and strong judges reduced agreement. However, three strong judges under unanimity routing reached a kappa of 0.79 with 89.7% auto‑coverage. Applied out‑of‑domain, the audit flagged 25.5% of BIRD‑financial’s expert‑authored gold SQLs.

The study’s code and pre‑registration are available on GitHub (https://github.com/JamesL404/synca-audit).

Read the original at arXiv cs.CL · by Haowei Liu, Hsin-Tai Wu, Yi Fang primary sourceOpen source ↗
Topics · follow one to build your own front page

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories