# Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement

Digest AI · Research · published 2026-09-28T04:00:00Z

Canonical: https://digestai.news/story/researchers-audit-llm-as-judge-in-text-to-sql-pipeline-find-low-agreem

## Summary

Researchers examined the agreement between an LLM-as-judge and human annotators in a production text-to-SQL pipeline. The deployed gpt-4o-mini judge achieved a Cohen's kappa of 0.04 on a disagreement‑enriched set and 0.42 on a uniform‑random spot‑check, over‑flagging 77.1% of human‑faithful cases. The over‑flags were largely due to a mechanism called GRADE‑HALLUCINATION.

A self‑hosted Qwen3.6‑27B replacement scored a kappa of 0.72, comparable to Claude Opus 4.7’s 0.71, while costing roughly 1/300 of the gpt‑4o‑mini per call. Ensembling did not improve results; pairing weak and strong judges reduced agreement. However, three strong judges under unanimity routing reached a kappa of 0.79 with 89.7% auto‑coverage. Applied out‑of‑domain, the audit flagged 25.5% of BIRD‑financial’s expert‑authored gold SQLs.

The study’s code and pre‑registration are available on GitHub (https://github.com/JamesL404/synca-audit).

## Key points

- gpt-4o-mini judge kappa 0.04 on enriched set, 0.42 on spot‑check, over‑flagging 77.1%
- Qwen3.6‑27B replacement kappa 0.72, costs 1/300 of gpt‑4o‑mini
- Three strong judges unanimity routing reach kappa 0.79 at 89.7% coverage

## Why it matters

Highlights reliability gaps in LLM‑based verification for database queries, affecting accuracy and cost in production systems.

## Sources

1. [Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline](https://arxiv.org/abs/2609.30290) (arXiv cs.CL, 2026-09-28, primary source)

## Cite

Digest AI, "Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement", 28 September 2026, https://digestai.news/story/researchers-audit-llm-as-judge-in-text-to-sql-pipeline-find-low-agreem

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/researchers-audit-llm-as-judge-in-text-to-sql-pipeline-find-low-agreem.json
