{"version":1,"type":"story","url":"https://digestai.news/story/researchers-audit-llm-as-judge-in-text-to-sql-pipeline-find-low-agreem","json":"https://digestai.news/story/researchers-audit-llm-as-judge-in-text-to-sql-pipeline-find-low-agreem.json","markdown":"https://digestai.news/story/researchers-audit-llm-as-judge-in-text-to-sql-pipeline-find-low-agreem.md","slug":"researchers-audit-llm-as-judge-in-text-to-sql-pipeline-find-low-agreem","headline":"Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement","summary":"Researchers examined the agreement between an LLM-as-judge and human annotators in a production text-to-SQL pipeline. The deployed gpt-4o-mini judge achieved a Cohen's kappa of 0.04 on a disagreement‑enriched set and 0.42 on a uniform‑random spot‑check, over‑flagging 77.1% of human‑faithful cases. The over‑flags were largely due to a mechanism called GRADE‑HALLUCINATION.\n\nA self‑hosted Qwen3.6‑27B replacement scored a kappa of 0.72, comparable to Claude Opus 4.7’s 0.71, while costing roughly 1/300 of the gpt‑4o‑mini per call. Ensembling did not improve results; pairing weak and strong judges reduced agreement. However, three strong judges under unanimity routing reached a kappa of 0.79 with 89.7% auto‑coverage. Applied out‑of‑domain, the audit flagged 25.5% of BIRD‑financial’s expert‑authored gold SQLs.\n\nThe study’s code and pre‑registration are available on GitHub (https://github.com/JamesL404/synca-audit).","keyPoints":["gpt-4o-mini judge kappa 0.04 on enriched set, 0.42 on spot‑check, over‑flagging 77.1%","Qwen3.6‑27B replacement kappa 0.72, costs 1/300 of gpt‑4o‑mini","Three strong judges unanimity routing reach kappa 0.79 at 89.7% coverage"],"whyItMatters":"Highlights reliability gaps in LLM‑based verification for database queries, affecting accuracy and cost in production systems.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["OpenAI","Qwen","Anthropic"],"models":["gpt-4o-mini","Qwen3.6-27B","Claude Opus 4.7"],"people":["James L404"]},"firstPublishedAt":"2026-09-28T04:00:00Z","updatedAt":"2026-09-28T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"Auditing and Repairing LLM-as-Judge Failures in a Production Text-to-SQL Pipeline","url":"https://arxiv.org/abs/2609.30290","publishedAt":"2026-09-28T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement\", 28 September 2026, https://digestai.news/story/researchers-audit-llm-as-judge-in-text-to-sql-pipeline-find-low-agreem","publisher":"Digest AI","title":"Researchers audit LLM-as-judge in text-to-SQL pipeline, find low agreement","datePublished":"2026-09-28T04:00:00Z","url":"https://digestai.news/story/researchers-audit-llm-as-judge-in-text-to-sql-pipeline-find-low-agreem"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}