{"version":1,"type":"story","url":"https://digestai.news/story/experts-use-llm-disagreement-to-speed-codebook-revision","json":"https://digestai.news/story/experts-use-llm-disagreement-to-speed-codebook-revision.json","markdown":"https://digestai.news/story/experts-use-llm-disagreement-to-speed-codebook-revision.md","slug":"experts-use-llm-disagreement-to-speed-codebook-revision","headline":"Experts use LLM disagreement to speed codebook revision","summary":"Large‑scale text annotation relies on codebooks that AI annotators follow, but creating a robust codebook can take months. Researchers explored whether large‑language models (LLMs) could accelerate this process by flagging cases where different LLMs disagree, then directing human experts to those cases.\n\nThey tested three feedback strategies: (i) editing LLM‑generated revisions driven by cross‑LLM disagreement (Codebook Verifying), (ii) answering questions about LLM disagreements (Question Answering), and (iii) labeling disagreement cases with rationales (Rationale Labeling). Experiments on thousands of tutoring‑session transcripts showed that Rationale Labeling produced the highest LLM‑labeling accuracy, 64.9%, compared with the expert‑revised codebook at 57.8%. The best Question Answering setting achieved 60.5%.\n\nThe study concludes that LLMs can strategically target expert effort, cutting the time required for codebook revision from months to days while maintaining labeling performance.","keyPoints":["Rationale Labeling achieved 64.9% accuracy versus 57.8% for expert‑revised codebook","Question Answering reached 60.5% accuracy","Experiments used thousands of tutoring‑session transcripts"],"whyItMatters":"LLMs can focus human effort on the most contentious cases, reducing codebook development from months to days and improving annotation efficiency for large‑scale datasets.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-09-24T04:00:00Z","updatedAt":"2026-09-24T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation","url":"https://arxiv.org/abs/2609.26926","publishedAt":"2026-09-24T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Experts use LLM disagreement to speed codebook revision\", 24 September 2026, https://digestai.news/story/experts-use-llm-disagreement-to-speed-codebook-revision","publisher":"Digest AI","title":"Experts use LLM disagreement to speed codebook revision","datePublished":"2026-09-24T04:00:00Z","url":"https://digestai.news/story/experts-use-llm-disagreement-to-speed-codebook-revision"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}