Experts use LLM disagreement to speed codebook revision
Large‑scale text annotation relies on codebooks that AI annotators follow, but creating a robust codebook can take months. Researchers explored whether large‑language models (LLMs) could accelerate this process by flagging cases where different LLMs disagree, then directing human experts to those cases.
Key points
- Rationale Labeling achieved 64.9% accuracy versus 57.8% for expert‑revised codebook
- Question Answering reached 60.5% accuracy
- Experiments used thousands of tutoring‑session transcripts
They tested three feedback strategies: (i) editing LLM‑generated revisions driven by cross‑LLM disagreement (Codebook Verifying), (ii) answering questions about LLM disagreements (Question Answering), and (iii) labeling disagreement cases with rationales (Rationale Labeling). Experiments on thousands of tutoring‑session transcripts showed that Rationale Labeling produced the highest LLM‑labeling accuracy, 64.9%, compared with the expert‑revised codebook at 57.8%. The best Question Answering setting achieved 60.5%.
The study concludes that LLMs can strategically target expert effort, cutting the time required for codebook revision from months to days while maintaining labeling performance.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Anthropic's Claude discovers new enzyme system ART resembling crispr · 10 src
- MIT researchers publish book on visual AI for urban studies · 1 src
- Research shows evidence quality boosts fact‑verification scores · 1 src
- AI transforms radiotherapy workflows while clinical adoption lags · 1 src
- SemiAnalysis releases ClusterMAX 3.0 rating 77 GPU cloud providers · 1 src
Comments
via GitHub Discussions