DigestAI news desk

Cut through the AI noise.

Research

Experts use LLM disagreement to speed codebook revision

Large‑scale text annotation relies on codebooks that AI annotators follow, but creating a robust codebook can take months. Researchers explored whether large‑language models (LLMs) could accelerate this process by flagging cases where different LLMs disagree, then directing human experts to those cases.

1 source primary source

Key points

  • Rationale Labeling achieved 64.9% accuracy versus 57.8% for expert‑revised codebook
  • Question Answering reached 60.5% accuracy
  • Experiments used thousands of tutoring‑session transcripts

They tested three feedback strategies: (i) editing LLM‑generated revisions driven by cross‑LLM disagreement (Codebook Verifying), (ii) answering questions about LLM disagreements (Question Answering), and (iii) labeling disagreement cases with rationales (Rationale Labeling). Experiments on thousands of tutoring‑session transcripts showed that Rationale Labeling produced the highest LLM‑labeling accuracy, 64.9%, compared with the expert‑revised codebook at 57.8%. The best Question Answering setting achieved 60.5%.

The study concludes that LLMs can strategically target expert effort, cutting the time required for codebook revision from months to days while maintaining labeling performance.

Read the original at arXiv cs.CL · by Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang primary sourceOpen source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories