DigestAI news desk
Research updated

Study Finds Deictic Ambiguity Can Undermine Draft‑Verify‑Revise LLM Pipelines

Researchers examined a common three‑stage LLM orchestration pattern—draft, verify, revise—and discovered that context‑dependent expressions like “previous” can shift meaning between stages, a problem they call a deictic shift. Using a synthetic dataset of ten base examples rendered in three conditions, they varied which stage resolved the expression correctly and how much independent reasoning…

1 source primary source

Key points

  • Deictic shifts cause “previous” to refer to different items across draft‑verify‑revise stages.
  • GPT‑5.2’s accuracy rose from 0.156 to 0.942 with increased reasoning effort.
  • Gemini 3 Pro maintained >0.94 accuracy, outperforming GPT‑5.2 at lower cost for 5% of trials.

Six models from three providers (including GPT‑5.2 and Gemini 3 Pro) were evaluated across 21 reasoning‑effort configurations. Balanced accuracy ranged from 0.156 (below chance) to near‑perfect, with GPT‑5.2 improving from 0.156 without reasoning to 0.942 at its highest effort, while Gemini 3 Pro stayed above 0.94 throughout. The study also found that when the meta‑evaluator erred, it relied on surface cues rather than deeper reasoning. The authors advise engineers to make referents explicit in each pipeline stage to avoid costly misinterpretations.

Read the original at arXiv cs.AI · by Obinna I. Ekekezie primary source Open source ↗
Topics · follow one to build your own front page
OpenAIGoogleGPT-5.2Gemini 3 Pro

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Research

All →

Related stories