{"version":1,"type":"story","url":"https://digestai.news/story/ai-measurement-5-you-don-t-need-to-read-that-i-ve-done-it","json":"https://digestai.news/story/ai-measurement-5-you-don-t-need-to-read-that-i-ve-done-it.json","markdown":"https://digestai.news/story/ai-measurement-5-you-don-t-need-to-read-that-i-ve-done-it.md","slug":"ai-measurement-5-you-don-t-need-to-read-that-i-ve-done-it","headline":"[AI Measurement #5] You don't need to read that \"I've done it\"","summary":"The author examined 135 instances where Claude‑based models (Opus 5, Sonnet 5, Haiku 4.5) answered “I’ve done it” after attempting to fix small Python bugs. The study found that typical textual cues—politeness, length, assertiveness, or the number of cited grounds—did not predict correctness. In fact, longer explanations were slightly more likely to be wrong, and the presence of strong language such as “everything” or “completely” showed no reliable pattern.\n\nThe key insight is that the content of the AI’s reply is not a trustworthy signal. The author recommends three practical steps: ignore the reply text, trust only explicit “I’m not done” messages (which were always correct in the sample), and create a single hidden verification condition before asking the AI to act. If multiple AIs give divergent answers, treat the result as unreliable. This approach can be applied to code fixes, document generation, or any task where an AI claims completion.\n\nThe findings suggest that users should shift focus from reading AI explanations to designing concrete, pre‑defined checks that can be evaluated independently of the AI’s output.","keyPoints":["Politeness, length, or assertive language did not correlate with success","All 45 \"I'm not done\" replies were accurate; hidden checks proved reliable"],"whyItMatters":"Understanding that AI self‑reports are unreliable helps developers and businesses avoid costly rework and builds safer workflows around AI‑generated code or content.","category":{"slug":"agents","name":"Agents & Tools","url":"https://digestai.news/category/agents"},"entities":{"companies":["Anthropic"],"models":["Claude Opus 5","Claude Sonnet 5","Claude Haiku 4.5"],"people":[]},"firstPublishedAt":"2026-09-20T21:23:00Z","updatedAt":"2026-09-20T21:23:00Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"note.com","title":"[AI Measurement #5] You don't need to read that \"I've done it\"","url":"https://note.com/generativeai_new/n/n6ffcd8136de4?hl=en","publishedAt":"2026-09-20T21:23:00Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"[AI Measurement #5] You don't need to read that \"I've done it\"\", 20 September 2026, https://digestai.news/story/ai-measurement-5-you-don-t-need-to-read-that-i-ve-done-it","publisher":"Digest AI","title":"[AI Measurement #5] You don't need to read that \"I've done it\"","datePublished":"2026-09-20T21:23:00Z","url":"https://digestai.news/story/ai-measurement-5-you-don-t-need-to-read-that-i-ve-done-it"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}