{"version":1,"type":"story","url":"https://digestai.news/story/researchers-propose-surrogate-log-probabilities-to-audit-llm-agents","json":"https://digestai.news/story/researchers-propose-surrogate-log-probabilities-to-audit-llm-agents.json","markdown":"https://digestai.news/story/researchers-propose-surrogate-log-probabilities-to-audit-llm-agents.md","slug":"researchers-propose-surrogate-log-probabilities-to-audit-llm-agents","headline":"Researchers propose surrogate log-probabilities to audit LLM agents","summary":"A new arXiv paper introduces a method to audit black-box LLM agents by using a low-cost, open-weight surrogate model. The approach addresses the problem that frontier chat APIs hide token probabilities, making it difficult to detect silent errors in tool calls or code before they execute. The surrogate runs in parallel, reading the same context and proposed action as the main agent, and scores the call using its own log-probabilities.\n\nThe method uses complementary readouts, including teacher forcing, request-PMI, and discriminative verdicts, to identify wrong argument values or holistically incorrect calls. It is training-free and requires no access to the agent's internals, costing only one prefill pass. On difficult coding tasks, the method achieved an AUROC of 0.825, significantly outperforming the actor's stated confidence, which was near chance at 0.598.\n\nThe paper demonstrates two deployment modes: a real-time gate that escalates low-confidence calls for review, and confidence feedback that allows the agent to adapt. These modes improved accepted-action accuracy and task success on live-execution benchmarks, with statistical significance reported at p <= 1e-4.","keyPoints":["Method uses open-weight surrogate log-probabilities to audit black-box LLM agents without accessing internals.","Achieved AUROC 0.825 on coding tasks, outperforming actor confidence (0.598) by +0.07 to +0.28.","Real-time gating and confidence feedback improved task success on live-execution benchmarks."],"whyItMatters":"This offers a practical, low-cost way to detect silent errors in deployed LLM agents, improving safety and reliability without requiring model transparency or retraining.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-10-06T04:00:00Z","updatedAt":"2026-10-06T04:00:00Z","sourceCount":3,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.AI","title":"Proxy Confidence: Auditing Black-Box LLM Agents with a Surrogate's Log-Probabilities","url":"https://arxiv.org/abs/2610.03894","publishedAt":"2026-10-06T04:00:00Z","type":"primary","primary":true,"lead":true},{"outlet":"arXiv cs.AI","title":"SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond Markdown","url":"https://arxiv.org/abs/2610.04008","publishedAt":"2026-10-06T04:00:00Z","type":"primary","primary":true,"lead":false},{"outlet":"arXiv cs.AI","title":"Teaching Agents to Code Reliably","url":"https://arxiv.org/abs/2610.03984","publishedAt":"2026-10-06T04:00:00Z","type":"primary","primary":true,"lead":false}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers propose surrogate log-probabilities to audit LLM agents\", 6 October 2026, https://digestai.news/story/researchers-propose-surrogate-log-probabilities-to-audit-llm-agents","publisher":"Digest AI","title":"Researchers propose surrogate log-probabilities to audit LLM agents","datePublished":"2026-10-06T04:00:00Z","url":"https://digestai.news/story/researchers-propose-surrogate-log-probabilities-to-audit-llm-agents"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}