Researchers propose surrogate log-probabilities to audit LLM agents
A new arXiv paper introduces a method to audit black-box LLM agents by using a low-cost, open-weight surrogate model. The approach addresses the problem that frontier chat APIs hide token probabilities, making it difficult to detect silent errors in tool calls or code before they execute. The surrogate runs in parallel, reading the same context and proposed action as the main agent, and scores…
Key points
- Method uses open-weight surrogate log-probabilities to audit black-box LLM agents without accessing internals.
- Achieved AUROC 0.825 on coding tasks, outperforming actor confidence (0.598) by +0.07 to +0.28.
- Real-time gating and confidence feedback improved task success on live-execution benchmarks.
The method uses complementary readouts, including teacher forcing, request-PMI, and discriminative verdicts, to identify wrong argument values or holistically incorrect calls. It is training-free and requires no access to the agent's internals, costing only one prefill pass. On difficult coding tasks, the method achieved an AUROC of 0.825, significantly outperforming the actor's stated confidence, which was near chance at 0.598.
The paper demonstrates two deployment modes: a real-time gate that escalates low-confidence calls for review, and confidence feedback that allows the agent to adapt. These modes improved accepted-action accuracy and task success on live-execution benchmarks, with statistical significance reported at p <= 1e-4.
Coverage and discussion
3sources- SkillScriptBench: Benchmarking Self-Evolution of Executable Agent Skill Packages Beyond MarkdownPrimary source · arXiv cs.AI ·
- Teaching Agents to Code ReliablyPrimary source · arXiv cs.AI ·
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Anthropic study: task understanding beats job title for AI success · 1 src
- Researchers test fixed token codes for language models at 100B-token scale · 1 src
- Baibaichuchu at the NTCIR-19 FinArg-3 Task: When Is Maximum Possible Profit Predictable from Investor Text? · 1 src
- Researchers introduce JEVal benchmark to test General decision models · 1 src
- Researchers release OncoNoteBERT for oncology note processing · 1 src
Comments
via GitHub Discussions