{"version":1,"type":"story","url":"https://digestai.news/story/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model","json":"https://digestai.news/story/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model.json","markdown":"https://digestai.news/story/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model.md","slug":"nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model","headline":"Nvidia researchers improve AI agent reliability with a judging model","summary":"Nvidia researchers introduced Mid-Harness, a method to improve AI agent reliability by having a generator model propose multiple actions and a separate judge model select the best one for execution. On the TerminalBench-Lite benchmark, the approach increased the first-try success rate from 50.00% to 68.03% when pairing a TMAX-9B generator with a GPT-5.6 Sol verifier. The method aims to reduce failures in stochastic environments like command-line interfaces where a single bad command can derail a task.\n\nThe paper, published on December 16, 2026, by authors including Minki Kang and Ehsan Hosseini-Asl, also found that using the same model as both generator and verifier yielded the best results. Mid-Harness outperformed trajectory scaling—repeating entire task attempts—in both effectiveness and token efficiency. Nvidia’s work follows earlier releases like the ACES framework (August 2026) and the Open Agent Safety Platform (September 2026), reflecting its focus on agent reliability. While the improvement is significant, real-world testing remains needed to validate performance across varied environments.","keyPoints":["Mid-Harness increases agent success rate from 50.00% to 68.03% on TerminalBench-Lite benchmark","Method uses a generator model to propose actions and a judge model to select the best one","Researchers found same-model verification (TMAX-9B) outperformed external judge models like GPT-5.6 Sol"],"whyItMatters":"Mid-Harness offers a cost-effective way to improve agent reliability by reducing failed commands without retraining. Lower token costs and potential for self-verification could accelerate adoption in production environments.","category":{"slug":"agents","name":"Agents & Tools","url":"https://digestai.news/category/agents"},"entities":{"companies":["Nvidia"],"models":["TMAX-9B","GPT-5.6 Sol"],"people":["Minki Kang","Ehsan Hosseini-Asl"]},"firstPublishedAt":"2026-10-01T10:56:00Z","updatedAt":"2026-10-01T10:56:00Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"cryptobriefing.com","title":"Nvidia researchers improve AI agent reliability with a judging model","url":"https://cryptobriefing.com/nvidia-mid-harness-ai-agent-reliability","publishedAt":"2026-10-01T10:56:00Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"New Methods Boost AI Agent Reliability","url":"https://digestai.news/thread/researchers-propose-twincheck-to-verify-ai-agent-tool-calls","storyCount":2},"cite":{"text":"Digest AI, \"Nvidia researchers improve AI agent reliability with a judging model\", 1 October 2026, https://digestai.news/story/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model","publisher":"Digest AI","title":"Nvidia researchers improve AI agent reliability with a judging model","datePublished":"2026-10-01T10:56:00Z","url":"https://digestai.news/story/nvidia-researchers-improve-ai-agent-reliability-with-a-judging-model"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}