Nvidia researchers improve AI agent reliability with a judging model
Nvidia researchers introduced Mid-Harness, a method to improve AI agent reliability by having a generator model propose multiple actions and a separate judge model select the best one for execution. On the TerminalBench-Lite benchmark, the approach increased the first-try success rate from 50.00% to 68.03% when pairing a TMAX-9B generator with a GPT-5.6 Sol verifier. The method aims to reduce…
Key points
- Mid-Harness increases agent success rate from 50.00% to 68.03% on TerminalBench-Lite benchmark
- Method uses a generator model to propose actions and a judge model to select the best one
- Researchers found same-model verification (TMAX-9B) outperformed external judge models like GPT-5.6 Sol
The paper, published on December 16, 2026, by authors including Minki Kang and Ehsan Hosseini-Asl, also found that using the same model as both generator and verifier yielded the best results. Mid-Harness outperformed trajectory scaling—repeating entire task attempts—in both effectiveness and token efficiency. Nvidia’s work follows earlier releases like the ACES framework (August 2026) and the Open Agent Safety Platform (September 2026), reflecting its focus on agent reliability. While the improvement is significant, real-world testing remains needed to validate performance across varied environments.
The story so far
2 episodes →- Nvidia researchers improve AI agent reliability with a judging modelthis story
Nvidia researchers improve AI agent reliability with a judging model
cryptobriefing.com · 1 October 2026
Loading the full article…
This text was published by cryptobriefing.com and written by Diego Almada Lopez. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- Amazon releases Strands Decider 2B, a Jev-inspired open-source decision model · 2 src
- Decagon launches Voice 3 agent with Chord speech model for multilingual calls · 2 src
- NVIDIA neom agent toolkit uses Amazon s3 vectors for agent memory · 1 src
- OpenAI launches dots agent to handle ongoing tasks without human input · 2 src
- Anthropic deploys Claude buying agent to handle inbound sales · 2 src
Comments
via GitHub Discussions