DeReAct improves weaker AI agents with external gating policies
Researchers introduced DeReAct, a modular agent architecture that separates action validation and task completion control from the main LLM policy. The approach uses a Critic to validate actions before execution and a Context Manager to verify task completion. Testing on GAIA and SWE-bench Verified showed DeReAct improved Pass@1 scores by 6.5–7.0 points for Qwen3-Coder-480B and 4.2–5.2 points…
Key points
- DeReAct uses a Critic to validate actions and a Context Manager to certify task completion
- Pass@1 improved by 6.5–7.0 points for Qwen3-Coder-480B on GAIA and SWE-bench Verified
- Gains were 4.2–5.2 points for Claude Sonnet 4.5 and diminished with stronger models
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- Glow reports AI coding agents exposed 13,000 internal images on GitHub · 1 src
- Researchers introduce XiangqiBench to evaluate closed-loop performance of LLM agents in Chinese chess · 1 src
- AWS, Google Cloud launch AI agent tools; FTC investigates OpenAI, Anthropic · 1 src
- Microsoft's voice ai starts listening in 0.12 seconds · 1 src
- GitHub Trending shows eight AI agent harness projects dominate · 1 src
Comments
via GitHub Discussions