Google Research open-sources RRSI for self-improving AI agents without overfitting
Google Research, alongside UNC-Chapel Hill, Stanford, and Washington University in St. Louis, has released RRSI (Regularized Recursive Self-Improvement), an open-source framework that lets AI agents refine their own harness—prompts, tools, memory, and workflows—without altering model weights. The method prevents overfitting by constraining the improvement loop, ensuring gains transfer beyond the…
Key points
- RRSI lets AI agents edit prompts, tools, and workflows without changing model weights, preventing overfitting
- Benchmarks show Terminal-Bench 2.1 improved from 74.2% to 80.2% and SWE-bench Verified from 82.0% to 83.8%
- Open-source framework uses Apache 2.0 license, requires Python 3.10+, and defaults to Claude Opus 4.8 on Vertex AI
RRSI introduces mechanisms like an annealed edit budget, evidence-aware credit, and a leakage critic to filter out memorization and noise. Benchmark results show improvements across eight datasets, including Terminal-Bench 2.1 (74.2% → 80.2%) and SWE-bench Verified (82.0% → 83.8%), with out-of-distribution gains on JobBench (+4.7), GDPval (+3.5), and APEX-Agents (+3.7). The framework reduces policy token usage by 30–36% compared to unregularized evolution. The code is Apache 2.0-licensed and works with any LiteLLM-compatible model, defaulting to Claude Opus 4.8 on Vertex AI.
Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
MarkTechPost · 29 September 2026
Loading the full article…
This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- Manus launches Manus 2.0 and Cue with standalone email, phone, wallet · 2 src
- Google to replace Gemini Gems with skills by November 17 · 4 src
- Instinct founder says personal agents could challenge Big Tech in months · 1 src
- H company launches Holo4 agentic models for desktop and API tasks · 3 src
- Researchers release EmailBench to test AI agents on email tasks · 1 src
Comments
via GitHub Discussions