DigestAI news desk

Cut through the AI noise.

Agents & Tools4 min read

Google Research open-sources RRSI for self-improving AI agents without overfitting

Google Research, alongside UNC-Chapel Hill, Stanford, and Washington University in St. Louis, has released RRSI (Regularized Recursive Self-Improvement), an open-source framework that lets AI agents refine their own harness—prompts, tools, memory, and workflows—without altering model weights. The method prevents overfitting by constraining the improvement loop, ensuring gains transfer beyond the…

1 source

Key points

  • RRSI lets AI agents edit prompts, tools, and workflows without changing model weights, preventing overfitting
  • Benchmarks show Terminal-Bench 2.1 improved from 74.2% to 80.2% and SWE-bench Verified from 82.0% to 83.8%
  • Open-source framework uses Apache 2.0 license, requires Python 3.10+, and defaults to Claude Opus 4.8 on Vertex AI

RRSI introduces mechanisms like an annealed edit budget, evidence-aware credit, and a leakage critic to filter out memorization and noise. Benchmark results show improvements across eight datasets, including Terminal-Bench 2.1 (74.2% → 80.2%) and SWE-bench Verified (82.0% → 83.8%), with out-of-distribution gains on JobBench (+4.7), GDPval (+3.5), and APEX-Agents (+3.7). The framework reduces policy token usage by 30–36% compared to unregularized evolution. The code is Apache 2.0-licensed and works with any LiteLLM-compatible model, defaulting to Claude Opus 4.8 on Vertex AI.

Full story from MarkTechPost · by Asif RazzaqOpen source ↗

Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

MarkTechPost · 29 September 2026

Loading the full article…

This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page
Google ResearchUNC-Chapel HillStanfordWashington University in St. LouisClaude Opus 4.8Gemini 3.5 FlashAsif Razzaq

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Agents & Tools

All →

Related stories