DigestAI news desk

Cut through the AI noise.

Agents & Tools4 min read

Google researchers find way to prevent self-improving AI agents from memorizing tests

Self‑improving AI agents risk memorizing the tests they run, which can hurt performance on new tasks. A new paper from Google Cloud AI Research and several universities proposes Regularized Recursive Self‑Improvement of Agent Harnesses (RRSI) to curb this effect while also cutting compute costs.

1 source

Key points

  • RRSI boosts unseen benchmark scores by up to 4.7 points while cutting token use by 30%
  • RRSI harness uses Claude Opus 4.8 frozen model and outperforms baseline on 8 benchmarks
  • Harness optimization with Gemini 3.5 Flash raised Gemini 3.1 Flash Lite accuracy from 11.2 to 14.6

RRSI adds a budget that limits how many edits a candidate can bundle, shrinks that budget over time, and uses a critic to reject benchmark‑specific changes. In experiments on eight benchmarks, the method raised scores on training tasks by up to 14.1 points and on unseen tasks by up to 4.7 points, while using about 30 % fewer tokens at runtime.

The study also shows that harnesses optimized on one frozen model can help weaker models; a Gemini 3.5 Flash harness improved Gemini 3.1 Flash Lite accuracy from 11.2 to 14.6. Nvidia’s SoL‑Pi and Google’s prior “dream” work illustrate similar gains, suggesting harness design is a key lever for safe, efficient agents.

The story so far

2 episodes →
  1. Google researchers find way to prevent self-improving AI agents from memorizing teststhis story
Full story from The Decoder · by Jonathan KemperOpen source ↗

Google researchers find a way to keep self-improving AI agents from memorizing their tests

The Decoder · 4 October 2026

Loading the full article…

This text was published by The Decoder and written by Jonathan Kemper. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page
GoogleGoogle Cloud AI ResearchNvidiaClaude Opus 4.8Gemini 3.5 FlashGemini 3.1 Flash LiteOpus 4.6

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Agents & Tools

All →

Related stories