Google researchers find a way to keep self-improving AI agents from memorizing their tests
Sonntag 14:40 · 04.10. · The Decoder EN

Plus AI research Copy the url to clipboard Share this article Go to comment section Google researchers find a way to keep self-improving AI agents from memorizing their tests Jonathan Kemper View the LinkedIn Profile of Jonathan Kemper Oct 4, 2026 Nano Banana Pro prompted by THE DECODER AI agents that keep optimizing their own working environment quickly tend to overspecialize on their test tasks. A new method from Google Cloud AI Research and several universities aims to prevent that while also cutting compute costs.
Modern AI agents wrap a fixed language model in a so-called harness , a framework of prompts, workflows, tools, memory, and logic that controls what the model sees at each step.
The harness decides whether an agent reads the right file before changing it, whether it recovers from a mistake, and whether it delivers its results cleanly. According to a new research paper, much of the recent progress in agents comes from work on the harness, not from new models.
Until recently, this was done by hand. People reviewed failed runs and patched the harness manually. Newer methods automate the loop by having a language model rewrite the harness itself, again and again, based on feedback from the test tasks.
The researchers call this a practical form of recursive self-improvement. The system produces feedback that it uses to optimize the harness, which in turn controls the system's own behavior.
The paper shows that this self-optimization comes with a catch. Because the agent keeps working on the same limited set of test tasks, it ends up memorizing them. Its scores on the training tasks go up, while gains on new, unseen tasks shrink or disappear entirely.
The researchers say this happens in several ways. The search memorizes patterns that only fit one particular benchmark, favors candidates that score well purely by chance, and piles on unnecessary complexity that raises the test score without making the agent any better.
Other methods mostly improve on the training tasks, but little of that carries over to unseen benchmarks. RRSI raises scores there in all three domains. | Image: Google
RRSI (Regularized Recursive Self-Improvement of Agent Harnesses) works on both ends of the optimization loop while leaving the harness fully editable. When the system proposes new changes, a budget caps how many independent edits a candidate can bundle at once.
That budget shrinks over time. Early rounds allow larger rewrites, while later rounds only permit small changes that can be clearly traced to a result. The system also keeps track of earlier attempts so it doesn't keep chasing the same failed ideas. When progress stalls, it deliberately experiments with parts of the harness it hasn't touched yet.
RRSI reins in self-optimization at two points, when proposing new changes and when deciding which of them become a permanent part of the harness. | Image: Google
When it comes to picking changes, a critic reviews every proposal and throws out any that hardcode task names, solutions, or other benchmark-specific tricks. Another rule only accepts higher compute costs if they come with a measurable performance gain. Components that no longer help get removed.
The researchers tested RRSI on eight benchmarks spanning coding, agentic office work, and engineering design. The underlying model, Claude Opus 4.8 , stayed frozen throughout. The team compared RRSI with the unmodified baseline harness and four recent optimization methods.
According to the paper, RRSI gains up to 14.1 points on the tasks it was trained on and up to 4.7 points on five benchmarks it never saw. It also uses about 30 percent fewer tokens at runtime than the unregularized version. Overall performance never fell below the baseline on any of the unseen benchmarks, which typically happens with a harness that has memorized its tasks.
The RRSI harness also improves on every task outside the training set, with the biggest gain of 4.7 points on JobBench. | Image: Google
Every method did well on the training tasks, but the results flipped on new ones. Two methods even ended up below the baseline harness. RRSI posted the smallest training gain of all the variants and was the only method that landed well above the baseline on unseen tasks. The guardrails are meant to produce exactly this tradeoff.
Among the optimized harnesses, RRSI needs the fewest tokens and steps and performs best on new tasks, though the unmodified baseline harness is even leaner. | Image: Google
A coding harness optimized with Gemini 3.5 Flash raised the accuracy of the much weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points without any modifications. The mechanisms the system found don't depend on the capability of the model used to discover them.
The authors note that their study only covers harnesses built around frozen models and doesn't address cases where the model weights change.
They conclude that self-improvement only makes AI agents reliably more capable when repeated feedback gets turned into lasting changes. The code is available on GitHub .
Manually designed harnesses often don't generalize to new tasks, as tests on ARC-AGI-3 have shown. With a purpose-built harness, Opus 4.6 scored 97.1 percent in a familiar environment and 0 percent in an unfamiliar one. Nvidia recently presented SoL-Pi , a related method in which a research agent automatically rebuilds the harness of coding agents. It cuts token use by up to 49 percent without a noticeable drop in performance.
Shortly before that, Google had agents "dream" about past search runs to improve their search strategy. That work also leaves the model itself unchanged.
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.