Google researchers have developed a method to stop self-improving AI agents from memorising their own tests. The new approach from Google Cloud AI Research and several universities prevents this overfitting while also reducing compute costs.
In this article
Modern AI agents wrap a fixed language model in a harness. This framework includes prompts, workflows, tools, memory, and logic that control what the model sees at each step. The harness decides whether an agent reads the right file before changing it, whether it recovers from a mistake, and whether it delivers its results cleanly. A new research paper suggests much of the recent progress in agents comes from work on the harness, not from new models.
Until recently, this was done by hand. People reviewed failed runs and patched the harness manually. Newer methods automate the loop by having a language model rewrite the harness itself, again and again, based on feedback from the test tasks. The researchers call this a practical form of recursive self-improvement. The system produces feedback that it uses to optimize the harness, which in turn controls the system’s own behavior.
Self-optimization leads agents to memorize their test tasks
The paper shows that this self-optimization comes with a catch. Because the agent keeps working on the same limited set of test tasks, it ends up memorising them. Its scores on the training tasks go up, while gains on new, unseen tasks shrink or disappear entirely.
The researchers say this happens in several ways. The search memorises patterns that only fit one particular benchmark, favours candidates that score well purely by chance, and piles on unnecessary complexity that raises the test score without making the agent any better.
Other methods mostly improve on the training tasks, but little of that carries over to unseen benchmarks. RRSI raises scores there in all three domains. | Image: Google
Shrinking edit budgets and a strict critic keep the harness general
RRSI (Regularized Recursive Self-Improvement of Agent Harnesses) works on both ends of the optimization loop while leaving the harness fully editable. When the system proposes new changes, a budget caps how many independent edits a candidate can bundle at once.
That budget shrinks over time. Early rounds allow larger rewrites, while later rounds only permit small changes that can be clearly traced to a result. The system also keeps track of earlier attempts so it doesn’t keep chasing the same failed ideas. When progress stalls, it deliberately experiments with parts of the harness it hasn’t touched yet.
RRSI reins in self-optimization at two points, when proposing new changes and when deciding which of them become a permanent part of the harness. | Image: Google
When it comes to picking changes, a critic reviews every proposal and throws out any that hardcode task names, solutions, or other benchmark-specific tricks. Another rule only accepts higher compute costs if they come with a measurable performance gain. Components that no longer help get removed.
Giving up training gains pays off on new tasks
The researchers tested RRSI on eight benchmarks spanning coding, agentic office work, and engineering design. The underlying model, Claude Opus 4.8, stayed frozen throughout. The team compared RRSI with the unmodified baseline harness and four recent optimization methods.
According to the paper, RRSI gains up to 14.1 points on the tasks it was trained on and up to 4.7 points on five benchmarks it never saw. It also uses about 30 percent fewer tokens at runtime than the unregularized version. Overall performance never fell below the baseline on any of the unseen benchmarks, which typically happens with a harness that has memorised its tasks.
The RRSI harness also improves on every task outside the training set, with the biggest gain of 4.7 points on JobBench. | Image: Google
Every method did well on the training tasks, but the results flipped on new ones. Two methods even ended up below the baseline harness. RRSI posted the smallest training gain of all the variants and was the only method that landed well above the baseline on unseen tasks. The guardrails are meant to produce exactly this tradeoff.
Among the optimized harnesses, RRSI needs the fewest tokens and steps and performs best on new tasks, though the unmodified baseline harness is even leaner. | Image: Google
Harnesses optimized on one model also help weaker ones
A coding harness optimized with Gemini 3.5 Flash raised the accuracy of the much weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points without any modifications. The mechanisms the system found don’t depend on the capability of the model used to discover them.
The authors note that their study only covers harnesses built around frozen models and doesn’t address cases where the model weights change.
They conclude that self-improvement only makes AI agents reliably more capable when repeated feedback gets turned into lasting changes. The code is available on GitHub.
Manually designed harnesses often don’t generalize to new tasks, as tests on ARC-AGI-3 have shown. With a purpose-built harness, Opus 4.6 scored 97.1 percent in a familiar environment and 0 percent in an unfamiliar one. Nvidia recently presented SoL-Pi, a related method in which a research agent automatically rebuilds the harness of coding agents. It cuts token use by up to 49 percent without a noticeable drop in performance.
Shortly before that, Google had agents “dream” about past search runs to improve their search strategy. That work also leaves the model itself unchanged.




