Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

Disclosure: Some links in this article are affiliate links. AI Maestro may earn a commission if you make a purchase, at no…

By Vane September 29, 2026 3 min read
Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting

Google Research, alongside UNC-Chapel Hill, Stanford, and Washington University in St. Louis, has released RRSI (Regularized Recursive Self-Improvement). The system allows an LLM agent to rewrite its own prompts, tools, memory, and control flow without altering model weights. It constrains the improvement loop so performance gains hold on benchmarks the agent never optimised against.

The code is available under Apache 2.0 and requires Python 3.10 or higher. It accepts any LiteLLM model string, though defaults assume Claude Opus 4.8 running on Vertex AI.

Why self-improving harnesses overfit

Standard harness evolution loops propose edits, score them on a fixed set, and keep the winner. Because the same tasks are reused every round, the loop can memorise them. The RRSI research identifies three failure modes: benchmark-specific fitting, noise chasing, and complexity accumulation. Each one widens the gap between evolve-set scores and real transfer.

How RRSI works

RRSI keeps every harness component editable. It regularises how the search moves instead.

Proposal side

  • Annealed edit budget: a cosine schedule lets early rounds bundle several edits. Late rounds allow a single attributable change.
  • Evidence-aware credit: each candidate is logged with its component, hypothesis, diff, score change and cost change. The proposer reads this ledger, so falsified ideas are not retried.
  • Structured exploration: when progress stalls inside the noise band, budget shifts to components the run never touched.

Selection side

  • Leakage critic: rejects task names, entities, answers or benchmark-specific logic before any scoring.
  • Noise-adjusted floor: gains must clear the variance measured on the unchanged base harness.
  • Cost rule: extra inference tokens must be paid for by measured gain.
  • Pruning: components that stop producing gains become deletion targets.

The research team frame these as analogies to classic regularisers. The edit budget maps to L0, pruning to Lasso (L1) and the cost rule to Ridge (L2).

Results across 8 benchmarks

All 6 held-out splits improved. With Gemini 3.5 Flash as the policy, Terminal-Bench 2.1 rose from 64.6 to 78.7. SWE-bench Verified rose from 76.8 to 79.0.

The harness is also lighter. On the agentic workspace instance, RRSI uses 2.42M policy tokens per trial. Unregularised evolution uses 3.80M. The abstract reports this as 30% fewer; the project page says 36%.

RRSI vs closest competitors

Scores come from Table 1 of the RRSI research paper. All methods share the same starting harness, policy, evolve split and candidate budget.

FeatureRRSIMeta-HarnessAHETTHEHarnessX
Core ideaRegularised proposal and selectionAgentic proposer over code, scores and traces of all prior candidatesObservability-driven loop; edits paired with verified predictionsEvolves harness during test time, no gold labelsModular typed primitives, trace-driven adaptation
Model weightsFrozenFrozenFrozenFrozenFrozen
Cost rule and pruningYesNo*No*No*No*
Harvey LAB evolve score90.593.090.791.191.8
OOD average (H0 = 39.7)43.640.639.238.039.7

*Per the RRSI research team. OOD average is the mean of JobBench, GDPval and APEX-Agents, computed from Table 1.

Meta-Harness leads on the Harvey LAB evolve split. RRSI has the smallest evolve gain but the only OOD average more than 1 point above H0.

What it means

For developers building agentic workspaces, the practical change is a shift from blind iteration to constrained optimisation. Previously, an agent might spend dozens of rounds refining a prompt that works only on its training data. RRSI forces that agent to prove a gain is real and efficient before accepting it. The cost rule specifically prevents the agent from swapping a cheap tool for a complex one just to boost a score. This reduces the risk of deploying systems that look smart in testing but fail or become prohibitively expensive in production.

Scroll to Top