THE CRUNCH

AI agents that rewrite their own working setup, known as a harness, tend to ace the tasks they practise on and stall on anything new. Researchers at Google Cloud AI Research and several universities propose a fix called RRSI, short for Regularized Recursive Self-Improvement of Agent Harnesses, which uses shrinking edit budgets and a strict critic to keep the optimisation honest. In tests across eight benchmarks, it delivered smaller training gains than rival methods but was the only one to hold up well on tasks it had never seen.

The catch is that self-optimising agents tend to memorise their test tasks. Because the loop keeps working on the same limited set of benchmarks, scores on training tasks climb while gains on new, unseen tasks shrink or vanish. The paper identifies several ways this happens: the search latches onto patterns that only fit one benchmark, favours candidates that score well by chance, and piles on complexity that raises test scores without making the agent genuinely better.

RRSI attacks the problem at both ends of the loop. A shrinking edit budget caps how many independent changes a candidate can bundle, allowing larger rewrites early on and only small, clearly traceable tweaks later. A critic then reviews every proposal and rejects anything that hardcodes task names, solutions or other benchmark-specific tricks, while a rule only accepts higher compute costs when they come with a measurable performance gain. The system also tracks earlier attempts to avoid chasing the same failed ideas, and deliberately experiments with untouched parts of the harness when progress stalls.

The researchers tested RRSI on eight benchmarks spanning coding, agentic office work and engineering design, with the underlying model, Claude Opus 4.8, kept frozen throughout. Against an unmodified baseline harness and four recent optimisation methods, RRSI gained up to 14.1 points on trained tasks and up to 4.7 points on five benchmarks it never saw, while using about 30 percent fewer tokens at runtime than the unregularised version. Performance never fell below the baseline on any unseen benchmark. The tradeoff was deliberate: every method did well on training tasks, but two methods ended up below the baseline harness on new tasks, and RRSI posted the smallest training gain of all variants while being the only one to land well above the baseline on unseen work.

The mechanisms also appear to transfer downwards: a coding harness optimised with Gemini 3.5 Flash lifted the much weaker Gemini 3.1 Flash Lite from 11.2 to 14.6 points without any modifications, which the authors say suggests the discovered mechanisms do not depend on the capability of the model used to find them.