In this article
Google Deepmind researchers have released a technique called Dream-RSI that allows AI agents to refine their search strategies by replaying past attempts instead of running new, expensive computations.
The goal for self-improving AI is to discover new algorithms, solve mathematical problems, or write faster code without human intervention. These systems follow a standard loop: propose a solution, evaluate the result, learn from it, and try again. Over thousands of iterations, they aim for a better outcome.
Complex tasks create enormous search spaces. The agent must decide which approaches to pursue, which to run in parallel, and which to discard. This exploration phase determines whether the search succeeds or wastes compute on dead ends.
Dream-RSI improves those decisions by changing how the agent searches, not the underlying model. Current methods handle exploration in two ways. A fixed strategy cannot learn from experience, leading to repeated dead ends. Adapting the strategy avoids rigidity but costs time. Finding out if a strategy works requires many attempts, and testing countless alternatives means repeating long, expensive runs.
Replaying past searches makes new strategies cheaper to test
The team proposes reusing data from a completed search to test alternative strategies within the space the agent has already explored. The agent records its attempts and results as it searches, providing the data needed to replay those decisions later.
Researchers compare this to finding your way through an unfamiliar area. On a first visit, you hit dead ends, double back, and struggle to find a route. Once you have a mental map, though, you can plan another route without visiting every spot again.
Dream-RSI applies that principle to recorded search histories. Rather than testing a new strategy in a live run, the agent runs it against stored results. This lets it check what would have happened if it had pursued other approaches first or abandoned some earlier. The system does not invent entirely new solutions during replay; it tests different decisions within the recorded search tree.
Because those results already exist, the agent does not need to generate or evaluate solutions again, avoiding the expensive computations a live run would require. That makes testing new search strategies much cheaper. The researchers call this process “dreaming.” The agent plays through thousands of variations and selects the best one before putting it to work in a live search.
The process repeats in a loop. After each search, the agent uses the recorded results to test better strategies, then applies the improved version to its next live run. Throughout this cycle, only the search strategy changes; the model generating the solutions remains untouched.
Dream-RSI finds better solutions with fewer attempts
The researchers tested Dream-RSI with Gemini 3.1 Pro and Gemini 3.7 Flash on eight tasks across three areas. Each comparison used a baseline with the same starting conditions but a fixed search strategy.
One task asked the system to write the fastest possible program for a statistical calculation commonly used in genomics and finance. Dream-RSI’s program ran faster than the established libraries sklearn and glmnet on all six test datasets.
With Gemini 3.1 Pro, average runtime fell from 3,587 to 2,931 milliseconds, while the number of attempts dropped from 550 to 317. Dream-RSI also outperformed a competing system called SimpleTES, which needed 51,200 runs, compared with Dream-RSI’s 317 attempts.
The same pattern held for math optimization tasks and efforts to write efficient GPU kernels, with comparable or better results at much lower computational cost. On two GPU tasks, Dream-RSI matched performance while cutting the number of runs by a factor of up to 2.43. On two others, it delivered up to 2.09 times the performance within the same budget.
Explicit instructions can limit exploration
In a follow-up analysis, the researchers tested another way to use search histories. Instead of replaying them to test strategies, they condensed them into instructions telling the agent where to search.
On one GPU task, the version with these instructions performed worse than the version without them. The researchers suggest that overly specific directions can narrow the search space too much, preventing the agent from exploring a broader range of approaches.
The same analysis showed how the learned strategy adjusted its effort. As performance improved, it initially reduced the number of attempts. When progress stalled, it increased the search effort again, which coincided with further gains. The researchers have shared code and more details on GitHub.
Recursive self-improvement has drawn growing attention lately. Developments in this field are part of why Anthropic CEO Dario Amodei recently warned about the pace of AI research.
Google Deepmind introduced AlphaEvolve in 2025, using the same basic principle. Gemini Flash generates code proposals, Gemini Pro analyzes them, and an evolutionary algorithm selects the best versions. Dream-RSI works one level above that process by optimizing the search strategy itself.
AutoTTS takes a related approach, using a coding agent to search for algorithms in a simulated environment. These algorithms decide when a language model should start, expand, or abandon reasoning paths. The resulting methods beat manually designed methods while using less compute.
Google Research recently presented a different way to reuse past runs with WikiSkill. That system records failures and successes in a wiki and turns them into reusable instructions for the agent. Dream-RSI’s follow-up analysis suggests that explicit instructions like these can restrict exploration on open-ended search tasks.
Meta goes further with Hyperagents, allowing agents to rewrite the mechanism that controls how they improve.
What it means
Developers building these systems can now test new search strategies without paying the full computational cost of a live run. The agent learns to be smarter about how it searches before it even starts a new task.




