Google, in collaboration with Google DeepMind, the University of Maryland, and the University of Virginia, published the preprint Dream-RSI: Recursive Self-Improvement through Evolving Worlds (arXiv:2609.14858) and a project page with an interactive demonstration. The method stores the history of past search agent runs as a tree of attempts and replays it as an exact simulator: thousands of alternative exploration strategies are evaluated offline, without a single new agent call. A separate LLM agent rewrites only the strategy code — the model, agent, and evaluator remain unchanged. On Gemini 3.1 Pro, this sped up solving the Lasso task from 3587 to 2931 ms with 317 agent calls instead of 550. Until the code is released, the material remains a preprint: the repository is announced, but the code itself is still being prepared for release.



What happened
The essence of the method: the history of past search agent runs is saved as a tree of attempts, where each node is a version of the solution, its diagnostics, and evaluation, and the entire tree works as an exact replay simulator. On it, thousands of alternative exploration strategies are run without a single new agent call — with a different order of branches, different degrees of parallelism, and different stopping points. A separate LLM agent repeatedly rewrites the strategy code, each version is evaluated on all accumulated trees, and since the current strategy also participates in the selection, the new one is guaranteed to be no worse than the old one. The model, the agent itself, and the evaluator do not change. In experiments on Gemini 3.1 Pro, the version with Dream-RSI sped up solving Lasso from 3587 to 2931 ms with 317 agent calls instead of 550; the baseline speed of VGG16 and LayerNorm from KernelBench was achieved in 2.43 and 1.79 times fewer attempts, respectively; for convolution+division and convolution+max combinations, code faster than the baseline was found — 2.09 times faster in the first case. The zhengkid/Dream-RSI repository has been announced, but the code is still being prepared for release.
Context
The method addresses a bottleneck at the meta-level in agent search. To check whether a new exploration strategy is really better than the previous one, it usually requires running the entire discovery cycle again, and such feedback costs as much as the research itself. Dream-RSI is based on the observation that logs of already completed runs contain everything necessary for cheap verification: if you record not only the result, but also each attempt with diagnostics and evaluation, they can be replayed as a deterministic simulator. The authors call this the agent's "dreaming" — the world is not launched again, but is reproduced from the recorded branches, hence Evolving Worlds in the title of the work. This approach also shifts the debate about recursive self-improvement: it is not the set of model weights that is improved, but the behavior strategy on top of a fixed model. In this light, the logs of agent search runs turn out to be an undervalued asset that can be turned into an exact offline simulator and optimize the strategy on it with thousands of runs.
Why this matters for the industry
For the industry, the main shift is economic. Meta-level feedback usually requires re-running the entire discovery cycle, and it is precisely this that makes recursive self-improvement expensive; Dream-RSI turns already paid run logs into a free exact simulator, so that one expensive online run pays for thousands of offline policy evaluations. The most expensive part of an agent product — iteration on orchestration — becomes sharply cheaper, and the advantage goes to teams that are already accumulating detailed run logs. Since only the strategy code is self-improving — which branches to develop, how many attempts to run in parallel, when to stop — the method remains an add-on on top of an unchanged model, agent, and evaluator, and therefore is potentially transferable to existing agent systems without retraining the base model. If transferability is confirmed, the "replay optimization" pattern may become a separate layer of agent stacks: an offline strategy optimizer as a service, a replay module in agent observability tools, or a "improve strategy on my logs" button in product assistants. This is a forecast, not a fixed fact, and it depends, among other things, on the release of open code.
Why this matters for users
For those who build agent search and work with test-time compute, the method directly hits the most expensive line item — meta-level feedback: new strategies can be checked on recorded attempt trees, not with new agent calls. The practical benefit for teams is already methodological today: it is worth starting to deterministically save the attempt trees of your agents — the version of the solution, diagnostics, and evaluation — so that by the time the open code appears, you have ready material for offline optimization, and the pattern itself can be reproduced on your own logs with standard tools. End users of agent services will benefit later and indirectly: if the approach takes root, products will be able to improve the search strategy without retraining the model and find solutions in fewer agent calls, which over time translates into faster and cheaper service operation. For now, this is a preprint, not a product: there is no API and pricing, and you can only use your own implementation of the pattern.
What is still unknown / limitations
The method currently exists only as a preprint: the code in the zhengkid/Dream-RSI repository is in the status of code release in preparation, there is no API and pricing, there are no independent replications, so the results cannot be reproduced immediately. The "no worse" guarantee is a guarantee on the replay simulator: it works on a composite metric — the best result found, a penalty for the number of attempts, a bonus for parallelism — and only within the recorded branches; in a live environment, there is no such guarantee, and the replay is exact only where the branches were fixed. Experiments are shown on Gemini 3.1 Pro in narrow domains — the Lasso solver and kernels from KernelBench — and the transferability of the method to other models and tasks is not confirmed. Finally, the price of the meta-cycle itself is not disclosed: how much it costs to store trees and offline evaluate thousands of strategies — this is a key question for practical implementation.
Sources
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds — project page with interactive demonstration
- Dream-RSI: Recursive Self-Improvement through Evolving Worlds — arXiv:2609.14858 preprint
- zhengkid/Dream-RSI — official project repository on GitHub
Author
Look at AI, editorial team
