I run a team of 6 AI agents, going 24 hours a day. It sounds cool, but once you’re actually doing it, you realize the most exhausting part isn’t that they’re not smart enough — it’s that they often don’t remember anything.
The same pitfall it fell into yesterday, it falls into again today on a different task. Last round it already figured out that a certain path doesn’t work, and the next round it charges right down that same path again, just as eagerly. Every single time, it’s like its first day on the job, having completely forgotten everything I taught it last week. You’re not managing an employee who accumulates experience — you’re managing an intern who’s permanently stuck on day one, just one that types very fast.
So when I saw Dream-RSI, released by Google’s research team, I stared at it for a long time. Because it addresses exactly the thing I wrestle with every single day.
Agents Are Starting to “Dream”
Let me be clear about what it actually does, because the name sounds poetic but the mechanism is pretty practical.
A typical agent works like this: it gets an instruction, tries one path, succeeds or fails, and in most cases, that experience just evaporates. Next time around, it doesn’t really remember why it got stuck last time, or which step turned out to be a waste of effort.
Dream-RSI adds one key step: after the agent finishes a round of a task, instead of rushing straight into the next one, it “dreams” first.
It stores every past attempt and its outcome — this one succeeded, that one failed, this step took a detour. Then it uses these old records to work things out: would trying a different path first have gone better? When should it give up earlier instead of burning more time? Should it run several approaches in parallel instead of committing to one path all the way through?
Here’s the part that made me go “oh, so that’s how it works”: because these outcomes already happened in real runs, this replaying doesn’t require re-executing the task at all. It’s replaying inside its own head, not actually redoing the work. It’s like someone who just finished a full day of chess, lying in bed that night replaying every move from the day — what if I’d gone the other way on move seven? He doesn’t need to actually play out a whole new game; in his head, he’s already walked through several possible branches.
After replaying, the agent picks out the approaches that performed better and applies them directly to the next round. And that new round leaves behind even more successes and failures, so it “dreams” again, and adjusts its strategy again. Round after round, it gets more accurate. That’s what the paper calls “recursive self-improvement” — it improves itself, and it does so continuously, round by round.
From 550 to 317 — Nearly 40% Less Work for the Same Result
Talking about the concept alone can feel abstract, so Google’s research team tested this across 8 tasks spanning algorithms, mathematical optimization, and GPU kernels. One example in particular really stuck with me.
Take one algorithm task with Gemini 3.1 Pro: compared to sticking with a fixed approach, once this dream-and-replay mechanism was applied, the agent’s call count dropped from 550 to 317 — and the program it ultimately found ran even faster.
Put that number into a real-world context. The same job, done with nearly 40% less work, and the result wasn’t just not worse — it was better. For someone like me who watches how many calls and tokens agents burn through every single day, that’s not an abstract academic number — it’s real cost and real time. Cutting 40% of the back-and-forth means 40% less waiting, 40% less on the bill, and a better answer to show for it.
A dreaming agent isn’t working harder — it’s better at picking its path.
What’s Really Impressive Is the Method Compounding
The real insight that made me sit up while reading this wasn’t “the agent got smarter” — it was how it got smarter.
Notice something easy to skim past: throughout this whole process, the underlying model itself doesn’t change. Google didn’t swap in a bigger brain, and there was no retraining. What got improved was the agent’s method, its strategy, for figuring out “what to do next.”
In other words, the progress isn’t about “which brain you’re using” — it’s about “how you go about doing things.”
That might sound like a minor technical detail to most people, but for anyone actually running an agent team, actually relying on agents to get work done, it’s almost a shift in how you think about the whole problem. Our corner of the world has a very strong reflex: if results aren’t good enough, swap in a bigger model. New model comes out? Upgrade right away. As if the only path to progress is chasing an ever more powerful brain.
Dream-RSI offers another path: let the method itself compound. The same brain, as long as it can remember, replay, and turn last round’s lessons into this round’s strategy, can keep getting better round after round — and that improvement accumulates, snowballs. That’s a much better deal than “swap in a bigger model,” because switching models is a one-time jump, while compounding method gets stacked, layer on layer, every single round.
The longer I spend running agents, the more convinced I am of one thing: an ordinary agent that remembers its lessons will, over time, beat an agent with a goldfish memory but a huge brain. Just like with people. The strongest person you know usually isn’t the smartest one — it’s the one who’s best at replaying and learning from what happened.
To Be Honest, It’s Not “Plug and Play” Yet
I don’t want to make it sound like something you can bolt onto your agent tomorrow afternoon — that wouldn’t be honest.
Right now, Dream-RSI is a research result, tested on specific tasks like algorithms, mathematical optimization, and GPU kernels — tasks that have clear right and wrong answers and produce measurable outcomes. These tasks share one thing in common: you can clearly log “did this succeed or fail, how fast did it run” — which is exactly what gives the agent something to “dream” about. It’s not yet something you can hand off and expect to work universally across every scenario.
So think of this as a direction well worth keeping an eye on, not a tutorial you can copy this afternoon. The direction being right doesn’t mean the road is already paved. There’s still distance to cover here, and the vaguer a task is — the less it has a clear right answer — the more research is still needed to figure out how this path should even work.
But direction, in itself, often matters more than whether something is usable right now. Because it tells you where the smart people are putting their effort.
One Idea Worth Taking With You
If you’re using agents too, or getting ready to have agents do work for you, what I want to leave you with isn’t a tool — it’s a shift in the question you ask.
We’re too used to asking: “Which more powerful model should I switch to?”
Dream-RSI makes me want to ask a different question instead: “How do I get my agent to remember the lesson from last time?”
These two questions take you to completely different places. Chase the model, and you’re forever waiting on the next release, forever spending money to upgrade, and it never quite feels like enough no matter how much you upgrade. Chase the method, and you’re building something that keeps getting better on its own — it might be pretty dumb today, but every round, it stacks a little more on top.
I don’t know how far the Dream-RSI path will ultimately go. But it at least helped me think through one thing clearly: instead of constantly swapping brains, it’s worth first figuring out how to stop the brain you already have from starting over at zero every time.
The one that dreams isn’t necessarily the stronger one. But it’s definitely the one less likely to walk the same path in vain.
Source: Google’s research team released Dream-RSI. Related coverage: https://m.theblockbeats.info/flash/367592