By Yunzhong, from Aofeisi
Quantum Bit | Official account QbitAI
Push aside a baffle, and the passage is sealed; bring over a ramp, and the high wall becomes possible to climb.
Every action changes the world, and the changed world becomes the starting point for the agent's next decision.
When such a world can be constructed from code and painted in real time by a diffusion model, AI gains a "training ground" it can explore over and over: act personally, observe consequences, verify hypotheses, and carry the experience into the next round.
This is AgentGarten, newly released by the MirroS team. It connects an executable code environment with a real-time neural renderer: code defines the physical rules, the model generates visual feedback, and with a throughput of over 30 fps, it lets agents continuously interact with the world.
In the classic hide-and-seek experiment, change happened quickly. The hider learned in round 4 to move the baffle to build a bunker, and the seeker mastered climbing over the wall via the ramp in round 10. After one failed jump, the seeker even pushed the ramp closer, adjusted its position, and tried again.
What drives the evolution of these strategies is the "experiment manual" that the agents write themselves.
At the end of each round, they review their attempts, record findings, and note open questions; in the next round, they continue exploring with this experience in hand.
Code makes the world runnable, real-time rendering makes the world visible, and continuous rehearsal keeps accumulating experience.
AgentGarten connects the three into a closed loop, exploring a bigger question:
When agents have a world in which they can act, make mistakes, and accumulate experience, how will intelligence continuously evolve within it?
Full blog post: https://mirros.ai/blog/worlds-for-evolving-agents
Technical report: https://mirros.ai/report/agent-garten.pdf
Code: https://github.com/MirroS-Lab/AgentGarten
Project homepage: https://mirros-lab.github.io/agent-garten
What agents lack is a world where they can practice over and over
If you want AI to truly understand the physical world, merely showing it massive amounts of video is far from enough. Just as you cannot learn to swim by only watching videos, agents must personally "get into the field", that is, take actions, observe the actual results, and then decide the next step.
This requires an environment with physical causal determinism: the box you push must stay where it was pushed, the passage you seal must become impassable, and all subsequent actions must be continuously constrained by these physical changes.
The two mainstream technical approaches each have their limitations:
Traditional code/game engine environmentsoffer precise states and transparent rules; every physical collision is clearly defined.
But the shortcoming is in the visuals: procedurally built scenes often feature only crude geometric gray-box models and repetitive textures. To give hundreds or thousands of open scenes realistic lighting and textures would require an astronomical amount of art modeling and engineering effort.
World models centered on video generation(such as various large video models) can certainly generate stunning high-definition imagery and produce subsequent frames in response to actions.
But the problem is that physical information is hidden entirely implicitly in the neural network's memory, with no explicit state that can be queried or intervened upon. If you want to check how much water is left in a cup, have a robotic arm tighten a screw, or set up a strict set of task referee rules, you are stuck in a pure video black box.
MirroS's solution is to:let each play its role: physics goes to code, and visuals go to the neural model.
Code drives the world forward, while the model takes care of 'how things look'.
In AgentGarten, a standard interaction loop works like this:
The agent issues an action command; the code environment immediately computes physical collisions and state updates, and exports a lightweight 'geometric sketch' (i.e., spatial depth or surface normals) from the agent's first-person camera perspective;
The neural renderer then reads this geometric sketch, combines it with prior visual memory, and instantly renders a photorealistic next frame to hand back to the agent.
△Figure 2 | The closed-loop flow between the code environment and the neural renderer. The agent issues action commands, the code world updates states and exports geometric conditions, and the neural renderer generates the next visual observation in real time.
This division of labor brings two major qualitative leaps:
First, physical rules are always "deterministic and controllable".
Scene layouts, collision detection, and game outcomes are all adjudicated by code, producing no "hallucinated physics".
Second, building entirely new worlds becomes extremely fast.
In the past, constructing 100 realistic simulation environments required a professional art team to model, texture, and light them, at a cost too high to bear. In AgentGarten, however, as long as simple 3D depth and normal outlines can be output, any minimalist code scene can directly apply this neural renderer, instantly gaining realistic lighting and materials. Wingsuit flying, robotic arm operation, kitchen cooking, multi-car racing—all share the same visual interface at the bottom layer. For the first time, the cost of expanding virtual environments has shifted from "art-labor-intensive" to "code-programmatic expansion".
△Figure 3 | Diverse physical worlds across tasks. Covering spatial navigation, object manipulation, camera re-shooting, and multi-car interaction. Left: simple white models exported from code; right: photorealistic visuals generated in real time by the renderer.
More crucially, this renderer isstreaming: for each action step, a segment of the scene is drawn; the agent sees the result clearly before making its next decision, making interaction flow as continuously as the real world.
Revisiting Hide-and-Seek: 25 million games versus 4 rounds
With the world built, the next question is the agent: by acting and reviewing repeatedly in such a world, can it improve on its own?
A classic reference is OpenAI's 2019 hide-and-seek experiment: hiders and seekers competed in an arena with boxes and ramps, and through large-scale self-play successively learned to build shelters and use ramps to climb walls.
The MirroS team redid this experiment in AgentGarten with a one-on-one setup: the hider arranges the scene first, then it is the seeker's turn to enter.
The rules of the contest are clear and pure: the hider first arranges the scene to block entrances, then the seeker comes in to search. The agents control movement and grabbing by writing Python code, making decisions solely from the images generated before their eyes, knowing nothing about spatial coordinates or the opponent's position.
Both sides start from completely blank manuals, each maintaining a policy library composed of individual skills. Each round consists of ten games; after the round, they review and consolidate, and in the next round a random subset of the skill library is applied.
Soon, those classic game strategies emerged one after another within a few rounds: moving boxes to block doors, carrying ramps to build bridges, and after failing to climb a wall, pulling the ramp closer to retry.
The 2019 research was famous for the emergence of these classic behaviors, and it provides an intuitive comparison (see Figure 4): in that study, the AI relied on reinforcement learning from scratch, taking about 25 million games to figure out shelter building, and about 100 million games to learn to use ramps to climb walls.
In AgentGarten, however, the agents, facing first-person visual frames, review and revise their textual manuals between rounds: the hider learned to build shelters in round 4, and the seeker mastered using ramps to climb walls in round 10.
△Figure 4 | Comparison of the scale of games played before representative strategies emerged. Self-play RL, trained from scratch on privileged physical states, required tens of millions to hundreds of millions of games; the MirroS agent, acting on first-person rendered frames, quickly mastered the corresponding strategies within a few rounds through round-by-round reflection on its manual.
Taking notes while playing: writing experience down
The rapid progress in Hide-and-Seek relies on a universal practice workflow.
MirroS's approach is tolet the agent take its own notes. Everything it learns is written into its manuals.
The practice is divided into four rigorous, progressive stages:
- Receiving the task: the agent receives a task document specifying the action goals and red lines of the rules, but with no standard answers.
- Blind-box trial and error: the agent acts autonomously under strict step and time limits. There is only the camera view—no god-mode coordinates, no minimap, and no score visible before the end of the game.
- Writing the manual: after a round ends, the agent reviews itself and honestly records what it tried, what phenomena it observed, which judgments are in doubt, and what new hypotheses should be verified next time.
- Archiving and passing on: the manual is archived and frozen, handed down directly as inherited prior knowledge to the agent of the next round.
Browsing through these manuals written by the agents, one finds them strongly colored by scientific inquiry. The models rigorously distinguish between "observed facts" and "made-up guesses," and even deliberately leave pitfall-avoidance guides for their successors:
- The hider distilled a practical truth: "The core of defense lies in blocking the passages, not merely blocking the line of sight", advising successors to use blind spots to reduce gaps, and to use small movements to test the blocking effect;
- The searcher worked out a control rule: "Before moving an object, first try pulling it; don't mistake stuck for gripped"—after approaching, press the micro-pull key once and see whether the object moves relative to the room's landmarks, thereby verifying whether it has truly been grasped;
- A bridge-crossing driver's warning is especially thought-provoking: "The target disc sliding into the blind spot at the front of the vehicle does not mean the car has safely arrived”。
Experience that withstood testing was adopted by successors, while failed lessons pushed them to try new strategy branches.
The same loop, run again across four worlds
Hide-and-seek is only a beginning. Because the code world is essentially a programmable software environment, the same "explore—review—consolidate" loop can easily be replicated in many more complex scenarios.
The team ran four rounds of drills in four other vastly different worlds:
- Puppy Companionship: Within 60 seconds, the agent must maintain the puppy's affection by reaching out to pet it and playing ball at the right moments. The score improved from 13 points in the first round to 19 points in the 4th round.
- Narrow Bridge Meeting: Two cars, from a first-person driving perspective, negotiate to yield to each other, passing each other on a narrow single lane and swapping ends in as short a time as possible. The total time for both cars to complete the pass was slashed from 71 seconds to 41 seconds.
- Cooperative Herding: Two sheepdogs, relying only on their own vision, work in tacit coordination to drive four sheep into the pen and keep them there for 5 seconds. From round 1, when time ran out with only three sheep penned, to the following three rounds of reliably penning all of them with "not one left out".
- Quarry Loader: A heavy loader must push rocks, feed material, and store it in the warehouse. In round 1 it barely managed to push the rocks, but by round 4 it could cleanly run through the entire workflow 31 seconds ahead of the 360-second time limit.
These vastly different tasks all illustrate the same thing: this experience-accumulation loop works just as well in entirely different worlds.
How the visuals keep up with the actions: how 30 fps on a single GPU was achieved
For an agent to "see and act" in a virtual world, the neural renderer must overcome three mountains that have long plagued generative video:keeping up with sudden movements, staying stable over ultra-long interactions, and running fast enough。
Starting from a base omnimodal model, the MirroS team proposed theAdversarial Forcingtraining method, and combined it with inference optimizations to open up this visual pipeline:
1. From generating entire clips offline to streaming block-by-block rendering
Video models used to generate an entire clip in one go; AgentGarten transformed this into streaming, block-by-block generation. When the agent makes a move, the model immediately produces the corresponding short block of visuals, so it can see the result before making the next decision.
2. No drift over long-horizon interactions (exact replay)
The longer generation runs continuously, the more easily tiny errors accumulate into appearance drift. Many models cut off the historical computation graph to save GPU memory, so gradients cannot flow back to historical memory; if the entire sequence is recomputed uniformly in a second pass, tiny differences in the underlying execution order introduce errors of several percentage points.
The team designed an exact replay mechanism: the first pass is a gradient-free forward rollout, and the second pass re-executes with gradients at exactly the same block-by-block pace, precisely reproducing the inference trajectory and stabilizing temporal consistency over long-sequence interactions.
3. Introducing real-video adversarial signals: avoiding visual degradation and artifacts
Relying solely on the model's own generated samples for score distillation, long-horizon inference easily becomes distorted and develops grid-like artifacts.
The team introduced a real-video discriminator during training to build an adversarial signal, anchoring on real physical image quality—both constraining the visual texture and making the generated results strictly conform to the geometric outlines given by the code.
4. Sustaining 30+ fps real-time interaction
The team hand-wrote custom Triton fused kernels, eliminating GPU memory round-trips and avoiding the minutes-long cold-start wait of traditional full-graph compilation; combined with CUDA Graph to eliminate CPU scheduling overhead, and swapping in an ultra-lightweight decoder with latency under 10 milliseconds, they finally stably supported closed-loop interaction at 480p resolution with throughput above 30 fps.
Preparing worlds that can be experienced for continuously evolving agents
Executable code makes training worlds endlessly programmatically generatable; the learned renderer gives these worlds perceptible, realistic physical feedback; and the accumulated experiment handbooks make every bout of fumbling a stepping stone for the next round of exploration.
Combining these three, the agent gains a space in which it can truly explore, err, and evolve.
This is the core idea behind MirroS's exploration of Physical RSI (Recursive Self-Improvement in the physical world).
Scaling for language models is in full swing, while the Scaling of intelligence in the physical world has only just begun.
True general embodied intelligence requires continuous growth through real back-and-forth interaction with the physical world:
Let every unknown become the starting point of the next round of evolution.
Full blog post: https://mirros.ai/blog/worlds-for-evolving-agents
*This article is republished by QbitAI with authorization; the views belong solely to the original authors.
