量子位官网 · 10/9/2026, 10:03:48
MirroS Releases AgentGarten: Combining Code-Defined Physics with Diffusion Rendering for Embodied Agent Evolution
The MirroS team has released and open-sourced AgentGarten, an embodied AI training framework that connects executable code environments with real-time neural renderers. By defining precise physical causal rules via code and generating realistic visual feedback at over 30 fps using diffusion models, it resolves the conflict between the crude visuals of traditional engines and the uncontrollable physics of video-based world models. In hide-and-seek experiments, agents utilized an 'experiment manual' mechanism to record trial-and-error experiences, successfully emerging complex strategies like moving barriers and building ramps, validating the feasibility of Physical Recursive Self-Improvement (RSI).
SOURCE COVERAGEOriginal coverage
Contents7 sections
By Yunzhong, from Ao Fei Si QbitAI | Official Account: QbitAI
Push aside a barrier, and the passage is sealed; bring in a ramp, and the high wall becomes surmountable.
Every action changes the world, and that changed world becomes the starting point for the agent’s next decision.
When such a world can be constructed by code and rendered in real-time by diffusion models, AI gains a "training ground" it can explore repeatedly: act personally, observe consequences, verify hypotheses, and carry experience into the next round.
This is AgentGarten, recently released by the MirroS team. It connects an executable code environment with a real-time neural renderer: code defines physical rules, while the model generates visual feedback, enabling agents to continuously interact with the world at throughputs exceeding 30 fps.
In the classic hide-and-seek experiment, changes happened quickly. By the 4th round, the hider learned to move barriers to build cover; by the 10th round, the seeker mastered using ramps to climb over walls. After a failed jump, the seeker would actively push the ramp closer, adjust its position, and try again.
Driving this evolution of strategies were the "experiment manuals" written by the agents themselves.
At the end of each round, they reviewed attempts, recorded findings, and noted questions; in the next round, they continued exploring armed with this accumulated experience.
Code makes the world runnable, real-time rendering makes the world visible, and continuous practice allows experience to accumulate.
AgentGarten closes the loop among these three elements, exploring a larger question:
When agents possess a world where they can act, make mistakes, and accumulate experience, how does intelligence continuously evolve within it?
Full Blog: https://mirros.ai/blog/worlds-for-evolving-agents Technical Report: https://mirros.ai/report/agent-garten.pdf Code: https://github.com/MirroS-Lab/AgentGarten

△ Figure 1 | Three typical game strategies emerged during matches. Moving barriers to seal covers, setting up ramps to scale high walls, and actively adjusting and retrying after an initial jump failure.
What Agents Lack Is a World Where They Can Practice Repeatedly
To truly understand the physical world, simply showing AI massive amounts of video is far from enough. Just as one cannot learn to swim merely by watching recordings, agents must "step onto the field" themselves—taking actions, observing actual results, and then deciding on the next step.
This requires an environment with deterministic physical causality: pushed boxes must remain in place, blocked passages must remain impassable, and all subsequent actions must be continuously constrained by these physical changes.
The two current mainstream technical paths each have limitations:
Traditional code/game engine environments offer precise states and transparent rules, with every physical collision clearly defined.
However, the shortcoming lies in visuals: procedurally built scenes often consist only of rough geometric white-boxes and repetitive textures. Achieving realistic lighting and material quality for hundreds or thousands of open scenes requires astronomical costs in artistic modeling and engineering effort.
World models centered on video generation (such as various large video models) can indeed generate stunning high-definition imagery and produce subsequent frames based on actions.
But the problem is that physical information is entirely implicit, hidden within the neural network's memory, failing to provide a queryable, intervenable explicit state. Checking how much water remains in a cup, tightening a screw with a robotic arm, or establishing strict task adjudication rules are impossible within a pure video black box.
MirroS’s solution is to let each component do what it does best: physics to code, and light/shadow to neural models.
Code Drives the World, Models Handle "How It Looks"
In AgentGarten, a standard interaction loop operates as follows:
The agent issues an action command; the code environment immediately calculates physical collisions and state updates, exporting a lightweight "geometric sketch" (spatial depth or surface normals) from the agent’s first-person camera perspective.
Subsequently, the neural renderer reads this geometric sketch, combines it with previous visual memory, and instantly renders a photorealistic next frame to present to the agent.

△ Figure 2 | The closed-loop flow between the code environment and the neural renderer. The agent issues action commands, the code world updates state and exports geometric conditions, and the neural renderer generates the next visual observation in real-time.
This division of labor brings two qualitative leaps:
First, physical rules remain "deterministic and controllable."
Scene layout, contact detection, and win/loss conditions are adjudicated by code, producing no "hallucinated physics."
Second, building new worlds becomes extremely fast.
Previously, constructing 100 realistic simulation environments required professional art teams for modeling, texturing, and lighting, resulting in prohibitive costs. In AgentGarten, as long as simple 3D depth and normal outlines can be output, any minimalist code scene can directly utilize this neural renderer to instantly gain realistic lighting and materials. Wingsuit flying, robotic arm manipulation, kitchen cooking, and multi-car racing share the same underlying visual interface—the cost of expanding virtual environments has shifted for the first time from "art-labor-intensive" to "code-procedural-expansion."
△ Figure 3 | Diverse physical worlds across tasks. Covering spatial navigation, object manipulation, camera re-shooting, and multi-car interactions. Left: Simple white-box models exported by code; Right: Photorealistic visuals generated in real-time by the renderer.
More critically, this renderer uses streaming generation: the action moves one step, the image draws a segment, the agent sees the result before making the next decision, and the interaction flows coherently like the real world.
Revisiting Hide-and-Seek: 25 Million Games vs. 4 Rounds
With the world built, the next focus is on the agents: Can they improve autonomously by acting and reviewing repeatedly in such a world?
A classic reference is OpenAI’s 2019 hide-and-seek experiment: hiders and seekers competed in arenas containing barriers and ramps, learning to build covers and use ramps to climb walls through large-scale self-play.
The MirroS team recreated this experiment in AgentGarten, adopting a one-on-one setup: the hider arranges the scene first, then the seeker enters.
The match rules are clear and pure: the hider first arranges the scene to block entrances, then the seeker enters to search. Agents control movement and grasping by writing Python code, making decisions solely based on the generated visuals, with no knowledge of spatial coordinates or opponent positions.
Both sides start with completely blank manuals, each maintaining a strategy library composed of individual skills. Each round consists of ten matches, followed by review and consolidation; in the next round, parts of the skill library are randomly selected for application.
Soon, those classic game strategies emerged within just a few rounds: moving barriers to seal doors, carrying ramps to build bridges, and pulling ramps closer to retry after failing to climb a wall.
The 2019 study became famous for the emergence of these classic behaviors and provides an intuitive comparison (see Figure 4): in that study, AI relied on reinforcement learning from scratch, taking approximately 25 million games to discover cover construction and about 100 million games to master climbing walls using ramps.
In AgentGarten, agents face first-person visual scenes, reviewing and revising textual manuals between rounds. The hider learned to build cover by round 4, while the seeker mastered wall-jumping via ramps by round 10.

△ Figure 4 | Comparison of game scales required before representative strategies emerge. Self-play RL trains from scratch on privileged physical states for tens of millions to hundreds of millions of games; MirroS agents act based on first-person rendered views, rapidly mastering corresponding strategies within just a few rounds through iterative manual reflection.
Learn as You Play: Writing Down Experience
The rapid progress in Hide-and-Seek relies on a general-purpose rehearsal pipeline.
MirroS’s approach is to let agents “take notes” themselves. All knowledge acquired is written into manuals.
The rehearsal process is divided into four rigorous progressive stages:
- Task Reception: Receive a task file defining action goals and rule constraints, but with no standard answers provided.
- Blind Trial-and-Error: Act autonomously under strict step and time limits. Agents see only camera feeds—no god-view coordinates, no minimaps, and no scores until the endgame.
- Manual Writing: After each round, conduct self-reflection and honestly record: what was tried, what phenomena were observed, which judgments are uncertain, and what new hypotheses should be tested next.
- Archiving & Inheritance: Freeze and archive the manual, passing it directly as prior legacy to the next round’s agent.

△ Figure 5 | Single-round rehearsal loop: Read task, act, write manual, archive. The next round’s agent continues exploration starting from the manuals accumulated by predecessors.
Reviewing these agent-written manuals reveals a strong scientific inquiry flavor. Models rigorously distinguish between “observed facts” and “speculative guesses,” and specifically annotate pitfalls for successors:
- The hider summarized a practical truth: “Defense hinges on blocking passages, not merely obstructing line-of-sight.” It advised successors to use blind spots to minimize gaps and test obstruction effectiveness with small movements.
- The seeker discovered a control principle: “Try pulling an object before moving it; don’t mistake jamming for grasping.” Upon approaching, press the micro-pull key briefly to check if the object moves relative to room landmarks, verifying true grasp.
- A bridge-crossing driver’s warning was particularly insightful: “The target disc sliding into the vehicle’s blind spot does not mean the car has safely arrived.”
Experiences that withstand scrutiny are adopted by successors, while failed lessons prompt them to explore new strategic branches.
Same Loop, Four Different Worlds
Hide-and-Seek was just the beginning. Since the code world is essentially a programmable software environment, the same “explore–reflect–consolidate” loop can be easily replicated across more complex scenarios.
The team ran four rounds of rehearsal in four distinctly different worlds:
- Companion Dog: Agents must maintain the dog’s affection within 60 seconds by reaching out to pet and playing ball at appropriate times. Scores improved from 13 in Round 1 to 19 in Round 4.
- Narrow Bridge Passing: Two cars negotiate yielding based on first-person driving views, swapping ends on a single-lane narrow bridge in minimal time. Total completion time for both cars dropped significantly from 71 seconds to 41 seconds.
- Cooperative Herding: Two sheepdogs coordinate solely via their own line-of-sight to drive four sheep into a pen and hold them there for 5 seconds. From timing out with only three sheep penned in Round 1, they stabilized to achieve “zero misses” (all sheep penned) in the subsequent three rounds.
- Quarry Loader: A heavy loader must push stones, feed materials, and store them. While Round 1 barely managed to push stones, by Round 4, the agent completed the entire workflow cleanly 31 seconds ahead of the 360-second limit.

△ Figure 6 | Four-round evolution comparison across four task worlds. Companion Dog, Narrow Bridge Passing, Cooperative Herding, and Quarry Loader all demonstrate robust behavioral evolution.
These vastly different tasks illustrate the same point: this experience-consolidation loop works equally well when transferred to entirely different worlds.
How Visuals Keep Up with Actions: Achieving 30 FPS on a Single Card
To enable agents to “see and act” in virtual worlds, neural renderers must overcome three long-standing challenges in generative video: keeping up with sudden actions, maintaining stability during ultra-long interactions, and running fast enough.
Starting from a foundational multimodal model, the MirroS team proposed the Adversarial Forcing training method, combined with inference optimizations, to establish this visual channel:
1. From Offline Whole-Sequence Generation to Streaming Chunk-by-Chunk Rendering
Traditional video models generate entire sequences at once. AgentGarten transforms this into streaming chunk-wise generation. When an agent performs an action, the model immediately produces the corresponding short visual chunk, allowing the agent to observe the result before making the next decision.
2. Preventing Drift in Long-Horizon Interactions (Precise Replay)
As continuous generation extends, minor errors accumulate into appearance drift. Many models cut off historical computation graphs to save VRAM, preventing gradients from backpropagating to memory history; recomputing the full sequence uniformly in a second pass introduces percentage-level errors due to subtle differences in underlying execution order.
The team designed a precise replay mechanism: the first pass performs gradient-free forward expansion, while the second pass re-executes with gradients following the exact same chunk-by-chunk rhythm. This precisely reproduces the rollout trajectory, stabilizing temporal consistency in long-sequence interactions.
3. Introducing Real Video Adversaries: Avoiding Degradation and Artifacts
Relying solely on score distillation using the model’s own generated samples leads to distortion and grid-like artifacts during long rollouts.
The team introduced real-video discriminators during training to construct adversarial signals. Anchored by real-world physical quality, this constrains visual texture while ensuring generated results strictly adhere to geometric contours defined by the code.
4. Maintaining 30+ FPS Real-Time Interaction
The team hand-crafted custom Triton fused operators to eliminate VRAM round-trips and avoid the minutes-long cold-start waits associated with traditional full-graph compilation. Combined with CUDA Graphs to remove CPU scheduling overhead and a super-lightweight decoder with sub-10ms latency, the system stably supports closed-loop interaction at 480p resolution with throughput exceeding 30 fps.
Preparing Experiencable Worlds for Continuously Evolving Agents
Executable code allows training worlds to be procedurally generated endlessly; learned renderers provide these worlds with perceivable, realistic physical feedback; consolidated experiment manuals turn every attempt into a stepping stone for the next round of exploration.
Combining these three elements gives agents a space where they can truly explore, trial-and-error, and evolve.
This is the core philosophy behind MirroS’s exploration of Physical RSI (Recursive Self-Improvement in the Physical World).
While Scaling Laws for language models are in full swing, the scaling of intelligence in the physical world is just beginning.
True general embodied intelligence requires continuous growth through genuine reciprocal interactions with the physical world:
Let every unknown become the starting point for the next round of evolution.
Full blog post: https://mirros.ai/blog/worlds-for-evolving-agents
This article is republished with permission from QbitAI. The views expressed are solely those of the original author.