Announced 8 Oct 2026 · Sources checked
What did MirroS open on 8 October?
GitHub’s API for MirroS-Lab/AgentGarten reports created_at 8 October 2026 at 03:53 UTC, pushed_at 9 October at 06:40 UTC, and Apache-2.0. The README title is “AgentGarten” with the subtitle “Code Worlds for Evolving Agents.” The News block dated 2026/10/08 lists the technical report, the MirroS blog, and renderer training and inference code.
The Hugging Face renderer repo was created 8 October at 11:15 UTC and last modified at 11:19 UTC, sha 893b595a52082402cc2a2a785fd9061732a03a1d. The card lists student.safetensors as “the few-step autoregressive generation tower (7.0 B parameters, BF16), distilled with Adversarial Forcing,” plus und.safetensors, “the frozen text tower of Cosmos3-Nano, unchanged,” and the Cosmos3-Nano text tokenizer. The Wan 2.2 VAE “is not included here.”
arXiv 2610.12374, submitted 8 October 2026, is titled “AgentGarten: Code Worlds for Evolving Agents,” in cs.CV. The HTML experimental page we opened says it was created on 8 October 2026. Named authors on the abs page begin Jiawei Chi, Shangchen Miao, Zhiyuan Shi and others; the README citation is “MirroS Team.” This is a renderer-and-paper drop, not a closed world-model launch like Odyssey-3 and not the same object as DreamTrue’s 8 October robot-video preprint.

How is the renderer supposed to work?
The README’s one-line split is: “Code determines how the world changes, and the renderer learns how those changes should look.” Simulators and game engines “maintain persistent state and execute program-defined rules.” A shared neural renderer “generates the agent’s observations from the depth and surface normals each world exports.”
The adaptation of NVIDIA Cosmos3-Nano is described in three stages: bidirectional rectified-flow training of whole clips from a first frame; autoregressive block-causal training with teacher forcing or diffusion forcing, sampled with a KV cache; and Adversarial Forcing, which “distills the autoregressive model into a few-step student.” The card says the student is four-step. Cosmos3Stream “runs the renderer block by block with a bounded KV cache.” Each block is four latent frames, denoised, then committed so the next block can depend on what was just generated.
Install notes on the README are Python 3.11+ and PyTorch 2.9+ on a CUDA host, plus separate downloads of nvidia/Cosmos3-Nano (transformer, text tokenizer, negative prompt) and Wan-AI/Wan2.2-TI2V-5B’s Wan2.2_VAE.pth. prepare_serving “serves the transformer with hand-written Triton kernels, cuBLAS matrix products, and CUDA graphs, without torch.compile.” That is a serving recipe, not a measurement we took. For other open world-action stacks that keep camera history, see NVIDIA Long-WAM. For what an agent is doing when it writes the next action, see AI agents explained.
What do the hide-and-seek numbers actually claim?
The 8 October blog, dated Oct 8, 2026, revisits “the game of OpenAI’s 2019 study of emergent tool use.” It says OpenAI’s report “reports shelter construction after roughly 25 million episodes, followed by seeker ramp use after another 75 million,” and later restates ramp use at “roughly 100 million.” In MirroS’s setting, “the hiders used a panel to build a shelter by round 4, and the seekers used a ramp to cross walls by round 10.” Those episode counts and round numbers are the blog’s.
The protocol on that page is not the 2019 self-play RL setup. “Our pretrained agents see neural-rendered first-person frames, act through short Python programs, and keep what they learn in a playbook of skill files.” Each role reviews its games after every round. The experiment “begins with empty playbooks, and each round consists of five games.” The arXiv abstract compresses the same claim to “agents learning from just 4 rounds compared with millions for a conventional reinforcement learning counterpart.” That is an authors’ comparison across two different learning stacks, not a rerun of OpenAI’s code.
The blog then says the same loop ran “for four rounds in each of four more worlds: a companion dog, a one-lane bridge, sheep herding, and a quarry loader.” Those are task descriptions on the blog. They are not evidence that those worlds ship in the GitHub tree we opened.
| Object | Status on pages we opened | What that is not |
|---|---|---|
| Renderer training and streaming inference | Checked TODO; code in MirroS-Lab/AgentGarten | A hosted playground |
| Renderer weights | MirroS-Lab/AgentGarten-renderer on 8 October | The Wan 2.2 VAE |
| Code worlds and practice loop | Unchecked TODO | A runnable hide-and-seek gym tonight |
| Real-time engine and frontend | Unchecked TODO | An interactive viewer |

What is still a TODO, and what did we not run?
The README TODO list marks four items done: training code, streaming inference, the report and blog, and the renderer checkpoints. Two remain open: “Real-time rendering engine: streaming server and interactive frontend” and “Code worlds and the agent practice loop: rounds of play, review, and playbooks.” The overview paragraph is consistent with that: “The neural renderer is available now… the code worlds and the agent practice loop will be released here as well.”
We did not clone the repo onto a CUDA host, download Cosmos3-Nano or the Wan VAE, run Cosmos3Stream, or render a validation clip. We did not play hide-and-seek, write a playbook, or open the project page as a substitute for the missing worlds. The blog’s clips of a panel, a ramp, and a second climb are the authors’ illustrations. They are not a package in the tree we listed.
What do the two licenses cover?
GitHub’s license key on AgentGarten is Apache-2.0. The LICENSE file we opened is the standard Apache License, Version 2.0, January 2004. That grant is for the repository’s code.
The renderer card says the weights “derive from Cosmos3-Nano and are released under the same OpenMDW License Agreement, version 1.1.” The OpenMDW-1.1 text we opened at openmdw.ai/license/1-1/, last updated 27 May 2026, is a permissive-style grant to deal in Model Materials without restriction, with notice retention on distribution and termination if you sue for patent or copyright infringement of those materials. It “does not impose any restrictions or obligations with respect to any use, modification, or sharing of any outputs.” It is not Apache, and it is not RAIL-M. The card says the text tower and tokenizer are NVIDIA’s files, redistributed unchanged under that license.
What should a team try if they already have Cosmos3-Nano?
The honest first step on the pages we opened is the README’s validation path: point WM_COSMOS3_NANO and WM_WAN22_VAE, download the AgentGarten-renderer artifact, and render the checkpoint’s validation clips with train.max_iterations=0. That tests whether the student loads. It does not test hide-and-seek.
If you need a playable world tonight, this drop does not yet advertise one. Wait for the unchecked TODOs, or keep the 8 October paper and blog in the research pile. We did not time a block, measure frames per second, or compare Adversarial Forcing to the cited DMD2, Self Forcing, or rCM recipes.
Common questions
Can I play hide-and-seek in AgentGarten tonight?
Not on the pages we opened. The renderer code and checkpoint are up. The README still lists code worlds and the agent practice loop as unchecked TODOs.
Are the renderer weights Apache-2.0?
No. The GitHub repo is Apache-2.0. The Hugging Face card puts the weights, including NVIDIA’s unchanged text tower, under OpenMDW-1.1.
Did the authors rerun OpenAI’s 2019 hide-and-seek?
Not in those words. The blog compares their pretrained agents and playbooks, seeing rendered frames and writing Python, with OpenAI 2019’s self-play RL episode counts. That is an authors’ cross-stack comparison.
What to remember
Use the 8 October README for an open Cosmos3-Nano renderer and two still-open TODOs. Use the blog and arXiv for the authors’ 4-round versus millions hide-and-seek claim. Keep Apache on the code and OpenMDW-1.1 on the weights.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





