Announced 8 Oct 2026 · Sources checked
What landed on 8 October?
The export API record for 2610.12459 lists published 2026-10-08T17:59:31Z. The abstract frames a gap: open-loop video models can look plausible without adapting to what they just generated, while many closed-loop systems keep a frozen executor that never learns the atomic actions the planner requests.
WorldGuide’s answer is joint training on the same step-level demonstrations. The ContextPlanner learns next-action and completion prediction from visual progress. The Executor is fine-tuned to realize those actions as short clips. Hierarchical visual memory keeps long-horizon context without unbounded tokens.
That sits next to other recent world-model drops we already covered, including NVIDIA’s Long-WAM and Odyssey-3. WorldGuide is a procedural video loop, not a robot latent planner like RoboJEPA.
Middle East AI News summarized the work on 10 October. The primary objects remain the arXiv HTML/PDF, the MBZUAI Oryx project page, the GitHub README, and the Hub checkpoint card we opened the same day.
How does the closed loop actually run?
From an initial image and a language goal, the planner sees the current visual state plus the latest three clip–action pairs. It emits a natural-language atomic instruction or a completion token. If the task is not done, the Executor—a HunyuanVideo-1.5 diffusion transformer conditioned on the frozen planner’s action embeddings—renders the next clip. That clip becomes the new state; memory updates; the loop continues until completion or a step budget.
The paper initializes the planner from Qwen2.5-VL-7B and trains it first with token-level cross-entropy on response tokens. The Executor then trains under flow matching while the planner stays frozen for action embeddings. At inference, action choice depends on generated results, not a prewritten script.
Hierarchical memory follows a YUME/FramePack-style schedule: newer latent frames keep finer spatial embeddings; older history is compressed. The authors report memory gains of 5.77% Task Success on WorldGuide Bench and 17.91% on Video-CraftBench in their ablations. Those percentages are theirs.
Ablations they publish also attribute +21.62% Task Success to closed-loop versus open-loop, +18.61% to visual feedback, and a 4.46% cut in repeat/skip errors when the planner sees generated clips. Treat every ablation delta as author-reported until you rerun the scripts.
Completion is a first-class planner output rather than a fixed clip budget. That matters for origami-length procedures where open-loop generators either stop early or keep fabricating folds after the crane is finished. The paper’s Figure 1 contrast—open-loop drift, frozen-executor closed loop, jointly trained loop—is the conceptual diagram to keep when skimming the HTML.

What is WorldGuide Bench, and what scores are claimed?
Because ordinary instructional datasets lack paired atomic actions, visual consequences, and completion signals, the authors introduce WorldGuide Bench: about 58,679–59K step-annotated videos across 245 tasks and 27 procedural categories (paper engineering, block building, origami, cooking, assembly, knotting, cloth folding, and more).
On that bench they report 33.33% Task Success for WorldGuide versus 29.90% for MiniMax-H3 even though MiniMax-H3 is given reference action plans. On VideoCraft-Bench under goal-only conditioning they report 47.69% versus 32.73% for MiniMax-H3. Aesthetic quality and overall consistency rows also favor WorldGuide in their tables.
Action-faithful robot video models such as DreamTrue attack a different failure mode—keeping robot motion honest under counterfactual post-training. WorldGuide’s claim is procedural completion in generated video space, not Franka control.
| Setting | WorldGuide | MiniMax-H3 (paper’s comparison) |
|---|---|---|
| WorldGuide-Bench Task Success | 33.33% | 29.90% (with reference plans) |
| VideoCraft-Bench (goal-only) | 47.69% | 32.73% |

What can you download today?
github.com/mbzuai-oryx/WorldGuide exists: created 2 October 2026, last push we recorded 9 October 06:52 UTC, with README, download_models.py, hyvideo/, scripts/inference/, trainer/, and worldcompass/. The GitHub license API returned 404 for a LICENSE file when we checked; treat reuse terms as unspecified until the authors add one. Acknowledgements credit HunyuanVideo-1.5 and Wan.
The README’s live Hub badge points at ankanmbz/WorldGuide-Ckpt. The model API lists createdAt 2026-10-08T21:10:14Z and lastModified 2026-10-09T06:25:12Z, pipeline_tag image-to-video, with text_encoder/llm/*.safetensors and transformer/diffusion_pytorch_model.safetensors among the siblings. MBZUAI/WorldGuide-Ckpt exists as a card/README stub without those weight shards on the copy we opened. The project page still labels the dataset badge “Coming Soon.”
Inference scripts document closed-loop rollouts with visual memory (mode2, last two steps) until an EOF token or about 70 steps, plus a text-only planner baseline without visual feedback. Training docs cover feature precomputation and multi-GPU runs. We did not download multi-gigabyte shards or execute the bash entrypoints.
The official MBZUAI/WorldGuide-Ckpt card we opened on 10 October is still a thin README stub that points readers at the method and GitHub; the weight siblings that download_models.py expects—planner under text_encoder/llm and the DiT under transformer/—were present on ankanmbz/WorldGuide-Ckpt instead. Until MBZUAI mirrors those shards under the org account, treat ankanmbz as the runnable checkpoint path named in the README badge, and re-check the Hub API before scripting a fleet download.
What should readers not conclude?
33% task success is still a hard problem statement, not a solved origami factory. MiniMax-H3’s comparison includes a privileged reference-plan condition the authors disclose. Aesthetic scores do not prove real-world utility. There is no public WorldGuide Bench dump yet under the Hub links we opened, so independent reproduction of Table rows is blocked on data access even if weights load.
Hierarchical JEPA planners such as AMI Labs’ H-JEPA optimize latent plans for control; WorldGuide optimizes generated video procedures. Do not merge their scoreboards.
This article is an evidence review of pages opened on 10 October 2026. It is not a video quality test, not a ranking against every closed-loop baseline, and not a claim that MBZUAI has shipped a consumer API.
Common questions
Is WorldGuide open source?
Training and inference code are public on GitHub. Checkpoints are on ankanmbz/WorldGuide-Ckpt. We found no LICENSE file via the GitHub license API on 10 October, and the dataset badge still says coming soon.
Does it control a real robot?
No. The paper’s loop generates video clips of procedural tasks. Robot foundation-model recipes are a separate line of work.
Why compare against MiniMax-H3 with reference plans?
That is the authors’ chosen strong baseline condition. They still report WorldGuide ahead on Task Success; the privilege is disclosed in the abstract.
What to remember
WorldGuide is an 8 October MBZUAI closed-loop video world model with public code, Hub weights under ankanmbz, and author-reported gains on a new procedural bench. Use it to study planner–executor coupling; do not treat 33% Task Success as product-ready procedural video.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





