What did NVIDIA publish, and when?

arXiv lists the paper as submitted on 7 October 2026. The abstract page we opened on 9 October carries citation_date 2026/10/07. Equal-contribution authors are Wei Huang and Bohan Zhang; the project page lists NVIDIA, MIT, HKU and UCSD.

The inspectable software landing is later than the preprint stamp. GitHub’s API for NVlabs/LongLive reports pushed_at 2026-10-08T02:08:47Z. The Long-WAM/ directory we listed on 9 October has its own README, Apache-2.0 LICENSE, configs, scripts, src/longwam, stage1 video pretraining, infra backends, and docs/VERIFICATION.md dated 2026-10-07. The parent repository is the existing LongLive 2.0 video-generation project, not a new empty repo.

There is still no NVIDIA Newsroom or Developer Blog item that announces Long-WAM as a product. The 8 October NVIDIA blog post we opened is an Omniverse how-to. That is a different story from Microsoft’s RTX Spark Windows PCs, which NVIDIA and Microsoft did blog. Long-WAM’s primary surfaces are the arXiv page, the NVlabs project page, the Long-WAM tree, and the Efficient-Large-Model checkpoints.

What is the paper’s central claim?

World-action models treat a video generator as the start of a robot policy: predict how the scene will change, then emit an action chunk. Long-WAM asks a narrower question that is easy to confuse with a generic context-window explainer. The authors vary the duration of real camera history in the causal prefix, keep the forecast and action horizons fixed, and train a separate checkpoint for each window. Access to a longer prefix, they write, is not the same as having learned to use it.

On RoboCasa GR-1, the abstract says success moves from 63.3% at 0.0 seconds of history — current observation only — to 78.7% at 19.2 seconds. Section 5.3 and the project page give 2.4-to-19.2-seconds as 66.3% to 78.7% for the robot-domain autoregressive start. After the same causal adaptation, a bidirectional Wan2.2 start is reported at 61.7% at 2.4 seconds and 61.6% at 19.2 seconds. The robot-domain AR lead over that start grows from 3.3 points with no history to 17.1 points at 19.2 seconds. Those comparisons are the authors’.

Table 4 in the HTML paper labels a “Long-WAM (2.4 s)” row 63.3%, which matches the abstract’s zero-history figure, not the 66.3% in the prose. We report the abstract and section 5.3 numbers and note the table label. On LIBERO-Long, the authors say 2.4 seconds of history lifts success from 94.5% to 99.5%. At 38.4 seconds on GR-1 they report a drop; they say 80.4% of sampled history frames at that length are padding because training trajectories average 12.1 seconds.

How does the system predict, then act?

Stage 1 continues LongLive-2.0’s 16-second autoregressive checkpoint on what the paper calls about 10,000 window-equivalent hours of robot and egocentric video, without action labels: RoVid-X, AgiBot World, EgoDex, EgoVerse and VITRA. Because this stage is video-only, the authors do not need a shared action space. That is not a claim the video model transfers to an unseen robot. The later policy is closer to multimodal control than to a text-only planner: one camera stream, language, proprioception, then actions.

Stage 2 couples a video expert and an action expert through an asymmetric Mixture-of-Transformers interface. Video queries never read action tokens. In the default inverse-dynamics (IDM) mode the video expert partially denoises future latents to noise level 0.9, never decodes pixels, and hands a visual key-value cache to the action expert. The action expert then denoises a chunk conditioned on observed history plus that predicted future. A second public mode, CodeDenoise, updates video and action jointly; the README ships IDM and COD bundles for LIBERO and RoboTwin 2.0.

The runtime is why they can talk about 19.2-second history on an edge GPU. They overlap inference with robot motion, encode incoming frames with a streaming VAE, quantize video to NVFP4 while keeping actions and KV in BF16, and publish device builds for GeForce RTX 5090, DGX Spark and Jetson AGX Thor. On RoboTwin 2.0 they report IDM falling only from 94.4% synchronous success to 94.2% asynchronous, against much larger drops they list for Fast-WAM and LingBot-VA. Those are their comparisons.

What else do the authors report on robots and planners?

The physical results are small-N author evaluations, 20 trials per policy and condition. On a Unitree G1 they report conveyor grasping of 100%, 100%, 95% and 90% at 3.0 to 7.5 cm/s, while π0.5 and Fast-WAM fall to 0% at the two highest speeds. Dynamic cup stacking at 3 cm/s is 19 of 20 for Long-WAM and 0 of 20 for both baselines. On a bimanual YAM they report an 81.7% mean on three tasks they say last more than 40 seconds. We have not seen those trials.

On RoboCasa365 they keep a 2.4-second checkpoint fixed and add GPT-6 Astra as a planner. Overall success is reported as 31.4% for the policy alone and 54.4% with the planner, versus 25.2% for the planner alone. Unseen composite tasks move from 6.1% to 35.0%. That is a hierarchical simulation result, not a shipped NVIDIA product. It sits nearer AWS’s Physical AI Toolchain write-up as another stack claim than as an API you can call without your own model access. The README’s agentic dry-run says live runs need credentials the repo does not ship.

Other author-best figures include 94.4% average on RoboTwin 2.0 and 34.9% success on DOMINO after dynamic-data fine-tuning. DOMINO is still a minority of tasks. We did not re-run the 24×10 GR-1 protocol. For a different NVIDIA preprint that also asks readers to wait for independent code, see the UNREAL retrieval paper.

What can you actually download?

The official checkpoint table in the README points at Efficient-Large-Model, not at gated personal Hub dumps that also appear in search. On 9 October, unauthenticated API calls showed gated=false for Long-WAM-RoboCasa-GR1-19.2s, Long-WAM-LIBERO-IDM, Long-WAM-YAM and LongLive2.0-Robot-M. Those cards’ lastModified stamps are 6 October 2026, the day before the arXiv date. docs/VERIFICATION.md, dated 7 October, says the authors confirmed 13 listed ELM repositories public and ungated and did not download weights in that check.

Named bundles include LIBERO IDM/COD, RoboTwin 2.0 IDM/COD, RoboCasa365, YAM, and five RoboCasa GR-1 history lengths from 0 to 19.2 seconds. GR-1 cards describe a 20 Hz window and a 16-step action chunk; keep model.pt, config.yaml and dataset_stats.json together. LongLive2.0-Robot-S and Robot-M are video checkpoints; S and M are sequence-length coverage, not model size.

Installation is a sparse clone of LongLive, then work from Long-WAM/ as its own root. Training and deployment extras pin different datasets versions; GR-1 and RoboCasa365 use incompatible robocasa packages. Device builds are longwam build rtx5090, spark or thor. The README says VAE/text assets, the GR-1 text cache and simulators are separate, and that “full benchmark reproduction remains unverified.” Verification.md adds that the 7 October repo check ran no GPU jobs, simulator rollouts or robot commands.

What should readers not assume?

78.7%, 99.5%, 95% and 107.4 ms are author measurements. The latency protocol excludes preprocessing, text encoding, controller and IPC overhead. Stretching history eightfold, they say, multiplies RTX 5090 chunk time by about 3.2×. After device tuning they report 107.4 ms on RTX 5090, 328.2 ms on DGX Spark and 378.7 ms on Jetson AGX Thor. Those are hardware-lab numbers, not a promise that a G1 on a conveyor meets a 100 ms budget at 19.2 seconds of history.

The code license is Apache 2.0 for Long-WAM sources; Wan2.2, simulators and the listed video datasets keep their own terms. YAM is labeled a pretraining release. The G1 and YAM demo GIFs on the project page are the authors’ recordings. We did not find a Long-WAM item on blogs.nvidia.com dated 7–9 October 2026.

This is an evidence review of pages opened on 9 October. It is not a first-hand training run or a claim that Long-WAM is generally better than the baselines outside the authors’ tables.

Common questions

Is Long-WAM a product I can buy from NVIDIA?

No. It is a research code tree and a set of Hugging Face checkpoints. The paper and project page are the announcement. NVIDIA’s 8 October corporate blog does not list Long-WAM.

Did anyone outside the authors rerun the 78.7% GR-1 score?

Not in the pages we opened. The README says full benchmark reproduction remains unverified. Verification.md records CPU tests and ungated Hub metadata, not simulator rollouts.

If I already have a bidirectional video model, will adding history help?

The authors say no on GR-1 after their causal adaptation: the Wan2.2 start shows no net gain from 2.4 to 19.2 seconds. The gain they report is tied to autoregressive video pretraining. That is their ablation, not a law of robotics.

THE TAKEAWAY

What to remember

Open the 7 October paper for the AR-versus-bidirectional plot, and the Long-WAM README for the sparse clone and the five GR-1 history bundles. Keep 78.7% and the G1 cup-stacking rates in the author column until someone else runs the same protocols.

Sources & further reading

  1. Long-WAM: Scaling the Context of World-Action Models ↗
  2. Long-WAM: Scaling the Context of World-Action Models (HTML) ↗
  3. Long-WAM README ↗
  4. Long-WAM project page ↗
  5. Long-WAM RoboCasa-GR1 19.2s ↗
  6. Long-WAM verification scope ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories