Announced 8 Oct 2026 · Sources checked
What was posted on 8 and 9 October?
The arXiv Atom record we opened lists 2610.12468v1 as published at 17:59:51 UTC on 8 October 2026, primary class cs.RO, with a cs.CV secondary. The title is “DreamTrue: Action-Faithful Robot World Model with Counterfactual Post-Training.” The HTML copy’s footnote says Xiangshuo Liu was an intern at CASIA during the work. The bibtex block on the GitHub README repeats the title, names the same nine authors, sets year 2026, and adds the note “Manuscript under review.”
The public GitHub repository brave-eai/DreamTrue did not exist on the paper’s posting day. GitHub’s repository JSON, retrieved 9 October, gives created_at 2026-10-09T02:35:37Z, pushed_at 2026-10-09T06:12:46Z, default branch main, language Python, stargazers_count 4, and license null. A GET of /brave-eai/DreamTrue/blob/main/LICENSE returned 404. The root README points to the paper, a project page at brave-eai.github.io/DreamTrue, and ModelScope dataset huoxingdawang/DreamTrue.
This is a video world-model paper, not a new robot product. For another 9 October latent-prediction write-up, see Meta FAIR’s RoboJEPA note. For a longer-horizon action-conditioned video model, see NVIDIA Long-WAM.
How do they say the generator is built?
Stage I converts each embodiment’s actions into a shared image-space bundle: rendered RGB, depth, amodal masks, gripper openings mapped to background intensity, and Plücker camera-ray maps. A VACE branch injects those features into a video DiT. The wmvideo README, opened the same day, names the public inference runtime as “a focused Wan2.1 V2V-VACE 14B runtime with a three-view launcher.”
Stage II builds counterfactual trajectories by applying an SE(3) perturbation to a recorded final pose, interpolating from the start, and solving inverse kinematics. A vision-language reward model, which the ModelScope card says is fine-tuned from a Qwen3.5 VLM, predicts defect probabilities on three axes. Those scores, plus a PSNR term on recorded actions, guide DiffusionNFT post-training. The paper’s own limitation paragraph says the action representation, generator, and VLM reward all operate in image space, so occlusion can still reward a visually tidy but physically wrong contact.
The public checkpoint is a dual-LoRA file, not a from-scratch 14B train. For what that adapter usually means, see LoRA research explained. For why a dataset license field is not the same as a complete source grant, see open weights versus open source.
What can you actually download?
The root README splits the public tree into calibration, wmvideo, and reward. The wmvideo README says the dual-LoRA file lives at ModelScope path wmvideo/checkpoint-wmvideo.safetensors, that launchers run offline and do not fetch weights, and that Wan2.1-VACE-14B base weights must come from Wan-AI on Hugging Face or ModelScope. It names a condition mini package, condition/condition-mini.tar.gz, and lists default three-view layouts for AgiBotWorld-Beta, DROID, RoboMIND 2.0, and RoboTwin 2.0. Training code is not included in that component.
The ModelScope markdown we opened on 9 October, last_updated 2026-10-08, lists those two files plus reward/checkpoint/ and reward/reward-data-mini.tar.gz. A “Coming soon” block — echoed by the calibration README — says the full calibration set, a fine-tuned SAM3 checkpoint, and the full reward annotations are still pending. The same page prints license “Apache License 2.0” and downloads 25. A download counter is not a quality signal.
The project page at brave-eai.github.io/DreamTrue returned a JavaScript shell when we fetched it. We used the arXiv HTML, the GitHub README files, and the ModelScope markdown for the facts above. We did not clone the repository or load the LoRA.
What did we not verify?
This is an evidence review of the 8 October abstract, Atom record, and HTML, the 9 October GitHub API record and README files, and the 8 October ModelScope dataset markdown. We did not attend the AgiBot World Challenge, did not rerun nDTW or the human defect protocol, and did not confirm that every listed baseline used the same views, horizon, or prompt. Treat the method as what the paper describes, the scores as what the authors printed, and the download as a mini-pack inference drop plus a dataset license field that the GitHub repo does not repeat.
Common questions
Is DreamTrue an open-source model you can retrain from scratch?
Not on the pages we opened. The public wmvideo tree is an inference runtime plus a dual-LoRA file that expects Wan2.1-VACE-14B base weights from Wan-AI. Training code is not in that component. ModelScope’s dataset listing prints Apache License 2.0. The GitHub repository has no LICENSE file.
Did the authors collect a new real-robot corpus?
They say they did not. The paper’s pitch is better use of existing recordings: AgiBotWorld-Beta, DROID, RoboMIND 2.0, and RoboTwin 2.0, plus counterfactual actions derived from those trajectories. WidowX250 and Piper scenes are described as qualitative transfer, not as the training mix.
Should we treat the AgiBot Challenge row as an official ranking?
Treat it as Table 2 in the authors’ HTML. They list three teams and put themselves first on three of four columns. We did not open a separate Challenge leaderboard or the judges’ protocol.
What to remember
Use 2610.12468 for the calibration-plus-counterfactual recipe and keep every AgiBot percentage in the authors’ column. Use the 9 October GitHub and ModelScope pages for what is actually fetchable: mini packs, a dual-LoRA, and a dataset Apache 2.0 field. Do not invent a GitHub license, and do not treat “coming soon” full data as released.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





