Announced 8 Oct 2026 · Sources checked
What did Odyssey announce on 8 October?
The dated launch object is Odyssey’s post “Meet Odyssey-3: Our Most Powerful Foundation World Model,” 8 October 2026, signed Oliver Cameron and Jeff Hawke. It says a research preview is open and that physical-AI developers should get in touch. A 15 September post, “Introducing Odyssey-3,” was the earlier teaser.
A Business Wire release at 10:30 AM EDT the same day, opened by us on FinancialContent, datelines Palo Alto and lists six domains: robot arms, Flexion humanoids, vehicles, indoor drones, agent-training environments, and Grand Theft Auto V. The CEO quote in that wire — “Odyssey-3 is a big step toward a single intelligence that can understand and operate in the world around us” — is Odyssey’s.
This is a closed world-model product, not an open checkpoint. For an 8 October robot video-model preprint with public mini packs, see CASIA and Amap’s DreamTrue note. For a longer-horizon action-conditioned video model, see NVIDIA Long-WAM.
What does Odyssey say the model is?
The launch post calls Odyssey-3 a learned dynamical system, implemented as an autoregressive diffusion transformer. It says the model learns physics, dynamics, and cause-and-effect from visual observations, then uses that knowledge to simulate environments and to train policies for different machines.
The preview generates an embodied environment from a prompt and continues it as a person or agent acts. Odyssey says it offers first-person and third-person navigation plus independent camera movement. The distilled few-step variant it points readers to is Odyssey-3 Flash. The wire says Flexion has already adopted the model. That is an adoption claim, not a contract we inspected.
Which Physics-IQ numbers are Odyssey’s, and what does the live board show?
Odyssey’s post says Odyssey-3 Pro sets a new state of the art on Physics-IQ Verified video-to-video at 66.1, “the highest reported score,” and that Pro also scores 54.7 on image-to-video. A caption says scores average four runs, best-of-8 uses one run with the same prompts, Odyssey-3 is 832×480, Pro is 1280×720, and the source is the Physics-IQ Verified leaderboard dated 7 October 2026. Those rows are Odyssey’s.
Physics-IQ Verified, which Odyssey attributes to Anates Labs and DeepMind, asks models to continue videos of real experiments. The live “Leaderboard All” view we opened on 9 October ranks by net improvement versus the track mean, not by a 66.1 headline. FLUX 3 large is first at +12.27 pp ± 0.44 pp; Odyssey-3 Pro is second at +11.15 pp ± 0.42 pp; Odyssey-3 is third at +9.99 pp ± 0.78 pp. A no-LLM Odyssey-3 row on best-practice prompts is eighth at +2.15 pp ± 0.84 pp.
The live spatial submetric lists FLUX 3 large at 64.36 ± 0.59 and Odyssey-3 Pro at 61.78 ± 0.40. Spatiotemporal lists an Odyssey-3 row first at 44.70 ± 1.86. We did not recover a 66.1 cell on that default view.
WorldMark is Odyssey’s own evaluation. It says that, using the benchmark’s captions and the mean of 13 metric scores, Odyssey-3 ranks first in first-person stylized, third-person real, and third-person stylized. The charts print 79.0 on first-person stylized, ahead of HY-World 1.5 at 76.9, and 76.3 on third-person stylized. Odyssey writes that these measure generated-world properties, not a physical-system test.
| Rank on the live default table | Model as printed | Net improvement vs track mean |
|---|---|---|
| 01 | FLUX 3 [large], Black Forest Labs | +12.27 pp ± 0.44 pp |
| 02 | Odyssey-3 Pro, Odyssey | +11.15 pp ± 0.42 pp |
| 03 | Odyssey-3, Odyssey | +9.99 pp ± 0.78 pp |
| 08 | Odyssey-3, no LLM, BPP | +2.15 pp ± 0.84 pp |
What did they show on robots, cars, and games?
The launch post says an action decoder or policy is trained on paired observations and actions. On robot arms, it claims that with tens of hours of demonstrations the model completed tasks and showed recoveries absent from those demos, including reorienting a gripper after a missed grasp. Labeled clips include “Pour the cereal into the bowl” and “Close the screwbox.”
On humanoids, Odyssey says Flexion built policies that beat the tested vision-language-action baselines under lighting changes that caused those baselines to fail. Those comparisons are Odyssey’s.
On vehicles, the 8 October post says the team drove on real roads in India with a policy trained on 20 hours of driving data and a frozen backbone. The 15 September post and the Business Wire reprint both write “20 hours of simulated driving data.” The September page adds that policies trained only in simulation traveled about 77% as far between safety-driver interventions as policies trained on real footage. Keep those sentences in the publishers’ columns. For other physical-AI write-ups, see AWS’s Physical AI toolchain and Meta FAIR’s RoboJEPA paper.
The wire also lists indoor drones trained in simulation, agent-training worlds, and play in Grand Theft Auto V. Those are demonstration claims.
How do they say they built it, and what can you try?
The build section lists three data families: internet video with schema-verified event annotations; gameplay with time-aligned keyboard and mouse; and simulated rigid-body interactions. Odyssey says it trains a multi-step video diffusion transformer, extends it autoregressively with teacher forcing and causal masking, then distills a few-step real-time variant with distribution-matching and adversarial losses.
The call to action is “Try Odyssey-3 Flash.” The pages we opened list no Hub repo, weight license, or parameter count. The Physics-IQ caption assumes $1 per MI355X GPU-hour; the live Anates page estimates Odyssey prompt rewriting at $0.01 per video. Those are published cost models, not invoices we paid. The Next Web’s $310 million June Series B line is prior company news, not this launch.
What did we not run?
This is an evidence review of the 8 October launch post, the 15 September teaser, the Business Wire reprint, the 9 October Physics-IQ Verified page, and The Next Web’s 16:21 UTC report. We did not open the preview, generate a video, or recompute Physics-IQ or WorldMark. Treat 66.1 and the WorldMark firsts as vendor evaluations, and the live Anates ranks as a different default sort.
Common questions
Can you download Odyssey-3 weights?
Not on the pages we opened. The public object is a research preview Odyssey labels Flash, plus a contact form for physical-AI developers. There is no Hub repo, no license file, and no parameter count in the launch post.
Is Odyssey-3 Pro first on the live Physics-IQ board?
Odyssey’s 8 October post says Pro’s 66.1 video-to-video best-of-8 score is the highest reported, citing a 7 October snapshot. The live default table we opened on 9 October ranks by net improvement versus the track mean and lists FLUX 3 large first. Those are different published views. We did not rerun the benchmark.
Did they train a new driving stack from scratch?
Odyssey says the backbone stayed frozen and that a waypoint policy used about 20 hours of data. The 15 September post and the Business Wire reprint call that data simulated. The 8 October post says the car then ran on real Indian roads. We did not inspect the logs.
What to remember
Use the 8 October Odyssey post for the diffusion-transformer recipe and the research-preview offer. Keep 66.1 and the WorldMark firsts in Odyssey’s column, and keep the live Anates net-improvement ranking — FLUX 3 large then Odyssey-3 Pro — in the board’s column. Do not invent open weights or an independent robot test.
Sources & further reading
- Meet Odyssey-3: Our Most Powerful Foundation World Model ↗
- Introducing Odyssey-3: A General-Purpose Physical Intelligence ↗
- Odyssey Launches Odyssey-3, Its Most Powerful Foundation World Model Built to Power AI Across the Physical World ↗
- Physics-IQ Verified leaderboard ↗
- Odyssey-3 drives a car in India and runs humanoids from one world model ↗
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





