What did Meta FAIR publish, and when?

The dated research object is the 7 October 2026 arXiv preprint, identifier 2610.10515. The abstract page we opened lists citation_date 2026/10/07. The API record names v1 at 17:54:42 UTC that day, from Artem Zholus. Authors listed on the HTML paper include Zholus, Nicolas Beltran-Velez, Jianhao Yuan, Sarath Chandar, Tushar Nagarajan, Daniel Severo, Koustuv Sinha, Michal Drozdzal, Adriana Romero Soriano, Jeannette Bohg, Nicolas Ballas and Mahmoud Assran. Affiliations span FAIR at Meta, Mila, Polytechnique Montréal and the Chandar Research Lab. Zholus’s author note says the work was done at Meta. Ballas and Assran are marked joint last authors.

The paper’s headline is not a new robot product. It is a family of action-conditioned latent world models, trained on a mixture the authors count as 23 public manipulation datasets, 12 embodiments, 26 action-space identifiers, 2.8713 million episodes, 15,022 video hours and 6,692 action hours. They call the 8B predictor the largest JEPA predictor trained to date. That ranking is theirs.

This is a robotics world-model paper, not a hosted policy API. The nearer open stack on this site is NVIDIA’s Long-WAM, which ships inspectable LongLive code and ungated Hugging Face checkpoints. RoboJEPA’s named code path did not.

How does the predictor work?

RoboJEPA does not generate pixels at training time. A frozen V-JEPA 2.1-G encoder maps up to several camera views into a 1664-dimensional patch grid. The trainable object is a feed-forward vision transformer that predicts the next latent frame in one pass, then can roll that prediction forward under a candidate action sequence. Training uses a plain ℓ1 loss in that frozen space, split between teacher-forced next-step prediction and a short autoregressive rollout so the model sees its own errors.

Planning is a single goal image, encoded once. At each step the authors search for an H-step action sequence whose imagined latent lands closest to that goal, using the cross-entropy method, then execute the first action and replan. They say they keep planning hyperparameters fixed across embodiments and sample actions from a uniform distribution rather than from a learned policy proposal. A diffusion decoder exists only to visualise rollouts; it is not the controller.

The encoder stays frozen on purpose. The scaling laws therefore describe a dynamics head on a fixed representation, not a jointly trained vision backbone. That is a different object from a generalist vision-language-action policy and from the simulation samples in AWS’s Physical AI Toolchain.

What scaling law do the authors fit?

They train predictors from 22 million to 8 billion parameters and from about 2×10^19 to 9.5×10^22 FLOPs. On held-out DROID (real Franka) and RoboCasa (simulation) scenes they measure nine-step latent rollout error, keep the best error at each compute budget, and fit four candidate curves on the 22M–2B frontier. The form that best extrapolates to the held-out 4B and 8B models is the second-order power law L(C)=E+A·C^(α−γ ln C), where C is training FLOPs and E is an irreducible loss.

That is a compute law for this recipe, not a new Chinchilla ratio for every robot dataset. For the language-model version of “balance parameters and tokens,” see our Chinchilla explainer. The RoboJEPA authors say their fits already sit close to data saturation because they train multiple epochs on a fixed corpus.

Downstream, they report a capability order. End-effector 3D reaching appears around 10^20 FLOPs. Moving while holding an object appears around 3×10^20. After they lengthen the cooldown rollout to K=10 steps, obstacle avoidance takes off around 10^21 FLOPs and saturates around 10^22. On the object-push task, no model in their plot records a positive success rate until about 10^22 FLOPs. Those thresholds are read from their figures and text, not from a third-party rerun.

Authors’ reported compute thresholds for RoboCasa planning capabilities
CapabilityApprox. training computeAuthors’ note
End-effector 3D reach~10^20 FLOPsGreedy pose reaching; saturates early
Hold object while moving~3×10^20 FLOPsStill greedy; 50M–100M models on the full mix
Obstacle reach~10^21 to ~10^22 FLOPsNeeds long-horizon (K=10) cooldown
Object pushafter ~10^22 FLOPsNo positive success before that compute

What do they report on a real Franka?

Section 3.4 deploys the same family on a DROID Franka with a left external camera and a wrist camera. The goal is one final image, not a language instruction. Tasks are Grasp, Object Lift, and Pick and Place. Table 3, which they say is a compressed view of Table 12, lists progress and success percentages for RoboJEPA-4B, RoboJEPA-8B, π0-FAST and π0.5.

RoboJEPA-8B is listed at Grasp 67/67, Object Lift 65/50 and Pick and Place 42/27. π0.5 is listed at 25/5, 12/0 and 69/53. The authors write that the VLA baselines differ in goal specification and training mix, that neither side received task-specific fine-tuning, and that the VLA rows are “contextual references rather than directly comparable baselines.” They hypothesise that π0.5’s 0% lift score may reflect DROID’s pick-and-place-heavy language, which a goal-image planner would not see.

Authors’ Table 3 real-robot success rates (%). Not an independent rerun.
ModelGoalGrasp successLift successPick and Place success
π0-FASTText221240
π0.5Text5053
RoboJEPA-4BImage603021
RoboJEPA-8BImage675027

Can you download the checkpoints?

The abstract, the footnote and the conclusion all say the authors release all model checkpoints plus training and robot-deployment code. The paper prints https://github.com/facebookresearch/robo_jepa and https://robojepa.github.io.

We opened both on 9 October. The GitHub HTML page and the GitHub API both returned HTTP 404. The project page is a stub: “RoboJEPA / Scaling Robotic Latent World Models / Project page coming soon.” Hugging Face’s papers page for 2610.10515 was reachable; it is a paper landing, not a weight repo. We did not find a public safetensors, GGUF or license file we could inspect.

Until that tree or another first-party download resolves, the release sentence in the PDF is a claim about intent, not an artifact we could clone. Do not treat a 404 as a temporary CDN glitch we verified later in this run; it was 404 at check time.

What is not established?

The authors list the limits themselves. The encoder is frozen, so the law is not a joint scaling result for representation plus dynamics. The corpus is fixed, so they expect more diverse interaction data rather than parameters alone if the frontier is to move. Planning uses one image goal and no language. Actions are sampled uniformly, then ranked by the world model. We did not count FLOPs, rerun CEM, or touch a Franka.

The “first scaling law for multi-embodiment robotic world models trained on real robot data” and “largest JEPA predictor” lines are the authors’ priority claims. The VLA comparison is theirs, with the mismatch they disclose. This is an evidence review of pages opened on 9 October. It is not a robot test and not a ranking of RoboJEPA against Long-WAM, Cosmos or π0.5 on a shared protocol.

Common questions

Are RoboJEPA weights public today?

Not on the pages we opened on 9 October. The paper names github.com/facebookresearch/robo_jepa. That repository and its API record returned 404. The project page has no files.

Is this the same model as V-JEPA 2?

No. RoboJEPA keeps a frozen V-JEPA 2.1-G encoder and trains a larger action-conditioned predictor on a multi-embodiment robot mix. V-JEPA 2-AC, as the authors describe it, was a smaller post-train on less robot data.

Did the 8B model beat π0.5?

Not as a single score. On the authors’ Table 3 it leads Grasp and Lift and trails Pick and Place. They say the comparison is contextual because goals and training mixes differ.

THE TAKEAWAY

What to remember

Use the 7 October preprint for the architecture, the second-order compute law and Table 3. Use the 9 October GitHub 404 and the stub project page for availability. Do not plan a download day until a first-party tree actually serves files.

Sources & further reading

  1. RoboJEPA: Scaling Robotic Latent World Models ↗
  2. RoboJEPA: Scaling Robotic Latent World Models (HTML) ↗
  3. arXiv API record 2610.10515 ↗
  4. RoboJEPA project page ↗
  5. facebookresearch/robo_jepa repository ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories