What was posted on 8 October, and what is the September page?

The arXiv Atom record we opened lists 2610.12424v1 as published at 17:55:12 UTC on 8 October 2026, primary class cs.RO, with cs.AI secondary. Authors start Zimo Wen, Yijin Chen, and Yuxuan Cao, each marked equal contribution in the HTML, then Wendi Chen, Yanwen Zou, Wenye Yu, Fuhang Kuang, Han Xue, Jun Lv, and corresponding authors Chuan Wen and Cewu Lu.

The paper’s abstract points to a project page and to github.com/nssmd/RoboRSI. The URL printed in the HTML, lab.noematrix.ai/blog/2-roborsi-research-preview/, redirected when we opened it to lab.noematrix.ai/blog/2-roborsi/. That page titles itself “RoboRSI Research Report,” names “Noematrix Team,” and prints “Published September 2026.” The README citation uses month = sep and the same Noematrix Team author. GitHub’s repo API, opened 9 October, lists created_at 2026-08-29T08:23:35Z, pushed_at 2026-10-09T12:56:26Z.

This is a code-as-policies self-improvement harness, not a new world model. For other 8 October physical-AI write-ups on this site, see ARC’s robot-reasoning recipe and CASIA and Amap’s DreamTrue note.

What is Top-Down Skill Refinement?

TSR, the paper’s organizing idea, decomposes a task into compound, atomic, and base skills with scoped responsibilities and input–output contracts. A failure is attributed to the responsible branch: a bad decomposition revises the task structure, a bad composition revises the compound skill, and a failed operation revises the atomic or base skill.

Four agents run each round. The Manager holds the hierarchy, skill versions, and release decisions. The Planner returns an executable plan for one atomic task. The Engineer calls tools and implements missing skills. The Reviewer judges postconditions from fresh observations and tool traces and writes a revision patch. The HTML says every atomic task has a postcondition checked against new observations before the agent proceeds.

Stable sequences are consolidated into compound skills so later episodes select and monitor a parameterized skill instead of reconstructing every tool decision. People, the authors say, still enter objectives, constraints, and corrections at the task level; those notes stay attached to the responsible skill.

Which Table 1 numbers are the authors’?

Section 4.2 says the main comparison is 5,974 evaluation episodes, plus 2,770 ablation episodes and 104 physical rounds. All simulated methods use GPT-5.6-SOL. On LIBERO, LIBERO-PRO, and RoboTwin, RoboRSI starts from an initial library and improves online during evaluation. On LIBERO-Plus the LIBERO library is held frozen. Maestro, “the most costly baseline,” is run only on LIBERO and LIBERO-PRO.

Table 1 prints success as successful/valid episodes and coverage as solved/total tasks, counting each task once after its first success. LIBERO is 30 tasks × 5 episodes; LIBERO-PRO is 120 × 5; LIBERO-Plus is 840 perturbation instances of 30 tasks; RoboTwin baselines are 50 × 3, while RoboRSI reports the first 154 episodes of its online run on the same 50 tasks.

The authors’ “highest success on all four benchmarks” sentence uses those Table 1 episode rates: +5.3 points on LIBERO versus OpenETA, +11.0 on LIBERO-PRO, +5.7 on frozen LIBERO-Plus, +2.7 on RoboTwin. About two thirds of OpenETA and Maestro failures on LIBERO-PRO are premature completions — the agent stops after a tool reports a release. For RoboRSI that share is 10% and 14% on LIBERO-PRO and LIBERO-Plus.

Author-reported Table 1 rows from the HTML, not an independent rerun. Success is successful/valid episodes.
BenchmarkOpenETA successRoboRSI successRoboRSI coverage
LIBERO (online)76/150 (50.7%)84/150 (56.0%)24/30 (80.0%)
LIBERO-PRO (online)231/600 (38.5%)297/600 (49.5%)102/120 (85.0%)
LIBERO-Plus (frozen library)306/840 (36.4%)354/840 (42.1%)28/30 (93.3%)
RoboTwin (online; 154 RoboRSI episodes)32/150 (21.3%)37/154 (24.0%)26/50 (52.0%)

How does the README’s 95/120 panel differ?

The README table we opened on 9 October lists “LIBERO cumulative task pass rate 95/120” and, in the same cell, “ten sequential rounds moved 32/120 → 83/120.” The Noematrix page says that with parallel execution, “cumulative task pass rate reaches 95/120 on LIBERO after approximately one day.” Paper §4.3.2 is a third panel: a one-day self-iteration on a 120-task LIBERO catalog in which coverage grows from 32 to 71 of 120 in five rounds, taking 7.0 hours of active execution, “separate from the evaluation in Table 1.”

Those cumulative counts mark a task passed at least once across evolving releases. Table 1’s 84/150 is successful episodes on a 30-task × 5-episode protocol with online revision. The README also prints LIBERO-Plus 398/840 adaptive Pass@2 against a “fixed release 261/840,” which is not Table 1’s frozen 354/840. RoboTwin in the README is 36/50 cumulative versus a 9/50 single-role baseline, not Table 1’s 37/154.

The repo’s reproduce_libero_pro.sh, the README says, “evaluates the current frozen release; it does not replay the cumulative results above.” That sentence is the authors’ own split between what the paper claims and what a fresh clone measures.

What did they show on the physical robot?

Appendix B names a WheelSingleArm M1: an Athena Pro Max mobile base, a RealMan arm, and a Zhiyuan gripper, with in-hand and chassis RGB-D cameras. The case study is household floor cleanup — search, approach, grasp, carry, place — across 104 physical development runs and 24 hours of cumulative operation. The navigation map stays fixed; objects and receptacles are partially randomized; on-site personnel restore a common base layout and confirm a placement only when they see the object enter its designated receptacle. For another 8 October latent world-model paper, see Meta FAIR’s RoboJEPA. For a vendor physical-AI packaging note, see AWS’s Physical AI toolchain.

Figure 4 follows ten selected runs. The earliest selected run placed no objects; later runs placed several and sometimes returned to the dock. Progress was not monotone. The 104 runs “trace development of an evolving system across varying task manifests, rather than repeated trials of a frozen policy on an identical task.” A separate pouf-in-the-wrong-bucket failure was revised from retained observations and was not given a matched post-repair physical trial.

The Noematrix page’s “first deployment of a Multi-Agent framework on a real-world mobile robot” and the 2024 “two engineers… about one month… 2,000–3,000 lines” contrast are that page’s claims. They are not Table 1.

What can you run, and what did we not run?

The public repo is nssmd/RoboRSI. The README offers a recursive clone, an OPENAI_API_KEY for any OpenAI-compatible Responses endpoint, and scripts/reproduce_libero_pro.sh, which it says creates an isolated environment, clones LIBERO-PRO, pulls zhouxueyang/LIBERO-Pro assets, and runs a frozen Pass-1 campaign. A web console is documented on port 8787. We did not clone the repo, spend API credits, or replay an episode.

This is an evidence review of the 8 October abs, HTML, and Atom record, the 9 October README and GitHub API, and the September Noematrix page. Treat every success rate as the authors’.

Common questions

Is 95/120 the same result as 84/150?

No. 84/150 is Table 1’s successful-episode rate on 30 LIBERO tasks × 5 episodes with online self-improvement. 95/120 is a cumulative task-pass count on a 120-task catalog across evolving releases, printed on the README and the September project page. Paper §4.3.2’s 32→71 of 120 in five rounds is a third panel.

Did they release weights for a new robot foundation model?

No. The inspectable artifact is a multi-agent harness and evaluation scripts. The simulated comparisons use GPT-5.6-SOL as the backbone. There is no Hub checkpoint in the pages we opened.

Is the physical 104-run study a controlled A/B test?

The paper presents it as a development trace: skills may change between runs, manifests vary, and placements are confirmed by people on site. It is not a frozen-policy evaluation on one repeated task, and one reported placement failure was not given a matched post-repair retry.

THE TAKEAWAY

What to remember

Use 2610.12424 for Top-Down Skill Refinement and for Table 1’s 84/150, 297/600, 354/840, and 37/154 rows. Use the README for the frozen reproduction path and keep 95/120 in the cumulative-coverage column. Do not invent a new robot foundation model or an independent LIBERO rerun.

Sources & further reading

  1. RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments ↗
  2. RoboRSI (HTML) ↗
  3. arXiv Atom record 2610.12424 ↗
  4. nssmd/RoboRSI README ↗
  5. nssmd/RoboRSI repository API ↗
  6. RoboRSI Research Report ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories