Announced 8 Oct 2026 · Sources checked
What did the Fudan team publish on 8 October?
The arXiv export API lists 2610.12457 published 2026-10-08T17:59:20Z. Corresponding author Jingjing Chen is marked at Fudan’s Institute of Trustworthy Embodied AI; coauthors span Fudan’s College of Computer Science and Artificial Intelligence and Singapore Management University.
The project website at emilia113.github.io/SpatialHarness restates the problem in plain language: multimodal foundation models can emit robot actions from vision and language, yet fine insertion, stacking, and precise placement still fail when fixed cameras hide the geometry that matters. The authors ask whether better observations at test time can help without retraining the policy or moving the physical cameras.
That framing sits next to other October robot-reasoning recipes such as Arc’s action-grounded traces and DreamTrue’s action-faithful world model. SpatialHarness is not another pretraining story. It is an embodied harness around a frozen controller.
We opened the HTML paper, the project site, the GitHub README, and the Hugging Face results dataset API on 10 October 2026. We did not download trial videos or reproduce the 120-trial table.
How does SpatialHarness change what the policy sees?
Before a trial, the system reads success conditions and works backward to the spatial relationships that must be judged—plug-to-socket alignment, peg-to-hole fit, gripper-to-handle contact. It then chooses fixed virtual viewpoints that reveal those relationships. Tasks that need local mating features get one local and one global view; whole-object relations get two global views.
During execution the harness keeps a simulated scene aligned with the real cell. Measured robot motion advances the simulation. Interaction-aware synchronization distinguishes static objects, continuously held objects, and grasp/release transitions, using visual correction, spatial-constraint programs, and hypothesis replay when contact outcomes diverge. When rendered views disagree with real observations, the harness marks them unreliable so the policy can fall back to physical cameras.
The frozen multimodal policy—GPT-6 Astra at low reasoning effort in the authors’ setup—receives the usual real observations plus rendered complementary views, object poses, and reliability flags. Actions still come from the same policy through tool calls. Nothing in the reported protocol fine-tunes Astra’s weights for these four tasks.
If you already follow OpenAI’s consumer GPT-6 / Intelligent UI rollout, treat Astra here as the authors’ robot-control configuration of that model family—not as a claim about ChatGPT’s Intelligent UI surface. See our GPT-6 Intelligent UI note for the product announcement path.

What success rates are claimed on the real robot?
The evaluation uses four tasks: block stacking, plug insertion, Tower of Hanoi, and block in drawer. Both baseline and SpatialHarness runs use the same physical cameras and control protocol. Each method gets 15 trials per task and 20 interaction rounds per trial—60 trials per method, 120 total.
The authors report average success rising from 18.33% to 83.33%. Plug insertion rises from 26.67% to 66.67%. Tower of Hanoi rises from 0% to 100%. The GitHub README’s caption for Figure 4 further breaks the counts as block stacking 7/15→15/15, plug insertion 4/15→10/15, Hanoi 0/15→15/15, and block in drawer 0/15→10/15 (overall 11/60→50/60). Those fractions match the percentages when rounded as the paper presents them.
| Task | Baseline (15 trials) | SpatialHarness (15 trials) |
|---|---|---|
| Block stacking | 7/15 (paper figure) | 15/15 |
| Plug insertion | 4/15 → 26.67% | 10/15 → 66.67% |
| Tower of Hanoi | 0/15 → 0% | 15/15 → 100% |
| Block in drawer | 0/15 | 10/15 |
| Average across four tasks | 18.33% | 83.33% |

What can you open today, and what is still missing?
The project site hosts method copy, demo excerpts, and links to the arXiv HTML/PDF. The GitHub repository emilia113/SpatialHarness currently serves the project documentation tree (README, docs/, posters, figures) and was pushed 2026-10-10T10:47:37Z on the copy we opened. The site’s Code button still reads “Coming soon.” The API license field is empty—do not invent an Apache or MIT grant for harness source that is not published yet.
Hugging Face dataset yzy05/SpatialHarness holds full-result videos (baseline and ours, external and wrist views) across the four tasks. The dataset API we opened lists createdAt 2026-10-09T08:31:12Z and lastModified 2026-10-10T10:52:19Z. That is useful for qualitative inspection of failure modes; it is not a pip-installable harness.
Compare that openness profile with Meta FAIR’s RoboJEPA drop, where a named facebookresearch tree now serves code and non-commercial weight URLs. SpatialHarness today is paper + demos + result videos, with reproducible harness code still pending the authors’ release.
What should readers not conclude?
Do not read 100% Hanoi as proof that every GPT-6 Astra deployment will clear peg puzzles. The figure is fifteen trials in the authors’ cell with their synchronization stack and low-reasoning Astra setting. Different cameras, lighting, object meshes, or simulator mismatch can erase the gain.
Do not treat the harness as free compute. Maintaining a synchronized scene, rendering extra views, and asking the multimodal policy to consume them adds latency and modeling work even when weights stay frozen. The paper’s point is that this cost can be cheaper than collecting new demonstrations or fine-tuning a VLA—not that the stack is zero-overhead.
Do not confuse illustrative robot photography with the evaluation cell. Cover and inline images in this article are archival Commons photographs; they do not depict Fudan’s setup or Astra tool calls.
Common questions
Does SpatialHarness fine-tune GPT-6 Astra?
Not in the reported protocol. The authors keep Astra frozen and change the observations and spatial context supplied at test time. Any later fine-tune would be a different experiment.
Is the harness code public?
Not as a complete installable package on the copy we opened. The GitHub repository hosts the project site assets and README; the website still labels Code as coming soon. Result videos are on Hugging Face.
Why might a strong policy still fail fine insertion?
The authors argue the failure is often insufficient spatial evidence in the fixed camera views—occlusion and poor viewing angles—rather than a total absence of manipulation skill. Their results support that hypothesis inside their four-task suite; they do not prove it for every contact-rich task.
What to remember
SpatialHarness is an 8 October test-time recipe: synchronize a simulator, render the views that reveal mating geometry, and keep GPT-6 Astra frozen. The author-reported jump from 18.33% to 83.33% average success is large enough to watch—and incomplete enough that you should wait for harness code before planning a production port.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





