What appeared on arXiv and GitHub on 8–9 October?

The arXiv Atom record we opened for 2610.12403v1 lists published 2026-10-08T17:44:19Z, primary class cs.CV, with authors Hongxing Li, Dingming Li, Yixin Li, Yong Du, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, and Yongliang Shen. The HTML affiliations we inspected place the team at Zhejiang University; the corresponding email on the page is hongxing.li@zju.edu.cn.

GitHub’s API object for ZJU-REAL/ViSkill reports created_at 2026-10-08T07:07:53Z, pushed_at 2026-10-09T03:21:55Z, default branch main, and license MIT. The LICENSE file we opened is titled MIT License with copyright line “Copyright (c) 2025 RAGEN.AI.” The README News block says code and Sokoban/FrozenLake models shipped on 8 October and that the paper was listed on arXiv on 9 October—treat those as the project’s own changelog dates next to the Atom timestamp.

This is a skill-library reinforcement loop for vision-language agents, not a new foundation model. For nearby 8 October skill and robot write-ups on this site, see RoboRSI’s top-down skill refinement and Tencent’s TeamAI Git harness for shared agent skills.

How does a visual skill card differ from a text skill?

The paper’s opening claim is that most skill-augmented agents linearize spatial layouts into language and lose geometry that later episodes need. ViSkill instead keeps a composite skill card: a rendered annotated trajectory image plus a short distilled strategy text, a geometric descriptor for retrieval, a utility score, and a use count.

At episode start the agent retrieves a card by blending geometric similarity with historical utility, gated by a retrieval threshold. Retrieved cards condition the policy during interaction and can add a skill-guided reward when the episode fails the environment’s binary success check. Successful trajectories that clear length and diversity gates are distilled into new library entries, so the library and the PPO policy co-evolve.

An optional cold-start path seeds the library with solver-derived cards before PPO. The README says cold start is off by default and needs an external API base URL and key in the YAML files under examples/train/. The paper names GPT-5.4-mini as the external model used for cold-start strategy distillation in their runs—that is a training-time dependency description, not a claim that GPT-5.4-mini ships inside the Hub checkpoints.

Diagram of ViSkill’s loop: skill retrieval, agent–environment interaction, skill-guided PPO optimization, and distillation into a visual skill library.
ViSkill methodology diagram from figures/methodology.png in ZJU-REAL/ViSkill (MIT License), opened 11 October 2026. The schematic is the authors’ illustration of retrieval, interaction, PPO, and distillation; it is not a measured training run. Photo: ZJU-REAL/ViSkill contributors. MIT License · Cropped and resized.

Which Table 1 numbers are the authors’?

Table 1 compares proprietary prompting baselines, open Qwen2.5-VL sizes, VLM-R1-3B, VAGEN, Atlas-VA, and the two ViSkill rows. ViSkill without cold-start is printed at 0.82 Sokoban, 0.84 FrozenLake, 1.00 on every PrimitiveSkill column (Place, Stack, Drawer, Align, Swap), and 0.89 overall. ViSkill + Cold-Start rises to 0.88 / 0.85 / 1.00 average / 0.91 overall.

Atlas-VA is the strongest non-ViSkill overall row at 0.87; VAGEN is 0.80. Among proprietary rows, GPT-5 is 0.70 overall and Claude 4.5 Sonnet is 0.58. The paper states that ViSkill and the RL baselines share a Qwen2.5-VL-3B backbone and environment seeds. Those are the authors’ evaluation settings. We did not re-prompt GPT-5 or retrain Atlas-VA.

Table 2’s representation swap is the sharpest ablation we opened for the visual-card claim: Text-Skill drops Sokoban to 0.63 and FrozenLake to 0.59; OCR-Skill is 0.68 / 0.62; full ViSkill stays at 0.82 / 0.84. Table 3 says removing optimization while keeping retrieval and distillation collapses Sokoban to 0.29 and FrozenLake to 0.33—the co-evolution claim rests on that row more than on the cold-start bump.

Author-reported Table 1 rows from the HTML we opened 11 October 2026. Not an independent rerun.
MethodSokobanFrozenLakePrimitiveSkill avgOverall
VAGEN0.790.740.880.80
Atlas-VA0.790.831.000.87
ViSkill0.820.841.000.89
ViSkill + Cold-Start0.880.851.000.91
Author-reported success-rate table comparing proprietary, open-source, and ViSkill rows on Sokoban, FrozenLake, and PrimitiveSkill.
Table graphic from figures/results.png in ZJU-REAL/ViSkill (MIT License), opened 11 October 2026. Numbers are the authors’ evaluation rows from the ViSkill paper; Ai Lookout did not rescore the environments. Photo: ZJU-REAL/ViSkill contributors. MIT License · Cropped and resized.

What can you download, and under which licenses?

The training and agent code live at github.com/ZJU-REAL/ViSkill under the MIT LICENSE we opened. The README’s install path is conda python=3.10, bash install.sh, optional bash install_render.sh for ManiSkill/PrimitiveSkill, then environment-specific train_ppo_qwen25vl3b_skill.sh scripts. Default backbone text in those scripts is Qwen2.5-VL-3B-Instruct.

Hugging Face hosts hongxingli/ViSkill-Sokoban (created 2026-10-08T11:14:08Z, lastModified 2026-10-09T06:29:15Z) and hongxingli/ViSkill-FrozenLake (created 2026-10-08T15:43:49Z, lastModified 2026-10-09T03:21:35Z). Both cards we opened list license apache-2.0, base_model Qwen/Qwen2.5-VL-3B-Instruct, and pipeline_tag image-text-to-text. Each ships model shards plus a skill_library/ directory of PNG skill cards. The Sokoban card’s transformers load snippet uses Qwen2_5_VLForConditionalGeneration; the card tells readers to use the GitHub repo for the full retrieval loop.

Keep license columns separate: MIT on the training tree, Apache-2.0 on the Hub fine-tunes, and whatever terms attach to Qwen2.5-VL-3B-Instruct as the base. For another open from-scratch component aimed at agent search stacks, see Waterloo’s Project Greenhouse Gaggle reranker. For test-time spatial scaffolding without retraining, see SpatialHarness.

What should a team verify before trusting the 0.89 headline?

Pin the paper version (2610.12403v1), the GitHub commit you clone, and the Hub revision SHAs we recorded for the Sokoban and FrozenLake cards. Re-run the authors’ YAML seeds before comparing to Atlas-VA or VAGEN; backbone, temperature, and environment server URLs all sit in those configs.

Treat PrimitiveSkill’s 1.00 average as a saturated sim column, not evidence for warehouse manipulation. The paper’s ethics note, as summarized in secondary coverage we opened, flags that experiments are simulated and that real deployment needs task-specific safety constraints—we quote that as a limitation, not as a lab we visited.

Cold-start needs an external model API in the authors’ scripts. If you only download the Hub fine-tunes, you get a policy checkpoint and skill PNGs, not an automatic library-growth service. Budget GPU time for PPO before you promise co-evolution on your own tasks.

We did not clone the repo, start ManiSkill, download multi-gigabyte safetensors, or reproduce Table 1. This article is an evidence review of the abs, HTML, Atom record, README, LICENSE, GitHub API object, and Hub cards opened on 11 October 2026.

Ivy-covered red-brick teaching building with a light stone central entrance and a blue shared bicycle on a stone path. No people appear.
Third Teaching Building, Zhejiang University Yuquan Campus, photographed 5 October 2017 by 猫猫的日记本. CC BY-SA 4.0 archival photograph via Wikimedia Commons. No identifiable people appear. The 2017 campus exterior does not depict ViSkill GPUs, skill cards, or Sokoban evaluation. Photo: 猫猫的日记本. CC BY-SA 4.0 · Cropped and resized.

Common questions

Are the ViSkill Hub weights MIT like the GitHub tree?

No. The ZJU-REAL/ViSkill LICENSE we opened is MIT. The hongxingli/ViSkill-Sokoban and ViSkill-FrozenLake cards we opened list Apache-2.0. Keep those columns separate from the Qwen base-model terms.

Does 0.91 mean ViSkill beats GPT-5 on Sokoban in production?

No. 0.91 is the authors’ overall average for ViSkill + Cold-Start on their three sim suites with a Qwen2.5-VL-3B RL policy. GPT-5’s 0.70 overall in Table 1 is a proprietary prompting baseline in the same paper table, not a production deployment study.

Is PrimitiveSkill a real robot benchmark?

In this paper PrimitiveSkill is a simulated, coordinate-grounded manipulation suite (Place, Stack, Drawer, Align, Swap) that the authors evaluate alongside 2D Sokoban and FrozenLake. The pages we opened do not present it as a physical Franka trial.

THE TAKEAWAY

What to remember

Use 2610.12403 and the MIT GitHub tree for the visual skill-card loop; use the Apache-2.0 Hub cards for downloadable Sokoban and FrozenLake fine-tunes; keep 0.89 / 0.91 in the authors’ Table 1 column.

Sources & further reading

  1. ViSkill: Reinforcing VLM Agents with Evolving Visual-Native Skills ↗
  2. ViSkill (HTML) ↗
  3. arXiv Atom record 2610.12403 ↗
  4. ZJU-REAL/ViSkill README ↗
  5. ZJU-REAL/ViSkill LICENSE (MIT) ↗
  6. ZJU-REAL/ViSkill repository API ↗
  7. hongxingli/ViSkill-Sokoban model card ↗
  8. hongxingli/ViSkill-FrozenLake model card ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories