Announced 8 Oct 2026 · Sources checked
What landed with OneSearch-VL?
The arXiv export API lists 2610.12419 published 2026-10-08T17:52:46Z. The README frames the product as a unified agent for single-image, multi-image, and video deep research: locate visual clues, retrieve external knowledge, and compose source-supported answers.
The GitHub API record we opened for appletea233/OneSearch-VL shows created_at 2026-10-08T12:52:29Z, license Apache-2.0, and description matching the paper. The README points to SFT/, RL/, and onesearch_vl/ packages, plus a Hugging Face collection at huggingface.co/OneSearch-VL.
That puts OneSearch-VL in the same practical lane as other open research-agent stacks we have covered, including IQuest’s SAIL scientific agent, while staying closer to visual grounding than text-only literature search.
We opened the HTML paper, README, GitHub API JSON, and Hub model/dataset API records on 10 October 2026. We did not download the 8B checkpoint or run prepare.py on the benches.
Authors listed on the HTML abstract include Hongyu Li, Manyuan Zhang, Kaituo Feng, Shu Chen, Dian Zheng, Hao Li, Hao Yu, Zhangquan Chen, Zoey Guo, Ray Zhang, Shaofei Huang, Tianrui Hui, Linjiang Huang and Si Liu. The paper’s contribution list is the VGEG-centered engine, the EVGR reward, the two OneSearch benches, and the joint SFT/RL recipe that produces OneSearch-VL-8B—not a claim that every hosted deep-research product is obsolete.
What is a Visually Grounded Evidence Graph?
VGEG is the authors’ shared task-level structure. It encodes dependencies among localized visual anchors, entity relations, source-supported facts, and the operations that produce an answer. The same graph shape is reused three ways: to construct and verify questions and expert trajectories, to derive EVGR rewards during RL, and to organize evaluation by research operation.
The data engine discovers visual anchors, attaches verified web facts, and builds compositional questions with explicit evidence dependencies. Expert trajectories are filtered for answer correctness and process quality. Together with adapted single-image data from OpenSearch-VL, that yields OneSearch-VL-SFT-110K (~110K trajectories) and OneSearch-VL-RL-10K (~10K RL tasks).
EVGR scores whether a rollout obtained the required facts (evidence traceability) and used the correct images, regions, or video frames (visual grounding). Combined with answer and query rewards, it is meant to punish fluent answers that skip the visual chain. The README’s ablation line says full EVGR adds 3.8 points over answer-and-query rewards on the paper’s six-benchmark ablation average—still an author-reported delta.
Operation labels on the benches matter for how you read the leaderboard. A single-anchor lookup that fails because OCR missed a logo is a different bug from a multi-anchor join that retrieves two correct facts and then subtracts the wrong years. The VGEG design is meant to make those failure modes separable during both training rewards and evaluation slices.

What benchmarks and scores are claimed?
OneSearch-MI-Bench holds 301 multi-image questions; OneSearch-Video-Bench holds 307 video questions. Both organize items around six operations: single-anchor lookup, multi-hop retrieval, knowledge-conditioned count, multi-anchor join, arithmetic, and comparison. Hub API lastModified stamps we opened are 2026-10-08T03:14:45Z for MI-Bench.
Versus Qwen3-VL-8B with tool access, the authors report +20.2 percentage points on OneSearch-MI-Bench, +17.6 on OneSearch-Video-Bench, and +27.0 on VideoDR. On seven single-image research benchmarks they report 58.3 average accuracy for OneSearch-VL-8B, +1.7 over OpenSearch-VL-8B. Joint SFT wins their data-mixture ablation average.
| Setting | Authors’ figure | Baseline they cite |
|---|---|---|
| OneSearch-MI-Bench | +20.2 points | Qwen3-VL-8B with tools |
| OneSearch-Video-Bench | +17.6 points | Qwen3-VL-8B with tools |
| VideoDR | +27.0 points | Qwen3-VL-8B with tools |
| Seven single-image benches (avg) | 58.3 accuracy (+1.7) | OpenSearch-VL-8B |
| EVGR ablation average | +3.8 points | answer + query rewards only |

How do you actually run it?
The README’s getting-started path is clone the Apache-2.0 repo, install from SFT/README.md (Python 3.11+), download SFT-110K or the RL-10K bundle, and point launchers at a Qwen3-VL-8B-Instruct base or the published OneSearch-VL-8B checkpoint. Inference wrappers evaluate single-image parquet benches, OneSearch-MI manifests that preserve image-index mapping, and video preparations via each dataset’s prepare.py.
The Hub model API for OneSearch-VL/OneSearch-VL-8B lists createdAt 2026-09-29T07:38:30Z and lastModified 2026-10-07T23:23:50Z, with model.safetensors among the sibling files. Treat the September create stamp as Hub provenance for the weight repo and the 8 October arXiv/GitHub pair as the public paper-and-code announcement we are covering.
If your stack already uses Qwen image models, keep licenses straight: Qwen-Image-2.1-Turbo ships under Qwen’s research terms, while this agent repo is Apache-2.0 and the Qwen3-VL base retains its own card terms. Read both before you ship a product path.
The repository layout separates concerns: SFT recipes under SFT/examples/onesearch, RL launchers that load RL/.env, and onesearch_vl inference/eval wrappers. That split is helpful if you only want to score the published 8B checkpoint without standing up Megatron training. It still assumes you can host vision tools and search backends compatible with the authors’ wrappers.
What should teams test before trusting the lifts?
Rerun the authors’ evaluation scripts against your tool environment. Search APIs, video frame samplers, and fetch limits differ; a 20-point gap measured with the paper’s tools can shrink when your tools refuse URLs or truncate frames.
Watch token and search budgets. The paper’s pitch includes structured evidence reuse; your deployment may still thrash the search API if termination conditions differ. Log tool calls per question type before you compare cost to a hosted deep-research SKU.
Inspect failure modes on multi-anchor join and arithmetic items. The README’s Ford/Brembo case study is the intended happy path—ground two brands, retrieve founding years, subtract—but production questions often mix OCR noise, stale web facts, and ambiguous crops.
For retrieval quality outside this agent, keep a separate embedder bake-off. Our Perplexity pplx-embed-v2 note is about late-interaction text/code embeddings, not visual deep research, but the same rule applies: vendor averages are not your corpus.
We did not measure latency, tool-call cost, or hallucination rates on OneSearch-VL-8B.
Common questions
Is OneSearch-VL only for videos?
No. It is trained as a shared policy across single-image, multi-image, and video research with a common tool environment. The new benches emphasize multi-image and video because those were the weaker open baselines.
What is open versus author-reported?
Open on the copies we opened: Apache-2.0 GitHub code, Hub 8B weights, SFT/RL datasets, and the two OneSearch benches. Author-reported: the percentage-point lifts versus Qwen3-VL-8B with tools and the EVGR ablation delta.
Do I need the VGEG at inference time?
The agent receives the visual input, question, and tools and gathers its own evidence. VGEGs are described as supervision and evaluation references during data/RL/benchmark construction, not as a user-supplied graph you must draft for every query.
What to remember
OneSearch-VL is an 8 October open multimodal deep-research agent: VGEG-centered data, EVGR rewards, Apache-2.0 code, and an 8B checkpoint with large author-reported gains on new multi-image and video benches. Clone it if you can rerun their evals; do not paste the +20-point headline into a vendor deck without that rerun.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





