Announced 9 Oct 2026 · Sources checked
What appeared on Hugging Face on 9 October?
The Hub API object we opened lists id tencent/Youtu-Parsing-Omni, createdAt 2026-10-09T06:42:18.000Z, lastModified 2026-10-09T09:43:48.000Z, sha 01f7e3e4916898c30bf9593735a9b3bd328492fa, pipeline_tag image-text-to-text, library_name transformers, and license other with license_name youtu-parsing. The safetensors block reports 5,334,472,192 BF16 parameters. The card’s own table says “5B.”
The README news block has two lines: “[2026-10] Technical report, model weights, vLLM plugin and inference examples released,” and “[TBD] Open-source evaluation code will be released.” The technical-report link points at a PDF inside TencentCloudADP/youtu-parsing/youtu_parsing_omni/paper/. The citation still prints “arXiv preprint arXiv:[TBD].”
This is a parser-weight drop, not a hosted document API. For a same-week Apache-2.0 OCR family that returns boxes, see LightOnOCR-3. For why transcription and document understanding are different jobs, see OCR versus document AI.
What does Tencent say one 5B model can parse?
The card’s one-sentence claim is that a single input — a document page, a natural image, a chart or flowchart, a geometry figure, an audio clip, or an audio-visual video — produces one structured JSON envelope. Perception fields listed are layout elements, text, tables, formulas, bounding boxes, timestamps, ASR, OCR, acoustic events, and camera motion. Cognition fields listed are captions, narratives, and reports. The output family is selected by a task prompt from prompts/youtu_parsing_omni.json.
Eight tasks are tabulated. document returns layout elements with boxes, text, LaTeX, OTSL tables, Markdown charts, Mermaid flowcharts, and reading order. natural_image returns entities and text with boxes, tags, captions, and a global description. graphics_chart and graphics_flowchart each return one element. graphics_geometric returns points, lines, arcs, shapes, relations, and measurements. audio returns vocal and non-vocal segments with timestamps, speakers, ASR, and acoustic events. natural_video and textrich_video return temporal segments; the text-rich path also emits a Markdown structured_report.
That is one model claiming several modalities, not proof that one checkpoint matches a specialist on every row. For the broader multimodal framing, see multimodal AI explained. For a hosted router that swaps parsers rather than shipping one 5B checkpoint, see LlamaIndex OpenDocRouter.
How does Tencent describe the encoder and the training loop?
The method section contrasts “existing omni-modal parsers” that keep separate vision and audio towers with a single Youtu-Omni-Encoder “initialized from a pre-trained text language model.” Thin stems map pixels and log-mel frames to tokens. A shared bidirectional Transformer packs images, audio, and video under one (t, h, w) encoding. Mergers feed a Youtu-LLM decoder that emits OmniSchema JSON.
For video, frames and audio chunks are packed in time. Fusion layers start a new attention window at every audio-to-frame boundary. Vision tokens take t from the frame index; audio tokens take t from their time offset with h = w = 0.
Post-training is named OmniSchema-Aware On-Policy Distillation (OSAD). After SFT and schema-routed RLVR, a student samples its own parses. Rollouts that pass JSON, schema, box, and repetition checks are scored by a frozen copy that also sees the reference parse. Content tokens use Jensen–Shannon divergence; structure tokens use forward KL. Those are Tencent’s training claims. We did not train a round.
Which scores are on the card, and what are they not?
OmniDocBench v1.6 is the document table. Youtu-Parsing-Omni is printed at 96.96 Overall, 0.0271 Text Edit, 96.80 Formula CDM, 96.79 Table TEDS, 98.20 Table TEDS-S, and 0.1104 Reading Order Edit. TeleOCR 1.2B is 96.91 Overall. Gemini-3-Pro is 92.91. GPT-5.2 is 86.59. The older Youtu-Parsing 2.5B row is 93.74. Those numbers are Tencent’s table, not a rerun.
OmniParsingBench is the omni table. Gemini-3-Pro is first at 77.44 average. Youtu-Parsing-Omni is 75.08, with a 95.02 chart score that the table marks as the column lead, 74.33 geometry, 62.84 natural image, 78.08 audio, 74.18 natural video, and 75.62 text-rich video. Logics-Parsing-Omni is 72.26. Qwen3-Omni-30B-A3B is 66.44. GPT-5.4 and Qwen3.5-397B-A17B have dashes on audio and video.
ChemOCR prints 74.6 average similarity and 54.2 Tani@1.0 against DeepSeek-OCR 2 at 74.9 and 49.6. PDMX-Synth, a music-score row, prints character error 22.73, which the table marks as the column lead, with symbol and line error behind LEGATO. Evaluation code is still TBD. We did not score a page.
| Bench | Youtu-Parsing-Omni figure | Comparison printed on the card |
|---|---|---|
| OmniDocBench v1.6 Overall | 96.96 | TeleOCR 96.91; Gemini-3-Pro 92.91 |
| OmniParsingBench average | 75.08 | Gemini-3-Pro 77.44; Logics-Parsing-Omni 72.26 |
| OmniParsingBench chart | 95.02 | Marked as the column lead |
| ChemOCR Tani@1.0 | 54.2 | DeepSeek-OCR 2 at 49.6; DeepSeek leads average similarity 74.9 to 74.6 |
| Parameters | 5B on the card; 5,334,472,192 BF16 on the Hub API | Not a third-party count |
How do you serve it, and what does the license actually say?
The GitHub repository TencentCloudADP/youtu-parsing existed on 21 January 2026 for the earlier Youtu-Parsing line. GitHub’s API listed updated_at 2026-10-09T16:48:30Z and pushed_at 2026-10-09T09:45:14Z when we opened it. The Omni path is youtu_parsing_omni. Setup wants Python 3.10 or newer and CUDA. vLLM serving pins vllm 0.19.0 and transformers 5.2.0 and uninstalls peft. Transformers inference wants transformers 5.10.2 in a different environment. ffmpeg is required for audio and video.
scripts/vllm.sh starts an OpenAI-compatible server on 127.0.0.1:8000. The documented engine flags include max-model-len 65536, max-num-seqs 1, gpu-memory-utilization 0.70, and a multimodal prompt limit of eight images, four videos, and eight audio clips. Decoding in the reported recipe is greedy. Media must sit under DATA_ROOT. The card tells Transformers users to set trust_remote_code=True and, for a Hub id, to pin a reviewed --revision.
The LICENSE file opens: “Youtu-Parsing IS NOT INTENDED FOR USE WITHIN THE EUROPEAN UNION.” Clause 0 says that sentence prevails in a conflict. Copyright is Tencent, 2026. The grant otherwise tracks a permissive, Apache-like text with a patent-termination clause. This is not Apache-2.0. The Hub tag is license:other. Read the file before you ship a parser in the EU.
What should a document team test before swapping a specialist OCR?
The useful trial is a page set you already score: multi-column reports, tiny text, tables, formulas, a chart, a short lecture clip, and one audio file. Compare the JSON against your current stack at the same render resolution. The card’s own OmniDocBench row is a vendor table. A 0.05 Overall gap on that table is not a production SLA.
Pin the commit. The custom modeling_youtu_vita.py files execute under trust_remote_code. The serving recipe listens on localhost only in the script we read; that is a default, not a deployment design. Evaluation code is still marked TBD, so you cannot replay Tencent’s benches from the repo as it stood on 9 October.
This is an evidence review of the Hub card, the Hub API object, the LICENSE file, and the GitHub repository metadata. We did not download the 10 GB of weights, start vLLM, or parse a page.
Common questions
Is this Apache-2.0, and can EU teams use it?
No on both as written. The file is a custom youtu-parsing license. The first clause says the work is not intended for use in the European Union and that this clause prevails in a conflict. Ask counsel before you treat that as a complete territorial ban or as a paper restriction.
Did Tencent beat Gemini-3-Pro on every parsing bench?
Not on the card we opened. OmniDocBench v1.6 Overall is printed above Gemini-3-Pro. OmniParsingBench average is printed below it, 75.08 against 77.44. Those are Tencent’s numbers.
Is the January GitHub repo the Omni launch?
No. TencentCloudADP/youtu-parsing was created on 21 January 2026 for the earlier Youtu-Parsing line. The Omni weights object we dated is the 9 October Hub repository.
What to remember
Use the 9 October Hub card for the eight-task JSON schema, the vLLM pins, and the EU territorial clause. Keep 96.96 and 75.08 in Tencent’s column. Do not load custom code from an unpinned revision, and do not treat a missing arXiv id as a missing weight file.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





