What did LightOn publish on 8 October?

The Hugging Face community article is dated 8 October 2026 and named as a LightOn release. The authors listed on the post are Said Taghadouini, Adrien Cavaillès and Baptiste Aubertin. It says LightOnOCR-3 is a new family of lightweight OCR models that keep transcription and add “bounding box coordinates with labels for all document regions, image descriptions as well as extract numerical data in figures or charts.” The same post says the models are Apache 2.0 for research and commercial use.

The undated product page on lighton.ai repeats the same family pitch: one model call for text, layout, image descriptions and chart data, with a “transcription-only” mode that matches LightOnOCR-2 when the prompt is empty. That page also carries LightOn’s cost line: in self-hosted deployments, “processing costs can fall below one cent per thousand pages, depending on hardware, utilization and workload.” That is a vendor estimate. We did not host a replica or measure a page cost.

This is a new open-weight reader, not a hosted swap API. It sits next to, and is not the same announcement as, LlamaIndex’s OpenDocRouter parse catalog, which bills third-party parsers by the token. LightOn’s artifacts are checkpoints you run.

What is actually on Hugging Face and GitHub?

We opened three LightOn cards on 9 October. lightonai/LightOnOCR-3-4B was created on 6 October and last modified at 20:29 UTC on 8 October; the API listed model.safetensors and license apache-2.0. The 1B and 0.8B cards were created on 7 October and also returned apache-2.0 plus a non-empty sibling file list. The 4B README says it “loads with the Qwen3.5 classes (Transformers ≥ 5.2, saved with 5.5.4),” is “trained and evaluated with thinking disabled,” and names Qwen/Qwen3.5-4B as the base model.

The GitHub repository lightonai/LightOnOCR describes “a minimal client, CLI and viewer for the LightOnOCR models (versions 1, 2 and 3) served with vLLM, plus the code to reproduce our benchmarks.” GitHub’s API listed the repo as created on 30 September 2026, last pushed at 15:58 UTC on 8 October, and licensed Apache-2.0. The README we fetched says grounding “works with LightOnOCR-3 only” and pins “vLLM 0.30.0 with transformers 5.16.1.” Those are LightOn’s compatibility notes. We did not install that extra or serve a model.

Apache 2.0 on a card is a grant of rights to that checkpoint. It is not the same as publishing training data or calling the stack open source in the stronger sense discussed in open weights versus open source. The 4B citation block still points at the earlier LightOnOCR-2 arXiv note (2601.14251), not a new LightOnOCR-3 paper we opened.

How do the two prompt modes work?

Both the blog and the 4B card describe two supported prompts. With the image only, the model “output[s] all textual elements of the page as markdown.” LightOn says that is the LightOnOCR-2 interface, so a v2 caller can switch weights without changing the request. With the text “grounding,” each block is prepended with a placeholder that names a type and a box in page coordinates normalized to 0–1000, for example a title, paragraph, table, image or chart. Image blocks get “a short description.” Chart blocks get “an HTML table of the data points extracted from the figure.” That is the jump from plain OCR to document understanding that LightOn is selling. Other instructions are “out of distribution.”

The 4B card lists the placeholder labels: text, title, list, header, footer, page_number, footnote, caption, formula, code, table, image, chart, header_image, footer_image and aside_text, plus a “+” suffix for continuation blocks. It says grounding adds about 25% more output tokens than plain transcription on its measurement. The blog’s formatting section says the compact markers produce 9–14% fewer output tokens on 512 olmOCR-Bench pages than Chandra-OCR-2 and Infinity-Parser2-Pro in grounding mode. Those token counts are LightOn’s.

Recommended rendering on the 4B card is “PDFs at 400 DPI” with “a target longest dimension of 2048px,” aspect ratio preserved. The README’s serving table uses a 2048-pixel page size for the 0.8B and 4B models and 1540 pixels for the 1B. The blog’s speed section says the 0.8B and 4B models “reach their best olmOCR-Bench scores at 400 DPI with a 5 MP cap,” about 4.8k image tokens per page, and that shrinking to 1540 pixels on the long side raises peak throughput. All of those figures are from LightOn’s vLLM 0.30.0, one-H100 setup. We did not repeat the load test.

What do LightOn’s benchmark tables actually say?

The blog’s olmOCR-Bench table puts Infinity Parser Pro (35.1B listed parameters) first at 87.6 overall and LightOnOCR-3-4B second at 86.3. At the same listed 4B size, LightOn’s 4B row is 0.5 above Chandra 2 at 85.8. Category leads on that table are mixed: the 4B leads ArXiv (91.1), the 1B leads multi-column pages (85.9), and the 0.8B leads long tiny text (94.1). Infinity Parser Pro remains ahead on old scans and headers/footers. LightOnOCR-2-1B is footnoted at 83.2 overall with headers/footers excluded, so that older overall is not comparable to the other rows.

On ParseBench, LightOn reports the 4B and 0.8B first and second on the five-category overall (75.1 and 74.6), ahead of Infinity Parser Pro at 74.3. On fr-bench-pdf2md, which LightOn describes as a French-document set, the 4B leads overall at 74.1, with its largest listed gains on handwritten pages, forms and multi-column layouts. Chandra 2 still leads some graphics and long-table columns on that table. The LightOn website’s equal-weight average of the three overall scores puts the 4B at 78.5 and the 0.8B at 76.9.

LightOn is explicit that these scores are not raw dumps. “A simple deterministic rewrite can therefore change a model’s score noticeably without changing what it actually extracted.” The team says it did not tailor training data to the benchmarks’ preferred markup, then applied “normalization functions to our raw inputs before scoring” so it could compare with other rows. It publishes those functions in the GitHub repo. That is the right way to read a vendor leaderboard; it is also why a third party that skips the rewrite will not match the table. For how marketing benches get ahead of what a reader can reproduce, see how to read AI benchmark marketing.

What should readers not assume?

A safetensors file and an Apache-2.0 badge are not a measured page cost, a hosted SLA, or a claim that LightOn beat every commercial OCR API on your invoices. The “below one cent per thousand pages” sentence is conditioned on hardware, utilization and workload. The speed numbers are one H100, one vLLM version and LightOn’s page render. The 4B card’s citation still names the LightOnOCR-2 paper. Community GGUF and MLX repos appeared on 8 October under other accounts; those are third-party conversions, not the cards we are dating.

LightOnOCR-3 is also not a compact Arabic-specialist reader. For TII’s 6 October Arabic OCR adaptation, see Falcon-OCR-Arabic. We have not compared those weights on the same pages, and LightOn’s French-benchmark lead does not transfer by announcement.

We did not download the weights for inference, reproduce olmOCR-Bench, or inspect training data. The inspectable facts on 9 October are the dated blog post, the three public cards, the Apache-2.0 repo, and LightOn’s own caveats about rewrite scripts and supported prompts.

Common questions

Is LightOnOCR-3 a hosted API I can call with a key?

The artifacts we opened are Hugging Face checkpoints, a vLLM-oriented GitHub client, and LightOn’s product page. That is self-hosted software. It is not OpenDocRouter or another token-billed parse endpoint.

Can a LightOnOCR-2 integration keep the same request?

LightOn says yes for plain transcription: an empty prompt still returns Markdown. Grounding is new and LightOnOCR-3-only. The README also pins specific vLLM and transformers versions.

Are the 86.3 and 75.1 scores raw model output?

No. LightOn says it applies published normalization functions before scoring so it can sit on the same tables as other models. A run that skips that rewrite is a different measurement.

THE TAKEAWAY

What to remember

Use the 8 October Hugging Face post for the launch date, the three cards for weights and licenses, and the GitHub README for how to serve them. Keep LightOn’s benches in the vendor-with-rewrite column until someone else reruns the scripts.

Sources & further reading

  1. LightOnOCR-3: High-Performance OCR and Layout Extraction in One Model ↗
  2. LightOnOCR-3 ↗
  3. lightonai/LightOnOCR-3-4B model card ↗
  4. lightonai/LightOnOCR README ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories