What did NVIDIA publish?

The Hugging Face article is credited to Aleksander Ficek, Igor Gitman, Sean Narenthiran, Mehrzad Samadi and Somshubra Majumdar, all listed with NVIDIA accounts, and is marked published 7 October 2026. The claim is that Nemotron is “a strong, adaptable foundation for building world-class specialist models,” shown by gold-medal-level systems at both IMO 2026 and IOI 2026.

The score table in the post is: IOI 2026, Nemotron-3-Ultra-CC with SFT and GenCorrect, 535.4/600, above a 361.12 gold threshold and a 498.27 top human score; IMO 2026, Nemotron 3 Ultra general, SFT and RL checkpoints in a generate-verify-refine system, 30/42, above an official gold threshold of 29. The IOI result “came from a live, prospective run under the same time, internet-access, and submission constraints as human contestants” and “was an unofficial, unsupervised benchmark and was not included in the official IOI ranking.” IMO “submitted proofs were graded by official IMO graders.”

What is the specialization recipe?

NVIDIA’s four-part recipe is: start with a Nemotron base; curate domain problems and reasoning traces; apply SFT and, where useful, RL; pair the specialist with an inference loop that generates, evaluates and improves answers. “We did not need to build a new foundation model for every challenge.” That is the point of the post, and it is also why this is not a Nemotron 4 launch. For why contest scores and marketing tables come apart, see how AI benchmark marketing works.

For competitive programming the team says it curated 22,000 problems and synthetic traces. Nemotron-3-Nano-CC (30 billion total / 3 billion active parameters) got SFT and RL. Nemotron-3-Ultra-CC (550 billion total / 55 billion active) got SFT. On IOI 2025, Nano is said to have moved from 130 points before post-training to 280 after SFT, 291 after RL, and 468 with GenCorrect, crossing a 438.3 gold line; Ultra-CC reached 502 with the same test-time strategy. Those 2025 numbers are NVIDIA’s development table. The 2026 Ultra-CC system is the 535.4/600 run.

How did the IMO system differ from the IOI system?

The IMO project trained one SFT specialist and one RL specialist from Nemotron 3 Ultra. The SFT corpus is described as 414,890 quality-filtered examples across 15,818 unique proof problems, covering generation, refinement, verification and meta-verification. RL used 9,597 problems “near the model's capability frontier.” The final system used both specialists plus the general model: generate candidates, score them, critique, refine, then a high-compute selection stage. It worked in natural language, “with no formal prover, external tools, or internet access,” scored 30/42 with full credit on four of six problems, and exceeded the official gold threshold. That is a different setup from Lean-checked manuscript pipelines such as OpenAI’s 722 math manuscripts.

NVIDIA says the medals “were not produced by fine-tuning alone, and they were not produced by brute-force sampling alone.” GenCorrect is credited for turning IOI fine-tuning gains into larger improvements across feedback rounds; complementary SFT and RL checkpoints are credited on IMO over drawing more samples from one checkpoint.

What did NVIDIA open on Hugging Face?

The Nemotron Labs IMO 2026 collection, which we opened, is described as checkpoints, training data and a benchmark from the IMO paper. Items listed on the collection page include Nemotron-3-Labs-Ultra-Math-SFT, Nemotron-3-Labs-Ultra-Math-RL, the general NVIDIA-Nemotron-3-Ultra-550B-A55B-BF16 card, Nemotron-Math-Proofs-v3-SFT and v3-RL datasets, and Nemotron-IMO-Bench (200 rows in the viewer metadata). The blog says the IMO paper describes training and the generate-verify-refine system, and that NeMo-Skills includes the IMO inference pipeline, prompts, submitted proofs and a quickstart.

For IOI, the blog says Nemotron-3-Ultra-CC is on Hugging Face and that the IOI paper plus NeMo-Skills cover GenCorrect and the evaluation pipeline. The article sidebar also names nvidia/NVIDIA-Nemotron-Labs-3-Competitive-Coding-550B-A55B-NVFP4. We did not download the weights or run the quickstart.

What should readers not assume?

A gold-level unofficial IOI score is not an official IOI medal. NVIDIA is explicit that the run was unsupervised and off the ranking. The IMO 30/42 used official graders, which is a stronger claim, and still a vendor-organized submission rather than an independent replication here.

These specialists are not a new general coding model in the sense of a fresh pretraining run, and they are not a small on-device coder. For a same-week open coding-model release with a different scope, see JetBrains Mellum 2.1. We have not timed a LiveCodeBench prompt or graded an IMO proof.

Common questions

Did Nemotron officially win IOI 2026?

No. NVIDIA says the 535.4/600 run used contest constraints but was unofficial, unsupervised, and not in the official ranking.

Was IMO 30/42 graded by the contest?

NVIDIA says the submitted proofs were graded by official IMO graders and that 30 points is above a 29-point gold threshold. We did not see the marked scripts.

Is this a new Nemotron base model?

The post presents it as fine-tuning and test-time systems on Nemotron 3, plus open checkpoints and recipes, not as a new foundation-model launch.

THE TAKEAWAY

What to remember

Keep three facts together: Nemotron 3 specialists, heavy inference loops, and NVIDIA’s own scoring rules — unofficial IOI, official IMO graders. Open weights and a collection are not a reproduced medal.

Sources & further reading

  1. One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO ↗
  2. Nemotron Labs IMO 2026 collection ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories