Announced 8 Oct 2026 · Sources checked
What did JetBrains publish on 8 October?
The JetBrains AI blog post “Mellum2.1 Gets to Work: A Fast Open Model for Coding Agents,” dated 8 October 2026 with a 13:00 timestamp and signed Bulat Salimzianov, says JetBrains is releasing Mellum2.1, the next version of the 12B mixture-of-experts model it open-sourced in June. The architecture has not changed since version 2: it remains a compact model with 2.5 billion active parameters, released under the Apache 2.0 license. The post says Mellum2 was fast but could not work inside a repository at the level JetBrains wanted, and that after a summer of reinforcement learning in real environments — “millions of sandboxed runs across thousands of environments” — Mellum2.1 “explores a codebase, edits files, and checks its own changes.”
The live Hub checkpoint is JetBrains/Mellum2.1-12B-A2.5B-Thinking. The Hugging Face API we opened lists that repository as created on 20 September 2026, last modified at 23:08 UTC on 7 October, license apache-2.0, base model JetBrains/Mellum2-12B-A2.5B-Base, and 12,149,923,072 BF16 parameters across five safetensors shards. The Mellum2.1 collection, last updated at 10:13 UTC on 8 October, contains only that Thinking card and JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF. There is no Mellum2.1 Instruct repository in that collection.
The 1 June post “Mellum2 Goes Open Source,” by Anton Semenkin and Nikita Pavlichenko, and a same-day Hugging Face blog, described Mellum2 as a from-scratch 12B model for routing, Q&A and sub-agents. The Mellum2 Technical Report on arXiv (2605.31268, submitted 29 May 2026) is the architecture paper the 2.1 card still cites. Mellum2.1 is a post-training increment, not a new backbone.
How is Mellum2.1 built, and what stayed the same?
The Thinking card says the architecture is unchanged: 12 billion parameters total, 2.5 billion active, 28 layers, hidden size 2,304, 64 experts with 8 active, grouped-query attention with 32 query heads and 4 key-value heads, a 131,072-token context, and sliding-window attention of 1,024 on three of every four layers. That is the same mixture-of-experts shape as Mellum2. For what an MoE model spends compute on, see our mixture-of-experts explainer.
JetBrains says almost all of the 2.1 work went into post-training, primarily reinforcement learning. The card lists three changes: RL moved from a short final stage to the main part of training; new RL tasks in math, competitive programming, science, tool use and software engineering, mixing open datasets with JetBrains-built tasks and filtering every source; and software-engineering training inside real repositories with a shell and file-editing tools, rewarded when tests pass. The blog repeats the “millions of sandboxes” line. Those are vendor descriptions of the training setup. We have not seen the RL datasets or the sandbox logs.
The card labels Mellum2.1 a thinking model for repository tasks, commands, tool calls, and hard coding, math and reasoning. The GGUF card says the model emits its chain of thought inside <think>...</think> blocks before the final answer. The May report described a separate Mellum2 Instruct variant that answers directly. That Instruct line is not in the 2.1 collection we opened.
What scores does JetBrains report?
The card and the blog say JetBrains evaluated Mellum2.1, Mellum2 Thinking, Gemma 4 E4B and Qwen3.5-9B with the same pipeline in thinking mode, and that every number is self-reported. Non-agentic rows use greedy decoding. Agentic rows use the open-source Pi v0.73.1 harness with shell and file tools, a 114K-token context, up to 16K tokens per turn, and each model’s default sampling — temperature 1.0 for Mellum2.1. Mellum2 Thinking was re-run on this pipeline, so its numbers differ slightly from the May technical report.
The largest advertised jump is agentic. SWE-bench Verified is 47.0 for Mellum2.1 against 2.0 for Mellum2 Thinking, 23.0 for Gemma 4 E4B and 50.0 for Qwen3.5-9B. Terminal-Bench 2.1 is 17.4 against 0.6, 3.4 and 21.7. SWE-bench Pro is 28.0 against 0.0, 4.0 and 38.0. JetBrains is ahead of its own prior checkpoint and of Gemma 4 E4B on those three rows, and behind Qwen3.5-9B on all three. LiveCodeBench v6 is 82.0, HumanEval+ 91.5 and MBPP+ 79.4. AIME 25/26 — the mean of AIME 2025 and AIME 2026, 30 questions each — is 83.3, below Qwen3.5-9B’s 86.7. WorkBench is 44.6, under Mellum2’s 45.1. HarmBench’s harmful rate falls to 8.5 from 21.5; Qwen3.5-9B is 6.6. XSTest safe compliance falls to 88.8 from 91.2.
Those figures are one vendor table. They are not an independent bake-off, not a submission to a public leaderboard we checked, and not a measure of your repository. That is the same rule as evaluating an AI agent on the task you actually run.
| Model | SWE-bench Verified | Terminal-Bench 2.1 | SWE-bench Pro |
|---|---|---|---|
| Mellum2.1 Thinking | 47.0 | 17.4 | 28.0 |
| Mellum2 Thinking | 2.0 | 0.6 | 0.0 |
| Gemma 4 E4B | 23.0 | 3.4 | 4.0 |
| Qwen3.5-9B | 50.0 | 21.7 | 38.0 |
How do you load it, and what licence applies?
The Thinking card shows a vLLM serve line with --reasoning-parser qwen3 and, for tools, --enable-auto-tool-choice and --tool-call-parser hermes, at --max-model-len 131072. The Python quickstart uses temperature 0.6, top_p 0.95, top_k 20 and max_tokens 81920. We did not start that server.
The same card, and the 8 October blog, still say “GGUF builds for llama.cpp, Ollama, and LM Studio, as well as the multi-token prediction (MTP) head for speculative decoding in vLLM, are coming soon.” The GGUF repository was created at 18:43 UTC on 7 October 2026. Its README lists five single-file quants: BF16 at 24.3 GB, Q8_0 at 12.9 GB, Q6_K at 10.9 GB, Q4_K_M at 8.1 GB (recommended) and MXFP4_MOE at 7.0 GB, with KL-divergence figures against BF16 on Wikitext-2. It documents llama-server -hf JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M and ollama run hf.co/JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF:Q4_K_M. That is a live GGUF drop next to a “coming soon” sentence. We found no matching MTP-head repository.
The Thinking card, the GGUF card and the official Apache 2.0 text we opened all say Apache 2.0. That is a permissive licence, not a source dump of the training data, and not the revenue-capped licence on some other recent open weights. For the difference, see open weights versus open source. Loading a Hub weight is still a trust decision; treat the files as you would any other third-party checkpoint. Our note on agent sandboxing is about a different class of risk — running tools the model calls — not about this licence.
The blog’s speed claims are vendor charts: under heavy load Mellum2.1 “serves almost twice as many tokens as Qwen3.5-9B,” and “for a single request, MTP makes it about 1.6 times faster.” Those sentences sit beside the “MTP … coming soon” line. We did not measure tokens per second.
What should readers not assume?
Mellum2.1 is not a new architecture, not an independently scored SWE-bench winner, and not a measured replacement for Qwen3.5-9B. Qwen3.5-9B leads the three agentic rows JetBrains published. WorkBench and XSTest are not uniformly up. There is no 2.1 Instruct card in the collection. GGUF is already on the Hub even though the announcement copy still says it is coming. MTP is advertised, not shipped in a repository we could open. We have not downloaded the safetensors or GGUF files, have not run Pi v0.73.1, and have not timed vLLM.
Common questions
Is there a Mellum2.1 Instruct checkpoint?
Not in the Mellum2.1 collection we opened on 8 October. That collection holds Thinking and Thinking-GGUF only. Mellum2 shipped an Instruct variant in June; we did not find a 2.1 Instruct sibling.
Can I run it in Ollama or llama.cpp today?
The GGUF repo’s README says yes, with Q4_K_M as the recommended file. The 8 October blog and the BF16 Thinking card still say those builds are coming soon. We did not run either command.
Did Mellum2.1 beat every comparable open model?
No. On JetBrains’s own agentic table, Qwen3.5-9B is ahead on SWE-bench Verified, Terminal-Bench 2.1 and SWE-bench Pro.
What to remember
Use Mellum2.1 if you want an Apache-licensed 12B/2.5B-active thinking checkpoint and can accept JetBrains’s own scores. Read the GGUF repo if you need llama.cpp, and do not wait for an MTP head the card still marks as coming soon. Do not treat 47.0 on SWE-bench Verified as an independent result.
Sources & further reading
- Mellum2.1 Gets to Work: A Fast Open Model for Coding Agents ↗
- JetBrains/Mellum2.1-12B-A2.5B-Thinking model card ↗
- JetBrains/Mellum2.1-12B-A2.5B-Thinking-GGUF model card ↗
- Apache License, Version 2.0 ↗
- Mellum2 Goes Open Source: A Fast Model for AI Workflows ↗
- Introducing Mellum2: A 12B Mixture-of-Experts Model by JetBrains ↗
- Mellum2.1 collection ↗
- Mellum2 Technical Report ↗
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





