What did Underdog release, and when?

The dated announcement is the 7 October 2026 page “Meet Underdog Saluki 27B,” signed “The Underdog team.” It says Qwen3.8-27B was squeezed into a 7.89 GB file that runs on standard llama.cpp, and that Saluki is “the best 2-bit build of this model we know of at picking the right tool.” That ranking is Underdog’s. The page says the weights are open source, Apache-licensed in spirit, and that users need not cite Underdog.

The Hub card ConwayResearch/Underdog-Saluki-27B-1.0 repeats the 7.89 GB size, names the GGUF `Underdog-Saluki-27B-1.0-IQ2-mix.gguf`, and states Apache 2.0. Credits go to the Qwen team and to ISTA-DASLab’s Qwen3.8-27B-GSQ-RCO-GGUF, both Apache 2.0. The launch page says this is a preview checkpoint and that full training notes will come later.

A 2-bit file is a quantization, not a new 27B trained from scratch. The usual memory and quality trade-offs are the ones in our quantization explainer. Apache on the GGUF also inherits the base model’s terms; that is the weaker “open weights” sense discussed in open weights versus open source, not a from-scratch training dump.

How is the 7.89 GB file built?

Underdog describes three layers of work. Qwen3.8-27B is the dense 27-billion-parameter teacher. ISTA-DASLab’s 2-bit GSQ-RCO release is the public quantization Saluki is built on; the launch page lists that ISTA file at 8.4 GB and 76 of 120 on Underdog Bench. Underdog’s own pass shrinks the file to 7.89 GB and, it says, targets tool calling. The card’s filename tag is IQ2-mix. The full recipe for that last pass is not on the pages we opened.

The launch page compares several small builds of the same teacher: Saluki 7.89 GB, a 7.9 GB “Main” build that scores higher on the same bench, Mia 2-bit EXL3 at 9.1 GB, ISTA IQ2_XS at 8.4 GB, and Ternary Bonsai 2 at 5.95 GB. Main is “ours” but not the launch file. Mia needs ExLlamaV3 and will not load in llama.cpp. Saluki is the llama.cpp-native file being published.

What do Underdog’s tool-calling benches say?

Underdog Bench is 120 tasks taken from the Berkeley Function Calling Leaderboard, 24 in each of five kinds of tool use. The page says the tasks were picked and frozen on 27 September, before these models were tested, and that every model ran the same harness in its own runtime. Thinking is off and temperature is 0 on the card’s tool-calling rows.

Saluki scores 88 of 120. The full-size Qwen3.8-27B scores 84. Main scores 91. ISTA’s 2-bit scores 76. Bonsai 2 scores 70. Underdog also reports that Saluki still solves 76 of the 84 tasks the full model got right. On the two “pick the right function” rows, Saluki is 47 of 48. The page is explicit: “We don’t claim they’re smarter than the original. 120 tasks is a modest test, and part of that lead is likely run-to-run variation.”

A second table covers 100 parallel-call tasks. With strict scoring Saluki is 42, the full model 35, and Mia 69. A “forgiving” parser that allows small formatting slips lifts Saluki to 55. The launch page says 23 Saluki replies had no call the checker could read, mostly formatting. Mia remains the leader on that slice and is a larger 9.1 GB ExLlama file.

Underdog Bench and parallel-call scores from Underdog’s 7 October page (vendor-run)
ModelSizeUnderdog Bench / 120Parallel calls / 100 (strict)
Saluki (launching)7.89 GB8842
Main (Underdog, not launching)7.9 GB91—
Qwen3.8-27B full size~54 GB8435
ISTA 2-bit IQ2_XS8.4 GB76—
Mia 2-bit EXL39.1 GB8569

Where does the compression show?

The launch page’s “good to know” list says Saluki keeps 82–85% of the full model on competition math (AIME) and that letter-level wordplay is its weakest instruction type. On 50 SWE-bench Verified GitHub issues, Saluki fixed 30 against 33 for the full model and 29 for Bonsai 2, including 27 of the 33 the full model fixed. WebWalkerQA browsing is 60 of 150 for Saluki; the full model was not run on that set.

The Hub card’s own table lists IFEval prompt-loose 93.5 against a public 91.5 for the full model, MBPP+ 78.0 against 83.9, MuSR 67.5 against 79.6, and AIME 2025/2026 avg@4 of 79.2 / 80.0 against public 96.7 / 94.6. The card warns that public scores come from a different harness. IFEval on the launch page is 93.5 loose / 90.9 strict, with unfinished prompts counted as failures.

Qwen3.8’s thinking switch still applies. The card’s default is thinking on, temperature 0.6; fast tool calls should set `enable_thinking` false and temperature 0. That is the same control discussed for Qwen’s thinking modes in our Qwen 3 thinking-modes note, not a new Underdog sampler.

How do you run it, and what about images?

The card’s quickstart is `huggingface-cli download ConwayResearch/Underdog-Saluki-27B-1.0 Underdog-Saluki-27B-1.0-IQ2-mix.gguf` and `llama-server -m … --jinja -ngl 99 -fa on -c 32768`. `--jinja` turns on the Qwen3.8 chat template for tool calls and thinking. The launch page says a 16 GB laptop is in scope. We did not time tokens per second.

Vision is inconsistent across the two first-party pages. The launch page’s “good to know” list says “Text in, text out. This release doesn’t take images.” The Hub card lists optional mmproj files of 928 MB (F16) or 629 MB (Q8_0) and shows a `--mmproj` llama-server example. Until Underdog reconciles those sentences, treat the main 7.89 GB file as text-only and the mmproj as an add-on documented only on the card.

What is not established?

We did not load the GGUF, rerun BFCL, or reproduce SWE-bench Verified. Underdog Bench is a 120-task slice the company froze, not the public BFCL leaderboard. The 88-versus-84 lead is inside the variation the authors themselves flag. The “best 2-bit” sentence is Underdog’s comparison with the small files they chose to run.

The page says Qwen3.8-27B “beats Claude Opus 4.6 Max… according to Artificial Analysis.” That is a claim about the teacher, attributed to a third-party board, not a Saluki result we opened on Artificial Analysis.

This is an evidence review of pages opened on 9 October. It is not a first-hand laptop test and not a recommendation to replace a full-precision Qwen3.8-27B deploy.

Common questions

Is Saluki a new 27B model?

No. It is a 2-bit GGUF of Qwen3.8-27B, further compressed from ISTA-DASLab’s GSQ-RCO release. The teacher and the ISTA file remain the upstream artifacts.

Does it beat the full model?

On Underdog’s 120-task tool-calling slice, Saluki scores 88 and the full model 84. Underdog says not to read that as “smarter,” and math and SWE-bench Verified still trail the teacher on the same pages.

Will it load in Ollama or LM Studio?

The pages we opened promise stock llama.cpp and apps built on it. They do not name an Ollama library tag. Check the GGUF in the runtime you actually use.

THE TAKEAWAY

What to remember

Use the 7 October Underdog page for the launch date, the 120-task bench and the math caveat. Use the Hugging Face card for the IQ2-mix filename, Apache-2.0 license and llama.cpp flags. Keep the 88-versus-84 tool-call lead in the vendor-with-caveat column.

Sources & further reading

  1. Meet Underdog Saluki 27B ↗
  2. ConwayResearch/Underdog-Saluki-27B-1.0 model card ↗
  3. Apache License, Version 2.0 ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories