What did mncai publish on 8 October?

mncai’s Hugging Face blog dated 8 October 2026 describes how the team post-trained Hunmin-397B-A17B-CUA for GUI grounding and multi-step desktop/browser tasks. The companion model card presents the same checkpoint as a general-purpose vision-language model based on Qwen3.5-397B-A17B.

The method narrative starts from a tension: Qwen-CUA is a stronger computer-use agent than the base model but, in mncai’s evaluations, worse on high-resolution GUI grounding and Korean-language benchmarks. The team transferred a low-rank portion of the weight difference on selected modules, then continued with SFT and GRPO agent RL.

Computer-use agents remain a crowded evaluation surface. For the category framing, see our computer-use agents guide.

The low-rank transfer step is the distinctive method claim: rather than full continued pretraining on Qwen-CUA’s recipe, mncai copied a low-rank slice of the weight difference on selected modules, then ran SFT and GRPO. Their blog says the transfer checkpoint alone already lifted OSWorld materially, with SFT/RL adding a smaller additional gain—stage numbers that are still mncai’s measurements on their hosts.

Yeouido skyscrapers along the Han River under autumn light. No identifiable people appear.
Yeouido, Seoul, November 2019. CC BY 2.0 archival photograph via Wikimedia Commons. Contextual Seoul finance-district view; it does not depict Hunmin’s GUI agent. Photo: Willem van Valkenburg / Wikimedia Commons. CC BY 2.0 · Cropped and resized.

Which scores are the authors’, and what should you not compare?

All agent evaluations in the card use image_max=5. mncai reports OSWorld (360) 70.5 and WindowsAgentArena (153) 50.9 for Hunmin versus 48.2 / 41.8 for the base and 77.1 / 56.7 for Qwen-CUA under the same reproduced setup. ScreenSpot-Pro (single) is listed at 75.6 for Hunmin versus 62.2 for Qwen-CUA and 72.7 for the base.

The authors warn not to compare these agent scores with results that use substantially larger visual histories such as image_max≈20. That warning is load-bearing if you are reading public OSWorld leaderboards.

Product computer-use betas quote different benches entirely; Pine’s SaaS-Bench private beta is not an OSWorld substitute—see Pine Computer.

Under the shared image_max=5 reproduction, Hunmin’s OSWorld 70.45 trails Qwen-CUA’s 77.12 while beating it on ScreenSpot-Pro (75.61 vs 62.20) and staying near the base on KMMLU-Redux (81.43 vs base 82.59 vs Qwen-CUA 78.07). Across eight Korean benchmarks, mncai says Hunmin stays within roughly two points of the base. Those rows are why the article title emphasizes not dumping Korean scores: the release is framed as a preservation trade, not a pure OSWorld chase.

Modern glass towers in Yeouido against a bright sky. No identifiable people appear.
Yeouido finance-district skyline, Seoul. CC BY 4.0 archival photograph via Wikimedia Commons. Illustrative Seoul context; it is not a ScreenSpot-Pro screenshot. Photo: S h y numis / Wikimedia Commons. CC BY 4.0 · Cropped and resized.

What is the practical deployment story?

Training reportedly ran on one node with 8×B200 GPUs using stacks named in the blog (MS-Swift, Ray, Megatron-Core, vLLM) and public post-training data/environments such as ScaleCUA. Serving a 397B MoE still implies serious GPU budget even with 17B active parameters.

The Hub tags include `license:apache-2.0` and bilingual `en`/`ko` markers. Apache-2.0 is a clearer commercial starting point than many research-only cards, but you still inherit Qwen upstream notices and must validate safety filters for desktop control.

Harnessed agent RL tooling is evolving quickly on the training side as well; Microsoft’s Agent Lightning v1 is a different stack aimed at RL in agent harnesses, not this checkpoint.

Training logistics in the blog are unusually concrete for a Hub write-up: one 8×B200 node; MS-Swift / Ray / Megatron-Core / vLLM; GRPO rollouts from two FP8 vLLM engines (TP4 each) alternating with a Megatron actor via sleep/wake on the same eight GPUs. That is a recipe note for labs with similar iron, not a promise that consumer GPUs can fine-tune the 397B MoE. Serving still means MoE-capable inference even with 17B active.

How should teams read the trade-off?

mncai’s own comparison says Hunmin and Qwen-CUA occupy different points on a specialization–preservation curve: Hunmin better on fine-grained grounding and several Korean benchmarks in their reproduction; Qwen-CUA stronger on long-horizon computer-use agent scores.

If your bottleneck is clicking the right 4K control in Korean enterprise software, the grounding story matters. If your bottleneck is 30-step OSWorld-like workflows, the card still puts Qwen-CUA ahead under the shared protocol.

Independent computer-use evaluations remain scarce relative to vendor tables. OpenAI’s Ironclad computer-use evaluation note is one external framing of how hard reliable measurement still is.

mncai also warns that merge completion is not effect preservation: they discuss FP8/merge experiments and ultimately kept the SFT adapter frozen, stacking another adapter for RL instead of assuming a merged BF16 base retained the skill. If you adapt the recipe, verify computer-use and Korean probes after every merge, not only after loss curves look healthy. Pick Hunmin when grounding and bilingual retention matter more than matching Qwen-CUA’s long-horizon OSWorld row under the same image_max=5 protocol.

What did we not test?

We did not download Hunmin’s shards, run OSWorld or ScreenSpot-Pro, or verify Korean benchmark deltas. This article reports the 8 October Hugging Face blog and the model card we opened on 10 October 2026.

Common questions

Is Hunmin stronger than Qwen-CUA overall?

Not on mncai’s own OSWorld comparison under image_max=5. Hunmin trails Qwen-CUA there while leading on ScreenSpot-Pro and several Korean rows in the same tables.

What license are the weights under?

The Hugging Face model tags list Apache-2.0 for mncai/Hunmin-397B-A17B-CUA. Confirm the card and any upstream Qwen notices before shipping.

Why does image_max matter?

mncai evaluates agents with image_max=5 and warns that scores are not directly comparable to setups that keep much longer visual histories.

THE TAKEAWAY

What to remember

Hunmin’s 8 October write-up is a concrete Apache-2.0 computer-use MoE release with honest vendor trade-offs. Match the image_max=5 protocol before you declare a winner against Qwen-CUA.

Sources & further reading

  1. Hunmin-CUA: Post-Training a 397B MoE for Computer Use ↗
  2. mncai/Hunmin-397B-A17B-CUA ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories