What shipped on 7 October?

The Hugging Face model API we opened lists 0arch-io/kisoku-1.6b created 2026-10-07T23:36:08.000Z, license apache-2.0, with a safetensors weight file and tokenizer artifacts. Companion hubs named in the GitHub README include kisoku-1.6b-chat (preview), kisoku-1.6b-gguf, intermediate checkpoint branches, and an evaluation dataset.

The GitHub repository 0arch-io/kisoku is Apache-2.0 and holds MaxText/JAX training configs, evaluation scripts, contamination audit code, synthetic long-context generators, and report/kisoku-report-draft.md (draft 4, 2026-10-07). The pretraining text itself is not redistributed; manifests list sources because Nemotron sets forbid redistribution.

That openness profile is closer to a research dump than a polished lab launch. Compare the documentation intensity with other small open releases such as Underdog’s Saluki 27B GGUF path and the license hygiene checklist in open weights versus open source.

GitHub activity we opened shows the repository created 1 October and pushed through 9 October 2026, so the Hub upload on the 7th is the cleanest “weights public” timestamp for this article’s eventDate.

How was Kisoku trained?

The report describes a Qwen3-style decoder in MaxText: 22 layers, 2048 embedding width, 16 query / 4 KV heads, tied embeddings, Llama 3.2 tokenizer (128,256 vocab), RoPE base 5,000,000, exact parameter count 1,600,749,056. Pretraining used a TPU v4-32 from Google’s TPU Research Cloud; evaluation used a single RTX 4090 (long-context RULER runs also used an RTX PRO 6000).

Token accounting in the report: stages 1–3 about 509.6B tokens; long-context Phases A–C add to a combined 524.4B. Stage mixes lean on Nemotron-CC, Ultra-FineWeb, StarCoder, math corpora, and later OpenThoughts3 / Nemotron-CC-Math. The author states no distillation loss was used, but “roughly a third or more of the documents it read were written or rewritten by larger models.”

Optimizer and schedule details in the report include Muon for two-dimensional attention/MLP matrices, AdamW for embeddings/norms/biases, peak learning rate 3e-4, and a warmup-stable-decay schedule written for 243,000 steps at about 2.1M tokens per step. The author logs many crash restarts during stage 1 and hand-copies stage-final checkpoints because only the latest five autosaves were retained.

Long-context training moved from 4K to 32K (Phase A, ending about 2026-10-01) then 64K (Phases B/C through about 2026-10-04), with YaRN ×2 used at inference for 128K RULER measurements. Chat fine-tuning is a separate preview with documented abstention/coverage tradeoffs.

Technical illustration of a Google TPU v4 board with labeled components. No people appear.
TPU v4 board illustration from Jouppi et al., released under CC BY 4.0 on Wikimedia Commons (File:TPU v4.png). No people appear. Diagram of the accelerator class Kisoku used via TRC; not a photograph of the author’s pod or exported weights. Photo: Norman P. Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Cliff Young, Xiang Zhou, Zongwei Zhou, and David Patterson. CC BY 4.0 · Cropped and resized.

What scores does the author report?

On lm-evaluation-harness 0.4.13 in bfloat16, the released base checkpoint’s short-context table (HF README) includes GSM8K 15.0, MMLU 34.2, ARC-Easy 64.4, ARC-Challenge 40.1, BBH 28.6, PIQA 73.0, WinoGrande 56.8, HellaSwag 57.8, HumanEval 14.6, TriviaQA 23.1. Against Llama 3.2 1B measured in the same harness, Kisoku wins five and Llama wins five.

The author repeatedly warns that SmolLM2 1.7B and Qwen2.5 1.5B beat both models on most tests, and that several Kisoku–Llama gaps sit inside bootstrap intervals. RULER averages with YaRN show Kisoku ahead of Llama 3.2 1B at 4K/16K/32K/64K, level at 128K, behind at 8K, and behind Granite 4.0 1B and Qwen3.5 models at every length.

Stages 2 and 3 (the math/reasoning-heavy mix) are where GSM8K and HumanEval move most; the long-context phases barely change short-context averages. That is useful when someone tries to read “64K trained” as a general quality upgrade—it is mainly a context story in this report.

Author-reported short-context scores (HF README / report draft 4; same harness)
BenchmarkKisoku 1.6B (released)Llama 3.2 1B
GSM8K (5-shot)15.05.8
MMLU (5-shot)34.231.3
ARC-Challenge40.136.9
HellaSwag57.864.2
HumanEval pass@114.618.9
TriviaQA (5-shot)23.140.7

What caveats matter before you quote it?

Contamination audit: the report describes verbatim 80-character probes against the pretraining corpus and says accuracy on the clean subset is the same or higher—but exact-match probes miss paraphrase leakage. HumanEval’s 164 problems make few-point gaps fragile.

Parameter and distillation caveats: Kisoku is 1.6B versus Llama 3.2 1B’s 1.24B, and Llama’s card claims distillation from larger models while Kisoku claims random init with heavy model-written text in the mix. Those are not the same training story even when the harness matchup is fair.

Data licensing: you get weights and recipes, not the Nemotron text shards. Redistribution limits are the author’s; verify your own compliance before commercial fine-tunes.

If you care about tiny local models more than solo TPU lore, also read our SmolLM2 local-models note—the Kisoku report itself says SmolLM2 leads on this suite.

Close-up of a green circuit board with chips, capacitors, and solder pads. No people appear.
Circuit-board close-up photographed 17 January 2018 by Jonathan Cutrer. CC BY 2.0 via Wikimedia Commons. No people appear. Contextual electronics photograph; not Kisoku’s evaluation GPU or Hugging Face hosting hardware. Photo: Jonathan Cutrer from San Angelo, Texas, United States. CC BY 2.0 · Cropped and resized.

Who should try the weights?

Researchers who want a reproducible small-model recipe with manifests and audit code will get more value than users hunting a chat assistant. The base model is a continuation model; the chat preview is explicitly unfinished.

For local inference experiments, the GGUF repo is the practical entry point; for training science, start with the MaxText configs and the contamination folder. If your goal is a stronger 1–2B daily driver, the author’s own tables point you toward other small open models such as Phi-class profiles rather than crowning Kisoku.

What did we not verify?

We did not run lm-evaluation-harness, download the full safetensors for inference, or audit the TRC logs. All benchmark numbers above are the author’s. Cover images are archival TPU and circuit photographs, not Rodriguez’s pod.

Common questions

Is Kisoku chat-tuned?

A chat preview exists at 0arch-io/kisoku-1.6b-chat. The main 1.6B repo is a base (continuation) model; the report documents known chat failures.

Does “18× less data than Llama 3.2 1B” make Kisoku better overall?

No. The author frames it as token efficiency on one suite with a 5–5 split, and notes stronger small models still win most rows.

What license applies?

Apache 2.0 on the GitHub repository and HF model card we opened. The Llama 3 tokenizer vocabulary note is called out on the model card.

THE TAKEAWAY

What to remember

Kisoku 1.6B is a rare fully opened solo TRC pretrain with weights, manifests, and a self-critical report. Quote the Llama token-efficiency story only with the author’s own caveats about distillation, synthetic data, and stronger small baselines.

Sources & further reading

  1. 0arch-io/kisoku README ↗
  2. Kisoku 1.6B technical report draft 4 ↗
  3. 0arch-io/kisoku-1.6b model card ↗
  4. 0arch-io/kisoku-1.6b model API ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories