Announced 7 Oct 2026 · Sources checked
What shipped on 7 October?
The Hugging Face model API we opened lists 0arch-io/kisoku-1.6b created 2026-10-07T23:36:08.000Z, license apache-2.0, with a safetensors weight file and tokenizer artifacts. Companion hubs named in the GitHub README include kisoku-1.6b-chat (preview), kisoku-1.6b-gguf, intermediate checkpoint branches, and an evaluation dataset.
The GitHub repository 0arch-io/kisoku is Apache-2.0 and holds MaxText/JAX training configs, evaluation scripts, contamination audit code, synthetic long-context generators, and report/kisoku-report-draft.md (draft 4, 2026-10-07). The pretraining text itself is not redistributed; manifests list sources because Nemotron sets forbid redistribution.
That openness profile is closer to a research dump than a polished lab launch. Compare the documentation intensity with other small open releases such as Underdog’s Saluki 27B GGUF path and the license hygiene checklist in open weights versus open source.
GitHub activity we opened shows the repository created 1 October and pushed through 9 October 2026, so the Hub upload on the 7th is the cleanest “weights public” timestamp for this article’s eventDate.
How was Kisoku trained?
The report describes a Qwen3-style decoder in MaxText: 22 layers, 2048 embedding width, 16 query / 4 KV heads, tied embeddings, Llama 3.2 tokenizer (128,256 vocab), RoPE base 5,000,000, exact parameter count 1,600,749,056. Pretraining used a TPU v4-32 from Google’s TPU Research Cloud; evaluation used a single RTX 4090 (long-context RULER runs also used an RTX PRO 6000).
Token accounting in the report: stages 1–3 about 509.6B tokens; long-context Phases A–C add to a combined 524.4B. Stage mixes lean on Nemotron-CC, Ultra-FineWeb, StarCoder, math corpora, and later OpenThoughts3 / Nemotron-CC-Math. The author states no distillation loss was used, but “roughly a third or more of the documents it read were written or rewritten by larger models.”
Optimizer and schedule details in the report include Muon for two-dimensional attention/MLP matrices, AdamW for embeddings/norms/biases, peak learning rate 3e-4, and a warmup-stable-decay schedule written for 243,000 steps at about 2.1M tokens per step. The author logs many crash restarts during stage 1 and hand-copies stage-final checkpoints because only the latest five autosaves were retained.
Long-context training moved from 4K to 32K (Phase A, ending about 2026-10-01) then 64K (Phases B/C through about 2026-10-04), with YaRN ×2 used at inference for 128K RULER measurements. Chat fine-tuning is a separate preview with documented abstention/coverage tradeoffs.

What caveats matter before you quote it?
Contamination audit: the report describes verbatim 80-character probes against the pretraining corpus and says accuracy on the clean subset is the same or higher—but exact-match probes miss paraphrase leakage. HumanEval’s 164 problems make few-point gaps fragile.
Parameter and distillation caveats: Kisoku is 1.6B versus Llama 3.2 1B’s 1.24B, and Llama’s card claims distillation from larger models while Kisoku claims random init with heavy model-written text in the mix. Those are not the same training story even when the harness matchup is fair.
Data licensing: you get weights and recipes, not the Nemotron text shards. Redistribution limits are the author’s; verify your own compliance before commercial fine-tunes.
If you care about tiny local models more than solo TPU lore, also read our SmolLM2 local-models note—the Kisoku report itself says SmolLM2 leads on this suite.

Who should try the weights?
Researchers who want a reproducible small-model recipe with manifests and audit code will get more value than users hunting a chat assistant. The base model is a continuation model; the chat preview is explicitly unfinished.
For local inference experiments, the GGUF repo is the practical entry point; for training science, start with the MaxText configs and the contamination folder. If your goal is a stronger 1–2B daily driver, the author’s own tables point you toward other small open models such as Phi-class profiles rather than crowning Kisoku.
What did we not verify?
We did not run lm-evaluation-harness, download the full safetensors for inference, or audit the TRC logs. All benchmark numbers above are the author’s. Cover images are archival TPU and circuit photographs, not Rodriguez’s pod.
Common questions
Is Kisoku chat-tuned?
A chat preview exists at 0arch-io/kisoku-1.6b-chat. The main 1.6B repo is a base (continuation) model; the report documents known chat failures.
Does “18× less data than Llama 3.2 1B” make Kisoku better overall?
No. The author frames it as token efficiency on one suite with a 5–5 split, and notes stronger small models still win most rows.
What license applies?
Apache 2.0 on the GitHub repository and HF model card we opened. The Llama 3 tokenizer vocabulary note is called out on the model card.
What to remember
Kisoku 1.6B is a rare fully opened solo TRC pretrain with weights, manifests, and a self-critical report. Quote the Llama token-efficiency story only with the author’s own caveats about distillation, synthetic data, and stronger small baselines.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





