What did Nature and Ai2 publish today?

Nature dated the article 7 October 2026 and marked it open access. Benjamin Minixhofer is the corresponding author, with a University of Cambridge email on the paper. Co-authors include Tyler Murray, Tomasz Limisiewicz, Anna Korhonen, Luke Zettlemoyer, Noah A. Smith, Edoardo M. Ponti, Luca Soldaini and Valentin Hofmann. They declare no competing interests. A same-day News & Views by Zhao Zhang and Yingfei Xiong at Peking University describes the work as giving token-based models access to individual characters.

Ai2’s 7 October blog treats the paper as the formal publication of Bolmo, introduced last December, and says it is releasing checkpoints beyond Olmo: Bwen 8B from Qwen 3 8B, Blama 8B from Llama 3 8B, and Stage 1 checkpoints that keep the original global model frozen.

That is a journal publication and a research release, not a consumer chatbot launch. For how subword pieces relate to the text you type, see our explainer on AI tokens. We did not download weights or generate from them.

What problem is byteification meant to solve?

Most large language models never see raw characters. They first split text into subword tokens from a fixed vocabulary. The paper says that hides information that matters for code, biological sequences, spelling and rare strings, and that it biases models toward English-centric vocabularies. Byte-level models read UTF-8 bytes instead. The authors argue that competitive byte-level systems usually had to be trained from scratch, which is expensive while subword recipes keep moving.

Byteification starts from an already trained subword model. The architecture is a latent tokenizer language model: a shallow local encoder over bytes, a boundary predictor that groups bytes into patches, the source model’s deep transformer, then a local decoder that predicts the next byte. Earlier systems in that family include DTP, BLT and H-Net. The authors say their design is built for the retrofit.

At prefill, the boundary predictor may look one byte ahead, so cuts can match how a subword tokenizer uses upcoming characters. At decode time it predicts the next byte and whether a boundary follows. We have not inspected a sampler trace.

How does the two-stage retrofit work?

Stage 1 trains the local encoder, decoder, boundary predictor and language-modelling head while the global transformer stays frozen. Losses copy subword cuts, match hidden states a few layers into the frozen model, and distill patch likelihoods. The paper used 9.8 billion tokens here, about 43 billion bytes.

Stage 2 unfreezes everything and trains on 39.3 billion tokens, about 173 billion bytes. The mix is about 172 billion Dolma 3 tokens plus about 75 million English character-level exercises in the style of the CUTE benchmark, sampled so they do not overlap the CUTE test words. Training ran for less than one epoch. Local layers are mLSTM blocks: one in the encoder, four in the decoder.

The starting checkpoints named in the paper are Olmo 3 7B after mid-training and long-context extension, OLMo 2 1B, Qwen3 8B Base and Llama 3 8B. Olmo is Ai2’s own open-training line; see our OLMo 2 training-transparency note for what “open training” usually includes and omits. The paper reports that Bolmo 1B ends up about 10 million parameters smaller than OLMo 2 1B, Bolmo 7B about 330 million larger than Olmo 3 7B, Blama 8B about 220 million larger than Llama 3 8B, and Bwen 8B about 120 million larger than Qwen3 8B.

What results does the paper claim?

Table 1 compares Bolmo 7B, Bwen 8B and Blama 8B with their sources, with Olmo 3 continued on the same mix, and with EvaByte 6.5B, TFree-HAT 7B and BLT 7B. The authors say the byteified models were close to their sources, that Bwen 8B statistically significantly beat the earlier public byte-level models, and that Bolmo 7B’s STEM score was 16.5 percentage points above BLT 7B. Character-understanding scores rose sharply versus the subword sources; an Olmo 3 continued-training control on the same mix did not match that gain.

On code they report higher pass@16 and lower pass@1 than the sources at temperature 0.6 and top-p 0.6, and warn that this may be a sampling interaction. They also show that raising bytes per patch can beat SuperBPE tokenizer transfer once a huge subword softmax becomes expensive.

A separate experiment merges Olmo 3’s RL-Zero instruction-following weights into Bolmo with task arithmetic, without extra training. The paper says Bolmo’s IFEval score then matches the post-trained Olmo 3 checkpoint. That is a merge on their harness, not a general recipe we tested. Read those tables the way our benchmark-marketing guide suggests: useful for the method, not as a ranking of today’s chat models.

Which checkpoints can you download, and under which licences?

Hugging Face hosts allenai/Bolmo-7B, allenai/Bolmo-1B, allenai/Bwen-8B and allenai/Llama-3-Blama-8B, with matching Stage 1 cards. The Bolmo-7B and Bwen-8B cards we opened say Apache 2.0 and name Ai2 as the developer. The Llama-3-Blama-8B card says the Llama 3 Community License. Apache 2.0 on one checkpoint is not a licence for the Llama-derived weights. That is the same distinction as in our open-weights versus open-source explainer.

The Hugging Face model API, checked on 7 October 2026, lists createdAt 26 August 2026 and lastModified 28 August 2026 for Bwen-8B and Llama-3-Blama-8B. Bolmo-7B’s createdAt is 13 December 2025; lastModified is 7 October 2026, 14:40 UTC, consistent with a card update on publication day. Ai2’s blog still says “we’re releasing new checkpoints.” We did not diff the weight files against August.

Code is at github.com/allenai/bolmo-core. Training data are at huggingface.co/datasets/allenai/bolmo_mix. The cards we opened recommend transformers 4.57.3 or later, Python 3.11, xlstm 2.0.4 and trust_remote_code=True; max_new_tokens counts bytes. Intended use is research and education under Ai2’s responsible-use guidelines. The cards warn that the base models have no safety filter.

What remains unproven?

The paper does not claim a chat match with GPT-6 or Claude. The suite is OlmoBaseEval plus CUTE and EXECUTE; we have not reproduced Table 1. Ai2 calls Bwen 8B its strongest byteified model and mentions a Bolmo 7B poetry fine-tune; that is anecdotal. For another recent Ai2 checkpoint aimed at scientists, see AstaBrief 8B, which is a report writer, not a byte-level base model.

Byteification still spends tens of billions of tokens and adds local modules. That is cheaper than a full 7B or 8B pretrain, not free. The authors say skipping stage 1 still works but leaves a bits-per-byte gap, and that task-arithmetic post-training needs embeddings that reset cleanly. This is not a hosted-assistant rollout.

Common questions

Did Ai2 upload Bwen and Blama for the first time today?

Not according to Hugging Face’s createdAt timestamps. The Bwen-8B and Llama-3-Blama-8B cards show 26 August 2026. Today’s event is the Nature paper and Ai2’s publication post. We did not prove the August files are bit-identical to today’s downloads.

Are all four models Apache 2.0?

No. The Bolmo and Bwen cards we opened state Apache 2.0. The Blama card states the Llama 3 Community License.

Did byteified models beat their source models overall?

The paper says they came close and sometimes surpassed the source on individual categories, with a large character-understanding gain. It does not claim a clean win on every task. We did not rerun the suite.

THE TAKEAWAY

What to remember

Nature has now peer-reviewed Ai2’s recipe for turning a subword LLM into a byte-level one without a full pretrain. If you want the method, read the paper and the cards, and match the licence to the backbone. If you want a 2026 chat ranking, this release does not provide one.

Sources & further reading

  1. Retrofitting language models to operate over bytes ↗
  2. Now in Nature: Retrofitting language models to operate over bytes ↗
  3. allenai/Bwen-8B model card ↗
  4. allenai/Llama-3-Blama-8B model card ↗
  5. allenai/Bolmo-7B model card ↗
  6. Retrofitted LLM can count the letter ‘i’s in ‘artificial intelligence’ ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories