What shipped, and when?

The dated artifact is GitHub release v1.15.0 of huggingface/trl. The GitHub API record we opened lists published_at as 8 October 2026 at 19:26:48 UTC and author qgallouedec, Quentin Gallouédec. The tag name and release name are both v1.15.0. We opened the release page, the v1.15.0 documentation index and the PyPI JSON document for trl 1.15.0 on 9 October.

The docs index and the PyPI description repeat the same lead line: “Up to 6.9× longer sequences on the same GPU,” with SFT, DPO, KTO, GRPO and RLOO scoring tokens through a fused LM head by default, and “52% to 82% less peak VRAM” at 8k tokens. The release notes add Distillation and publish the tables behind that rounded figure.

TRL is Hugging Face’s post-training library: supervised fine-tuning, preference methods and online RL trainers on top of Transformers. The 1.15 news is a scoring-path change inside those trainers, not a new base model. For the preference-method background, the useful prior on this site is direct preference optimization, which is one of the trainers that now uses the fused head.

What does the fused LM head actually do?

The release describes a memory problem after the decoder. The hot path TRL owns is the loss head: projecting hidden states through the language-model head to score tokens. Materializing the full logits tensor has shape [batch, sequence, vocabulary]. On a 262k-vocabulary model that tensor dominates activations.

The fused head replaces that materialization. A Triton kernel “projects the hidden states through the LM head in tiles and reduces to per-token log-probs and entropy directly.” The release says the path is on by default: “there is nothing to turn on.” Scoring needs Triton on a GPU. The platforms named are Linux with CUDA, ROCm or XPU. A Torch implementation is the fallback when that kernel does not apply.

SFT’s previous default already took a different shortcut. loss_type="chunked_nll" computes cross-entropy in token chunks so the full logits tensor is never held. The 1.15 table shows that default path unchanged at 107,520 tokens. The fused head is the new default scoring path for the trainers that still needed full-logits or Liger-chunked scoring, including SFT when you opt into loss_type="nll." That is an implementation change inside fine-tuning, not a new alignment algorithm.

What numbers did Hugging Face publish?

The max-length table uses gemma-3-1b, a 262k vocabulary, 79 GiB of GPU memory, bf16, batch size 1 except KTO at batch size 2, gradient checkpointing and sdpa. Those axes are Hugging Face’s. The headline 6.89× is KTO. DPO is 5.80×. SFT with nll is 5.25×. GRPO and RLOO are 4.00× and 4.26×, and both rows are labeled “scoring only.”

A second table holds peak memory and tokens per second at 8,192 tokens. Peak GiB falls from 48.97 to 12.00 on DPO (−75.5%), 64.91 to 11.98 on KTO (−81.5%), 24.15 to 10.30 on GRPO (−57.4%), 29.28 to 14.02 on RLOO (−52.1%) and 30.33 to 8.87 on SFT with nll (−70.7%). Step throughput in that table is 1.023× to 1.045×. The release’s prose summary is “52% to 82%” less peak memory and “2.3% to 10.9%” faster steps.

Hugging Face max trainable sequence length on gemma-3-1b, 262k vocabulary, 79 GiB cap, bf16, gradient checkpointing, sdpa. GRPO and RLOO are scoring only.
Trainerv1.14.2 tokensv1.15.0 tokensVendor ratio
DPO10,24059,3925.80×
KTO (batch size 2)9,21663,4886.89×
GRPO (scoring only)28,672114,6884.00×
RLOO (scoring only)23,552100,3524.26×
SFT (loss_type="nll")20,480107,5205.25×
SFT (default chunked_nll)107,520107,5201.00×

How were those charts run?

The release’s benchmark note is unusually specific, and it is still a vendor note. The job used one B300 with the PyTorch allocator capped at 79 GiB, which the authors describe as about an H100 80GB’s usable memory. The model is google/gemma-3-1b-pt built from its config with random bf16 weights, learning rate 1e-6, AdamW and synthetic token ids at exact length. Max length used a fresh process per attempt, two training steps, doubling from 8k then bisection to 1,024 tokens. Throughput used 12 steps, discarded the first two, and took the median of three A/B/A/B repeats.

The same note lists what was not measured: H100 hardware itself, flash-attention (sdpa only), real generation and multi-GPU. GRPO and RLOO replace generation with fixed random completions, so only scoring is timed. The authors put the A-versus-A noise floor at 0.1% to 0.8%. Those constraints belong next to the 6.89× cell. We did not rerun the job.

What breaks if you upgrade?

The fused path removes the old full-logits scoring path, the use_liger_kernel chunked path and _forward_redirection from GRPO, RLOO, DPO and KTO. A PEFT adapter on lm_head now raises. The documented replacement is modules_to_save=["lm_head"]. That is a real upgrade foot-gun for anyone who was LoRA-tuning the head.

use_liger_kernel=True is deprecated in those four trainers and scheduled for removal in v2.0.0. Liger’s layer kernels still apply, but fused_linear_cross_entropy is forced to False because it would replace the fused head, and setting it explicitly raises. The replacement named in the notes is model_init_kwargs={"use_kernels": True}. In DPO and KTO, compute_metrics and compute_loss(..., return_outputs=True) still receive full logits, from an extra forward pass taken only when they are used; both warn that this is deprecated. For adapter training on limited GPUs the usual prior remains QLoRA; the 1.15 change does not replace that recipe.

The release also drops Python 3.10 support and refuses nn.DataParallel. The PyPI record we opened requires Python >=3.11, names Apache-2.0, and depends on transformers >=4.56.2. Those are package facts, not performance claims.

What else is in 1.15.0, and what is not established?

The same tag adds selective activation checkpointing in SFT, assistant-only loss on vision datasets once transformers 5.18 is present, conversation-list logging in the completions table and AsyncGRPO. Those are listed features, not the fused-head charts. Preference training in TRL still sits on the same research line as InstructGPT-style RLHF; 1.15 does not change that method, only how tokens are scored.

We did not install TRL 1.15.0, compile the Triton kernel, or reproduce the gemma-3-1b tables on a B300 or an H100. We did not measure a PEFT-on-lm_head failure beyond the sentence in the notes. The 6.89× and 52–82% figures are Hugging Face’s, on random-weight gemma-3-1b under a 79 GiB cap, with generation excluded from the GRPO and RLOO rows.

This is an evidence review of the 8 October release and the v1.15.0 docs and PyPI records opened on 9 October. It is not a first-hand training run and not a ranking of TRL against Unsloth, Liger or other post-training stacks.

Common questions

Do I have to opt into the fused LM head?

No. The v1.15.0 release says it is on by default for SFT, DPO, KTO, GRPO, RLOO and Distillation. Scoring still needs Triton on CUDA, ROCm or XPU.

Why is default SFT unchanged in the max-length table?

SFT’s default loss_type is still chunked_nll, which already avoided materializing full logits. That row stays at 107,520 tokens. The 5.25× SFT jump is the loss_type="nll" row.

Will my LoRA-on-lm_head run keep working?

Not as a PEFT adapter on lm_head. The release says that configuration now raises. The documented replacement is modules_to_save=["lm_head"].

THE TAKEAWAY

What to remember

Install trl 1.15.0 for the default fused LM head if you have Triton on CUDA, ROCm or XPU. Read the PEFT-on-lm_head and use_liger_kernel notes before you upgrade a preference or RL trainer. Treat 6.89× and the 52–82% VRAM drops as Hugging Face’s gemma-3-1b numbers under a 79 GiB cap, not as a measurement we ran.

Sources & further reading

  1. trl v1.15.0 release notes ↗
  2. TRL documentation index v1.15.0 ↗
  3. trl 1.15.0 package metadata ↗
How this story was made

Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.

Our editorial standards
Back to all stories