Announced 8 Oct 2026 · Sources checked
What is Project Greenhouse’s first milestone?
The authors define agentic search as models using search tools in loops and argue that fully open, sovereign components can be built with modest compute. Their first delivered piece is a pointwise decoder-only reranker—score each (query, document) pair, then reorder a first-stage list.
They deliberately skip embedding-model pretraining and full agent RL as a first step, citing faster iteration when candidates are fixed. Commercial rerankers from cloud vendors are cited as evidence the component is useful alone.
Pointwise reranking sits after first-stage retrieval; for embedding-side context see vector search explained and our note on Perplexity’s late-interaction embedders.
The paper’s framing matters for how you cite it: Greenhouse is presented as a multi-year programme toward fully open agentic search stacks, and Gaggle is explicitly milestone one—a reranker—not a finished tool-using agent. That is why the Hub card and the arXiv abstract both talk about pointwise relevance scoring rather than planner loops or browser tools. Treat “sovereign agentic search” as the project banner and Gaggle as the downloadable component you can actually score today.

How was Gaggle trained?
Pretraining used Andrej Karpathy’s nanochat depth-34 recipe on ClimbMix to produce gaggle-nanochat-pretrained-climbmix-20260924 (~3.29B). Fine-tuning kept two LM head rows for the tokens “ true” and “ false,” yielding about 3.22B parameters in the reranker.
Relevance training used RLHN-250K with localized contrastive estimation for one epoch. The released headline model is a parameter soup of four runs (seeds/hardware variants) averaged at weight 0.25—checkpoint name gaggle-reranker-20261005.
The paper stresses the bulk of experiments used a handful of GPUs; SFT typically one or two GPUs, with pretraining on an 8×H100 server for the heavier phase. Those resource lines are the authors’.
ClimbMix pretraining followed by RLHN-250K fine-tuning is the authors’ recipe; the soup of four runs is their variance-reduction step for the headline checkpoint name. If you train a sibling model, expect to re-pin ClimbMix revision hashes, nanochat depth settings, and the exact LCE schedule before claiming parity. The 3.22B active reranker head that keeps only the “ true” / “ false” vocabulary rows is a deliberate compression of a larger pretrained LM into a binary relevance scorer—useful when you want open weights without importing a third-party backbone’s license stack, still subject to ClimbMix’s own data terms.

What does “fully open and sovereign” mean here?
The paper models training as a property hypergraph so every artifact and command can be cited. The authors argue that starting from third-party open weights still leaves upstream corpora and recipes opaque.
They scope out general coding and frontier reasoning. Greenhouse is framed as a long-running project; Gaggle is an existence proof for one retrieval component, not the finished agentic search stack.
Contrastive ablations fine-tune other backbones (Gemma-4-E2B, Qwen3.5-4B-Base, MiniCPM, nanochat-d34) with the same SFT recipe; those tables are in the paper for readers who want backbone comparisons.
What should a retrieval team try?
Load castorini/gaggle-reranker-20261005 with the card’s transformers example, score a BM25 top-100 from your corpus, and compare nDCG to your current cross-encoder before changing production.
Respect the 128/512 training truncation if you want numbers comparable to the card; longer contexts are untested per the model card.
If sovereignty is the goal, budget for ClimbMix’s NC license review—not only the MIT weight license.
A practical bake-off is three columns on the same corpus: your production cross-encoder, Gaggle soup, and one of the parent seed checkpoints if you care about soup variance. Log not only nDCG@10 but also latency at your batch size and the fraction of ties the pointwise head produces. If you already follow Waterloo’s Anserini/Pyserini tooling, wire the Hub example before you invent a new serving path—the card’s transformers snippet is the authors’ intended smoke test.
What did we not test?
We did not download the 12.9 GB shards, run trec_eval, or reproduce Table scores. This article reports arXiv 2610.11922 and the castorini Hub cards we opened on 10 October 2026.
Common questions
Is Gaggle an agent that calls search tools?
No. It is a pointwise reranker component intended for agentic search stacks. The paper’s broader agent goals remain future work.
Which checkpoint should I download?
castorini/gaggle-reranker-20261005 is the souped model behind the headline tables. Parent run checkpoints are also listed under the castorini org.
Did AiLookout verify the 0.641 DL mean?
No. That mean is reported on the Hub card and paper materials we opened.
What to remember
Project Greenhouse’s first public milestone is a from-scratch open reranker with MIT weights and author-reported TREC/BEIR gains. Use the Hub card for the recipe; rerun nDCG on your corpus before swapping production rankers.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





