Announced 8 Oct 2026 · Sources checked
What was posted on 8 and 9 October?
The arXiv Atom record we opened lists 2610.11451v1 as published at 08:04:22 UTC on 8 October 2026, primary class cs.CL, with cs.DL and cs.IR secondaries. The comment field says “16 pages, technical report.” Authors start with SAIL Model Team; the HTML marks Bryan Dai as corresponding author and Chi Liu as technical lead.
The public weight repo did not exist on the paper’s posting day. Hugging Face’s model API for IQuestLab/SAIL, retrieved 9 October, gives createdAt 2026-10-09T06:21:20Z, lastModified 2026-10-09T06:51:04Z, sha c3c8d2fbaa928e3f43fe3dfeae998eb07da1948f, gated false, license apache-2.0, and tags that include sail, science, mixture-of-experts, and arxiv:2610.11451. The dataset API for IQuestLab/SAIL-Training-Data gives createdAt 2026-10-09T06:21:40Z and lastModified 2026-10-09T06:51:44Z.
This is a scientific-agent post-train, not a new sequencer or a new literature index. For another Ai2 scientific-writing drop on this site, see AstaBrief 8B. AstaBench, which SAIL uses for most of its tables, is the same Allen Institute suite that paper cites.
What is the science-aware loop?
The abstract says SAIL is developed through a science-aware improvement loop: agents built on frontier AI models analyze the student model’s task failures and construct training tasks that address the gaps. Diagnosis looks at search and evidence selection on literature tasks, scientific assumptions and reasoning on coding tasks, and planning and revision on longer investigations. The agents draw on paper collections and scientific code repositories to build problems, interaction trajectories, and executable tasks with the required environments and tools.
Section 2 repeats that loop over multiple development cycles. Literature failures are split into exploration versus selection. Coding failures are traced to concepts, assumptions, and whether the model notices a bad computational output — the paper’s Figure 3 is a forward-Euler orbit that runs without errors while energy drifts. End-to-end failures are lost objectives, repeated dead ends, or unused intermediate findings. The authors say they validate generated material against sources and task-specific checks, including execution where applicable.
The technical-contributions list is three items: the open 35B / 3B model, the loop, and the training-plus-infrastructure recipe. “We release the model and most of its training data” is the paper’s grant sentence. The dataset card is more specific about what “most” excludes.
How do they train the student?
Section 3 starts from Qwen3.6-35B-A3B. Supervised fine-tuning mixes general instruction data with the scientific tasks from the loop and produces the shared checkpoint. Specialists then train on focused distributions with extra SFT, RL, or both. Multi-teacher on-policy distillation, citing Ma et al. 2026, has the student generate its own trajectories while a routed specialist scores every student token under the same context, including tool observations. Agentic RL then uses task verifiers on executable families. For the MoE shape this starts from, see our mixture-of-experts explainer.
Section 4 says the on-policy stack is built on verl: a training engine, a rollout engine, Uni-Agent adapters into sandboxes, and a gateway that records token-in/token-out IDs. Long episodes that exceed one context window are split when the harness compacts history, but they keep one episode identity and one terminal reward. RL and MOPD share that collection path.
The Hub README’s serve recipes match a language-model-only deployment: vLLM 0.19 or newer, or SGLang 0.5.10 or newer, eight GPUs, a 262,144-token context, the Qwen3 reasoning parser, and the qwen3_coder tool-call parser. The vLLM command passes --language-model-only. Search services, code execution, and research environments, the card says, are supplied by the application.
What can you download, and under what terms?
The model README we opened points to the paper, the training-data repo, Apache License 2.0, and the vLLM and SGLang commands above. Fifteen model-*.safetensors shards plus tokenizer and processor files are listed on the model API. The card’s base_model field is Qwen/Qwen3.6-35B-A3B. That is an Apache-2.0 post-train of an existing MoE, not a from-scratch public pretrain. For that distinction, see open weights versus open source.
The dataset README states 171,586 scientific training examples for supervised fine-tuning and prints the five-way split: 157,742 scientific-coding, 4,966 research-reasoning, 3,702 literature-analysis, 3,477 research-synthesis, and 1,699 end-to-end-discovery. Records are Parquet under sail-code, sail-search, and sail-e2e. Each row has an id, a messages list, optional tool schemas, and metadata. The same card says this repository provides text conversations and tool definitions; executable tools and task environments are not included.
usedStorage on the dataset API is 1,260,454,735 bytes. The model API’s usedStorage is 70,236,189,701 bytes. Those are Hub accounting figures, not a load we performed. The paper’s “most of its training data” sentence and the dataset’s missing environments are different statements. Do not collapse them.
What did we not run?
This is an evidence review of the 8 October abstract, Atom record, and HTML, and of the 9 October Hugging Face model and dataset API records and README files. We did not download the 70 GB shards, load the 171,586 Parquet rows, stand up vLLM, or rerun SciCode, AstaBench, or DeepResearch Bench II. Treat the loop as what the PDF describes, the percentages as what the authors printed, and the Hub drop as an Apache-2.0 weight-plus-SFT-text release without the execution environments.
Common questions
Is SAIL a from-scratch 35B model?
No. The paper and the Hub card say it is post-trained from Qwen3.6-35B-A3B, a 35B-total / 3B-active mixture-of-experts. The Apache-2.0 grant covers the SAIL checkpoint and the released SFT text. It is not a new pretrain.
Does the training-data repo include the tools SAIL used?
The dataset card says it includes text conversations and tool definitions. It also says executable tools and task environments are not included. Search, code execution, and research sandboxes have to come from your application, which is the same note on the model README.
Did SAIL beat GLM-5.2 overall?
Not on the authors’ own twelve-task aggregate. They print 59.76 for SAIL and 63.30 for GLM-5.2. They do print a SciCode lead, 50.35 against 47.57, and an ArxivDIGESTables lead, 35.24 against 35.13 for DeepSeek-V4-Flash-0731. Those are their tables.
What to remember
Use 2610.11451 for the science-aware loop and keep every AstaBench and SciCode percentage in the authors’ column. Use the 9 October Hub repos for what is actually fetchable: Apache-2.0 weights, 171,586 SFT rows, and no bundled execution environments. Do not invent an independent SciCode win or a from-scratch 35B pretrain.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





