Announced 2 Oct 2026 · Sources checked
What did Ai2 release?
AstaBrief 8B is a report generator, not a general chatbot. Given a user question and retrieved literature snippets, it writes a cited scientific report in a single pass. In Asta it is Fast mode, sitting beside a Claude-powered Thinking mode that still summarises and clusters snippets and writes section by section.
Ai2’s 2 October blog says Fast mode averages 51.1 seconds per report across the full Asta pipeline, against 178.5 seconds for Thinking mode, about 3.5 times faster. Separately, it claims nearly an order-of-magnitude cut versus the proprietary models it tracked during development. Those are vendor timings from the period of the project, not measurements we ran.
Open weights are the other half of the pitch: institutions can run the model on their own hardware when a query would reveal unpublished work. That is open weights, not a fully open training stack. Our open-weights versus open-source explainer is the distinction to keep while you read the cards.
How was AstaBrief trained?
The base is Qwen3-8B. Ai2 considered a reinforcement-learning recipe in the style of its DR Tulu work, then chose supervised fine-tuning plus direct preference optimisation because RL was more expensive and harder to debug.
Queries came from real OpenScholar and Asta ScholarQA logs, filtered for quality, language, scientific relevance and personal information. That left about 90,000 research-focused prompts. SFT targets were full reports from the multi-step ScholarQA pipeline, written by a mix of Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini and GPT-4.1, then quality-filtered to about 47,000 examples.
DPO used a held-out query pool. One report per query came from ScholarQA, usually Claude 3.5 or 3.7; the competitor was o3, o4-mini, DeepSeek-V3 or DeepSeek-R1 writing from the same retrieved excerpts. GPT-4.1 and DeepSeek-R1 judged the pair. Ai2 says those judges agreed with humans on 95% of a calibration set and kept only pairs where both judges agreed, leaving about 6,000 examples (the dataset card lists 6,622 rows). DPO ran in open-instruct on 8×H100 GPUs for seven epochs at a 5e-6 linear learning rate, DPO-norm loss, beta 10, and a 16,000-token maximum length.
What do the published scores show?
The development target was SQABench-CS2: 200 user-written computer-science questions in the blog, with a 100-question test split on the model card. Ai2 tracked rubric coverage, paragraph relevance (answer precision), citation precision and citation recall. First SFT runs improved coverage but lagged the Claude pipeline on precision and citations. The filter that helped most was dropping synthetic reports with low citation density.
On the ScholarQA-CS2 test set the card reports Qwen3-8B at 77.3 average, the SFT checkpoint at 83.7, and AstaBrief 8B at 87.0, with citation precision 90.5 and citation recall 78.2. Against Asta ScholarQA (86.2 test, 60.25 on DeepScholarBench) and DR-Tulu-8B (88.8 test, 56.26 DeepScholarBench), AstaBrief sits at 87.0 and 53.50. Its LLM-judged win rate versus Asta ScholarQA is 72% on the test split in that table.
A 14-question human study with three researchers favoured DR Tulu on overall preference; two of the three preferred AstaBrief on citation accuracy. All of this is Ai2’s evaluation, using 2025-era backing models. It is not an independent leaderboard. Read the tables the way our benchmark-marketing guide suggests: useful for the recipe, not for a 2026 ranking.
| System | Average | Citation precision | Citation recall |
|---|---|---|---|
| Qwen3-8B | 77.3 | 76.2 | 64.6 |
| AstaBrief-8B-SFT | 83.7 | 87.7 | 71.3 |
| AstaBrief-8B | 87.0 | 90.5 | 78.2 |
What does Fast mode change in Asta?
Thinking mode still retrieves, summarises, clusters and writes section by section. Fast mode skips those intermediate stages and asks AstaBrief to draft the whole report from the query plus snippets. Ai2 says that cut latency without sacrificing the development metrics it cared about. The model is meant to live inside that retrieval-augmented scaffold, not as a bare chat checkpoint.
Early product telemetry, from Ai2: 374 Asta users tried Fast mode; 29.1% used it on two or more days; they generated 3.67 report threads on average; 23% stayed on Fast and did not return to Thinking; another 18% switched, using Fast for about 40% of threads. Positive-feedback rates were 84.2% (Fast) and 85.2% (Thinking). Ai2 itself says feedback is sparse.
The blog is also clear about a metric it did not optimise: preserving the scope of a cited claim. A report can attach the right paper and still widen a sample finding into a population claim, shift past tense into a timeless rule, or turn a description into a recommendation. That gap is on the authors’ roadmap, not a result they claim to have closed.
What can you download, and under which licence?
The Hugging Face card licenses AstaBrief 8B under Apache 2.0 and names Qwen3-8B as the base. An SFT checkpoint and the DPO mix are published beside it, plus an example workflow for reports from local PDFs. Apache 2.0 on the weights is not a licence to redistribute the preference data. The DPO dataset card uses CC BY-NC 4.0 and warns that synthetic third-party outputs remain under those providers’ terms. Check both cards before you ship a fine-tune.
Intended use on the card is research and education under Ai2’s responsible-use guidelines. Recommended inference follows the training prompt, temperature 0.7, top-p 0.95, and a 4,096-token generation cap in the published snippet.
What should readers not infer?
Do not treat AstaBrief as a 2026 replacement for the current Claude or GPT report stacks. Ai2 says it has not rerun the full evaluation against today’s frontier. Do not treat citation precision as faithfulness to a paper’s actual claim. Do not treat Apache 2.0 weights as permission to republish the CC BY-NC DPO mix.
The useful fact is narrower: a Qwen3-8B post-train, filtered for citation density and aligned with DPO, can sit close to a heavier ScholarQA pipeline on Ai2’s 2025 tables and cut Asta latency. That is a systems result. Independent labs still have to test it on their own corpora.
Common questions
Is AstaBrief trained with reinforcement learning?
No. Ai2 considered an RL path similar to DR Tulu and instead used supervised fine-tuning plus offline DPO.
Are the weights and the training data under the same licence?
No. The model card states Apache 2.0 for AstaBrief 8B. The DPO dataset card states CC BY-NC 4.0.
Did Ai2 beat today’s Claude models on report quality?
It does not claim that. The comparisons use 2025-era ScholarQA and judge models, and the blog says the suite has not been rerun against current frontier systems.
What to remember
AstaBrief is a downloadable, one-pass report writer with Apache 2.0 weights and a narrower DPO-data licence. Use it for local scientific drafts. Treat the published scores as a 2025 methods result, and keep a human on claim scope.
Sources & further reading
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





