In this article
AstaBrief is now open source
The new AstaBrief model generates cited scientific reports in an average of 51.1 seconds, compared to 178.5 seconds for the previous proprietary Thinking mode.
Scientific work requires answers grounded in evidence. Models must preserve what the data actually supports rather than quietly broadening a study’s conclusions. Researchers also need to verify final outputs.
Scientists using Asta often bring substantial context and many constraints. They might ask the platform to compare approaches across a body of literature while accounting for a specific method, population, or setting. Many users return to generated reports later, treating them as working research artifacts rather than one-off answers.
The team wanted to help scientists generate cited reports faster with a model they could download and run themselves. They tested whether a small, open model trained specifically for scientific report generation could match the quality of proprietary models while reducing generation time and serving costs.
AstaBrief 8B turns a research question and retrieved literature excerpts into a cited report. It is available in Asta’s Generate a report feature today as Fast mode alongside Claude-powered Thinking mode. The team is also open-sourcing the model and training data so others can study, reproduce, and build on their approach.
Developing AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned report-generation pipeline that writes the full report in one pass rather than section by section. The result is nearly an order-of-magnitude reduction in report generation time compared to the proprietary models they tracked.
Across the full Asta pipeline, Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, about 3.5 times faster.
Those efficiency gains made AstaBrief a useful test case for a broader goal: building open language models that can be adapted to the specific demands of scientific work.
Open weights will also let institutions run AstaBrief on their own infrastructure. This is necessary when research questions reveal sensitive or unpublished work. Alongside the model weights, the team is releasing an example workflow that researchers can adapt to create reports from their own PDFs. This provides a starting point for local report generation.
This post covers how they trained AstaBrief, what they learned about grounding it in scientific evidence, and which parts of their approach they think can carry forward to future models for science. Most of the training and evaluation described was completed in 2025. The proprietary models used to generate training data and as comparison points reflect the frontier at the time. The team has not rerun the full evaluation against today’s frontier models. The results below are best read as evidence about the particular training and system design choices they tested.
Training the model
The goal with AstaBrief was to build an open-weights model with all the qualities that matter most for long-form scientific synthesis: answer quality, relevance, structure, and citation grounding. They started from Qwen3-8B and focused most of their effort on the post-training data, evaluation, and surrounding report-generation scaffolding.
Adapting general-purpose models for scientific work and training new scientific models from scratch is something the team is exploring broadly across Ai2. Through NSF OMAI, a U.S. national initiative led by Ai2 to build fully open AI infrastructure and models for scientific discovery, their researchers are working directly with scientific communities to understand what they need from future open models and where today’s general-purpose models fall short. That includes studying how needs differ across scientific fields and workflows, with more findings from that research to share in the future.
Recent work, including their DR Tulu, has shown that reinforcement-learning-based methods can improve long-form report generation for open-weights models, especially when judge models are involved in the training loop. They considered that path for AstaBrief, but ultimately focused on a simpler recipe built around supervised fine-tuning and direct preference optimization.
RL-based training can be unstable and expensive. They wanted to see how far they could push report generation quality with a cheaper, more operationally manageable setup. This approach is also easier to debug and iterate on.
That made the quality of the training data especially important. Rather than relying on a more complex optimization method to compensate for noisy examples, they spent much of the project figuring out how to generate, select, and filter examples that actually demonstrated the report-writing behaviour they wanted.
They also wanted AstaBrief to be faster so that users could get preliminary reports quickly that they could then iterate over in subsequent turns. For speed improvements, they decided to train AstaBrief to directly generate the final report in one pass given a user query and relevant retrieved snippets. This bypassed the expensive snippet summarization and clustering stages their Claude-based Thinking mode uses and did not write out the answer section-by-section. Interestingly, they found it was possible to do so without sacrificing performance.
Collecting SFT training data
The training pipeline began with real user queries submitted through the system described in their paper Synthesizing scientific literature with retrieval-augmented LMs and ScholarQA, the framework that now underpins Asta’s Generate a report feature. Rather than training only on synthetic prompts or benchmark-style tasks, they wanted AstaBrief to learn from real queries from real scientists.
Research suggests that scientists often ask different things of language models than users do of general-purpose chatbots or traditional search tools. In their analysis of hundreds of thousands of Asta queries, expert researchers frequently supplied substantial context, multiple constraints, and relationships between concepts rather than relying on short, keyword-style prompts.
More recent Asta user studies have also surfaced differences in how researchers want AI involved in their work. Some are comfortable using models for ideation or experimentation, while others prefer a narrower role in synthesis, literature surveillance, or pattern-finding. Across those differences, participants want clearer source traceability, more visibility into what a model is doing, and greater control over the context it uses.
They filtered the user logs they collected for quality, relevance, and privacy. They stripped out beta-tester and bot traffic, dropped queries that were too short to be meaningful, and used an LLM-based filtering pass to catch non-English queries, non-scientific requests, and prompts containing personal information. That left a pool of 90K research-focused queries.
For SFT, they generated full-report target outputs from the filtered queries using the multi-step ScholarQA pipeline behind Asta’s report generation. The pipeline retrieved relevant literature, organized the material into sections, and used a backing report-generating model to synthesize the evidence into a cited report. They drew on a mix of proprietary systems: Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1. After quality filtering, this yielded 47K usable training examples.
Creating DPO pairs
DPO required a different kind of training data. Instead of a single target report per query, they needed pairs of reports with one preferred over the other.
They built those pairs from a separate subset of queries not used during SFT data generation. One report per query came from the existing ScholarQA pipeline, typically backed by Claude 3.5 Sonnet or 3.7 Sonnet. The competing report was generated by feeding ScholarQA’s retrieved literature excerpts to a different model: o3, o4-mini, DeepSeek-V3, or DeepSeek-R1, depending on the example.
Two judge models, GPT-4.1 and DeepSeek-R1, compared each pair and picked a winner. They ensured that LLM judges were aligned with human preferences. They only kept pairs where both judges agreed, which gave them a cleaner preference set and cut much of the noise that typically shows up in preference data generated at scale.
After quality filtering, the final DPO dataset came to about 6K examples.
Using multiple generators and requiring agreement between two judges gave them a relatively simple way to construct preference data without treating any single model’s output or judgment as ground truth.
Filtering data for better attribution
Their main evaluation target was SQABench-CS2, a set of 200 user-written computer science research questions. They tracked four metrics throughout the development of AstaBrief:
- Rubric score, which measures how much necessary content is covered by the report.
- Answer precision, which measures whether each paragraph is relevant to the question.
- Citation precision, which measures whether each citation supports the claim it is attached to.
- Citation recall, which measures whether the report’s claims are fully supported by the citations provided.
For their final model, they also ran secondary evaluations: DeepScholarBench, a 63-query benchmark for long-form research synthesis built from recent ArXiv papers, and two separate pairwise evaluations against reports generated by the Claude-powered pipeline. This included an LLM-judged comparison on SQABench-CS2 and a small human study.
A report can sound polished and complete while meandering from the question or attaching citations to claims from which the underlying evidence does not follow. For scientific synthesis, they needed to measure those behaviours separately. But citation support is only part of scientific faithfulness. A model can cite the right study and still make a stronger claim than the study itself supports. This can happen in subtle ways, for example, turning a finding about a particular sample into a generic claim about an entire population, shifting a result reported in the past tense into a present-tense statement that sounds more universally true, or turning a descriptive finding into a recommendation for what clinicians, policymakers, or researchers should do.
Those kinds of generalizations are especially important for scientific report generation because each step can broaden the apparent scope of the evidence without introducing an obviously false statement. A cited sentence may therefore be technically related to its source while still overstating what researchers actually established. Their development metrics focused primarily on relevance, coverage, and citation grounding. A richer evaluation of scientific report writers should also test whether they preserve the scope and strength of the claims in their sources.
Their first SFT runs improved overall content quality, but they still lagged behind their Claude-powered report generation pipeline on answer precision and citation quality. In other words, the model got better at writing reports, but it was not grounded in evidence as consistently as they needed for scientific synthesis.
That pushed them to spend more time on data quality. They tested four statistics-based filters to identify weaker synthetic training examples:
- Output-to-input token ratio. Answers with very high ratios were often noisy because they were generating a lot of text from




