Open Source Hugging Face Blog

Open-sourcing AstaBrief, the fast report-generation model in Asta

AstaBriefAI2report generationopen weights

Language models can already help researchers search the literature, synthesize evidence, and work through complex questions, but scientific work places particular demands on them: answers must stay grounded in evidence, the models must preserve what the evidence actually supports rather than quietly broadening a study's conclusions, and researchers need to verify final outputs. In practice, Asta users bring substantial context and many constraints—for example asking it to compare approaches across a body of literature while accounting for a particular method, population, or setting—and many return to generated reports later, treating them as working research artifacts rather than one-off answers.

The goal was to let scientists generate cited reports faster with a model they could download and run themselves. To do that, the team tested whether a small, open model trained specifically for scientific report generation could match the report quality of the proprietary models they were using while reducing generation time and serving costs. The result is AstaBrief 8B, which turns a research question and retrieved literature excerpts into a cited report.

Developing AstaBrief required tens of thousands of real research queries, citation-focused filtering, preference data, and a redesigned report-generation pipeline that writes the full report in one pass rather than section by section. The payoff is nearly an order-of-magnitude reduction in report generation time compared to the proprietary models tracked: across the full Asta pipeline, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5× faster.

AstaBrief is available today in Asta's Generate a report feature as Fast mode, alongside the Claude-powered Thinking mode, and the weights and training data are being open-sourced so others can study, reproduce, and build on the approach. Open weights also let institutions run AstaBrief on their own infrastructure, which matters when research questions reveal sensitive or unpublished work. Alongside the model weights, an example workflow is being released so researchers can adapt it to create reports from their own PDFs as a starting point for local report generation. The post itself covers how AstaBrief was trained, what was learned about grounding it in scientific evidence, and which parts of the approach could carry forward to future models for science; most of the training and evaluation was completed in 2025, so the proprietary models used to generate training data and as comparison points reflect the frontier at the time, and the full evaluation has not been rerun.

Read original →

← Back to home