Briefing

AstaBrief: an open, fast model for cited scientific reports
A small open-weights model, trained on real researcher questions, now writes grounded literature reports in about a third of the time of a proprietary pipeline, and anyone can download it and run it themselves.
At a glance
AstaBrief 8B takes a research question plus retrieved literature excerpts and returns a cited report in a single pass. It is offered as "Fast mode" in Asta's Generate a report feature, alongside the Claude-powered "Thinking mode."
Model size and base8 billion parameters, fine-tuned from Qwen3-8B.
SpeedFast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, roughly 3.5x faster across the full pipeline.
RecipeSupervised fine-tuning followed by direct preference optimization, with no reinforcement learning.
ReleaseModel weights, training data, and an example workflow for building reports from your own PDFs.
Why scientific reports are a hard problem
General-purpose chat answers are not enough for research work. Scientists need outputs they can trace, check, and return to later.
Evidence-grounded answersEvery claim should be tied to a source that actually supports it.
Faithful scopeModels must not quietly broaden what a study concluded.
VerifiabilityResearchers need clear source traceability to review results without wading through filler.
Context-rich queriesExpert users supply detailed context, constraints, and relationships between concepts rather than short keywords.
Reports as lasting artifactsMany users revisit generated reports later as working research documents.
Training approach
The team focused on post-training data, evaluation, and the surrounding report-generation scaffolding rather than a heavier optimization method.
SFT plus DPO over RLReinforcement learning can be unstable and costly. A simpler recipe is cheaper, easier to debug, and faster to iterate on.
Data quality firstEffort went into generating, selecting, and filtering examples that demonstrate the desired report-writing behavior.
One-pass generationThe model writes the whole report directly from the query and retrieved snippets, skipping snippet summarization, clustering, and section-by-section writing.
No quality sacrificeRemoving those stages sped things up without a measurable loss in performance.
Built for iterationQuick preliminary reports let users refine their question in follow-up turns.
Building the training data
Training used real researcher queries instead of purely synthetic or benchmark-style prompts.
Query collection and cleaningUser logs were filtered to remove beta-tester and bot traffic, very short queries, non-English and non-scientific requests, and prompts with personal information. This left 90K research-focused queries.
SFT targetsFull reports were produced with the multi-step ScholarQA pipeline using several proprietary backing models (Claude 3.5 and 3.7 Sonnet, o3, o4-mini, GPT-4.1). After quality filtering, 47K examples remained.
DPO preference pairsEach pair compared a ScholarQA report against a report from another model (o3, o4-mini, DeepSeek-V3, or DeepSeek-R1) built from the same retrieved excerpts. About 6K pairs remained after filtering.
Dual-judge agreementGPT-4.1 and DeepSeek-R1 both judged each pair, and only pairs where they agreed were kept. The judges matched human preferences 95% of the time.
No single source of truthMixing generators and requiring two agreeing judges avoids treating any one model's output or opinion as ground truth.
Filtering for better attribution
Early SFT runs improved content quality but trailed the Claude-powered pipeline on answer precision and citation quality. Four statistical filters were tested to remove weak training examples.
Output-to-input token ratioVery high ratios often signal noisy answers that generate lots of text from too little evidence.
Citation relevanceLow average retrieval scores for cited papers suggest over-reliance on lower-ranked evidence.
Citation densityThe share of statements with at least one citation. Low density means long stretches of unsupported text.
Citation diversityThe share of retrieved papers actually cited. Low scores mean the report leans on only a few sources.
The winner: citation densityFiltering out low-density reports gave the strongest gains. More aggressive filtering, filter combinations, and learning-rate sweeps added little.
The lessonA simple "does it cite its claims?" signal beat more elaborate combinations. Composition and quality of post-training data can matter more than adding scientific text.
How it was evaluated
The main benchmark was SQABench-CS2, a set of 200 user-written computer science research questions. Four metrics were tracked throughout development.
Rubric scoreHow much of the necessary content the report covers.
Answer precisionWhether each paragraph is relevant to the question.
Citation precisionWhether each citation supports the claim it is attached to.
Citation recallWhether the report's claims are fully supported by the citations provided.
DeepScholarBenchA 63-query benchmark for long-form research synthesis built from recent ArXiv papers, used as a secondary check.
Pairwise comparisonsAn LLM-judged comparison against Thinking mode on SQABench-CS2, plus a small human study.
Results
After DPO, AstaBrief reached a level competitive with the Claude-powered pipeline and with DR Tulu across several answer and citation measures.
Competitive qualityWithin range of Thinking mode and DR Tulu on the development metrics.
Large speed gainNearly an order-of-magnitude faster generation than the proprietary models tracked, and 51.1 versus 178.5 seconds end to end.
Human studyThree researchers ranked reports on 14 questions. DR Tulu won on overall preference, while two of the three researchers preferred AstaBrief on citation accuracy.
Reading the numbersUnlike DR Tulu, AstaBrief was optimized for pairwise report ranking during DPO, so the comparison is not perfectly like-for-like.
Faithfulness goes beyond citations
A sentence can cite the right study and still overstate what that study showed. These subtle shifts are especially hard to catch because none produces an obviously false statement.
Sample to populationA finding about one sample becomes a claim about everyone.
Past to present tenseA reported result becomes a statement that sounds universally true.
Description to recommendationA descriptive finding becomes advice for clinicians, policymakers, or researchers.
A gap in current metricsThe development metrics covered relevance, coverage, and citation grounding. Future evaluations should also test whether scope and strength of claims are preserved.
Early usage in Asta
Early production data comes from 374 users who tried Fast mode. The team cautions that feedback is sparse.
Repeat use29.1% used Fast mode on two or more days, averaging 3.67 report threads each.
Committed users23% kept using Fast mode and never switched back to Thinking mode.
Mixed workflowsAnother 18% switched between modes, using Fast mode for about 40% of their threads.
SatisfactionPositive feedback was similar for both modes: 84.2% for Fast and 85.2% for Thinking.
Why open weights matter
Data privacyInstitutions can run the model on their own hardware, including behind a firewall, which matters when queries reveal sensitive or unpublished work.
No API dependencyReport generation no longer relies on a proprietary model API.
ReproducibilityReleased weights and training data let others study, reproduce, and extend the approach.
Local workflowAn example workflow shows how to generate reports from your own PDFs.
Mode choiceThinking mode remains available for more compute-intensive tasks.
Limitations to keep in mind
Dated comparisonsMost training and evaluation took place in 2025, and the full evaluation has not been rerun against today's frontier models.
Narrow benchmark domainThe main benchmark covers computer science questions only.
Small human studyThree researchers and 14 questions is useful signal, not a statistical conclusion.
Sparse feedbackUsage and satisfaction figures come from a small early cohort.
Teacher dependenceTraining targets came from proprietary models, so the student inherits some of their strengths and weaknesses.
What's next
Finer-grained preference learningMoving beyond whole-report preferences.
RAG plus RLStronger retrieval-augmented approaches combined with reinforcement learning.
Multi-turn and multi-tool abilitiesSupporting longer research conversations and richer tool use.
More data sources and query decompositionBroader scientific coverage and better handling of complex questions.
Scope-preserving evaluationMeasuring concision, organization, and whether evidence scope is kept intact.
Broader science modelsPart of a longer line of work, from ScholarQA and DR Tulu to future Olmo versions, and the NSF OMAI initiative for open scientific AI.
Key takeaways for practitioners
Filter on a simple grounding signalCitation density was the most effective data filter tested.
Use multiple generators and judgesCross-model agreement yields cleaner preference data without trusting a single model.
Simplify the pipelineSingle-pass generation can cut latency and cost without losing quality.
Measure scope, not only supportPlan evaluations that check whether claims are broadened beyond the evidence.
Train on real queriesResearcher questions are richer than benchmark prompts and better reflect real use.