The Allen Institute (Ai2) has open-sourced the weights of AstaBrief 8B under the Apache 2.0 license — a fine-tune of Qwen3-8B that, given a research query and provided fragments of scientific papers, writes a complete report with citations in a single pass. In the Asta service, the model became the fast "Generate a report" mode and prepares such a report approximately 3.5 times faster than the multi-step Thinking mode on Claude. The weights are already available on Hugging Face, so the "research query, paper fragments, finished report with citations" pipeline can be deployed on your own infrastructure.

image
image
image

What happened

Ai2 has released the weights of AstaBrief 8B in open access under the Apache 2.0 license. The model was trained on 47,000 SFT examples from real user queries in ScholarQA, where the target reports were generated by Claude 3.5/3.7 Sonnet, o3, o4-mini, and GPT-4.1, plus about 6,000 DPO pairs selected by two judges, GPT-4.1 and DeepSeek-R1: their evaluations agreed with human judgments in 95% of cases. RL was deliberately not used in this recipe. In the Asta service, the model appeared as the fast "Generate a report" mode and outputs a report in an average of 51.1 seconds versus 178.5 seconds for the Thinking mode on Claude. On the ScholarQA-CS2 set of 200 computer science questions, AstaBrief scored an average of 87 points: recall of key elements 90.2, answer accuracy 89, citation accuracy 90.5, citation completeness 78.2. In pairwise comparisons, it outperforms its own multi-step Asta ScholarQA pipeline on Claude in 72% of cases.

Context

This is not a finished research agent, but a step carved out of the pipeline. Asta's main mode, ScholarQA, works as a multi-step agent, still relying on expensive front-end models; Ai2 compressed the final report writing into a compact model, showing that for such a narrow structured task, distilling the behavior of strong teachers without RL is sufficient. Essentially, this is a transfer of the behavior of closed front-end models into open weights, which means the quality ceiling of AstaBrief is limited by what the teachers themselves can do. The approach changes the familiar form of a research agent: instead of a chain of steps with multiple generation passes — a single pass over pre-assembled materials for the question.

Why this matters for the industry

For the industry, the matter is primarily about economics. Open weights under Apache 2.0 turn the generation of reviews with citations from an expensive service function of commercial agents into a reusable artifact that institutions can keep on their own infrastructure — including over sensitive and unpublished data. A new type of component emerges, "fast report over my own corpus": a retrieval-agnostic RAG-writer, embeddable in any pipeline instead of expensive front-end model calls at the final generation step. Ai2's recipe will likely be replicated for other domains, and science-platform products, institutional repositories, and corporate assistants will start including "report with citations" as a low-latency button. The differentiator shifts from text generation to data, the retrieval stage, and citation checking: services whose value rests only on the smoothness of the report face pressure.

Why this matters for users

The practical part is available now. The weights allenai/AstaBrief_8B are on Hugging Face: the 8B model can be run locally via vLLM or transformers with recommended parameters temperature 0.7, top_p 0.95, max_tokens 4096, and a ready-made example for reports on your own PDFs is in the ai2-scholarqa-lib repository in the api/scholarqa/lite folder — an MVP on your own documents can realistically be built in an evening. Without installation, the model can be tried in Fast mode at asta.allen.ai. Two mandatory conditions: the model does not search for anything itself, paper fragments must be provided along with the question, and the prompt must follow the format from sft_prompt.txt, otherwise quality drops. What is written still needs to be checked: citation completeness is the weakest side, a smooth report may miss key sources.

What is still unknown / limitations

The key figures — 72% pairwise wins over the Asta ScholarQA pipeline on Claude and an average score of 87 on ScholarQA-CS2 — were obtained on Ai2's own benchmarks and test set, and independent replication has not yet been done. Human data diverges from the metrics: in a study with three scientists and 14 questions, participants generally preferred DR Tulu, although two of the three considered AstaBrief's citations more accurate, so the benchmark measurement and human perception currently say different things, and the sample is too small to resolve this. It is also unknown how replication behaves outside computer science and on which combinations of retrieval systems and prompts the claimed quality level is maintained.

Sources

Author

Look at AI, editorial team