hraness
Theme
Appearance

saved

Open-sourcing AstaBrief

by Ai2Ai2published

Hraness wrote this summary from a saved copy of the source. Quotations are taken word for word from the source.

gist

Ai2 is open-sourcing AstaBrief 8B, a Qwen3-8B model trained with supervised fine-tuning and direct preference optimization to write a cited scientific report in one pass from a question and retrieved excerpts. In Asta, Fast mode averages 51.1 seconds a report against 178.5 seconds for the Claude-powered Thinking mode. The practical lesson is the data: filtering synthetic reports for citation density mattered more than fancier filters, and the 2025 comparisons describe that recipe rather than today’s frontier.

ideas

  • A small open model can write the whole report. AstaBrief turns a question and retrieved excerpts into one cited report, skipping the section-by-section Claude pipeline.
  • The recipe is SFT and DPO, not reinforcement learning. Ai2 trained on real scientist queries, filtered synthetic reports, and preference pairs that two judges agreed on.
  • Citation density was the filter that worked. Dropping low-citation synthetic examples improved grounding more than stacking several more elaborate statistics.
  • The speed result is about that system, not a new leaderboard. Fast mode is about 3.5 times quicker in the Asta pipeline, and the comparisons were run against 2025 models.

quotes

“Fast mode averages 51.1 seconds per report compared with 178.5 seconds for Thinking mode, about 3.5× faster.”

Ai2, comparing Fast mode with Thinking mode.

“The strongest gains came from filtering out synthetic reports with low citation density”

Ai2, on which training filter helped.

“Open weights will also let institutions run AstaBrief on their own infrastructure”

Ai2, on why the weights are open.

“RL-based training can be unstable and expensive.”

Ai2, explaining why they avoided reinforcement learning.