AstaBrief Is Open Source: Run Your Own Report Engine

You write the same report every week: messy notes in, structured brief out, a person rereading every line before it ships. A model built for exactly that job just went open-weights, and the useful question isn't "is the benchmark good?" It's whether the output survives your sources, with your numbers, in your format.
Here's what shipped, what it was actually built for, and the one-afternoon test I'd run before trusting it with anything a client sees.
What Ai2 actually released
AstaBrief 8B is an open-weights model from Ai2 (the Allen Institute for AI) that turns a research question plus retrieved literature excerpts into a cited report. Ai2 announced it on October 2, 2026. It already powers the "Fast" mode of the Generate a report feature in Asta, sitting next to a Claude-powered "Thinking" mode, and you can also download the weights and run it on your own infrastructure (Ai2 announcement).
The facts worth knowing before you get excited:
- License and base: the model card lists Apache 2.0 and a Qwen3-8B base (Hugging Face model card).
- Training recipe: supervised fine-tuning followed by offline DPO. No reinforcement learning. Ai2 said RL can be unstable and expensive, and it wanted something easier to debug (Unite.AI).
- Teacher data: the target reports for SFT came from Ai2's multi-step ScholarQA pipeline, backed by proprietary models including Claude 3.5 Sonnet, Claude 3.7 Sonnet, o3, o4-mini, and GPT-4.1 (Ai2).
- It does not retrieve anything. You supply the query and the retrieved excerpts in the recommended prompt format. The card warns a different format may give degraded or inconsistent behavior (model card).
That last point matters more than it looks. This is a writer, not a research agent. Retrieval is your job.
Why it's fast, and what that number does and doesn't mean
AstaBrief is fast because the pipeline got shorter, not because the model is magic. It writes the final report in one pass from the query and retrieved snippets. It skips the snippet summarization and clustering stages and doesn't write section by section. Ai2 says it found this possible without sacrificing performance (Ai2).
Ai2's reported timings across the full hosted Asta pipeline: Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5× faster. Two caveats, both important:
- These are Ai2's own averages for the hosted pipeline. Nobody has established what a self-hosted deployment on your hardware will do, so don't plan around 51 seconds.
- Faster than Thinking mode is not a quality claim. It's a latency claim.
I haven't benchmarked it yet, so I'm not quoting speed numbers I didn't measure. Part of the test below is measuring your own.
| Fact | Source |
|---|---|
| Fast mode avg 51.1 s/report (hosted pipeline) | Ai2 blog |
| Thinking mode avg 178.5 s/report | Ai2 blog |
| Single-pass writing, no section-by-section stage | Ai2 blog |
| Local latency on your hardware | Unknown, measure it |
What the benchmarks cover, and the gap that matters to you
The published numbers are for scientific literature tasks, not for business reports. On the model card, AstaBrief-8B scores 86.3 on SQA-CS2 (Dev) and 87.0 on SQA-CS2 (Test), with win rates against Asta SQA of 55% (Dev) and 72% (Test). The card calls it competitive with Asta ScholarQA and DR-Tulu (model card).
The sub-scores are more useful than the headline:
| Metric | Score |
|---|---|
| Ingredient recall | 90.2 |
| Answer precision | 89 |
| Citation precision | 90.5 |
| Citation recall | 78.2 |
| Average | 87 |
Citation recall at 78.2 is the weakest number. Read that as: the model tends to be right about what it cites, but it may leave claims uncited. For anyone building a "every statement traces to a source" workflow, that's the failure mode to watch.
Now the honest gap. Nothing I found tests this model on notes, support tickets, invoices, or client updates. It was trained for scientific papers with citations. Applying it to your weekly client digest is extrapolation, and quality outside science is unconfirmed. Also, Ai2 hasn't rerun its full evaluation against current frontier models, so don't read these scores as "beats the big general models." They don't say that.
That gap is the whole reason for the test in the next section.
The one-afternoon test
Pick one report you already produce by hand and run five past examples through the model. You need five cases where you have the source material and the final version you actually sent, because then you know the correct answer. Don't score "overall quality." Count specific failures.
Failure types to tally per report:
- Wrong number: a figure in the output that differs from the source.
- Missing item: something your final version included that the draft dropped.
- Invented claim: a statement the source never made.
- Merged details: facts from two sources or two clients blended together.
Write down the count out of five for each type. Then the decision rule:
- Low failures, same type every time: you can build a check around them. A rule that flags any number in the output that doesn't appear in the source catches the most dangerous class automatically.
- Random failures: keep a human in the loop and use the model for first drafts only.
Either way, you now have a real number instead of a vibe.
Here's a minimal version of that number check. It extracts numeric tokens from the draft and flags any that never appear in the source:
import re
NUM = re.compile(r"\$?\d[\d,]*\.?\d*%?")
def numbers(text: str) -> set[str]:
# normalize: strip $ and commas so "$1,200" == "1200"
return {n.replace("quot;, "").replace(",", "") for n in NUM.findall(text)}
def unsupported_numbers(source: str, draft: str) -> list[str]:
return sorted(numbers(draft) - numbers(source))
source = open("source_notes.txt").read()
draft = open("model_draft.txt").read()
flags = unsupported_numbers(source, draft)
if flags:
print("REVIEW BEFORE SENDING. Numbers not found in source:", flags)
else:
print("No unsupported numbers (still read it).")
This is deliberately dumb. It won't catch a wrong figure that happens to also appear elsewhere in the source, and it can't catch an invented non-numeric claim. It catches the "looks great, wrong figure in paragraph three" case, which is the one that actually burns you. Treat it as a floor, not a verdict.
For timing, wrap each run so you record seconds per report on your own box. That's the only latency number that counts for you.
Running it yourself: hardware and prompt reality
Plan around the weights, then add headroom. Community quantizations list the original BF16 weights at about 16.4 GB, an FP8 version at about 9.4 GB, and a 4-bit GGUF (Q4_K_M) at roughly 4.7 GB. Those are third-party conversions, not Ai2 releases (FP8 repo example).
Treat those as weights-only sizes. Long-context memory comes on top, and claims that it "fits comfortably on a 24 GB card" come from third-party estimates, not an Ai2 spec. Test with your real input lengths before you commit hardware.
Two practical notes from the model card and release:
- Stick to the recommended prompt format. Ai2 warns that deviating may degrade or destabilize behavior. If your pipeline wraps tickets in a custom template, expect to spend time matching the format, not just pasting text in.
- Ai2 shipped an example workflow for building reports from your own PDFs alongside the weights (Ai2). If your sources are documents rather than tickets, start there instead of writing the scaffolding from scratch.
One thing to read before using this commercially: the license is Apache 2.0, but the model card also says it's intended for "research and educational use" under Ai2's Responsible Use Guidelines. How far client-report work is covered wasn't clarified in the sources I found. Read the guidelines yourself before you build a paid service on it.
Where this fits in a real pipeline
If the five-report test passes, the pipeline is short:
trigger: weekly schedule OR new batch of inputs
steps:
- collect_sources: # tickets, notes, PDFs -> excerpts
- build_prompt: # AstaBrief's recommended format, query + excerpts
- run_model: # self-hosted, one pass
- number_check: # flag numbers not present in source
- human_review: # 2-minute read, not a 2-hour write
- send
The model drafts; the check catches the dangerous class of error; a person reads before anything leaves. That's where the hours come back, and it's why I treat any report generator as a drafter, not a sender.
My view, and it's an opinion, not a measured result: for boring recurring work like this, a small open model you host and verify often beats paying per token on a giant general model for every client every week. The case rests on three things: cost becomes fixed instead of per-token, confidential text stays on your machine, and open weights don't change underneath you. Whether AstaBrief specifically is that model for business reports is exactly what the test answers. The reason to run it on your own data rather than trust a scientific-literature benchmark is that fast and clean-looking is not the same as correct.
Early signal from Ai2's own usage data is mildly encouraging but thin: of 374 users who tried Fast mode, 29.1% used it on two or more days, and 23% never switched back to Thinking mode. Ai2 itself treats that as early evidence. I'd treat it the same way.
Why bizflowai.io helps with this
At bizflowai.io, the report pipelines I build for clients follow the shape above: scheduled or batch triggers, a drafting step, a programmatic check that flags figures not present in the source, and a short human review before anything is sent. What I've found is that the drafting model is the easy part to swap; the verification step and the input plumbing are where the reliability comes from, so that's where the build time goes.
Want more like this?
I publish practical AI automation, GenAI engineering, and faceless content workflows on YouTube every week.
Subscribe to bizflowai.io on YouTube — never miss a new tutorial.
Planning an AI automation project or need a second opinion on your architecture?
Connect with me on LinkedIn — Lazar Milicevic, GenAI Engineer & bizflowai.io Founder.
Visit bizflowai.io for our services, case studies, and AI consulting.
Frequently asked questions
What is AstaBrief?
AstaBrief is a model released on Hugging Face by the Allen Institute for AI. It powers the report-generation feature in Asta, the institute's research assistant. Per the announcement, it is built to do one job quickly: take a pile of source material and produce a structured brief. It is now open, so you can download it and run it yourself instead of using it only through another company's app.
Why does an open report-generation model matter for small businesses?
Open weights offer three advantages. Cost: a self-hosted specialized model turns per-token fees on high-volume, repetitive reporting into a fixed cost. Data: confidential material like invoices and contracts stays on your own server. Control: hosted models change over time, while open weights don't, so a report format that works today keeps working next month.
How do I test whether a report model is accurate enough for my business?
Pick one report you already produce by hand and gather five past examples where you know the correct answer, with the source material and the final version you sent. Run each source through the model and compare line by line. Count specific failures: wrong numbers, missing items, and invented claims. Record the failure count out of five rather than scoring overall quality. The whole test takes about an afternoon.
Can I trust AI-generated reports without checking them?
No. Fast doesn't mean accurate. A report model can produce a clean-looking brief that quietly drops a number, merges two clients' details, or states something the source never said. The danger is output that looks great but contains one wrong figure. Treat any report generator as a drafter, not a sender, and have a person review every report before it goes out.
When should I automate report checks versus keep a human reviewing every draft?
Look at your test failures. If the failure count is low and the same type of error repeats, build an automated check around it, such as a rule that flags any number in the output that doesn't appear in the source. If the failures are random, keep a human in the loop and use the model only for the first draft. Even with checks, a person should do a quick final review.