Synthesizing scientific literature with retrieval-augmented language models
Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang et al. · Nature · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41586-025-10072-4
Methodology & findings
Study design
The study combines computational model development with benchmark creation and human evaluation.
Main result
OpenScholar-8B outperforms GPT-4o by 6.1% and PaperQA2 by 5.5% in correctness on multi-paper synthesis tasks. "Although GPT-4o hallucinates citations 78-90% of the time, OpenScholar achieves citation accuracy on par with human experts." In human evaluations, "experts preferred OpenScholar-8B and OpenScholar-GPT-4o responses over expert-written ones 51% and 70% of the time, respectively, compared with 32% for GPT-4o."
Research paradigm
Empirical/computational
Author conclusions
"To further research on LM-based systems that can help scientists navigate the complex, ever-growing task of scientific literature review, we introduce OpenScholar and ScholarQABench. OpenScholar, the first fully open retrieval-augmented system, uses open-weight LLMs and trained retrieval models to iteratively refine scientific output, addressing challenges such as hallucinations and citation accuracy." The authors emphasize: "Our expert evaluation across three scientific disciplines reveals that OpenScholar generates answers that are more helpful than those produced by expert annotators, who required an hour per annotation. Specifically, OpenScholar, using our trained 8B and GPT-4o achieves a 51% and 70% win rate against human-generated answers, respectively."
Risk of bias
Annotator expertise bias - evaluators may not have captured nuanced differences for questions outside their immediate areas of expertise; Small evaluation set sizes (110 computer science, 108 multi-disciplinary) introducing variance; Potential confounding from output length differences - OpenScholar-GPT-4o and OpenScholar-8B produce 2.4x and 2.0x longer responses than expert-written answers; Inter-annotator agreement relatively modest (0.68 pairwise, 0.70 relaxed) suggesting some disagreement; Recency bias - annotations reflect specific time points (July 2024 for Scholar-CS, September 2024 for Scholar-Multi); Domain limitation - evaluation primarily covers computer science, biomedicine, physics; excludes social sciences and other disciplines; Annotator-expertise bias due to small expert-annotated evaluation sets; Potential future contamination risks from public benchmark availability; Selection bias in expert annotators (limited to PhD-level researchers); Limited generalizability beyond computer science, biomedicine, physics, and neuroscience domains; Potential confounding factor of output length differences between model and human responses; Raters may have focused more on writing quality than factual correctness during usefulness assessments; Annotator expertise bias - evaluations conducted by 16 experts may not capture nuanced differences outside their immediate areas of expertise; Small sample size for long-form human evaluation (110 computer science, 108 expert answers) introduces variance; Confounding factor of output length - OpenScholar responses are 2.0-2.4x longer than human answers, potentially influencing preference judgments; Selection bias in question curation - expert writers may have introduced implicit biases in question selection; Potential contamination risk for public benchmark - static public benchmark may be exposed during model training; Inter-annotator agreement moderate (0.68-0.70) suggests some subjectivity in pairwise comparison with ties
Limitations
- "Expert annotation is costly and time-consuming, so our human-written evaluation sets are small (for example, 110 for computer science long-form question answering
- 108 expert answers), which may introduce variance and annotator-expertise bias." Additionally, "our automatic evaluation may not perfectly capture quality
- In Scholar-CS, we combine length, excerpts and rubric items with heuristic weights." The authors note that "ScholarQABench primarily focuses on computer science, biomedicine and physics, with no instances from social sciences or other engineering and scientific disciplines
- We recognize that our findings may not fully generalize to other domains." Furthermore, "OpenScholar does not consistently retrieve the most representative or relevant papers for certain queries" and outputs "may contain factual inaccuracies or unsupported information, particularly in versions based on our 8B model."
Open questions raised
- Challenges in retrieving most representative papers - enhanced retrieval methodologies incorporating citation networks and metadata could improve performance
- Improving factual accuracy in 8B model versions - limited capacity for instruction-following and scientific knowledge
- Handling of copyright-protected content - currently uses only open-access papers; fair data use in retrieval-augmented LMs remains an open question
- Fine-grained human analysis of citation accuracy and validity - authors leave more detailed analysis for future work
- Generalization across scientific domains - particularly for social sciences and restricted-access disciplines
- Multi-turn interactions - current work focuses on single-turn evaluation; multi-turn LM-human interactions remain challenging
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations