12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Synthesizing scientific literature with retrieval-augmented language models

Akari Asai, Jacqueline He, Rulin Shao, Weijia Shi, Amanpreet Singh, Joseph Chee Chang et al. · Nature · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
2/4
Quality (LMQS)
C
Evidence
14
Citations
86.21
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41586-025-10072-4

Methodology & findings

Study design

The study combines computational model development with benchmark creation and human evaluation.

Main result

OpenScholar-8B outperforms GPT-4o by 6.1% and PaperQA2 by 5.5% in correctness on multi-paper synthesis tasks. "Although GPT-4o hallucinates citations 78-90% of the time, OpenScholar achieves citation accuracy on par with human experts." In human evaluations, "experts preferred OpenScholar-8B and OpenScholar-GPT-4o responses over expert-written ones 51% and 70% of the time, respectively, compared with 32% for GPT-4o."

Research paradigm

Empirical/computational

Author conclusions

"To further research on LM-based systems that can help scientists navigate the complex, ever-growing task of scientific literature review, we introduce OpenScholar and ScholarQABench. OpenScholar, the first fully open retrieval-augmented system, uses open-weight LLMs and trained retrieval models to iteratively refine scientific output, addressing challenges such as hallucinations and citation accuracy." The authors emphasize: "Our expert evaluation across three scientific disciplines reveals that OpenScholar generates answers that are more helpful than those produced by expert annotators, who required an hour per annotation. Specifically, OpenScholar, using our trained 8B and GPT-4o achieves a 51% and 70% win rate against human-generated answers, respectively."

Risk of bias

Annotator expertise bias - evaluators may not have captured nuanced differences for questions outside their immediate areas of expertise; Small evaluation set sizes (110 computer science, 108 multi-disciplinary) introducing variance; Potential confounding from output length differences - OpenScholar-GPT-4o and OpenScholar-8B produce 2.4x and 2.0x longer responses than expert-written answers; Inter-annotator agreement relatively modest (0.68 pairwise, 0.70 relaxed) suggesting some disagreement; Recency bias - annotations reflect specific time points (July 2024 for Scholar-CS, September 2024 for Scholar-Multi); Domain limitation - evaluation primarily covers computer science, biomedicine, physics; excludes social sciences and other disciplines; Annotator-expertise bias due to small expert-annotated evaluation sets; Potential future contamination risks from public benchmark availability; Selection bias in expert annotators (limited to PhD-level researchers); Limited generalizability beyond computer science, biomedicine, physics, and neuroscience domains; Potential confounding factor of output length differences between model and human responses; Raters may have focused more on writing quality than factual correctness during usefulness assessments; Annotator expertise bias - evaluations conducted by 16 experts may not capture nuanced differences outside their immediate areas of expertise; Small sample size for long-form human evaluation (110 computer science, 108 expert answers) introduces variance; Confounding factor of output length - OpenScholar responses are 2.0-2.4x longer than human answers, potentially influencing preference judgments; Selection bias in question curation - expert writers may have introduced implicit biases in question selection; Potential contamination risk for public benchmark - static public benchmark may be exposed during model training; Inter-annotator agreement moderate (0.68-0.70) suggests some subjectivity in pairwise comparison with ties

Limitations

  • "Expert annotation is costly and time-consuming, so our human-written evaluation sets are small (for example, 110 for computer science long-form question answering
  • 108 expert answers), which may introduce variance and annotator-expertise bias." Additionally, "our automatic evaluation may not perfectly capture quality
  • In Scholar-CS, we combine length, excerpts and rubric items with heuristic weights." The authors note that "ScholarQABench primarily focuses on computer science, biomedicine and physics, with no instances from social sciences or other engineering and scientific disciplines
  • We recognize that our findings may not fully generalize to other domains." Furthermore, "OpenScholar does not consistently retrieve the most representative or relevant papers for certain queries" and outputs "may contain factual inaccuracies or unsupported information, particularly in versions based on our 8B model."

Open questions raised

  • Challenges in retrieving most representative papers - enhanced retrieval methodologies incorporating citation networks and metadata could improve performance
  • Improving factual accuracy in 8B model versions - limited capacity for instruction-following and scientific knowledge
  • Handling of copyright-protected content - currently uses only open-access papers; fair data use in retrieval-augmented LMs remains an open question
  • Fine-grained human analysis of citation accuracy and validity - authors leave more detailed analysis for future work
  • Generalization across scientific domains - particularly for social sciences and restricted-access disciplines
  • Multi-turn interactions - current work focuses on single-turn evaluation; multi-turn LM-human interactions remain challenging
Data: ScholarQABench - multi-domain benchmark with 2,967 expert-written queries and 208 long-form answers (stated as 3,000 research questions and 250 expert-written answers in some sections); OpenScholar DataStore (OSDS) - 45 million scientific papers with 236 million passage embeddings, constructed from peS2o v3 (papers up to October 2024) for inference, with peS2o v2 used for evaluations (papers up to January 2023); Synthetic training data - 130,000 training instances generated through the OpenScholar inference pipeline; OpenScholar DataStore (OSDS): 45 million open-access papers with 236 million passage embeddings; ScholarQABench: 2,967 expert-written queries and 208 long-form answers across computer science, physics, neuroscience, and biomedicine; peS2o: Scientific paper collection (v3 with papers up to October 2024; v2 with papers up to January 2023); Public demo deployed with 30,000+ users and 90,000 user queries collected; ScholarQABench benchmark with 2,967 expert-written queries and 208 long-form answers; peS2o v3 corpus (45 million papers, 236 million passage embeddings) available for OSDS data store. OpenScholar DataStore (OSDS) consists of 45 million open-access scientific papers with pre-computed dense embeddings from peS2o. Authors state: "We open-source all artefacts, including our code, models, data store, datasets and a public demo."Code: Public demo released at launch with over 30,000 users and nearly 90,000 collected user queries; Authors state: "We open-source all artefacts, including our code, models, data store, datasets and a public demo."; All artifacts available at https://doi.org/10.1038/s41586-025-10072-4; OpenScholar code, models, and data store are open-sourced; Public demo available for scientific literature synthesis; All artifacts including code, models, data store, datasets, and demo are open-sourced (specific GitHub/GitLab URLs not provided in text); Open-source release mentioned: "We open-source all artefacts, including our code, models, data store, datasets and a public demo" and "We released the first public demo for scientific literature synthesis, powered by OpenScholar-8B." Specific GitHub/GitLab URLs not provided in the paper text. Code available at https://doi.org/10.1038/s41586-025-10072-4 (Nature article methods section).Extracted from: pdfAgreement 42%

Explore related topics

Related papers