12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

SciRAG: Adaptive, Citation-Aware, and Outline-Guided Retrieval and Synthesis for Scientific Literature

Hang Ding, Yilun Zhao, Tiansheng Hu, Manasi Patwardhan, Arman Cohan · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18653/v1/2026.eacl-long.303

Methodology & findings

Study design

Mixed methods including automated evaluation (Correctness Score, Citation F1), human evaluation with three expert annotators using 1-5 Likert scales, hallucination analysis using three independent LLM judges, case studies, and ablation studies on reasoning modules.

Primary method

design science research with iterative system design, component ablation testing, and empirical evaluation

Main result

SciRAG consistently outperforms strong baselines across diverse benchmarks. "SciRAG consistently delivers top-tier answer quality, ranking first in correctness score on 4 out of 5 datasets." Specifically, "SciRAG outperforms strong baselines such as OpenScholar-OS-GPT4o and PaperQA2, with clear gains on SciFact (+2.8), PubMedQA (+9.3), ScholarQA-CS (+11.3), and ScholarQA-MULTI (+0.28) compared to OpenScholar-OS-GPT4o." In human evaluation, "SCIRAG leads in Organization, Coverage, Usefulness, demonstrating SCIRAG could generate a well-organized, comprehensive and useful result to scientific literature queries." Hallucination analysis shows low uncited statement rates: "The proportion of sentences judged unsupported was 6.0% for GPT-4o, 5.3% for DeepSeek-R1, and 6.6% for Gemini 2.5 Pro."

Research paradigm

pragmatist/design science

Author conclusions

The authors conclude: "We presented SciRAG, a novel retrieval-augmented generation framework designed specifically for scientific literature exploration." They emphasize that "By integrating adaptive retrieval with citation centric symbolic reasoning, SCIRAG establishes a new paradigm for trustworthy and scalable scientific knowledge synthesis." The work demonstrates that "extensive experiments on diverse open retrieval benchmarks, such as ScholarQA and PubMedQA, demonstrate that SCIRAG consistently outperforms strong baselines including OpenScholar (Asai et al., 2024) and Pa-perQA2 (Skarlinski et al., 2024), achieving higher factual accuracy and overall relevance."

Risk of bias

Limited human evaluation sample (30 queries from ScholarQA-CS only); Small annotator pool (three expert annotators only); Potential disciplinary bias (annotators may not represent all fields adequately); Model-dependent evaluation (uses GPT-4o for both system and some evaluation); Limited corpus access for PaperQA2 baseline comparison due to proprietary data restrictions; Limited number of expert annotators (only 3) may not capture disciplinary variance; Use of general-purpose models (GPT-4o, Llama-3.1) not fine-tuned for scientific domains introduces potential citation accuracy bias; PaperQA2 baseline evaluation limited by lack of access to private/license-protected papers; Case study evaluation based on 1 representative hard case; Limited human evaluation sample (30 queries from ScholarQA-CS dataset only); Potential model-specific bias from relying on GPT-4o as primary LLM; Limited annotator diversity (only 3 experts in CS); Citation F1 metric may undercount implicitly supported statements

Limitations

  • "Despite its strong performance, SciRAG has certain limitations
  • First, it relies on general-purpose language models such as GPT-4o and Llama-3.1, which are not fine-tuned for scientific citation accuracy and may miss precise attribution
  • Second, the symbolic reasoning and outline-based synthesis introduce non-trivial computational overhead, which may affect real-time applicability in large-scale deployments
  • Also, the human evaluation was conducted with a limited number of expert annotators, which may not capture disciplinary variance."

Open questions raised

  • Future work could explore domain-specific model tuning and lightweight alternatives to improve both precision and efficiency. The paper identifies need for better handling of real-time applicability in large-scale deployments and improved disciplinary representation in evaluation.
  • Future work should explore domain-specific model tuning to improve citation accuracy, lightweight alternatives to reduce computational overhead for real-time applicability in large-scale deployments, and evaluation with broader disciplinary representation.
  • Need for domain-specific model tuning for improved citation accuracy
  • Lightweight alternatives to reduce computational overhead for real-time applicability
  • Better capture of disciplinary variance in human evaluation
  • Fine-tuning for scientific citation accuracy compared to general-purpose LLMs
Data: ScholarQA benchmark (ScholarQA-CS, ScholarQA-BIO, ScholarQA-NEURO, ScholarQA-MULTI); SciFact benchmark; PubMedQA benchmark; QASA benchmark; OpenScholar Datastore (45+ million papers, 200+ million snippets); ScholarQA-CS benchmark (100 queries); ScholarQA-BIO benchmark (1,451 queries); ScholarQA-NEURO benchmark (1,308 queries); ScholarQA-MULTI benchmark (108 queries); SciFact (1,375 claim verification examples); PubMedQA (208 yes/no biomedical questions); QASA (843 reasoning-heavy questions from AI/ML papers); OpenScholar Datastore (over 45 million papers and 200 million snippets); ScholarQA (ScholarQA-CS, ScholarQA-BIO, ScholarQA-NEURO, ScholarQA-MULTI); SciFact; PubMedQA; QASA; OpenScholar Datastore (45 million papers, 200+ million snippets)Extracted from: pdfAgreement 59%

Explore related topics

Related papers