A large language model-driven scientific literature surveys generation framework based on multi-agents and retrieval-augmented generation
Ruihua Qi, Haobo Lyu, Weilong Li, Shuqin Chen, Xu Guo · PeerJ Computer Science · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.7717/peerj-cs.3882
Methodology & findings
Study design
Benchmark evaluation with comparative analysis against multiple baseline models (QFiD, FiD, BigBird, GPT-4o-mini, Llama4-17B-16E) on the SciReviewGen dataset; human-aligned evaluation using LLM-as-judge scoring; ablation studies to assess component contributions..
Primary method
ROUGE metrics (ROUGE-1/2/L); LLM-as-judge evaluation; ablation studies measuring component contributions (RAG removal impact: ROUGE-2/L reduction of 2.28/3.19; single-agent ablation impact: ROUGE-1 decrease of 1.15)
Main result
The framework achieves ROUGE-1/2/L scores of "42.84/16.57/18.25" on the SciReviewGen benchmark, "outperforming strong baselines including Query-weighted Fusion-in-Decoder (QFiD), Fusion-in-Decoder (FiD), BigBird, GPT-4o-mini, and Llama4-17B-16E." Additionally, "the framework also excels in human-aligned evaluation, attaining an LLM-as-judge score of 0.70, surpassing QFiD by +0.29."
Reports effect sizes.
Research paradigm
Empirical-computational (artifact development with benchmarking)
Author conclusions
The authors conclude that "Our work not only advances the state of the art in automated literature synthesis but also offers a scalable solution to mounting scholarly information overload." Additionally, they state that "Ablation studies confirm the critical role of both RAG and multi-agent design: removing RAG reduces ROUGE-2/L by 2.28/3.19, while single-agent ablation further decreases ROUGE-1 by 1.15."
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations