12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A large language model-driven scientific literature surveys generation framework based on multi-agents and retrieval-augmented generation

Ruihua Qi, Haobo Lyu, Weilong Li, Shuqin Chen, Xu Guo · PeerJ Computer Science · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.7717/peerj-cs.3882

Methodology & findings

Study design

Benchmark evaluation with comparative analysis against multiple baseline models (QFiD, FiD, BigBird, GPT-4o-mini, Llama4-17B-16E) on the SciReviewGen dataset; human-aligned evaluation using LLM-as-judge scoring; ablation studies to assess component contributions..

Primary method

ROUGE metrics (ROUGE-1/2/L); LLM-as-judge evaluation; ablation studies measuring component contributions (RAG removal impact: ROUGE-2/L reduction of 2.28/3.19; single-agent ablation impact: ROUGE-1 decrease of 1.15)

Main result

The framework achieves ROUGE-1/2/L scores of "42.84/16.57/18.25" on the SciReviewGen benchmark, "outperforming strong baselines including Query-weighted Fusion-in-Decoder (QFiD), Fusion-in-Decoder (FiD), BigBird, GPT-4o-mini, and Llama4-17B-16E." Additionally, "the framework also excels in human-aligned evaluation, attaining an LLM-as-judge score of 0.70, surpassing QFiD by +0.29."

Reports effect sizes.

Research paradigm

Empirical-computational (artifact development with benchmarking)

Author conclusions

The authors conclude that "Our work not only advances the state of the art in automated literature synthesis but also offers a scalable solution to mounting scholarly information overload." Additionally, they state that "Ablation studies confirm the critical role of both RAG and multi-agent design: removing RAG reduces ROUGE-2/L by 2.28/3.19, while single-agent ablation further decreases ROUGE-1 by 1.15."

Data: not_statedCode: not_statedExtracted from: pdfAgreement 65%

Explore related topics

Related papers