SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models
Tong Bao, Mir Tafseer Nayeem, Davood Rafiei, Chengzhi Zhang · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18653/v1/2025.emnlp-main.136
Methodology & findings
Study design
Empirical evaluation study with three task-based experiments (Fully LLM-based survey generation, RAG-based survey generation, and Human-guided survey generation).
Sample
N = 120, 8 groups
Primary method
One-way ANOVA tests were conducted for cross-disciplinary comparison with results at p-level significance. Kolmogorov-Smirnov (KS) tests performed for reference distribution comparisons. Semantic similarity computed using cosine distance on embedding vectors. Text similarity computed using threshold-based matching (0.95 for citations, 0.8 for structural overlap). Rank-based aggregation for multi-criteria reranking.
Main result
The study found that "QUAL-SG achieves the highest citation quality (F1 score of 16.73%), outperforming Naive-RAG and Fully-LLMGen by 10.80% and 8.97%, respectively." Additionally, "QUAL-SG also surpasses both baselines in content quality (Similarity +0.73%, ROUGE-L +1.40%, KPR +3.66%) and structural consistency (LLM evaluation +0.16 on a 5-point scale, semantic overlap +12.54%)."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/Empiricist
Author conclusions
"While LLMs have demonstrated efficiency and the ability to generate content considered useful by human evaluators, our human evaluation results reveal that, despite strong topical relevance, LLM-generated surveys exhibit limited information coverage and in-depth analysis, both essential for high-quality scientific surveys. Therefore, while LLMs can assist in survey generation, they are still unable to independently craft surveys that meet academic standards at the current stage."
Risk of bias
Data contamination: Some surveys and key references may have been included in LLM training data, potentially affecting performance estimates; Selection bias: Only 120 highly-cited surveys selected from 4,205 total surveys (29 surveys per domain), potentially biasing toward well-known topics; Evaluator bias: Human evaluators were PhD students in computer science; domain-specific expertise may vary across Biology, Medicine, Psychology, and Computer Science; Self-evaluation bias: LLM-as-judge used GPT-4o for automatic evaluation, which may introduce consistency biases; Data contamination: Possibility that surveys or key references used are open access and may have been included in training data of LLMs; Selection bias: Evaluation sample limited to 120 relatively short surveys from 4 domains; Evaluation bias: Use of same LLMs (GPT-4o) for both generation and evaluation poses potential self-evaluation bias; Annotation bias: Human evaluation conducted by only three second-year PhD students from computer science domain; Selection bias: Only 120 surveys selected from 4,205 available (sample may not be representative); Data contamination: Open-access surveys and key references may have been in LLM training data; Evaluator bias: Human evaluators were PhD students in computer science only; limited diversity of evaluators; Threshold selection bias: Citation matching uses 0.95 similarity threshold; structural overlap uses 0.8 threshold (chosen via preliminary experiments without rigorous validation)
Limitations
- "For copyright reasons, our approach is restricted to using only abstracts and bibliographic metadata of the retrieved papers, without access to full-text content
- This limitation may hinder the LLM's ability to capture finer-grained details and structural elements that are often present in full-length papers." Additionally, "To reduce API costs, we did not perform post-generation refinement, such as language polishing, citation formatting, structural adjustments, and PDF/Latex export." The authors also note: "While our empirical evaluation focuses on a subset of 120 relatively short surveys spanning multiple disciplines...we expect similar performance trends to hold across the full dataset."
Open questions raised
- Citation network analysis can be used to capture global relationships among papers and identify influential studies
- Analyzing human citation behavior (intent, frequency, location) can inform better reference selection mechanisms
- Training reference selection models on human-annotated datasets
- Leveraging full-text information instead of abstracts to enable more comprehensive contextual understanding
- Human-in-the-loop discourse control
- Factual consistency verification
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations