12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models

Tong Bao, Mir Tafseer Nayeem, Davood Rafiei, Chengzhi Zhang · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
2
Citations
6.13
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18653/v1/2025.emnlp-main.136

Methodology & findings

Study design

Empirical evaluation study with three task-based experiments (Fully LLM-based survey generation, RAG-based survey generation, and Human-guided survey generation).

Sample

N = 120, 8 groups

Primary method

One-way ANOVA tests were conducted for cross-disciplinary comparison with results at p-level significance. Kolmogorov-Smirnov (KS) tests performed for reference distribution comparisons. Semantic similarity computed using cosine distance on embedding vectors. Text similarity computed using threshold-based matching (0.95 for citations, 0.8 for structural overlap). Rank-based aggregation for multi-criteria reranking.

Main result

The study found that "QUAL-SG achieves the highest citation quality (F1 score of 16.73%), outperforming Naive-RAG and Fully-LLMGen by 10.80% and 8.97%, respectively." Additionally, "QUAL-SG also surpasses both baselines in content quality (Similarity +0.73%, ROUGE-L +1.40%, KPR +3.66%) and structural consistency (LLM evaluation +0.16 on a 5-point scale, semantic overlap +12.54%)."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/Empiricist

Author conclusions

"While LLMs have demonstrated efficiency and the ability to generate content considered useful by human evaluators, our human evaluation results reveal that, despite strong topical relevance, LLM-generated surveys exhibit limited information coverage and in-depth analysis, both essential for high-quality scientific surveys. Therefore, while LLMs can assist in survey generation, they are still unable to independently craft surveys that meet academic standards at the current stage."

Risk of bias

Data contamination: Some surveys and key references may have been included in LLM training data, potentially affecting performance estimates; Selection bias: Only 120 highly-cited surveys selected from 4,205 total surveys (29 surveys per domain), potentially biasing toward well-known topics; Evaluator bias: Human evaluators were PhD students in computer science; domain-specific expertise may vary across Biology, Medicine, Psychology, and Computer Science; Self-evaluation bias: LLM-as-judge used GPT-4o for automatic evaluation, which may introduce consistency biases; Data contamination: Possibility that surveys or key references used are open access and may have been included in training data of LLMs; Selection bias: Evaluation sample limited to 120 relatively short surveys from 4 domains; Evaluation bias: Use of same LLMs (GPT-4o) for both generation and evaluation poses potential self-evaluation bias; Annotation bias: Human evaluation conducted by only three second-year PhD students from computer science domain; Selection bias: Only 120 surveys selected from 4,205 available (sample may not be representative); Data contamination: Open-access surveys and key references may have been in LLM training data; Evaluator bias: Human evaluators were PhD students in computer science only; limited diversity of evaluators; Threshold selection bias: Citation matching uses 0.95 similarity threshold; structural overlap uses 0.8 threshold (chosen via preliminary experiments without rigorous validation)

Limitations

  • "For copyright reasons, our approach is restricted to using only abstracts and bibliographic metadata of the retrieved papers, without access to full-text content
  • This limitation may hinder the LLM's ability to capture finer-grained details and structural elements that are often present in full-length papers." Additionally, "To reduce API costs, we did not perform post-generation refinement, such as language polishing, citation formatting, structural adjustments, and PDF/Latex export." The authors also note: "While our empirical evaluation focuses on a subset of 120 relatively short surveys spanning multiple disciplines...we expect similar performance trends to hold across the full dataset."

Open questions raised

  • Citation network analysis can be used to capture global relationships among papers and identify influential studies
  • Analyzing human citation behavior (intent, frequency, location) can inform better reference selection mechanisms
  • Training reference selection models on human-annotated datasets
  • Leveraging full-text information instead of abstracts to enable more comprehensive contextual understanding
  • Human-in-the-loop discourse control
  • Factual consistency verification
Data: SurveyGen; S2ORC (Semantic Scholar Open Research Corpus); SurveyGen dataset (https://github.com/[repository]) - comprising 4,205 human-written surveys with 242,143 cited references, section-level structures, and rich metadata from S2ORC and OpenAlex; S2ORCExtracted from: pdfAgreement 60%

Explore related topics

Related papers