12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

SurveyGen: Quality-Aware Scientific Survey Generation with Large Language Models

Tong Bao, Mir Tafseer Nayeem, Davood Rafiei, Chengzhi Zhang · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
E
Evidence
2
Citations
6.13
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18653/v1/2025.emnlp-main.136

Methodology & findings

Study design

Empirical evaluation study with three task-based experiments (Fully LLM-based survey generation, RAG-based survey generation, and Human-guided survey generation).

Sample

N = 120, 4 groups

Primary method

One-way ANOVA tests were conducted for cross-disciplinary comparison with results at p-level significance. Kolmogorov-Smirnov (KS) tests performed for reference distribution comparisons. Semantic similarity computed using cosine distance on embedding vectors. Text similarity computed using threshold-based matching (0.95 for citations, 0.8 for structural overlap). Rank-based aggregation for multi-criteria reranking.

Main result

The study found that "QUAL-SG achieves the highest citation quality (F1 score of 16.73%), outperforming Naive-RAG and Fully-LLMGen by 10.80% and 8.97%, respectively." Additionally, "QUAL-SG also surpasses both baselines in content quality (Similarity +0.73%, ROUGE-L +1.40%, KPR +3.66%) and structural consistency (LLM evaluation +0.16 on a 5-point scale, semantic overlap +12.54%)."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/Empiricist

Author conclusions

"While LLMs have demonstrated efficiency and the ability to generate content considered useful by human evaluators, our human evaluation results reveal that, despite strong topical relevance, LLM-generated surveys exhibit limited information coverage and in-depth analysis, both essential for high-quality scientific surveys. Therefore, while LLMs can assist in survey generation, they are still unable to independently craft surveys that meet academic standards at the current stage."

Risk of bias

Data contamination: Some surveys and key references may have been included in LLM training data, potentially affecting performance estimates; Selection bias: Only 120 highly-cited surveys selected from 4,205 total surveys (29 surveys per domain), potentially biasing toward well-known topics; Evaluator bias: Human evaluators were PhD students in computer science; domain-specific expertise may vary across Biology, Medicine, Psychology, and Computer Science; Self-evaluation bias: LLM-as-judge used GPT-4o for automatic evaluation, which may introduce consistency biases; Selection bias: Evaluation sample limited to 120 relatively short surveys from 4 domains; Evaluation bias: Use of same LLMs (GPT-4o) for both generation and evaluation poses potential self-evaluation bias; Threshold selection bias: Citation matching uses 0.95 similarity threshold; structural overlap uses 0.8 threshold (chosen via preliminary experiments without rigorous validation)

Limitations

  • "For copyright reasons, our approach is restricted to using only abstracts and bibliographic metadata of the retrieved papers, without access to full-text content
  • This limitation may hinder the LLM's ability to capture finer-grained details and structural elements that are often present in full-length papers." Additionally, "To reduce API costs, we did not perform post-generation refinement, such as language polishing, citation formatting, structural adjustments, and PDF/Latex export." The authors also note: "While our empirical evaluation focuses on a subset of 120 relatively short surveys spanning multiple disciplines...we expect similar performance trends to hold across the full dataset."

Open questions raised

  • Citation network analysis can be used to capture global relationships among papers and identify influential studies
  • Analyzing human citation behavior (intent, frequency, location) can inform better reference selection mechanisms
  • Training reference selection models on human-annotated datasets
  • Leveraging full-text information instead of abstracts to enable more comprehensive contextual understanding
  • Human-in-the-loop discourse control
  • Factual consistency verification
Data: SurveyGen; S2ORC (Semantic Scholar Open Research Corpus)Extracted from: pdf

Explore related topics

Related papers