12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

SurveyGen-I: Consistent Scientific Survey Generation with Evolving Plans and Memory-Guided Writing

Jing Chen, Zhiheng Yang, Yixian Shen, Jie Liu, Adam Belloum, Chrysa Papagainni et al. · ArXiv.org · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
D
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2508.14317

Methodology & findings

Study design

Comparative system evaluation using LLM-as-Judge scoring across five content quality dimensions (coverage, relevance, structure, synthesis, consistency) and three reference quality metrics (number of references, citation density, recency ratio).

Primary method

Design science research with iterative modular architecture (component-based systems engineering)

Main result

SurveyGen-I yields "an 8.5% improvement in content quality, a 27% increase in citation density, and more than twice as many distinct references, while also demonstrating significantly better citation recency" compared to the strongest baseline. The system achieves "281 unique works per survey on average" with "citation density also rises substantially (17.28), exceeding SurveyForge (5.52) by around 3 times" and demonstrates that "89.1% of all citations are published within the past 5 years (RR@5), compared to 66.7% in SurveyX and SurveyForge."

Research paradigm

Design science / Systems engineering

Author conclusions

The authors conclude: "We present SurveyGen-I, a fully automated framework for generating academic surveys with high consistency, citation coverage, and structural coherence. By integrating multi-level retrieval, adaptive planning, and memory-guided writing, SurveyGen-I effectively captures complex literature landscapes and produces high-quality surveys without manual intervention. Extensive evaluations across six scientific domains demonstrate its effectiveness over existing methods, marking a step forward in reliable and scalable scientific synthesis."

Risk of bias

Reliance on third-party APIs and online retrieval may introduce coverage bias; LLM-as-Judge evaluation may reflect subjective preferences rather than universal standards; Topic selection bias: evaluation limited to six major scientific domains; Potential licensing bias in full-text access restrictions; LLM-as-Judge evaluation bias: GPT-4o-mini used for all content quality scoring may introduce systematic bias toward outputs similar to GPT-4o-mini's own generation style; Evaluation metric selection bias: Heavy weighting on citation density and recency may not reflect all aspects of survey quality; Comparison fairness: AutoSurvey baseline regenerated by authors rather than using official outputs; Dataset bias: Benchmark constructed by authors may not be representative of all survey topics; Thresholds not justified: Cosine similarity threshold (0.3) and LLM relevance score threshold (≥70) chosen empirically without validation; LLM-as-Judge evaluation bias: subjective preferences of the rating model (GPT4o-mini) may not align with universal writing standards; Comparison methodology bias: AutoSurvey baseline required reimplementation with different model version than original; Citation availability bias: system performance constrained by Semantic Scholar API coverage and licensing access

Limitations

  • The authors state: "our framework adopts an online retrieval strategy to ensure access to up-to-date literature
  • However, this design introduces network sensitivity, variable latency, and reliance on third-party APIs, which may restrict full-text access due to licensing constraints." Additionally, "for niche or emerging topics with limited source material, the achievable survey length and depth are naturally constrained
  • This shows a general challenge in automatic survey generation: content quality is ultimately bounded by the availability and granularity of the source literature." They also note that "some evaluation signals may reflect subjective preferences rather than universal writing standards."

Open questions raised

  • Citation expansion and tracing of indirect references in RAG-based systems
  • Cross-subsection consistency in parallel section generation
  • Dynamic outline adaptation based on intermediate content generation
  • Comprehensive literature coverage beyond surface-level embedding similarity
  • Need for broader community adoption and feedback for future enhancements
  • Current LLM-based survey generation frameworks remain limited in literature retrieval scope and depth, relying on embedding-based similarity search over fixed local databases that often fails to identify important papers with different terminology
Extracted from: pdfAgreement 70%

Explore related topics

Related papers