SurveyGen-I: Consistent Scientific Survey Generation with Evolving Plans and Memory-Guided Writing
Jing Chen, Zhiheng Yang, Yixian Shen, Jie Liu, Adam Belloum, Chrysa Papagainni et al. · ArXiv.org · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2508.14317
Methodology & findings
Study design
Comparative system evaluation using LLM-as-Judge scoring across five content quality dimensions (coverage, relevance, structure, synthesis, consistency) and three reference quality metrics (number of references, citation density, recency ratio).
Primary method
Design science research with iterative modular architecture (component-based systems engineering)
Main result
SurveyGen-I yields "an 8.5% improvement in content quality, a 27% increase in citation density, and more than twice as many distinct references, while also demonstrating significantly better citation recency" compared to the strongest baseline. The system achieves "281 unique works per survey on average" with "citation density also rises substantially (17.28), exceeding SurveyForge (5.52) by around 3 times" and demonstrates that "89.1% of all citations are published within the past 5 years (RR@5), compared to 66.7% in SurveyX and SurveyForge."
Research paradigm
Design science / Systems engineering
Author conclusions
The authors conclude: "We present SurveyGen-I, a fully automated framework for generating academic surveys with high consistency, citation coverage, and structural coherence. By integrating multi-level retrieval, adaptive planning, and memory-guided writing, SurveyGen-I effectively captures complex literature landscapes and produces high-quality surveys without manual intervention. Extensive evaluations across six scientific domains demonstrate its effectiveness over existing methods, marking a step forward in reliable and scalable scientific synthesis."
Risk of bias
Reliance on third-party APIs and online retrieval may introduce coverage bias; LLM-as-Judge evaluation may reflect subjective preferences rather than universal standards; Topic selection bias: evaluation limited to six major scientific domains; Potential licensing bias in full-text access restrictions; LLM-as-Judge evaluation bias: GPT-4o-mini used for all content quality scoring may introduce systematic bias toward outputs similar to GPT-4o-mini's own generation style; Evaluation metric selection bias: Heavy weighting on citation density and recency may not reflect all aspects of survey quality; Comparison fairness: AutoSurvey baseline regenerated by authors rather than using official outputs; Dataset bias: Benchmark constructed by authors may not be representative of all survey topics; Thresholds not justified: Cosine similarity threshold (0.3) and LLM relevance score threshold (≥70) chosen empirically without validation; LLM-as-Judge evaluation bias: subjective preferences of the rating model (GPT4o-mini) may not align with universal writing standards; Comparison methodology bias: AutoSurvey baseline required reimplementation with different model version than original; Citation availability bias: system performance constrained by Semantic Scholar API coverage and licensing access
Limitations
- The authors state: "our framework adopts an online retrieval strategy to ensure access to up-to-date literature
- However, this design introduces network sensitivity, variable latency, and reliance on third-party APIs, which may restrict full-text access due to licensing constraints." Additionally, "for niche or emerging topics with limited source material, the achievable survey length and depth are naturally constrained
- This shows a general challenge in automatic survey generation: content quality is ultimately bounded by the availability and granularity of the source literature." They also note that "some evaluation signals may reflect subjective preferences rather than universal writing standards."
Open questions raised
- Citation expansion and tracing of indirect references in RAG-based systems
- Cross-subsection consistency in parallel section generation
- Dynamic outline adaptation based on intermediate content generation
- Comprehensive literature coverage beyond surface-level embedding similarity
- Need for broader community adoption and feedback for future enhancements
- Current LLM-based survey generation frameworks remain limited in literature retrieval scope and depth, relying on embedding-based similarity search over fixed local databases that often fails to identify important papers with different terminology
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations