12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

MVSS: A Unified Framework for Multi-View Structured Survey Generation

Yinqi Liu, Yueqi Zhu, Yongkang Zhang, Xinfeng Li, Yufei Sun, Xin Wang et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale experimental evaluation on 76 computer science topics with multiple survey lengths (8k, 16k, 32k, 64k tokens).

Primary method

Design science research following structure-first paradigm with iterative refinement; multi-view structural alignment optimization; LLM-based generation and evaluation

Main result

MVSS consistently achieves the highest overall survey quality across all target lengths. As stated in the results: "At 16k tokens, MVSS attains an average score of 4.90, exceeding AutoSurvey (4.60) and HiReview (3.65). This performance gap further widens at longer lengths, indicating that MVSS frequently approaches, and in some cases matches, expert-level organization under direct human comparison." Furthermore, "at 64k tokens, MVSS reaches 5.00 on Coverage, Structure, and Relevance, matching human-written surveys under the same evaluation protocol."

Research paradigm

Design science / computational systems design

Author conclusions

The authors conclude: "We presented MVSS, a unified framework for multi-view structured survey generation that elevates conceptual structure from a secondary byproduct to a first-class optimization objective. By jointly constructing citation-grounded hierarchical knowledge trees, schema-driven comparison tables, and evidence-aware narrative text, MVSS enforces structural coherence and semantic alignment across survey views." They further state: "Beyond its empirical gains, MVSS reframes automated survey generation as a structure-centric synthesis problem, highlighting the role of explicit hierarchies and comparisons in literature understanding. We believe MVSS represents a step toward scalable systems that go beyond summarization to actively organize scientific knowledge."

Risk of bias

LLM-judge bias: Heavy reliance on frontier LLMs (deepseek-chat, gpt-4o, gemini-2.5-pro) for both generation and evaluation introduces potential model-specific biases; Domain representation bias: Evaluation limited to computer science topics from arXiv corpus; domains underrepresented in LLM pretraining may perform poorly; Corpus bias: 530,000 arXiv papers (2018-2024) may not represent all research or research traditions equally; Selection bias in reference surveys: 76 topics selected from Google Scholar with balancing for citation counts and coverage, but selection criteria not fully transparent; Evaluator bias: Double-blind human evaluation on only 30 sampled topics; modest sample size for human preference assessment; Hyperparameter tuning bias: Additional hyperparameters introduced by structural and alignment objectives not fully characterized across domains; LLM judge bias and calibration concerns; Limited domain representation (CS only, ArXiv corpus); Potential selection bias in 76 surveyed topics; Underrepresentation in LLM pretraining data for certain domains; LLM bias in generation and judgment; Underrepresentation of non-English domains in pretraining data; Selection bias in 76 computer science topics from arXiv; Potential bias in expert-written survey selection for reference; LLM judge calibration may not fully align with human evaluations across all dimensions

Limitations

  • The system still depends on frontier LLMs for both generation and judgment, which raises cost, reproducibility, and bias concerns, especially when extending to domains underrepresented in pretraining data
  • The authors also note: "Our structural and alignment objectives introduce additional hyperparameters whose robustness across domains and retrieval settings has not been fully characterized
  • Moreover, our evaluation focuses on 76 CS topics using an arXiv-based corpus, limiting generalizability to other disciplines, formats, or argumentative norms
  • Finally, MVSS models a static snapshot of a field and does not capture temporal evolution or uncertainty in conflicting evidence."

Open questions raised

  • Need for methods that extend to domains underrepresented in LLM pretraining data
  • Robustness of structural and alignment hyperparameters across diverse domains and retrieval settings
  • Generalizability beyond computer science to other academic disciplines with different argumentative norms and formats
  • Temporal evolution modeling: MVSS currently models static snapshots and does not capture field evolution over time
  • Uncertainty representation: No mechanism to capture conflicting evidence or uncertainty in taxonomic structures
  • Time-aware and uncertainty-aware structures for future work
Data: arXiv paper corpus: 530,000 papers (2018-2024); 76 computer science topics with reference surveys from Google Scholar; 30 topics used for human evaluation (subset of 76); 530,000 arXiv papers (2018-2024); 76 computer science survey topics with reference surveys; 530,000 arXiv papers (2018-2024) used as retrieval corpus; 76 computer science survey topics with reference expert-written surveys from Google ScholarExtracted from: pdfAgreement 62%

Explore related topics

Related papers