Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models
Tobias Schreieder, Tim Schopf, Michael Färber · ArXiv.org · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2508.15396
Methodology & findings
Study design
Systematic mapping study following guidelines by Petersen et al.
Main result
The survey identified 134 relevant publications on evidence-based text generation with LLMs. The analysis revealed that "86% of studies were published after 2023, highlighting limitations of earlier surveys like Li et al. (2023) that omits a substantial portion of the recent literature." Additionally, "research efforts in this area are highly fragmented, with studies typically aiming to cite (used in 75% of papers), attribute (62%), or quote (13%) evidence." The most common task is Question Answering (QA), accounting for "approximately 71% of all papers." The survey identified 300 distinct evaluation metrics, 17 frameworks, 231 datasets, and 11 benchmarks for evidence-based text generation.
Research paradigm
Systematic mapping study following established guidelines for literature reviews
Author conclusions
The authors conclude that "The field is still fragmented with varied terminology and evaluation standards. Our synthesis offers a clear reference point for advancing the reliability and verifiability of LLMs." They emphasize that "evidence-based text generation with LLMs is rapidly evolving, with 86% of papers emerging after 2023 and continued growth projected." They highlight the critical need for future work, stating: "future research should explore hybrid approaches that combine the strengths of both parametric and non-parametric attribution" and that "Future work should ensure these frameworks are adaptable to the diverse tasks within evidence-based text generation and comprehensively cover all evaluation dimensions."
Risk of bias
Publication bias: Inclusion limited to English-language publications with electronically accessible full texts, potentially excluding relevant non-English work; Database coverage bias: Literature search across specific databases may miss relevant studies in other venues; Temporal bias: Data cutoff in February 2025 limits 2025 publications to only 10 papers, potentially underrepresenting emerging work; Screening bias: While two independent reviewers screened publications, joint decision-making without reported inter-rater reliability metrics; Selection bias: Inclusion criteria focusing on explicit evidence references may miss related work on implicit attribution or grounding; Selection bias: Search limited to nine specific databases; unpublished work or work in non-English languages excluded; Publication bias: All included studies must be published with electronically accessible full texts; Screening bias: Despite dual independent review, potential for missed papers during title/abstract screening of 805 publications to 134 final papers; Language bias: Only English-language publications included; Temporal bias: Data cutoff February 2025 excludes recent publications; earlier surveys outdated; Database bias: Search limited to nine specific databases; grey literature potentially missed; Publication recency bias: Over 86% of papers published after 2023, skewing toward recent trends; Selection bias: Inclusion criteria required explicit evidence incorporation, potentially missing related work using implicit grounding; Two-author screening without inter-rater reliability statistics reported
Limitations
- The authors state that "Prior literature on evidence-based text generation remains limited in scope" and note that earlier surveys "lack a systematic literature search, and are already outdated." They identify that "only two frameworks and benchmarks [are] in common use" despite identifying "over 300 evaluation metrics," indicating "the urgent need for standardized evaluation frameworks." Additionally, they note that "many studies introducing unique and non-standardized metrics" for human evaluation underscores "the need for automated and scalable evaluation approaches." The paper's data cutoff in February 2025 means "only 10 publications from 2025" were included, limiting coverage of the most recent work.
Open questions raised
- Limited exploration of parametric attribution approaches (only 7 studies found): Need for methods that trace outputs to model training data rather than relying solely on external retrieval
- Lack of standardized evaluation frameworks and benchmarks: Only 2 frameworks and benchmarks in common use despite 300+ metrics identified
- Heterogeneity in human evaluation metrics: Many studies use custom, nonstandardized evaluation criteria limiting comparability
- Unexplored citation behavior and biases of LLMs: Citation behavior remains largely underexplored, including potential citation-related biases similar to human authorship
- Limited explainability of citation reasoning: Users unable to understand why LLMs select particular sources from multiple candidates
- Need for automated and scalable evaluation approaches: Current reliance on human evaluation limits standardization and reusability
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations