12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models

Tobias Schreieder, Tim Schopf, Michael Färber · ArXiv.org · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
I
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2508.15396

Methodology & findings

Study design

Systematic mapping study following guidelines by Petersen et al.

Main result

The survey identified 134 relevant publications on evidence-based text generation with LLMs. The analysis revealed that "86% of studies were published after 2023, highlighting limitations of earlier surveys like Li et al. (2023) that omits a substantial portion of the recent literature." Additionally, "research efforts in this area are highly fragmented, with studies typically aiming to cite (used in 75% of papers), attribute (62%), or quote (13%) evidence." The most common task is Question Answering (QA), accounting for "approximately 71% of all papers." The survey identified 300 distinct evaluation metrics, 17 frameworks, 231 datasets, and 11 benchmarks for evidence-based text generation.

Research paradigm

Systematic mapping study following established guidelines for literature reviews

Author conclusions

The authors conclude that "The field is still fragmented with varied terminology and evaluation standards. Our synthesis offers a clear reference point for advancing the reliability and verifiability of LLMs." They emphasize that "evidence-based text generation with LLMs is rapidly evolving, with 86% of papers emerging after 2023 and continued growth projected." They highlight the critical need for future work, stating: "future research should explore hybrid approaches that combine the strengths of both parametric and non-parametric attribution" and that "Future work should ensure these frameworks are adaptable to the diverse tasks within evidence-based text generation and comprehensively cover all evaluation dimensions."

Risk of bias

Publication bias: Inclusion limited to English-language publications with electronically accessible full texts, potentially excluding relevant non-English work; Database coverage bias: Literature search across specific databases may miss relevant studies in other venues; Temporal bias: Data cutoff in February 2025 limits 2025 publications to only 10 papers, potentially underrepresenting emerging work; Screening bias: While two independent reviewers screened publications, joint decision-making without reported inter-rater reliability metrics; Selection bias: Inclusion criteria focusing on explicit evidence references may miss related work on implicit attribution or grounding; Selection bias: Search limited to nine specific databases; unpublished work or work in non-English languages excluded; Publication bias: All included studies must be published with electronically accessible full texts; Screening bias: Despite dual independent review, potential for missed papers during title/abstract screening of 805 publications to 134 final papers; Language bias: Only English-language publications included; Temporal bias: Data cutoff February 2025 excludes recent publications; earlier surveys outdated; Database bias: Search limited to nine specific databases; grey literature potentially missed; Publication recency bias: Over 86% of papers published after 2023, skewing toward recent trends; Selection bias: Inclusion criteria required explicit evidence incorporation, potentially missing related work using implicit grounding; Two-author screening without inter-rater reliability statistics reported

Limitations

  • The authors state that "Prior literature on evidence-based text generation remains limited in scope" and note that earlier surveys "lack a systematic literature search, and are already outdated." They identify that "only two frameworks and benchmarks [are] in common use" despite identifying "over 300 evaluation metrics," indicating "the urgent need for standardized evaluation frameworks." Additionally, they note that "many studies introducing unique and non-standardized metrics" for human evaluation underscores "the need for automated and scalable evaluation approaches." The paper's data cutoff in February 2025 means "only 10 publications from 2025" were included, limiting coverage of the most recent work.

Open questions raised

  • Limited exploration of parametric attribution approaches (only 7 studies found): Need for methods that trace outputs to model training data rather than relying solely on external retrieval
  • Lack of standardized evaluation frameworks and benchmarks: Only 2 frameworks and benchmarks in common use despite 300+ metrics identified
  • Heterogeneity in human evaluation metrics: Many studies use custom, nonstandardized evaluation criteria limiting comparability
  • Unexplored citation behavior and biases of LLMs: Citation behavior remains largely underexplored, including potential citation-related biases similar to human authorship
  • Limited explainability of citation reasoning: Users unable to understand why LLMs select particular sources from multiple candidates
  • Need for automated and scalable evaluation approaches: Current reliance on human evaluation limits standardization and reusability
Data: Annotated dataset for this survey; Annotated dataset made available in public Git repository at https://github.com/faerber-lab/AttributeCiteQuote. Full list of 231 datasets provided in repository.; Annotated dataset in public Git repository at https://github.com/faerber-lab/AttributeCiteQuoteCode: AttributeCiteQuote; https://github.com/faerber-lab/AttributeCiteQuoteExtracted from: pdfAgreement 64%

Explore related topics

Related papers