12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Attribution in Scientific Literature: New Benchmark and Methods

Yash Saxena, Deepa Tilwani, Ali Mohammadi, Edward Raff, Sheth Amit, Srinivasan Parthasarathy et al. · arXiv (Cornell University) · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
2
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2405.02228

Methodology & findings

Study design

Computational benchmark evaluation using multiple large language models (LLMs) tested on the REASONS dataset across four prompting strategies: Direct Querying, Direct Querying with Metadata, Indirect Querying, and Sequential Indirect and Direct Prompting (SID Prompting).

Main result

The study found that "zero-shot direct prompting results reveal significant performance variations, with G1 achieving the lowest HRs (32.3%) and highest F-1 scores (0.40) across domains" and that "metadata provision creates performance convergence among proprietary models (G1-G3) while maintaining a significant gap with RAG models." Additionally, "indirect prompting results show substantially higher HRs across all models, with even G1 reaching 67.7% HR compared to 32.3% with direct prompting," demonstrating that "current models primarily succeed through information extraction rather than a deep understanding of scientific relationships."

Research paradigm

Positivist/empiricist - computational evaluation of LLM capabilities against a benchmark

Author conclusions

The authors conclude: "The REASONS benchmark forms the foundation for developing more trustworthy AI systems for scientific writing assistance, literature review, and knowledge synthesis that appropriately credit original sources. Standardized evaluation across different prompting strategies and domains enables researchers to identify specific attribution weaknesses that must be addressed before deploying AI assistants in high-stakes scientific contexts." They further note that "Future research need to focus on improving attribution through explicit reasoning mechanisms similar to the Toulmin model within retrieval-augmented frameworks. More sophisticated adversarial testing approaches including partial abstract modifications and misleading term insertion would provide deeper insights into model robustness."

Risk of bias

Dataset selection bias: exclusion of mathematics, statistics, and physics papers limits generalizability; Domain representation bias: QC (smallest domain with 53.0% HR for G1) versus CV (largest at 5,488 papers) shows domain-dependent performance; Model selection bias: evaluation emphasizes proprietary OpenAI and popular open-source models; Evaluation bias: zero-shot indirect prompting shows 67.7% HR for G1 but indirect queries may not reflect real-world attribution tasks; Adversarial test design bias: similarity threshold (0.70) for substitutions may not capture all failure modes; Domain representation bias: Smaller domains (QC with fewest papers) show highest hallucination rates, suggesting models may have less training data; Selection bias: Papers from arXiv only, limited to IEEE-formatted papers with specific licenses; Evaluation bias: Direct querying may favor models with strong memorization capabilities rather than reasoning; Model selection bias: Mix of proprietary and open-source models with varying sizes and architectures; Selection bias: Dataset restricted to IEEE-formatted arXiv papers with CC-compatible licenses, excluding mathematics, statistics, and physics; Domain representation bias: Uneven coverage across 12 scientific domains (CV: 5,488 papers vs QC: smaller representation); License-based selection bias: Exclusion of CC BY-NC-ND licensed papers; Model selection bias: Different models have different context window sizes and training cutoffs

Limitations

  • The authors state that "Our study deliberately excluded mathematics, statistics, and physics papers due to equation prevalence in their related work sections, which the theoremKb crawling method couldn't effectively process." Additionally, they note that "these studies has two main limitations: it primarily focuses on general-purpose content rather than specialized domains, and it typically provides attribution at too high a granularity (with the exception of the recent SelfCite)." Furthermore, the paper acknowledges that "Domain representation (as shown in Figure 7) significantly impacts attribution accuracy, with QC (smallest representation) showing the highest HRs (53.0% for G1)" suggesting limitations in specialized domain coverage.

Open questions raised

  • Lack of sentence-level attribution benchmarks for scientific domains
  • Need for domain-specific training in LLMs to meet specialized field requirements
  • Integration of knowledge graph representations and graph-theoretic retrieval approaches for more reliable source attribution
  • Evaluation of LLMs on mathematics, statistics, and physics papers (excluded from current study)
  • Development of explicit reasoning mechanisms similar to the Toulmin model within retrieval-augmented frameworks
  • More sophisticated adversarial testing approaches including partial abstract modifications and misleading term insertion
Data: REASONS dataset: https://github.com/YashSaxena21/REASONS (mentioned as "available in the GitHub repository"); REASONS dataset; REASONS benchmark dataset (GitHub: https://github.com/YashSaxena21/REASONS); Dataset comprises sentences from 12 scientific domains on arXiv (2017-2024)Code: https://github.com/YashSaxena21/REASONS (complete dataset and all experimental code); REASONS GitHub repository; https://github.com/YashSaxena21/REASONSExtracted from: pdfAgreement 47%

Explore related topics

Related papers