Attribution in Scientific Literature: New Benchmark and Methods
Yash Saxena, Deepa Tilwani, Ali Mohammadi, Edward Raff, Sheth Amit, Srinivasan Parthasarathy et al. · arXiv (Cornell University) · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2405.02228
Methodology & findings
Study design
Computational benchmark evaluation using multiple large language models (LLMs) tested on the REASONS dataset across four prompting strategies: Direct Querying, Direct Querying with Metadata, Indirect Querying, and Sequential Indirect and Direct Prompting (SID Prompting).
Main result
The study found that "zero-shot direct prompting results reveal significant performance variations, with G1 achieving the lowest HRs (32.3%) and highest F-1 scores (0.40) across domains" and that "metadata provision creates performance convergence among proprietary models (G1-G3) while maintaining a significant gap with RAG models." Additionally, "indirect prompting results show substantially higher HRs across all models, with even G1 reaching 67.7% HR compared to 32.3% with direct prompting," demonstrating that "current models primarily succeed through information extraction rather than a deep understanding of scientific relationships."
Research paradigm
Positivist/empiricist - computational evaluation of LLM capabilities against a benchmark
Author conclusions
The authors conclude: "The REASONS benchmark forms the foundation for developing more trustworthy AI systems for scientific writing assistance, literature review, and knowledge synthesis that appropriately credit original sources. Standardized evaluation across different prompting strategies and domains enables researchers to identify specific attribution weaknesses that must be addressed before deploying AI assistants in high-stakes scientific contexts." They further note that "Future research need to focus on improving attribution through explicit reasoning mechanisms similar to the Toulmin model within retrieval-augmented frameworks. More sophisticated adversarial testing approaches including partial abstract modifications and misleading term insertion would provide deeper insights into model robustness."
Risk of bias
Dataset selection bias: exclusion of mathematics, statistics, and physics papers limits generalizability; Domain representation bias: QC (smallest domain with 53.0% HR for G1) versus CV (largest at 5,488 papers) shows domain-dependent performance; Model selection bias: evaluation emphasizes proprietary OpenAI and popular open-source models; Evaluation bias: zero-shot indirect prompting shows 67.7% HR for G1 but indirect queries may not reflect real-world attribution tasks; Adversarial test design bias: similarity threshold (0.70) for substitutions may not capture all failure modes; Domain representation bias: Smaller domains (QC with fewest papers) show highest hallucination rates, suggesting models may have less training data; Selection bias: Papers from arXiv only, limited to IEEE-formatted papers with specific licenses; Evaluation bias: Direct querying may favor models with strong memorization capabilities rather than reasoning; Model selection bias: Mix of proprietary and open-source models with varying sizes and architectures; Selection bias: Dataset restricted to IEEE-formatted arXiv papers with CC-compatible licenses, excluding mathematics, statistics, and physics; Domain representation bias: Uneven coverage across 12 scientific domains (CV: 5,488 papers vs QC: smaller representation); License-based selection bias: Exclusion of CC BY-NC-ND licensed papers; Model selection bias: Different models have different context window sizes and training cutoffs
Limitations
- The authors state that "Our study deliberately excluded mathematics, statistics, and physics papers due to equation prevalence in their related work sections, which the theoremKb crawling method couldn't effectively process." Additionally, they note that "these studies has two main limitations: it primarily focuses on general-purpose content rather than specialized domains, and it typically provides attribution at too high a granularity (with the exception of the recent SelfCite)." Furthermore, the paper acknowledges that "Domain representation (as shown in Figure 7) significantly impacts attribution accuracy, with QC (smallest representation) showing the highest HRs (53.0% for G1)" suggesting limitations in specialized domain coverage.
Open questions raised
- Lack of sentence-level attribution benchmarks for scientific domains
- Need for domain-specific training in LLMs to meet specialized field requirements
- Integration of knowledge graph representations and graph-theoretic retrieval approaches for more reliable source attribution
- Evaluation of LLMs on mathematics, statistics, and physics papers (excluded from current study)
- Development of explicit reasoning mechanisms similar to the Toulmin model within retrieval-augmented frameworks
- More sophisticated adversarial testing approaches including partial abstract modifications and misleading term insertion
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations