12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Unmediated AI-Assisted Scholarly Citations

Stefan H. Szeider · Open Conference Proceedings · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.52825/ocp.v8i.3161

Methodology & findings

Study design

Comparative evaluation across three independent experiments using 104 obfuscated academic citations sampled from DBLP with stratified sampling (50% from 2020-2025, 25% from 2015-2019, 25% from 2010-2014).

Primary method

Design science research with implementation-based evaluation

Main result

The evaluation demonstrates that MCP-DBLP achieves significantly higher citation accuracy compared to web search alone. Specifically, "Evaluation on 104 obfuscated citations (averaged across three experiments) shows 82.7% perfect match rate for MCP-DBLP with unmediated export versus 28.2% for standard web search, with zero cases of metadata corruption." The unmediated export approach eliminates all corrupted metadata errors (0% CM rate), compared to 6.7% for web search baseline.

Research paradigm

Design science and engineering

Author conclusions

The authors conclude that "an architectural approach for connecting language models with bibliographic databases that combines conversational interaction with verified accuracy" is achievable through separation of concerns. They state that "MCP-DBLP implements this approach through the Model Context Protocol, providing conversational access to DBLP with unmediated BibTeX export. Evaluation across three independent experiments with 104 obfuscated citations each shows 82.7% PM for MCP-U versus 28.2% for Web, a 2.9× improvement, with zero metadata corruption." They emphasize that "the MCP architecture provides standardization, statefulness, and composability, making the approach broadly applicable across bibliographic databases and research disciplines."

Risk of bias

Selection bias: citations sampled from DBLP only, may not represent other databases; Evaluation bias: non-interactive mode may not reflect real-world usage with user clarification; Ambiguity bias: 15-19% Wrong Paper rate indicates citation ambiguity rather than method failure; Selection bias in citation sampling (only DBLP-indexed papers, may not represent all research domains); Potential lack of representativeness for non-computer science fields; Non-interactive evaluation mode may not reflect actual usage patterns where users disambiguate; Limited diversity in citation obfuscation difficulty levels; Selection bias: Citations were sampled from DBLP only, potentially favoring computer science papers; Confounding: Different obfuscation difficulty levels may introduce variance; Agent variability: Single model configuration (Claude Sonnet 4.5) may not generalize across different LLMs

Limitations

  • The paper identifies several limitations in the evaluation context: "Our experiments used non-interactive mode, where the agent processed citations autonomously without user feedback
  • In real-world applications, users would typically instruct the language model to ask for clarification when facing ambiguous references." Additionally, errors related to citation ambiguity (WP) remained at 15-19% across all methods, indicating limitations tied to input clarity rather than retrieval method
  • The remaining NF errors occur when papers lack DBLP indexing or input citations are too vague.

Open questions raised

  • The paper identifies that language models remain unreliable producers of bibliographic data despite multiple existing approaches. It notes the gap between current citation quality in long-form QA systems and publication standards, and the lack of MCP servers specifically designed for unmediated bibliographic export with metadata integrity guarantees.
  • Extension to other bibliographic databases (PubMed, arXiv, Semantic Scholar, institutional repositories)
  • Integration with multi-agent research workflows combining literature analysis, writing assistance, and fact-checking
  • Interactive disambiguation workflows to reduce Wrong Paper and Not Found rates
  • Support for non-computer science domains and specialized databases
  • The paper notes potential for interactive disambiguation to reduce Wrong Paper (WP) and Not Found (NF) rates when users can provide clarification on ambiguous references
Data: 104 obfuscated academic citations sampled from DBLP (ground truth: 50% from 2020-2025, 25% from 2015-2019, 25% from 2010-2014); MCP-DBLP source code: https://github.com/szeider/mcp-dblp; MCP-DBLP package on PyPI: https://pypi.org/project/mcp-dblp/; DBLP Computer Science BibliographyCode: https://github.com/szeider/mcp-dblp; https://pypi.org/project/mcp-dblp/; MCP-DBLPExtracted from: pdfAgreement 58%

Explore related topics

Related papers