12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

citecheck: An MCP Server for Automated Bibliographic Verification and Repair in Scholarly Manuscripts

Junhyeok Lee · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

System design and implementation with fixture-based evaluation and regression testing.

Primary method

Design science research with systems engineering approach; safety-oriented design treating bibliography repair as a safety problem

Main result

The paper presents citecheck, a system that "combines workspace-aware extraction, multi-source metadata retrieval, manifestation-aware matching, and policy-gated rewrite planning" to address the problem of reference inaccuracies in scholarly manuscripts. The key finding is that "citation hallucinations are now well documented: models can fabricate references, blend metadata from different papers, and emit plausible-looking but invalid identifiers," and the system addresses this through automated bibliographic verification.

Research paradigm

Design science / Systems engineering

Author conclusions

The authors conclude: "We presented citecheck, an MCP server for automated bibliographic verification and repair in scholarly manuscripts and paper-like project folders. The system combines workspace-aware extraction, multi-source metadata retrieval, manifestation-aware matching, and policy-gated rewrite planning. Its current repository already provides a functional prototype with structured outputs and automated tests, while its limitations remain explicit. We view citecheck as infrastructure for dependable scholarly editing systems and as a practical safeguard against both traditional reference errors and newer LLM-driven citation failures."

Open questions raised

  • Promising next steps include: expanding fixture-backed evaluations, building a realistic benchmark of malformed bibliographies, measuring citation-key rewrite accuracy on real LaTeX projects, and studying the tradeoff between review-only and replacement-safe operating modes in downstream author workflows.
  • "Promising next steps include expanding fixture-backed evaluations, building a realistic benchmark of malformed bibliographies, measuring citation-key rewrite accuracy on real LATEX projects, and studying the tradeoff between review-only and replacement-safe operating modes in downstream author workflows."
  • The authors identify the following future directions: "Promising next steps include expanding fixture-backed evaluations, building a realistic benchmark of malformed bibliographies, measuring citation-key rewrite accuracy on real LATEX projects, and studying the tradeoff between review-only and replacement-safe operating modes in downstream author workflows."
Code: Repository containing approximately 6.8k lines of TypeScript code with test suite (location not explicitly specified but referenced as publicly available); Repository available (exact URL not specified in paper; described as containing approximately 6.8k lines of TypeScript code with test fixtures); The authors mention "The current repository is a TypeScript workspace with approximately 6.8k lines across application and library code." Repository details are referenced but specific GitHub/GitLab URLs are not provided in the paper text.Extracted from: pdfAgreement 75%

Explore related topics

Related papers