Citation Hallucination Determines Success: An Empirical Comparison of Six Medical AI Research Systems
Xuefei Shi, Zhanxiao Tian · medRxiv · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.04.02.26350091
Methodology & findings
Study design
Comparative empirical evaluation study.
Sample
N = 7, 8 groups
Primary method
Descriptive statistics (means, ranges, standard deviations) across seven system configurations. Authors explicitly state: "Given the small number of systems -a constraint inherent to the field, as only a limited number of publicly available AI research systems exist -we report descriptive statistics (means, ranges, standard deviations) rather than formal hypothesis testing." No formal inferential statistics reported. For ablation comparison, paired differences reported across three tasks with relative improvement percentages. Programmatic citation verification used exact title matching with fuzzy fallback (Levenshtein distance ≤ 3). Regex-based extraction and exact matching with defined tolerances for numerical fidelity (OR/HR: ±0.01, P-values: exact, sample sizes: exact, percentages: ±0.5%). Analysis performed using Python 3.11.
Main result
Citation integrity emerged as the critical bottleneck in AI-generated medical research. The study found that "four of seven evaluated AI research systems produced manuscripts with unacceptably high citation hallucination rates, rendering them unreliable as scientific documents despite high scores in writing quality and structural completeness." AI Research Army's pipeline improved the weighted total score from 68.9 to 81.8 (+18.7%), with citation hallucination reducing from 7.2% to 2.9%. A complete ranking reversal occurred between single-model evaluation (v1) and the three-tier evaluation framework (v2), demonstrating that "a system that produces beautiful but dishonest manuscripts will score highly under subjective evaluation but fail under rigorous programmatic verification."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empirical - systematic evaluation of AI systems against objective criteria
Author conclusions
The authors conclude that "citation integrity is the critical bottleneck in AI-generated medical research manuscripts" and that "our six-dimension evaluation framework revealed that four of seven evaluated AI research systems produced manuscripts with unacceptably high citation hallucination rates, rendering them unreliable as scientific documents despite high scores in writing quality and structural completeness." They emphasize: "The difference between a 'beautiful paper' and a 'reliable paper' is not stylistic -it is the difference between contributing to scientific knowledge and polluting it." They recommend mandatory citation verification, multi-dimensional hybrid evaluation, and hard rules for scientific integrity in future evaluations.
Risk of bias
Self-evaluation bias: Authors developed both the evaluated system (AI Research Army) and the evaluation framework; Single time-point evaluation: Results may not generalize to other time periods or model versions; Limited system population: Only 6 systems evaluated; authors note 'only a limited number of publicly available AI research systems exist'; Citation verification coverage bias: CrossRef and PubMed APIs may under-detect legitimate references (false negatives); Subjective LLM judging: High inter-judge variability in clinical interpretation dimension (SD up to 21.0); Domain specificity: Evaluation limited to observational epidemiology; results may not generalize to RCTs or qualitative research; Hard-rule cap bias: Four systems triggered hard-rule threshold (D1 < 30), potentially masking performance improvements in other dimensions; Self-evaluation bias: Authors developed both the system (AI Research Army) and the evaluation framework; Selection bias in system comparison: Only six publicly available systems could be evaluated, not a random sample; LLM judge subjectivity in dimensions D5 and D6, with high inter-judge variability (SD up to 21.0); Citation verification false negatives: CrossRef and PubMed APIs may not index recent publications, preprints, or non-English literature; Temporal dependency: Evaluation conducted at single time point (March-April 2026); system performance may change with model updates; Evaluation methodology bias demonstrated: Single-model evaluation (v1) completely reversed rankings compared to three-tier evaluation (v2); Single timepoint evaluation: Conducted at one time point (March-April 2026), limiting generalizability across temporal changes; Database coverage bias: Citation verification relies on CrossRef and PubMed APIs which may not index all legitimate references; Citation verification coverage gaps: Missing recent publications, preprints, non-English literature, book chapters; LLM judge subjectivity: High inter-judge variability (SD up to 21.0) for clinical interpretation dimension; Model version dependency: Results tied to specific model versions evaluated at specific time
Limitations
- The authors acknowledged six major limitations: (1) "Limited task scope
- Our evaluation focused on observational epidemiology using NHANES data, comprising three clinical tasks across three domains
- While these tasks represent diverse analytical approaches (logistic regression, mediation analysis, stratified analysis), they cover a narrow slice of medical research
- Results may differ for randomized controlled trials, qualitative research, or non-NHANES datasets." (2) Citation verification coverage limited to CrossRef and PubMed APIs, which may not index all legitimate references, particularly recent publications or non-English literature
- (3) Temporal dependency - single time point evaluation (March-April 2026)
- (4) Numerical fidelity gap - the quality assurance pipeline "did not improve D2 (Numerical Fidelity), which remained at 53.3 for both pipeline and single-prompt versions." (5) LLM judge subjectivity with "high inter-judge variability (SD up to 21.0)" for clinical interpretation
Open questions raised
- Need to extend MedResearchBench to additional study designs (randomized controlled trials, qualitative research) and non-NHANES datasets
- Incorporating additional citation databases beyond CrossRef and PubMed (Semantic Scholar, Google Scholar API, preprint servers)
- Longitudinal evaluation tracking system performance over time
- Improving D2 (Numerical Fidelity) - authors identify this as an explicit gap: 'Adding a dedicated numerical accuracy checking agent that cross-references manuscript statistics against the results package'
- Developing more granular, rubric-based clinical interpretation scoring to reduce inter-judge variability
- Independent replication by third-party evaluators to address self-evaluation bias
Explore related topics
Related papers
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations