12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Citation Hallucination Determines Success: An Empirical Comparison of Six Medical AI Research Systems

Xuefei Shi, Zhanxiao Tian · medRxiv · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.04.02.26350091

Methodology & findings

Study design

Comparative empirical evaluation study.

Sample

N = 7, 8 groups

Primary method

Descriptive statistics (means, ranges, standard deviations) across seven system configurations. Authors explicitly state: "Given the small number of systems -a constraint inherent to the field, as only a limited number of publicly available AI research systems exist -we report descriptive statistics (means, ranges, standard deviations) rather than formal hypothesis testing." No formal inferential statistics reported. For ablation comparison, paired differences reported across three tasks with relative improvement percentages. Programmatic citation verification used exact title matching with fuzzy fallback (Levenshtein distance ≤ 3). Regex-based extraction and exact matching with defined tolerances for numerical fidelity (OR/HR: ±0.01, P-values: exact, sample sizes: exact, percentages: ±0.5%). Analysis performed using Python 3.11.

Main result

Citation integrity emerged as the critical bottleneck in AI-generated medical research. The study found that "four of seven evaluated AI research systems produced manuscripts with unacceptably high citation hallucination rates, rendering them unreliable as scientific documents despite high scores in writing quality and structural completeness." AI Research Army's pipeline improved the weighted total score from 68.9 to 81.8 (+18.7%), with citation hallucination reducing from 7.2% to 2.9%. A complete ranking reversal occurred between single-model evaluation (v1) and the three-tier evaluation framework (v2), demonstrating that "a system that produces beautiful but dishonest manuscripts will score highly under subjective evaluation but fail under rigorous programmatic verification."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empirical - systematic evaluation of AI systems against objective criteria

Author conclusions

The authors conclude that "citation integrity is the critical bottleneck in AI-generated medical research manuscripts" and that "our six-dimension evaluation framework revealed that four of seven evaluated AI research systems produced manuscripts with unacceptably high citation hallucination rates, rendering them unreliable as scientific documents despite high scores in writing quality and structural completeness." They emphasize: "The difference between a 'beautiful paper' and a 'reliable paper' is not stylistic -it is the difference between contributing to scientific knowledge and polluting it." They recommend mandatory citation verification, multi-dimensional hybrid evaluation, and hard rules for scientific integrity in future evaluations.

Risk of bias

Self-evaluation bias: Authors developed both the evaluated system (AI Research Army) and the evaluation framework; Single time-point evaluation: Results may not generalize to other time periods or model versions; Limited system population: Only 6 systems evaluated; authors note 'only a limited number of publicly available AI research systems exist'; Citation verification coverage bias: CrossRef and PubMed APIs may under-detect legitimate references (false negatives); Subjective LLM judging: High inter-judge variability in clinical interpretation dimension (SD up to 21.0); Domain specificity: Evaluation limited to observational epidemiology; results may not generalize to RCTs or qualitative research; Hard-rule cap bias: Four systems triggered hard-rule threshold (D1 < 30), potentially masking performance improvements in other dimensions; Self-evaluation bias: Authors developed both the system (AI Research Army) and the evaluation framework; Selection bias in system comparison: Only six publicly available systems could be evaluated, not a random sample; LLM judge subjectivity in dimensions D5 and D6, with high inter-judge variability (SD up to 21.0); Citation verification false negatives: CrossRef and PubMed APIs may not index recent publications, preprints, or non-English literature; Temporal dependency: Evaluation conducted at single time point (March-April 2026); system performance may change with model updates; Evaluation methodology bias demonstrated: Single-model evaluation (v1) completely reversed rankings compared to three-tier evaluation (v2); Single timepoint evaluation: Conducted at one time point (March-April 2026), limiting generalizability across temporal changes; Database coverage bias: Citation verification relies on CrossRef and PubMed APIs which may not index all legitimate references; Citation verification coverage gaps: Missing recent publications, preprints, non-English literature, book chapters; LLM judge subjectivity: High inter-judge variability (SD up to 21.0) for clinical interpretation dimension; Model version dependency: Results tied to specific model versions evaluated at specific time

Limitations

  • The authors acknowledged six major limitations: (1) "Limited task scope
  • Our evaluation focused on observational epidemiology using NHANES data, comprising three clinical tasks across three domains
  • While these tasks represent diverse analytical approaches (logistic regression, mediation analysis, stratified analysis), they cover a narrow slice of medical research
  • Results may differ for randomized controlled trials, qualitative research, or non-NHANES datasets." (2) Citation verification coverage limited to CrossRef and PubMed APIs, which may not index all legitimate references, particularly recent publications or non-English literature
  • (3) Temporal dependency - single time point evaluation (March-April 2026)
  • (4) Numerical fidelity gap - the quality assurance pipeline "did not improve D2 (Numerical Fidelity), which remained at 53.3 for both pipeline and single-prompt versions." (5) LLM judge subjectivity with "high inter-judge variability (SD up to 21.0)" for clinical interpretation

Open questions raised

  • Need to extend MedResearchBench to additional study designs (randomized controlled trials, qualitative research) and non-NHANES datasets
  • Incorporating additional citation databases beyond CrossRef and PubMed (Semantic Scholar, Google Scholar API, preprint servers)
  • Longitudinal evaluation tracking system performance over time
  • Improving D2 (Numerical Fidelity) - authors identify this as an explicit gap: 'Adding a dedicated numerical accuracy checking agent that cross-references manuscript statistics against the results package'
  • Developing more granular, rubric-based clinical interpretation scoring to reduce inter-judge variability
  • Independent replication by third-party evaluators to address self-evaluation bias
Data: NHANES data (publicly available) - https://www.cdc.gov/nchs/nhanes/index.htm (implied, not explicitly linked); NHANES data (publicly available): Used for three clinical tasks - Cardio 000 (NHANES 2017-2018), Mental 000 (NHANES 2017-2020), Metabolic 002 (NHANES 2007-2012); NHANES data (publicly available): https://www.cdc.gov/nchs/nhanes/; Evaluation scripts and raw scores: https://github.com/TerryFYL/ai-research-armyCode: https://github.com/TerryFYL/ai-research-army (evaluation scripts, evaluation code, raw scores, and generated manuscripts); https://github.com/TerryFYL/ai-research-army (Evaluation scripts and evaluation output JSON files); GitHub repository: https://github.com/TerryFYL/ai-research-army (evaluation scripts and prompts)Extracted from: pdfAgreement 48%

Explore related topics

Related papers