12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Comparing the Accuracy of 4 Artificial Intelligence Models in PubMed Citation Generation for Glaucoma Research

Mustafa Civelekler, Mehmet Citirik · Journal of Glaucoma · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1097/ijg.0000000000002716

Methodology & findings

Study design

Comparative accuracy assessment study using standardized test cases.

Sample

N = 35, 6 groups

Primary method

Descriptive statistics reporting accuracy percentages for each model. Expert review classification system used for validation (categorical: 'Fully Cited,' 'Partially Cited,' 'Not Cited'). No inferential statistical tests or confidence intervals reported.

Main result

The study found that "DeepSeek, a biomedically enriched model, outperformed the others, with an accuracy of 92.0%. Copilot and Gemini achieved moderate accuracies of 66.7% and 25.8%, respectively, while ChatGPT achieved the lowest citation accuracy at 19.4%." Additionally, "Expert review confirmed that even the best model produced citation errors, emphasizing the need for human oversight."

Reports effect sizes.

Research paradigm

Positivist/Empiricist

Author conclusions

The authors conclude that "AI models-particularly biomedically enriched tools such as DeepSeek-can accelerate citation drafting, but citation hallucinations and metadata errors remain common. AI should serve as a decision support tool for reference retrieval and formatting, not a substitute for rigorous manual verification before submission."

Risk of bias

Selection bias: Only 35 paragraphs from a single source (Review of Ophthalmology, fourth edition) may not represent the full diversity of glaucoma research citations; Temporal bias: Model performance may change with updates to underlying data and model architecture; Expert reviewer bias: Expert classification of citations could vary based on individual judgment without explicit inter-rater reliability testing reported; Selection bias in choice of test paragraphs (all from single source); potential temporal bias as AI models are subject to updates and changes; possible evaluator bias in expert review process; limited to glaucoma research domain.; Selection bias in choice of test paragraphs (limited to 35 paragraphs from a single textbook source); Temporal bias: model performance may vary with updates and changes to underlying data; Potential evaluator bias in expert review classification; Single domain expertise (glaucoma research) may not generalize to other ophthalmology subspecialties

Limitations

  • The authors note they "interpret this apparent advantage cautiously, as model details, updates, and changes in underlying data may influence performance." Additionally, while the study demonstrated performance differences, the authors acknowledge that "citation hallucinations and metadata errors remain common" even in the best-performing model.

Open questions raised

  • The authors identify the need for continued development of specialized biomedically enriched AI models and the importance of human oversight in citation generation. They highlight the limitations of current general-purpose models in specialized ophthalmology fields and suggest that domain-specific optimization improves performance.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 72%

Explore related topics

Related papers