Comparing the Accuracy of 4 Artificial Intelligence Models in PubMed Citation Generation for Glaucoma Research
Mustafa Civelekler, Mehmet Citirik · Journal of Glaucoma · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1097/ijg.0000000000002716
Methodology & findings
Study design
Comparative accuracy assessment study using standardized test cases.
Sample
N = 35, 6 groups
Primary method
Descriptive statistics reporting accuracy percentages for each model. Expert review classification system used for validation (categorical: 'Fully Cited,' 'Partially Cited,' 'Not Cited'). No inferential statistical tests or confidence intervals reported.
Main result
The study found that "DeepSeek, a biomedically enriched model, outperformed the others, with an accuracy of 92.0%. Copilot and Gemini achieved moderate accuracies of 66.7% and 25.8%, respectively, while ChatGPT achieved the lowest citation accuracy at 19.4%." Additionally, "Expert review confirmed that even the best model produced citation errors, emphasizing the need for human oversight."
Reports effect sizes.
Research paradigm
Positivist/Empiricist
Author conclusions
The authors conclude that "AI models-particularly biomedically enriched tools such as DeepSeek-can accelerate citation drafting, but citation hallucinations and metadata errors remain common. AI should serve as a decision support tool for reference retrieval and formatting, not a substitute for rigorous manual verification before submission."
Risk of bias
Selection bias: Only 35 paragraphs from a single source (Review of Ophthalmology, fourth edition) may not represent the full diversity of glaucoma research citations; Temporal bias: Model performance may change with updates to underlying data and model architecture; Expert reviewer bias: Expert classification of citations could vary based on individual judgment without explicit inter-rater reliability testing reported; Selection bias in choice of test paragraphs (all from single source); potential temporal bias as AI models are subject to updates and changes; possible evaluator bias in expert review process; limited to glaucoma research domain.; Selection bias in choice of test paragraphs (limited to 35 paragraphs from a single textbook source); Temporal bias: model performance may vary with updates and changes to underlying data; Potential evaluator bias in expert review classification; Single domain expertise (glaucoma research) may not generalize to other ophthalmology subspecialties
Limitations
- The authors note they "interpret this apparent advantage cautiously, as model details, updates, and changes in underlying data may influence performance." Additionally, while the study demonstrated performance differences, the authors acknowledge that "citation hallucinations and metadata errors remain common" even in the best-performing model.
Open questions raised
- The authors identify the need for continued development of specialized biomedically enriched AI models and the importance of human oversight in citation generation. They highlight the limitations of current general-purpose models in specialized ophthalmology fields and suggest that domain-specific optimization improves performance.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations