2112 Fact or Fiction: Exploring Hallucinated Sources and Outcomes in AI-Generated Research
Aditi Choudhary, Maria Eckmann, Ainsley Anderson, Rayan Khan, Rohit Prem Kumar, Ayesha A. Waheed et al. · Neurosurgery · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1227/neu.0000000000003964_2112
Methodology & findings
Study design
Experimental validation study.
Sample
N = 1100, 2 groups
Primary method
Manual categorical verification and frequency analysis. Percentages calculated for real (41.5%), false (46.7%), and partially false (11.8%) citations. Stratified analysis by medical specialty.
Main result
The study found that "Out of the 1,100 citations, only 41.5% were real, while 58.5% were fabricated." Within the fabricated citations, "46.7% were false and 11.8% were partially false," with significant variation by medical specialty. For example, "Oncology had the highest rate of real citations (80%), Family Medicine had the least number of false papers (16%), and Pediatrics showed the highest proportion (29%) of partially false articles."
Reports effect sizes.
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude that "Built with mathematical algorithms for general-purpose use, ChatGPT lacks the precision and domain-specific depth required for academic research." They emphasize that "These findings emphasize the need for strong oversight in AI-assisted research and urge developers to enhance and check training data for misinformation."
Risk of bias
Selection bias: Only ChatGPT 4.0 was tested; results may not generalize to other LLMs; Evaluator bias: Manual verification by reviewers could introduce subjective categorization errors; Specialty selection bias: Choice of ten medical specialties may not represent all fields equally; Prompt engineering: Specific phrasing of prompts may affect hallucination rates; Manual verification bias: single-pass manual verification without inter-rater reliability assessment; Selection bias: only tested ChatGPT 4.0; results may not generalize to other LLMs; Verification source bias: used only PubMed for verification; other databases not consulted; Prompt design bias: citation generation prompted in an unspecified manner that may influence hallucination rates; Selection bias: only ChatGPT 4.0 tested, no comparison with other LLMs; Manual verification bias: single-rater bias in citation categorization not addressed; Verification bias: reliance on PubMed search may miss citations in other databases; Specialty selection bias: ten specialties chosen may not be representative
Limitations
- The authors state that "ChatGPT lacks the precision and domain-specific depth required for academic research" and note that "Its narrative focus and limited citation accuracy make it unreliable for unsupervised literature searches." The study was limited to one language model (ChatGPT 4.0) and focused on citation generation rather than other research tasks.
Open questions raised
- The authors identify the need for stronger oversight mechanisms in AI-assisted research, improved training data quality and verification in LLMs, and development of domain-specific safeguards before deploying AI in academic research contexts.
- The paper identifies the need for improved oversight of AI-assisted research and calls for developers to enhance training data quality and implement misinformation detection mechanisms.
- The authors identify the need for enhanced training data verification in LLMs, improved domain-specific capabilities for medical research applications, and development of oversight mechanisms for AI-assisted research.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations