Assessing the Performance of 8 AI Chatbots in Bibliographic Reference Retrieval: Grok and DeepSeek Outperform ChatGPT, but None are Entirely Accurate
Álvaro Cabezas-Clavijo, Pavel Sidorenko-Bautista · Journal of Data and Information Science · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1515/jdis-2025-0326
Methodology & findings
Study design
Comparative analysis of eight generative AI chatbots (ChatGPT, Claude, Copilot, DeepSeek, Gemini, Grok, Le Chat, and Perplexity) using standardized prompts across five academic disciplines.
Main result
The study found that "26.5% were real and fully accurate (i.e., all five bibliographic elements were correct), while 33.8% were real but only partially correct (e.g., containing errors in the publication year or locating data). In contrast, 39.8% of the references were either incorrect or entirely fabricated by the AI systems." Additionally, "The chatbots that yielded the highest percentage of completely correct references were Grok (60%) and DeepSeek (48%). On the other hand, the chatbots most prone to fabricating references were Copilot (100%), Perplexity (72%), and Claude (64%)." Notably, "only two AIs -Grok and DeepSeek-did not fabricate any of the 50 references requested."
Research paradigm
positivist/empiricist
Author conclusions
The authors conclude that "while certain tools such as Grok and DeepSeek demonstrate promising performance, the high rate of fabricated references found in most models highlights the risks associated with the uncritical use of AI-generated information by students. These findings underscore the need to strengthen pedagogical strategies focused on information literacy, particularly within the evolving educational framework shaped by artificial intelligence." They further state that "only about one in four references generated was entirely correct—including authorship, year, title, publication venue, and locating data. This underscores that, while AI systems are becoming increasingly capable of generating legitimate bibliographic references, they are still far from being reliable tools for this fundamental academic task."
Risk of bias
Selection bias: Only free-access versions of chatbots were tested, excluding premium versions that may perform better; Temporal bias: Study conducted February 7-9, 2025; LLM models evolve rapidly and findings may quickly become outdated; Domain-specific bias: Testing limited to five disciplines; results may not generalize to specialized fields; Verification bias: Manual verification conducted by researchers; inter-rater reliability not reported; Prompt bias: Single standardized prompt per discipline may not capture all usage patterns by students; Selection bias: Only free-access versions tested; premium versions excluded, which may underestimate capabilities of paid tiers; Temporal bias: Study conducted February 7-9, 2025; rapid evolution of LLMs means findings may quickly become outdated; Prompt design bias: Single standardized prompt used; variation in prompt engineering could affect results; Verification bias: Manual verification via Google/Google Scholar may not catch all fabrications, especially sophisticated ones; Domain specificity bias: Five specific disciplines selected may not represent all academic areas; Language bias: English-language queries used; performance may differ in other languages; Selection bias in choice of five disciplines may not represent full academic landscape; Testing conducted in February 2025 only; findings may be outdated due to rapid model updates; Limited to free-access versions; findings not generalizable to premium chatbot versions; Manual verification method may introduce human error in reference validation; Single standardized prompt used; variation in prompt phrasing could affect results; Potential bias in selection of verification method (Google/Google Scholar searches)
Limitations
- The authors acknowledge that "the sample size -400 references-is necessarily small, though sufficient to draw meaningful conclusions about the performance of the chatbots in bibliographic reference generation." They also note that "the rapid pace at which large language models (LLMs) are evolving, and the speed at which they are being deployed by major technology companies, mean that some of these findings may quickly become outdated." Additionally, "this experimental study was conducted in a realistic scenario for university students, but not the only one," and the study was limited to free-access versions of chatbots, which "tend to be more limited and less reliable" compared to premium versions.
Open questions raised
- Expansion of analysis to more specific disciplinary contexts beyond the five major areas tested
- Evaluation of premium/paid versions of chatbots rather than free-access versions only
- Exploration of specialized academic discovery tools (e.g., Elicit, ResearchRabbit)
- Investigation of whether AIs are drawing from copyrighted sources without licensing agreements
- Development of safer, more ethical, and more effective learning ecosystems for university students
- Need to expand analysis to more specific disciplinary contexts beyond the five major areas
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- Artificial intelligence and the conduct of literature reviewsGerit Wagner · 2021 · 275 citations
- Artificial intelligence for literature reviews: opportunities and challengesF. J. Bolaños · 2024 · 188 citations