Hallucinations in Scholarly LLMs
Open Conference Proceedings · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.52825/ocp.v8i.3175
Methodology & findings
Study design
Selective review method using iterative search across multiple databases (Google Scholar, Scopus, ACM Digital Library, SpringerLink, arXiv).
Primary method
Qualitative synthesis and thematic analysis; no quantitative statistical methods employed
Main result
The paper identifies four main types of academic hallucinations: "Factual hallucination: It is a case where an LLM produces inaccurate, verifiable or fabricated data"; "Citation Hallucination: This is when a bibliographic element, like an author, article title, journal title, conference venue, publication date or DOI, is made up or modified"; "Interpretive hallucination: A hallucination that is caused by a model malrepresenting, overgeneralizing, or misinterpreting the results of a source article"; and "Contextual Hallucination: This occurs when an LLM uses old, domain-irrelevant, or lost context to make predictions." The research reveals that "30-40 percent of the citations produced by models like ChatGPT and Bard are wrong or completely fake" and that hallucinations pose serious threats to academic integrity and reproducibility.
Reports effect sizes and confidence intervals.
Research paradigm
Interpretivist/Conceptual synthesis
Author conclusions
The authors conclude: "In this paper, the conceptual perspective of the LLM-based hallucinations was presented in the framework of scholarly communication. It categorized core forms of hallucinations, explored the causes of such hallucinations, and discussed the consequences of such hallucinations on the academic integrity, academic reproducibility and academic trust." They further state that "To meet the criterion of verifiability and transparency, they will need domain targeted datasets, systems mindful of hallucinations, and explainable detection models. Finally, the missing component to turn LLMs into reliable research assistants thinkable, transparent, verifiable, and trustworthy AI helpers is the ability to see these technologies as more than mere text generators."
Risk of bias
Non-systematic review methodology may introduce selection bias in literature inclusion; Fragmented and rapidly evolving literature base increases risk of incomplete coverage; Scarcity of domain-specific benchmarks limits comparability across sources; Selection bias: Literature search limited to specific databases and timeframe (2019-2025); potential publication bias towards studies addressing hallucination problems; Interpretation bias: Qualitative synthesis without systematic protocol increases subjective interpretation risk; Coverage bias: Acknowledged fragmentation of literature in rapidly evolving field may result in incomplete representation
Limitations
- The authors explicitly state: "Although such a methodology provides systematic knowledge, it is not a systematic review
- The dynamic character of the research in the field of the LLM, as well as the scarcity of domain-specific benchmarks, entail the fact that the literature is still fragmented." Additionally, they note limitations of current scholarly knowledge graphs: "they do not always have complete coverage, they are slow to index new works, and they would usually only record the metadata, but not the detailed methodological or interpretive information."
Open questions raised
- Need for domain-specific datasets and evaluation benchmarks beyond existing work like SciHal25
- Requirement for evaluation of hallucination behaviors in non-scientific areas
- Need for explainability processes that explain why outputs are unreliable, not just detect hallucinations
- Incorporation of hallucination-detection algorithms into academic peer-review and publication pipelines
- Development of systems that are transparent and reduce the gap between black-box neural generation and human reasoning
- Requirement for more general and specific datasets to evaluate hallucination behaviors in non-scientific areas
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations