Hallucinations in Scholarly LLMs: A Conceptual Overview and Practical Implications
Manas Gaur · Maryland Shared Open Access Repository (USMAI Consortium) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.13016/m2vgrz-oust
Methodology & findings
Study design
Conceptual overview and literature synthesis.
Main result
The paper identifies four major types of hallucinations in scholarly LLMs: "Factual hallucination: When an LLM states either incorrect or non-existent information, it would be considered a factual hallucination," "Citation Hallucination: It is defined as invented bibliographic elements, including author names, article titles, journal names, or even DOIs," "Interpretive hallucination: It occurs when models either misrepresent or overgeneralize the findings of a source paper," and "Contextual Hallucination: This occurs when an LLM uses outdated or domain-irrelevant information to make unwarranted transfers from one domain to the other." The paper further notes that "up to 30–40% of the citations ChatGPT and Bard produce are either incorrect or completely fabricated" based on comparative analyses.
Research paradigm
Interpretivist/Conceptual Analysis
Author conclusions
The paper concludes that "This paper developed a conceptual overview of Large Language Models-based hallucination within the context of scholarly communication; thus, it classified the main types of hallucination, analyzed the causes, and examined the implications for academic integrity, reproducibility, and trust. The important mitigation approaches discussed in the paper include rag, post-generation verification, and neurosymbolic integration to improve factual reliability." The authors further assert that "Eventually, ethically led AI companions that are trustworthy and transparently designed, can change LLMs from simple text generators into trusted research assistants that will enhance, not compromise, the pillars of scientific knowledge."
Risk of bias
Selection bias: Non-systematic literature review with no transparent search strategy; Publication bias: Only published sources cited; no gray literature search methodology stated; Author bias: Anonymous authorship prevents positionality transparency, though single-author anonymity raises questions about institutional affiliation verification; Scope bias: Limited to English-language academic literature; international or non-English scholarship may be excluded
Open questions raised
- The authors identify several future research directions: (1) Development of specific datasets and benchmarks to evaluate hallucinations in academic contexts, noting that "While there are initiatives such as SciHal25, which attempts to define shared tasks for detecting hallucination in scientific content, more datasets from different domains are urgently needed"; (2) Development of explainable hallucination detection models that "does not only flag inaccuracies but also explains why any given output is potentially unreliable"; (3) Implementation of hallucination mitigation techniques within academic peer-review and publication pipelines where "Automatic reviewers could support human editors through cross-validation of references, the detection of paraphrased distortions, or unsubstantiated claims prior to publication."
- Domain-specific datasets and benchmarks for evaluating hallucinations in academic contexts beyond SciHal25
- Development of explainable hallucination detection models that explain why outputs are unreliable
- Implementation of hallucination mitigation techniques within academic peer-review and publication pipelines
- Integration of neurosymbolic paradigms embedding interpretability and verifiability into model architecture
- Datasets from different disciplines to capture various ways hallucinations occur across fields
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations