Structural Hallucination in Large Language Models: A Network-Based Evaluation of Knowledge Organization and Citation Integrity
Moses Boudourides · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Network-based hallucination stress test combining knowledge graph extraction, graph similarity analysis (Jaccard index), centrality comparison (degree, betweenness, PageRank), and citation integrity verification.
Main result
Across all three empirical domains, substantial structural divergence is observed. "In the lexical benchmark, macro-averaged F1 scores fall below 0.05; in the biographical benchmark, hallucination rates exceed 93%; and in the bibliometric benchmark, citation omission reaches 91.9%." Network-level comparison in the Roget reconstruction further reveals "node-set Jaccard similarity of 0.028 and fabrication rates above 94%." These findings demonstrate that "parametric language models generate semantically plausible approximations without reliably reconstructing externally verifiable relational structures."
Research paradigm
Empirical evaluation of computational systems using network-analytic and information quality frameworks
Author conclusions
"Structural hallucination represents a qualitatively different failure from isolated factual error: it affects how knowledge is organized rather than whether a single statement is true." The authors conclude that "structural fidelity cannot be inferred from local fluency alone" and that "the proposed stress test provides a reproducible instrument for evaluating the structural integrity of LLM-generated knowledge representations within knowledge organization and information quality research." They further argue that "the most consequential failures of LLMs in scholarly contexts are not isolated factual errors but structural hallucinations: distortions of conceptual organization, relational architecture, and bibliographic grounding that remain invisible to local (sentence-level) evaluation."
Risk of bias
Training data distributional bias favoring contemporary vocabulary over historical terms; Parametric memory constraints for database-structured relations; Selection of specific LLM model (not explicitly named in excerpt) may affect generalizability; Ground truth datasets (Roget 1911, Wikidata, Dimensions.ai) may have their own biases in representation; Selection of a single LLM model without comparison to other models or versions; Training data temporal bias toward contemporary vocabulary over historical terminology (acknowledged for Roget benchmark); Potential sampling bias in the selection of 50 COVID-19 publications and associated citations; Ground-truth databases (Wikidata, Dimensions.ai) may themselves contain biases or incompleteness; Selection bias: Choice of three specific domains (lexical, biographical, bibliographic) may not represent all types of knowledge organization systems; Dataset representativeness: Roget 1911 edition is historically specific; Wikidata philosophers limited to 1800-1850 birth cohort; COVID-19 publications may reflect recent training data bias; Temporal bias: Authors acknowledge that LLMs exhibit "temporal distributional bias in training data" leading to anachronistic reconstruction; Reference system selection bias: Reliance on single authoritative reference systems (Wikidata, Dimensions.ai) as ground truth may introduce selection bias
Open questions raised
- Inadequacy of sentence-level accuracy metrics for evaluating structural knowledge preservation in scholarly contexts
- Gap between conventional fact-checking and structural integrity assessment
- Need for methodology to evaluate knowledge organization at the relational/network level rather than propositional level
- Limited existing frameworks for assessing structural fidelity in LLM-generated knowledge representations
- The paper identifies gaps in LLM evaluation methodology: current benchmarks emphasize question-answering accuracy and factual recall but do not assess whether models preserve the structure of ontologies, authority files, or citation records when generating extended academic text. It also notes that evaluation of LLM outputs remains largely confined to sentence-level accuracy and semantic plausibility rather than structural integrity. Future work should extend the stress test across additional domains and compare performance across multiple LLM architectures and versions.
- Standard LLM benchmarks emphasize question-answering accuracy and factual recall but do not assess whether models preserve the structure of ontologies, authority files, or citation records when generating extended academic text
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations