12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Structural Hallucination in Large Language Models: A Network-Based Evaluation of Knowledge Organization and Citation Integrity

Moses Boudourides · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Network-based hallucination stress test combining knowledge graph extraction, graph similarity analysis (Jaccard index), centrality comparison (degree, betweenness, PageRank), and citation integrity verification.

Main result

Across all three empirical domains, substantial structural divergence is observed. "In the lexical benchmark, macro-averaged F1 scores fall below 0.05; in the biographical benchmark, hallucination rates exceed 93%; and in the bibliometric benchmark, citation omission reaches 91.9%." Network-level comparison in the Roget reconstruction further reveals "node-set Jaccard similarity of 0.028 and fabrication rates above 94%." These findings demonstrate that "parametric language models generate semantically plausible approximations without reliably reconstructing externally verifiable relational structures."

Research paradigm

Empirical evaluation of computational systems using network-analytic and information quality frameworks

Author conclusions

"Structural hallucination represents a qualitatively different failure from isolated factual error: it affects how knowledge is organized rather than whether a single statement is true." The authors conclude that "structural fidelity cannot be inferred from local fluency alone" and that "the proposed stress test provides a reproducible instrument for evaluating the structural integrity of LLM-generated knowledge representations within knowledge organization and information quality research." They further argue that "the most consequential failures of LLMs in scholarly contexts are not isolated factual errors but structural hallucinations: distortions of conceptual organization, relational architecture, and bibliographic grounding that remain invisible to local (sentence-level) evaluation."

Risk of bias

Training data distributional bias favoring contemporary vocabulary over historical terms; Parametric memory constraints for database-structured relations; Selection of specific LLM model (not explicitly named in excerpt) may affect generalizability; Ground truth datasets (Roget 1911, Wikidata, Dimensions.ai) may have their own biases in representation; Selection of a single LLM model without comparison to other models or versions; Training data temporal bias toward contemporary vocabulary over historical terminology (acknowledged for Roget benchmark); Potential sampling bias in the selection of 50 COVID-19 publications and associated citations; Ground-truth databases (Wikidata, Dimensions.ai) may themselves contain biases or incompleteness; Selection bias: Choice of three specific domains (lexical, biographical, bibliographic) may not represent all types of knowledge organization systems; Dataset representativeness: Roget 1911 edition is historically specific; Wikidata philosophers limited to 1800-1850 birth cohort; COVID-19 publications may reflect recent training data bias; Temporal bias: Authors acknowledge that LLMs exhibit "temporal distributional bias in training data" leading to anachronistic reconstruction; Reference system selection bias: Reliance on single authoritative reference systems (Wikidata, Dimensions.ai) as ground truth may introduce selection bias

Open questions raised

  • Inadequacy of sentence-level accuracy metrics for evaluating structural knowledge preservation in scholarly contexts
  • Gap between conventional fact-checking and structural integrity assessment
  • Need for methodology to evaluate knowledge organization at the relational/network level rather than propositional level
  • Limited existing frameworks for assessing structural fidelity in LLM-generated knowledge representations
  • The paper identifies gaps in LLM evaluation methodology: current benchmarks emphasize question-answering accuracy and factual recall but do not assess whether models preserve the structure of ontologies, authority files, or citation records when generating extended academic text. It also notes that evaluation of LLM outputs remains largely confined to sentence-level accuracy and semantic plausibility rather than structural integrity. Future work should extend the stress test across additional domains and compare performance across multiple LLM architectures and versions.
  • Standard LLM benchmarks emphasize question-answering accuracy and factual recall but do not assess whether models preserve the structure of ontologies, authority files, or citation records when generating extended academic text
Data: Roget's Thesaurus (1911 edition); Wikidata philosophers (1,532 records for philosophers born 1800-1850; 1,527 matched records); Dimensions.ai COVID-19 publications and citations (50 publications with 654 citations); Roget's Thesaurus 1911 edition (historical lexical ontology); Wikidata philosophers dataset (1,532 philosophers born 1800-1850, 1,527 matched records); Dimensions.ai bibliographic database (50 COVID-19 publications with 654 citations); Roget's Thesaurus (1911 edition) - referenced as ground truth but source repository not explicitly specified; Wikidata philosophers dataset (1,527 matched records of philosophers born 1800-1850) - accessible via Wikidata API; Dimensions.ai COVID-19 bibliographic data (50 publications with 654 citations) - sourced via dimcli Python libraryExtracted from: pdfAgreement 62%

Explore related topics

Related papers