12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

TopoGuard: Training-Free Hallucination Detection in Graph-Grounded Question Answering; LLM Hallucinations in Academic Research

Chongshan Lin · Libra · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18130/xmj4-m547

Methodology & findings

Study design

Mixed methods combining: (1) Technical system development and experimental evaluation of TopoGuard framework with GraphRAG stress testing using NovelQA and LightRAG datasets; (2) Qualitative document analysis of technical reports, system cards, academic literature, and mitigation strategies; (3) Preliminary interview case study with a university-affiliated researcher using LLMs in teaching and research..

Sample

N = 1, 2 groups

Main result

The study found that "graph topology provides a useful structural signal, but topology alone is incomplete" and that "LLM-based semantic reasoning performs best overall because it can interpret whether the retrieved graph path actually entails the claim." Additionally, "hallucination risk is task-dependent" with researchers feeling more comfortable using LLMs for structural tasks that can be checked against provided material, but facing higher risk when models are asked to retrieve facts or generate citations.

Reports effect sizes.

Research paradigm

Mixed methods (technical empiricism + interpretive sociotechnical analysis)

Author conclusions

The authors conclude that "trustworthy LLM use requires both better detection systems and stronger academic practices of verification" and that "LLMs can remain useful academic assistants, but their outputs need to be traceable with appropriate caution when evidence is weak." They emphasize that the thesis "contributes to the broader goal of making LLMs more reliable in research settings by combining algorithmic grounding with sociotechnical awareness."

Risk of bias

Selection bias in interview sampling (single university-affiliated researcher); Potential researcher bias in qualitative document analysis interpretation; Dependence on graph construction quality affecting verifier performance; Limited scope of GraphRAG stress testing datasets; Limited interview sample (single case study); Document analysis may be subject to selection bias in which documents were examined; No description of blinding or inter-rater reliability for document analysis; Potential confirmation bias in selecting documents supporting the theoretical framework; Graph construction quality dependency not controlled in technical experiments; Single case study in STS component may not represent broader researcher populations; Document analysis selection bias depending on which documents were chosen; Technical evaluation may be influenced by choice of datasets (NovelQA, LightRAG); Potential confirmation bias in selecting which hallucination types to test; Limited interview data may not capture diverse disciplinary perspectives

Limitations

  • The authors note that "when the graph is automatically built from long text, important event-level relations may be missing, which limits what any graph-based verifier can check." Additionally, they acknowledge that "embedding-based detection improves some cases, but it can confuse semantic relatedness with factual equivalence." The STS component relies on "qualitative document analysis and preliminary interview analysis," suggesting limited empirical scope for the social science component.

Open questions raised

  • The paper identifies the need for integration of technical hallucination detection with academic practice; further investigation of how different academic disciplines evaluate evidence differently; deeper understanding of verification labor distribution between humans and tools; and continued development of interpretable, training-free hallucination detection methods.
  • The paper identifies the need for: (1) better hallucination detection systems beyond topology-only approaches; (2) stronger academic verification practices; (3) understanding of task-dependent hallucination risk across different academic disciplines; (4) investigation of how verification labor is distributed between humans and tools; (5) improved graph construction methods from long text to capture event-level relations.
  • The thesis identifies the need for integration of technical and social dimensions of hallucination, discipline-specific verification standards, improved graph construction methods for automatically-built knowledge graphs, and stronger academic practices around citation verification and evidence traceability in LLM-assisted research
Data: not_statedCode: not_statedExtracted from: pdfAgreement 55%

Explore related topics

Related papers