GIScholarBench: Benchmarking LLM Overconfidence in GIS Research
Zongrng Li, Mingzheng Yang, Lei Zou, Hongxu Ma, Hao Tian, Siqi Zhou et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational benchmark evaluation using automated browser-based response collection from three deployed LLM systems (Claude Sonnet 4.5, Gemini 3, ChatGPT 5.3) accessed through native web interfaces.
Main result
The study found that "all three models tend to produce complete and confident bibliographic answers even when those answers are incorrect" in metadata retrieval tasks. Across all three tasks, "overconfidence persists across different academic workflows, but its form changes with task structure." In literature linking, "models are generally able to retrieve at least one relevant citation for most seed papers, but performance declines substantially when generating longer citation lists," with "Precision@20 values below 0.55" indicating "a large fraction of generated references remain unverifiable despite being presented in fluent and authoritative formats." In research direction generation, "model-generated directions are substantially more concentrated than real research trajectories," with models failing to capture "approximately 91% of the topic space represented by real downstream research."
Research paradigm
Empirical evaluation of computational systems under simulated and real-world conditions
Author conclusions
"This paper presents GIScholarBench, a large-scale benchmark for evaluating LLM academic capabilities and overconfidence in GI-Science scholarly workflows... The results show that LLMs often produce fluent and authoritative outputs even when their answers are incorrect, incomplete, or poorly calibrated." The authors conclude that "LLMs are useful as orientation and brainstorming tools, but should not be treated as authoritative sources for scholarly knowledge" and that "more productive AI-assisted GIScience workflows should position LLMs as auxiliary tools for orientation, filtering, and idea structuring, while relying on human researchers to verify factual claims, trace citations systematically, and actively search for directions that the model may overlook."
Risk of bias
Model selection bias: Only three deployed commercial LLM systems evaluated; no open-weight or domain-specialized models included; Domain specificity bias: Benchmark limited to GIScience literature; findings may not generalize to other disciplines; Temporal bias: Web-deployed models continuously updated, making exact replication impossible; Annotation bias: Ground truth constructed from Scopus metadata only; coverage bounded by 10,865-paper corpus; Collection bias: Browser-based automated collection may trigger anti-automation mechanisms or rate limiting; Prompt design bias: Strong/Weak conditions may not equally reflect real-world usage patterns; Selection bias: Benchmark limited to 25 core GIScience journals, may not represent periphery or emerging venues; Temporal bias: Web interface collection captures specific model versions; continuous updates not documented; Measurement bias: Fuzzy matching threshold (τ=0.20 for TF-IDF similarity, δ=0.15 for cluster assignment) could affect cross-task comparisons; Corpus limitation bias: Ground-truth for literature linking bounded by 10,865-paper corpus; out-of-corpus references not captured; Domain specificity bias: GIScience-specific characteristics (terminology, spatial reasoning, interdisciplinary structure) may not generalize to other scientific fields; Selection bias: Corpus limited to 25 core GIScience journals, potentially underrepresenting emerging venues or interdisciplinary outlets; Temporal bias: Web-deployed models continuously updated; behaviors observed during collection period may not be reproducible; Domain bias: GIScience-specific disciplinary structure may not generalize to other scientific fields; Evaluation bias: TF-IDF similarity threshold (τ=0.20 for Task 2, δ=0.15 for Task 3) choices not fully justified; fuzzy matching may introduce systematic matching biases; Measurement bias: K-Means clustering with k=20 selected via elbow method may oversimplify topic space or create artificial topic boundaries; Platform bias: Web interface behaviors differ from API-based deployments; anti-automation safeguards may affect response patterns
Limitations
- The authors acknowledge several limitations: "the browser-based collection strategy improves ecological validity by evaluating models through real-world user-facing web interfaces, but it also introduces reproducibility challenges
- Web-deployed LLM systems are continuously updated through model revisions, interface changes, and platform-level adjustments that are not always publicly documented." Additionally, "the benchmark is constructed entirely from GIScience literature, and the extent to which these findings generalize to other scientific domains remains uncertain." Finally, "our analysis characterizes overconfidence as a behavioral property of model outputs and does not measure confidence calibration in the strict sense
- We do not elicit explicit confidence estimates from the models, so we cannot quantify the gap between a model's self-reported certainty and its empirical accuracy."
Open questions raised
- Generalization to other scientific domains beyond GIScience (medicine, law, social sciences)
- Evaluation of open-weight, domain-finetuned, and retrieval-augmented LLM systems
- Explicit confidence calibration measurement via confidence scores and reliability diagrams
- Investigation of whether specialized corpora, retrieval mechanisms, confidence scoring, and explicit 'Not Found' constraints can improve factual reliability
- Assessment of whether overconfidence is a general LLM architecture property or domain-specific
- Extensions to additional model families beyond the three evaluated commercial systems
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations