12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

Zongrng Li, Mingzheng Yang, Lei Zou, Hongxu Ma, Hao Tian, Siqi Zhou et al. · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Computational benchmark evaluation using automated browser-based response collection from three deployed LLM systems (Claude Sonnet 4.5, Gemini 3, ChatGPT 5.3) accessed through native web interfaces.

Main result

The study found that "all three models tend to produce complete and confident bibliographic answers even when those answers are incorrect" in metadata retrieval tasks. Across all three tasks, "overconfidence persists across different academic workflows, but its form changes with task structure." In literature linking, "models are generally able to retrieve at least one relevant citation for most seed papers, but performance declines substantially when generating longer citation lists," with "Precision@20 values below 0.55" indicating "a large fraction of generated references remain unverifiable despite being presented in fluent and authoritative formats." In research direction generation, "model-generated directions are substantially more concentrated than real research trajectories," with models failing to capture "approximately 91% of the topic space represented by real downstream research."

Research paradigm

Empirical evaluation of computational systems under simulated and real-world conditions

Author conclusions

"This paper presents GIScholarBench, a large-scale benchmark for evaluating LLM academic capabilities and overconfidence in GI-Science scholarly workflows... The results show that LLMs often produce fluent and authoritative outputs even when their answers are incorrect, incomplete, or poorly calibrated." The authors conclude that "LLMs are useful as orientation and brainstorming tools, but should not be treated as authoritative sources for scholarly knowledge" and that "more productive AI-assisted GIScience workflows should position LLMs as auxiliary tools for orientation, filtering, and idea structuring, while relying on human researchers to verify factual claims, trace citations systematically, and actively search for directions that the model may overlook."

Risk of bias

Model selection bias: Only three deployed commercial LLM systems evaluated; no open-weight or domain-specialized models included; Domain specificity bias: Benchmark limited to GIScience literature; findings may not generalize to other disciplines; Temporal bias: Web-deployed models continuously updated, making exact replication impossible; Annotation bias: Ground truth constructed from Scopus metadata only; coverage bounded by 10,865-paper corpus; Collection bias: Browser-based automated collection may trigger anti-automation mechanisms or rate limiting; Prompt design bias: Strong/Weak conditions may not equally reflect real-world usage patterns; Selection bias: Benchmark limited to 25 core GIScience journals, may not represent periphery or emerging venues; Temporal bias: Web interface collection captures specific model versions; continuous updates not documented; Measurement bias: Fuzzy matching threshold (τ=0.20 for TF-IDF similarity, δ=0.15 for cluster assignment) could affect cross-task comparisons; Corpus limitation bias: Ground-truth for literature linking bounded by 10,865-paper corpus; out-of-corpus references not captured; Domain specificity bias: GIScience-specific characteristics (terminology, spatial reasoning, interdisciplinary structure) may not generalize to other scientific fields; Selection bias: Corpus limited to 25 core GIScience journals, potentially underrepresenting emerging venues or interdisciplinary outlets; Temporal bias: Web-deployed models continuously updated; behaviors observed during collection period may not be reproducible; Domain bias: GIScience-specific disciplinary structure may not generalize to other scientific fields; Evaluation bias: TF-IDF similarity threshold (τ=0.20 for Task 2, δ=0.15 for Task 3) choices not fully justified; fuzzy matching may introduce systematic matching biases; Measurement bias: K-Means clustering with k=20 selected via elbow method may oversimplify topic space or create artificial topic boundaries; Platform bias: Web interface behaviors differ from API-based deployments; anti-automation safeguards may affect response patterns

Limitations

  • The authors acknowledge several limitations: "the browser-based collection strategy improves ecological validity by evaluating models through real-world user-facing web interfaces, but it also introduces reproducibility challenges
  • Web-deployed LLM systems are continuously updated through model revisions, interface changes, and platform-level adjustments that are not always publicly documented." Additionally, "the benchmark is constructed entirely from GIScience literature, and the extent to which these findings generalize to other scientific domains remains uncertain." Finally, "our analysis characterizes overconfidence as a behavioral property of model outputs and does not measure confidence calibration in the strict sense
  • We do not elicit explicit confidence estimates from the models, so we cannot quantify the gap between a model's self-reported certainty and its empirical accuracy."

Open questions raised

  • Generalization to other scientific domains beyond GIScience (medicine, law, social sciences)
  • Evaluation of open-weight, domain-finetuned, and retrieval-augmented LLM systems
  • Explicit confidence calibration measurement via confidence scores and reliability diagrams
  • Investigation of whether specialized corpora, retrieval mechanisms, confidence scoring, and explicit 'Not Found' constraints can improve factual reliability
  • Assessment of whether overconfidence is a general LLM architecture property or domain-specific
  • Extensions to additional model families beyond the three evaluated commercial systems
Data: Scopus database records (10,865 papers from 25 GIScience journals, 2020-2025); Ground-truth metadata constructed from Scopus including titles, authors, DOIs, publication years, citation counts, keywords, abstracts, and reference lists; 10,865 GIScience papers from Scopus (25 core GIScience journals, 2020-2025). No explicit repository URL provided in the paper for dataset release.; GIScholarBench corpus: 10,865 papers from 25 core GIScience journals (2020-2025), constructed from Scopus database. Journals and abbreviations listed in Table 1. Specific URLs for data availability not provided in manuscript.Code: JSONL Batch Sender (v6.3) browser extension used for automated prompt collectionExtracted from: pdfAgreement 55%

Explore related topics

Related papers