Large language models for scientometric mapping of scientific controversy: A validated hybrid AI–Human framework
Teo Sušnjak, Cole Palffy, Tatiana Zimina, Nazgul Altynbekova, Kunal Garg, Leona Gilbert · Scientometrics · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-026-05681-3
Methodology & findings
Study design
Hybrid AI-human computational workflow with four stages: (1) data acquisition from academic databases (2000-2024) using Publish or Perish software; (2) automated classification using large language models (GPT-4o-mini as primary model, validated against six alternative LLMs); (3) structured expert validation via inter-rater reliability analysis with two biomedical scientist co-authors on a random sample of 150 abstracts; (4) thematic analysis combining LLM-driven extraction with expert review.
Sample
N = 1033, 8 groups
Primary method
Cohen's Kappa (pairwise comparisons of inter-rater reliability); Fleiss' Kappa (multiple raters); descriptive statistics for stance distribution by year, venue, and citation impact; journal stance-gap scores (percentage-point difference between PTLDS- and CLD-supporting studies); time-adjusted observed-minus-expected stance shares (year-specific corpus distribution weighted by journal publication-year mix); citation impact analysis using Scopus citation counts
Main result
The study found that "the hybrid AI-human framework reproduced expert judgement in complex stance-classification tasks at levels broadly comparable to expert-expert agreement, suggesting sufficient reliability for corpus-level scientometric analysis." Additionally, "PTLDS-supporting papers made up 34% of studies but received 52% of citations, whereas CLD-supporting papers comprised 24% of studies and received 22% of citations," demonstrating asymmetrical citation attention across competing stances in the Lyme disease literature.
Reports effect sizes.
Research paradigm
Pragmatist/Mixed-methods (hybrid AI-human framework with empirical validation)
Author conclusions
"This study introduced and validated a hybrid AI-human method for large-scale analysis of contested scientific discourse... The framework reproduced expert judgement at levels broadly comparable to expert-expert agreement, supporting its use for corpus-level scientometric analysis... work aligned with the mainstream post-treatment syndrome framing (PTLDS), which explains persistent symptoms without positing ongoing active infection after standard therapy, was more prevalent in higher-impact journals and received a disproportionate share of citations. Work aligned with the chronic-infection framing (CLD), which argues for persistent infection and often supports extended antimicrobial treatment, was concentrated in a different segment of the publication landscape... It offers a scalable and reliable framework for literature-based scientometrics and makes the structure of scientific disagreement measurable rather than anecdotal."
Risk of bias
Selection bias from incomplete corpus (7,528 records excluded due to irretrievable abstracts, non-English text, missing search terms); Database coverage bias (38,000+ records removed due to missing DOIs); Temporal bias from shifting terminology that may reflect changes in terminology rather than substantive claims; Anchoring bias from validation interface showing only two candidate justifications, preventing raters from rejecting both or providing free-text alternatives; Model selection bias (primary reliance on GPT-4o-mini rather than other LLMs); Abstract availability bias affecting observed stance prevalence; Venue publication patterns and citation asymmetries reflecting epistemic gatekeeping rather than truth; Selection bias: Missing or irretrievable abstracts (7,528 excluded) may not be random; corpus completeness and stance prevalence estimates may be affected; Language bias: English-language coverage only (non-English texts excluded); Database coverage bias: Records sourced from specific academic databases with potential under-representation of certain publication types; Anchoring bias: Validation interface presented two candidate justifications; raters could not reject both or provide free-text alternatives; Model selection bias: GPT-4o-mini as primary classifier; varying performance across alternative LLMs (Kappa range 0.458-0.717 with alternative models); Attrition: Initial dataset of 84,140 reduced to 1,033 for final analysis (98.8% attrition through filtering steps); Temporal terminology bias: "[S]ome temporal change may reflect shifting terminology rather than changing substantive claims"; Abstract availability bias: more than 38,000 records were excluded due to missing/duplicated DOIs; 27,180 excluded due to missing publication names, titles, or abstracts; English-language bias: non-English texts excluded; Database coverage bias: corpus represents accessible English-language literature rather than complete census; Anchoring bias in validation interface: raters saw two candidate justifications and could not reject both or provide free-text alternatives; Model selection bias: choice of GPT-4o-mini as primary classifier; agreement varied significantly with alternative LLMs; Temporal terminology bias: 'some temporal change may reflect shifting terminology rather than changing substantive claims'
Limitations
- "The framework is limited by abstract-only evidence, English-language coverage, residual interpretive ambiguity even after expert validation, and the compression inherent in thematic labels
- Abstract-only analysis can miss methodological detail and rhetorical structure in full texts
- Database coverage and abstract availability shape the corpus and can bias observed stance prevalence..
- Because the structure of irretrievable abstracts was not preserved at sufficient granularity for retrospective audit, our Lyme-specific findings should be interpreted as patterns within the retrieved corpus rather than exact prevalence estimates for the full underlying literature
- The validation interface also imposed constraints: raters saw two candidate justifications during the task and could not reject both or provide unconstrained free-text alternatives
- This design kept the evaluation structured, but it may have introduced anchoring effects..
Open questions raised
- Need for validated, scalable methods for literature-based scientometric analysis in contested domains
- LLM-assisted scientometrics lacks well-validated workflows and structured evaluation protocols that integrate automated outputs with human oversight
- Broader generalization beyond Lyme disease to controversies with weaker boundaries, more than two salient positions, or greater conceptual drift requiring adapted schemas
- Prompt stability and robustness across model variants and inference-time perturbations
- Temporal change attribution (distinguishing terminology shifts from substantive claim changes)
- Limited by abstract-only analysis; methodology could be extended to full texts for methodological detail and rhetorical structure
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations