12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Large language models for scientometric mapping of scientific controversy: A validated hybrid AI–Human framework

Teo Sušnjak, Cole Palffy, Tatiana Zimina, Nazgul Altynbekova, Kunal Garg, Leona Gilbert · Scientometrics · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-026-05681-3

Methodology & findings

Study design

Hybrid AI-human computational workflow with four stages: (1) data acquisition from academic databases (2000-2024) using Publish or Perish software; (2) automated classification using large language models (GPT-4o-mini as primary model, validated against six alternative LLMs); (3) structured expert validation via inter-rater reliability analysis with two biomedical scientist co-authors on a random sample of 150 abstracts; (4) thematic analysis combining LLM-driven extraction with expert review.

Sample

N = 1033, 8 groups

Primary method

Cohen's Kappa (pairwise comparisons of inter-rater reliability); Fleiss' Kappa (multiple raters); descriptive statistics for stance distribution by year, venue, and citation impact; journal stance-gap scores (percentage-point difference between PTLDS- and CLD-supporting studies); time-adjusted observed-minus-expected stance shares (year-specific corpus distribution weighted by journal publication-year mix); citation impact analysis using Scopus citation counts

Main result

The study found that "the hybrid AI-human framework reproduced expert judgement in complex stance-classification tasks at levels broadly comparable to expert-expert agreement, suggesting sufficient reliability for corpus-level scientometric analysis." Additionally, "PTLDS-supporting papers made up 34% of studies but received 52% of citations, whereas CLD-supporting papers comprised 24% of studies and received 22% of citations," demonstrating asymmetrical citation attention across competing stances in the Lyme disease literature.

Reports effect sizes.

Research paradigm

Pragmatist/Mixed-methods (hybrid AI-human framework with empirical validation)

Author conclusions

"This study introduced and validated a hybrid AI-human method for large-scale analysis of contested scientific discourse... The framework reproduced expert judgement at levels broadly comparable to expert-expert agreement, supporting its use for corpus-level scientometric analysis... work aligned with the mainstream post-treatment syndrome framing (PTLDS), which explains persistent symptoms without positing ongoing active infection after standard therapy, was more prevalent in higher-impact journals and received a disproportionate share of citations. Work aligned with the chronic-infection framing (CLD), which argues for persistent infection and often supports extended antimicrobial treatment, was concentrated in a different segment of the publication landscape... It offers a scalable and reliable framework for literature-based scientometrics and makes the structure of scientific disagreement measurable rather than anecdotal."

Risk of bias

Selection bias from incomplete corpus (7,528 records excluded due to irretrievable abstracts, non-English text, missing search terms); Database coverage bias (38,000+ records removed due to missing DOIs); Temporal bias from shifting terminology that may reflect changes in terminology rather than substantive claims; Anchoring bias from validation interface showing only two candidate justifications, preventing raters from rejecting both or providing free-text alternatives; Model selection bias (primary reliance on GPT-4o-mini rather than other LLMs); Abstract availability bias affecting observed stance prevalence; Venue publication patterns and citation asymmetries reflecting epistemic gatekeeping rather than truth; Selection bias: Missing or irretrievable abstracts (7,528 excluded) may not be random; corpus completeness and stance prevalence estimates may be affected; Language bias: English-language coverage only (non-English texts excluded); Database coverage bias: Records sourced from specific academic databases with potential under-representation of certain publication types; Anchoring bias: Validation interface presented two candidate justifications; raters could not reject both or provide free-text alternatives; Model selection bias: GPT-4o-mini as primary classifier; varying performance across alternative LLMs (Kappa range 0.458-0.717 with alternative models); Attrition: Initial dataset of 84,140 reduced to 1,033 for final analysis (98.8% attrition through filtering steps); Temporal terminology bias: "[S]ome temporal change may reflect shifting terminology rather than changing substantive claims"; Abstract availability bias: more than 38,000 records were excluded due to missing/duplicated DOIs; 27,180 excluded due to missing publication names, titles, or abstracts; English-language bias: non-English texts excluded; Database coverage bias: corpus represents accessible English-language literature rather than complete census; Anchoring bias in validation interface: raters saw two candidate justifications and could not reject both or provide free-text alternatives; Model selection bias: choice of GPT-4o-mini as primary classifier; agreement varied significantly with alternative LLMs; Temporal terminology bias: 'some temporal change may reflect shifting terminology rather than changing substantive claims'

Limitations

  • "The framework is limited by abstract-only evidence, English-language coverage, residual interpretive ambiguity even after expert validation, and the compression inherent in thematic labels
  • Abstract-only analysis can miss methodological detail and rhetorical structure in full texts
  • Database coverage and abstract availability shape the corpus and can bias observed stance prevalence..
  • Because the structure of irretrievable abstracts was not preserved at sufficient granularity for retrospective audit, our Lyme-specific findings should be interpreted as patterns within the retrieved corpus rather than exact prevalence estimates for the full underlying literature
  • The validation interface also imposed constraints: raters saw two candidate justifications during the task and could not reject both or provide unconstrained free-text alternatives
  • This design kept the evaluation structured, but it may have introduced anchoring effects..

Open questions raised

  • Need for validated, scalable methods for literature-based scientometric analysis in contested domains
  • LLM-assisted scientometrics lacks well-validated workflows and structured evaluation protocols that integrate automated outputs with human oversight
  • Broader generalization beyond Lyme disease to controversies with weaker boundaries, more than two salient positions, or greater conceptual drift requiring adapted schemas
  • Prompt stability and robustness across model variants and inference-time perturbations
  • Temporal change attribution (distinguishing terminology shifts from substantive claim changes)
  • Limited by abstract-only analysis; methodology could be extended to full texts for methodological detail and rhetorical structure
Data: Filtered corpus of abstract details (excluding abstract text), prompts, and classification outputs available on GitHub (https://github.com/teosusnjak/Lyme-disease-controversy); Filtered corpus of abstract details (excluding abstract text due to copyright), prompts, and classification outputs are available via GitHub repository; Filtered corpus of abstract details (excluding abstract text due to copyright), prompt templates, and classification outputs openly available via GitHub repository: https://github.com/teosuSnjak/Lyme-disease-controversyCode: https://github.com/teosusnjak/Lyme-disease-controversy; https://github.com/teosusjnak/Lyme-disease-controversyExtracted from: pdfAgreement 46%

Explore related topics

Related papers