12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Semantic stability protocol: intercoder reliability for zero-shot classification via AI coders

Hung-Yen Hsu, Mei‐Ling Hsu · Quality & Quantity · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11135-026-02832-9

Methodology & findings

Study design

Computational simulation study using zero-shot classification on 424 Chinese news articles.

Main result

The study found that "high intercoder reliability (mean α > 0.94) is achievable with two AI coders (RQ1)" and that "Krippendorff's α, calculated by treating each of the 100 classification runs as an independent observer in a run-level reliability matrix (100 runs × 424 documents), was 0.8485, surpassing the conventional content analysis threshold of α ≥ 0.80." The research demonstrates that "systematic aggregation yields high procedural reliability despite individual-run instability, consistent with the vector consistency hypothesis."

Research paradigm

Computational empiricism with methodological innovation

Author conclusions

The authors conclude: "This study has developed and empirically tested a methodological framework that converts zero-shot classification into transparent, auditable procedures for social science research. By reimagining repeated LLM outputs as AI coders, the study illustrates that traditional reliability standards can be met with minimal computational investment. Empirically, two AI coders (20 runs) sufficed to achieve mean intercoder α > 0.94, indicating that systematic aggregation yields high procedural reliability despite individual-run instability, consistent with the vector consistency hypothesis." They further state that the Semantic Stability Protocol "provides a deployable workflow: researchers can achieve high procedural reliability with as few as two AI coders, applying stratified aggregation strategies based on MR and ConfGap diagnostics and reserving human review for the small fraction of genuinely ambiguous cases."

Risk of bias

Model-specific findings: Results limited to DeepSeek Reasoner; generalization to other LLMs unclear; Language-specific: Testing only on Chinese text; may not generalize to other languages; Domain-specific: Net-zero emissions news; results may not apply to other classification tasks or text types; Non-independence of repeated samples: All 100 runs from same model with same weights; not statistically independent like truly distinct coders; Semantic center circularity: PA metric uses majority vote to define semantic center, then measures convergence to it (mathematical certainty for Vote strategy); Single-model evaluation: No comparison across different LLM architectures or providers; No external ground truth: No comparison against human expert coding across all stability strata; Model-specific bias: findings limited to single LLM (DeepSeek Reasoner); Language bias: Chinese-only evaluation; Domain bias: net-zero emissions news only; Parameter bias: temperature fixed at 0 (does not represent typical usage); Circular logic in PA metric: Vote strategy's convergence to PA=1.0 is mathematically predetermined by semantic center definition; Non-independence of repeated outputs: all runs share same model weights and prompt; Selection bias in dataset: 424 articles from 15 specific media organizations (April 2021-September 2024); Model-specific findings not generalizable to other LLMs; Language-specific (Chinese) - may not transfer to other languages; Domain-specific (net-zero emissions news) - classification task characteristics may differ across domains; Single temperature setting (zero) - stochasticity patterns may differ at other temperatures; Confidence scores are not calibrated probabilities but prompted self-assessments

Limitations

  • "All findings are specific to the tested configuration: a single model (DeepSeek Reasoner), a single language (Chinese), a single domain (net-zero news), and fixed parameters (temperature = 0)
  • The zero-temperature setting means the observed stochasticity likely represents a lower bound of system-level variance
  • stability profiles may differ across other configurations
  • Researchers should conduct pilot tests before applying the protocol to new settings." Additionally, "The repeated outputs share model weights and prompt, making them not statistically independent
  • reported α values measure output stability under stochastic perturbation rather than agreement among genuinely independent observers." Furthermore, "procedural reliability does not imply substantive validity
  • High α indicates consistent, not correct, classification."

Open questions raised

  • Extension to diverse LLMs: "Future research should extend this framework to diverse LLMs, languages, and classification schemes"
  • External validity assessment: "External validity assessment comparing protocol outputs against expert coding across all three stability strata is a necessary next step"
  • Controlled experiments on prompt design: "Whether prompt design independently affects single-output variability requires further controlled experiments"
  • Multi-model configurations: Potential for AI coders from different models to serve as convergent validity evidence
  • Generalization across configurations: Need for pilot tests when applying protocol to new model-language-domain combinations
  • Lack of conceptual framework for organizing repeated outputs as 'AI coders'
Data: Complete dataset available as Supplementary Dataset (424 Chinese news articles on net-zero emissions from 15 media organizations, April 2021 to September 2024); Supplementary Implementation Template provided; Complete dataset available as Supplementary Dataset (location not specified in paper); Complete dataset available as Supplementary Dataset (424 Chinese medium-to-long-form news articles concerning net-zero emissions from 15 media organizations, April 2021 to September 2024)Code: Ready-to-use implementation template provided as Supplementary Implementation TemplateExtracted from: pdfAgreement 52%

Explore related topics

Related papers