12,637 papers · updated 18 Sept 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Semantic stability protocol: intercoder reliability for zero-shot classification via AI coders

Hung-Yen Hsu, Mei‐Ling Hsu · Quality & Quantity · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
C
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11135-026-02832-9

Methodology & findings

Study design

Computational simulation study using zero-shot classification on 424 Chinese news articles.

Main result

The study found that "high intercoder reliability (mean α > 0.94) is achievable with two AI coders (RQ1)" and that "Krippendorff's α, calculated by treating each of the 100 classification runs as an independent observer in a run-level reliability matrix (100 runs × 424 documents), was 0.8485, surpassing the conventional content analysis threshold of α ≥ 0.80." The research demonstrates that "systematic aggregation yields high procedural reliability despite individual-run instability, consistent with the vector consistency hypothesis."

Research paradigm

Computational empiricism with methodological innovation

Author conclusions

The authors conclude: "This study has developed and empirically tested a methodological framework that converts zero-shot classification into transparent, auditable procedures for social science research. By reimagining repeated LLM outputs as AI coders, the study illustrates that traditional reliability standards can be met with minimal computational investment. Empirically, two AI coders (20 runs) sufficed to achieve mean intercoder α > 0.94, indicating that systematic aggregation yields high procedural reliability despite individual-run instability, consistent with the vector consistency hypothesis." They further state that the Semantic Stability Protocol "provides a deployable workflow: researchers can achieve high procedural reliability with as few as two AI coders, applying stratified aggregation strategies based on MR and ConfGap diagnostics and reserving human review for the small fraction of genuinely ambiguous cases."

Risk of bias

Model-specific findings: Results limited to DeepSeek Reasoner; generalization to other LLMs unclear; Language-specific: Testing only on Chinese text; may not generalize to other languages; Domain-specific: Net-zero emissions news; results may not apply to other classification tasks or text types; Non-independence of repeated samples: All 100 runs from same model with same weights; not statistically independent like truly distinct coders; Semantic center circularity: PA metric uses majority vote to define semantic center, then measures convergence to it (mathematical certainty for Vote strategy); Single-model evaluation: No comparison across different LLM architectures or providers; No external ground truth: No comparison against human expert coding across all stability strata; Parameter bias: temperature fixed at 0 (does not represent typical usage); Selection bias in dataset: 424 articles from 15 specific media organizations (April 2021-September 2024); Single temperature setting (zero) - stochasticity patterns may differ at other temperatures; Confidence scores are not calibrated probabilities but prompted self-assessments

Limitations

  • "All findings are specific to the tested configuration: a single model (DeepSeek Reasoner), a single language (Chinese), a single domain (net-zero news), and fixed parameters (temperature = 0)
  • The zero-temperature setting means the observed stochasticity likely represents a lower bound of system-level variance
  • stability profiles may differ across other configurations
  • Researchers should conduct pilot tests before applying the protocol to new settings." Additionally, "The repeated outputs share model weights and prompt, making them not statistically independent
  • reported α values measure output stability under stochastic perturbation rather than agreement among genuinely independent observers." Furthermore, "procedural reliability does not imply substantive validity
  • High α indicates consistent, not correct, classification."

Open questions raised

  • Extension to diverse LLMs: "Future research should extend this framework to diverse LLMs, languages, and classification schemes"
  • External validity assessment: "External validity assessment comparing protocol outputs against expert coding across all three stability strata is a necessary next step"
  • Controlled experiments on prompt design: "Whether prompt design independently affects single-output variability requires further controlled experiments"
  • Multi-model configurations: Potential for AI coders from different models to serve as convergent validity evidence
  • Generalization across configurations: Need for pilot tests when applying protocol to new model-language-domain combinations
  • Lack of conceptual framework for organizing repeated outputs as 'AI coders'
Data: Complete dataset available as Supplementary Dataset (424 Chinese news articles on net-zero emissions from 15 media organizations, April 2021 to September 2024); Supplementary Implementation Template provided; Complete dataset available as Supplementary Dataset (location not specified in paper)Code: Ready-to-use implementation template provided as Supplementary Implementation TemplateExtracted from: pdf

Explore related topics

Related papers