SPIRIT-CONSORT-ELM: Element-Level Assessment of Randomized Controlled Trial Reporting Using Large Language Models
Lan Jiang, Xiangji Ying, Andrew W. Brown, Mengfei Lan, Wenxuan Song, Joseph Menke et al. · medRxiv · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.06.06.26354746
Methodology & findings
Study design
This is a dataset development and computational modeling study.
Sample
N = 200, 10 groups
Primary method
Gwet's AC1 (inter-annotator agreement, preferred over Cohen's kappa for label imbalance handling); Percent agreement calculation; F1 score (harmonic mean of precision and recall); Accuracy (proportion of model predictions matching expert annotations); Micro-averaged F1 and averaged Gwet's AC1 for multi-answer questions; Bootstrap resampling (10,000 iterations) for confidence interval estimation; Precision, recall, specificity, and negative predictive value (NPV) as secondary metrics; Semantic similarity-based example selection using text embeddings (SPIRIT-CONSORT-TM encoder, Qwen-8B-Embedding decoder)
Main result
The study found that "GPT-5 achieves higher scores than Qwen-2.5" with an overall F1 score of 0.822 (95% CI: 0.794-0.847), accuracy of 0.852 (95% CI: 0.826-0.876), and Gwet's AC1 of 0.796 (95% CI: 0.760-0.829) on element-level assessment of RCT reporting. When expert-annotated text snippets were provided, "GPT-5 achieves 0.893 F1, 0.912 accuracy, 0.877 Gwet's AC1" representing a substantial improvement of "+7.1 pp in F1, +6.0 pp in accuracy, and +8.1 pp in Gwet's AC1" compared to automated evidence retrieval.
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/Empiricist - computational validation of automated assessment systems
Author conclusions
"In this study, we developed a set of 119 questions to support element-level assessment of reporting completeness in RCT publications... The dataset provides a fine-grained benchmark for evaluating RCT reporting at the element level. LLM performance indicates that automating element-level reporting assessment is feasible, although further improvements may be needed to achieve higher reliability and practical use. Decomposing checklist items into their constituent elements more faithfully operationalizes the intent of reporting guidelines and enables a more granular characterization of RCT reporting."
Risk of bias
Limited sample diversity: dataset comprises 200 articles (100 pairs), potentially limiting generalizability of reporting patterns; Annotation bias: single annotator (XY) for 150 of 200 articles after double annotation of 50 articles only; inter-annotator agreement issues particularly on interpretation-requiring questions (lowest agreement 0.33); Evidence retrieval bias: LLM performance depends on upstream SPIRIT-CONSORT-TM model performance (F1=0.74 at sentence level); errors in retrieval cascade to downstream model predictions; Model selection bias: prompt development and hyperparameter tuning performed on GPT-5 specifically using small validation subset; Text representation bias: current framework does not address figures and tables, which contain results information; particularly impacts CONSORT-specific questions (average snippet length 195 tokens vs 68 for SPIRIT); Conditional question dependency bias: errors in responses to preceding questions propagate to dependent conditional questions; Limited dataset diversity (200 articles only); potential selection bias in article selection from SPIRIT-CONSORT-TM corpus; inter-annotator agreement issues on questions requiring interpretation (lowest agreement 0.33); model performance dependency on evidence retrieval quality; CONSORT-specific questions underperform due to diffuse and lengthy results descriptions.; Limited dataset diversity - relatively modest number of articles (200) may not represent all RCT reporting patterns; Annotator expertise bias - three expert annotators with prior involvement in SPIRIT-CONSORT-TM may share similar interpretive biases; Model selection bias - prompt development used GPT-5 on validation subset, potentially optimizing for that model; Evidence retrieval bias - automated SPIRIT-CONSORT-TM retrieval model (F1=0.74 at sentence level) introduces errors that propagate to downstream task; Task formulation bias - conversion of multi-choice questions to binary sub-items may lose nuance in complex reporting elements
Limitations
- "Although the corpus contains 19,500 element-level judgments, it is based on a relatively modest number of articles, which may limit the diversity of reporting patterns represented in the dataset." Additionally, "the best-performing model achieved an overall F1 score of 0.822
- Although this performance is encouraging, higher performance may be necessary for reliable real-world adoption." Furthermore, "the computational cost might be substantial in practice, as 119 questions must be answered for a single article," and "the current framework is based on the SPIRIT 2013 and CONSORT 2010 reporting guidelines," requiring updates for newer guideline versions.
Open questions raised
- Higher LLM performance needed for reliable real-world adoption (current F1=0.822)
- Improved evidence retrieval methods to close the gap between SPIRIT-CONSORT-TM-based retrieval (F1=0.822) and expert-annotated text (F1=0.893)
- Methods to integrate information from figures and tables, particularly for CONSORT-specific questions
- Fine-tuning of open-weight models such as Qwen-2.5 as promising direction for improvement
- Deeper contextual reasoning capabilities for nuanced interpretation of trial results
- Automatic generation of updated questions aligned with evolving SPIRIT 2025 and CONSORT 2025 guidelines
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations