12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

SPIRIT-CONSORT-ELM: Element-Level Assessment of Randomized Controlled Trial Reporting Using Large Language Models

Lan Jiang, Xiangji Ying, Andrew W. Brown, Mengfei Lan, Wenxuan Song, Joseph Menke et al. · medRxiv · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.06.06.26354746

Methodology & findings

Study design

This is a dataset development and computational modeling study.

Sample

N = 200, 10 groups

Primary method

Gwet's AC1 (inter-annotator agreement, preferred over Cohen's kappa for label imbalance handling); Percent agreement calculation; F1 score (harmonic mean of precision and recall); Accuracy (proportion of model predictions matching expert annotations); Micro-averaged F1 and averaged Gwet's AC1 for multi-answer questions; Bootstrap resampling (10,000 iterations) for confidence interval estimation; Precision, recall, specificity, and negative predictive value (NPV) as secondary metrics; Semantic similarity-based example selection using text embeddings (SPIRIT-CONSORT-TM encoder, Qwen-8B-Embedding decoder)

Main result

The study found that "GPT-5 achieves higher scores than Qwen-2.5" with an overall F1 score of 0.822 (95% CI: 0.794-0.847), accuracy of 0.852 (95% CI: 0.826-0.876), and Gwet's AC1 of 0.796 (95% CI: 0.760-0.829) on element-level assessment of RCT reporting. When expert-annotated text snippets were provided, "GPT-5 achieves 0.893 F1, 0.912 accuracy, 0.877 Gwet's AC1" representing a substantial improvement of "+7.1 pp in F1, +6.0 pp in accuracy, and +8.1 pp in Gwet's AC1" compared to automated evidence retrieval.

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/Empiricist - computational validation of automated assessment systems

Author conclusions

"In this study, we developed a set of 119 questions to support element-level assessment of reporting completeness in RCT publications... The dataset provides a fine-grained benchmark for evaluating RCT reporting at the element level. LLM performance indicates that automating element-level reporting assessment is feasible, although further improvements may be needed to achieve higher reliability and practical use. Decomposing checklist items into their constituent elements more faithfully operationalizes the intent of reporting guidelines and enables a more granular characterization of RCT reporting."

Risk of bias

Limited sample diversity: dataset comprises 200 articles (100 pairs), potentially limiting generalizability of reporting patterns; Annotation bias: single annotator (XY) for 150 of 200 articles after double annotation of 50 articles only; inter-annotator agreement issues particularly on interpretation-requiring questions (lowest agreement 0.33); Evidence retrieval bias: LLM performance depends on upstream SPIRIT-CONSORT-TM model performance (F1=0.74 at sentence level); errors in retrieval cascade to downstream model predictions; Model selection bias: prompt development and hyperparameter tuning performed on GPT-5 specifically using small validation subset; Text representation bias: current framework does not address figures and tables, which contain results information; particularly impacts CONSORT-specific questions (average snippet length 195 tokens vs 68 for SPIRIT); Conditional question dependency bias: errors in responses to preceding questions propagate to dependent conditional questions; Limited dataset diversity (200 articles only); potential selection bias in article selection from SPIRIT-CONSORT-TM corpus; inter-annotator agreement issues on questions requiring interpretation (lowest agreement 0.33); model performance dependency on evidence retrieval quality; CONSORT-specific questions underperform due to diffuse and lengthy results descriptions.; Limited dataset diversity - relatively modest number of articles (200) may not represent all RCT reporting patterns; Annotator expertise bias - three expert annotators with prior involvement in SPIRIT-CONSORT-TM may share similar interpretive biases; Model selection bias - prompt development used GPT-5 on validation subset, potentially optimizing for that model; Evidence retrieval bias - automated SPIRIT-CONSORT-TM retrieval model (F1=0.74 at sentence level) introduces errors that propagate to downstream task; Task formulation bias - conversion of multi-choice questions to binary sub-items may lose nuance in complex reporting elements

Limitations

  • "Although the corpus contains 19,500 element-level judgments, it is based on a relatively modest number of articles, which may limit the diversity of reporting patterns represented in the dataset." Additionally, "the best-performing model achieved an overall F1 score of 0.822
  • Although this performance is encouraging, higher performance may be necessary for reliable real-world adoption." Furthermore, "the computational cost might be substantial in practice, as 119 questions must be answered for a single article," and "the current framework is based on the SPIRIT 2013 and CONSORT 2010 reporting guidelines," requiring updates for newer guideline versions.

Open questions raised

  • Higher LLM performance needed for reliable real-world adoption (current F1=0.822)
  • Improved evidence retrieval methods to close the gap between SPIRIT-CONSORT-TM-based retrieval (F1=0.822) and expert-annotated text (F1=0.893)
  • Methods to integrate information from figures and tables, particularly for CONSORT-specific questions
  • Fine-tuning of open-weight models such as Qwen-2.5 as promising direction for improvement
  • Deeper contextual reasoning capabilities for nuanced interpretation of trial results
  • Automatic generation of updated questions aligned with evolving SPIRIT 2025 and CONSORT 2025 guidelines
Data: SPIRIT-CONSORT-ELM; SPIRIT-CONSORT-TM; SPIRIT-CONSORT-ELM (Element-level dataset): 200 articles (100 protocol-results pairs) with 119 questions answered for 9,892 explicit responses out of 19,500 potential questions; SPIRIT-CONSORT-TM: prior annotated corpus of 100 paired RCT protocols and results publications; SPIRIT-CONSORT-ELM dataset (200 articles with 119 questions answered; 19,500 element-level judgments) - availability status not explicitly stated in paper; SPIRIT-CONSORT-TM corpus (100 paired RCT protocols and results publications used as basis) - referenced as prior workCode: llama.cppExtracted from: pdfAgreement 50%

Explore related topics

Related papers