12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Compliance of systematic reviews and meta-analyses in ophthalmology with the PRISMA statement: an AI-based assessment and longitudinal comparison with 2017 data

Seon Young Lee, Jae Seon Hong, Sang Hyeok Lee, Rajen K. Gupta · BMC Medical Research Methodology · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
1
Citations
11.28
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s12874-026-02825-0

Methodology & findings

Study design

Systematic assessment study using dual human review and AI evaluation.

Sample

N = 207, 2 groups

Primary method

Cohen's kappa statistic for inter-observer agreement assessment; Mann-Whitney U test for comparing findings with the 2017 study; compliance scores calculated from PRISMA 2020 checklist adherence (42-point scale).

Main result

The study found that "the mean compliance score, as assessed by human reviewers, was 36.28 out of 42 points (86.37%), indicating a substantial improvement in adherence to the PRISMA checklist compared with the level reported in the 2017 study (p < 0.00001)." Additionally, "compliance scores generated by the AI platforms demonstrated a moderate level of agreement with human assessments (Cohen's κ = 0.63 for ChatGPT, 0.54 for Gemini)."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude that "this study demonstrates a marked improvement in the reporting quality of systematic reviews and meta-analyses in ophthalmology following adoption of the 2020 PRISMA statement. Nonetheless, persistent deficiencies remain, particularly in the reporting of bias, sensitivity analyses, and research registration. The application of AI models offers promising potential for enhancing the efficiency and effectiveness of reporting quality assessments; however, further refinement is required to ensure consistency and accuracy."

Risk of bias

Potential reviewer bias in manual assessments despite dual review; AI platform variability (different platforms showed different agreement levels: ChatGPT κ=0.63 vs Gemini κ=0.54); Temporal confounding: comparison with 2017 data may reflect changes in journal practices rather than improved compliance alone; Selection bias: articles limited to 11 major ophthalmology journals; Potential assessor bias in human reviews despite dual assessment; AI platform variability (Cohen's kappa 0.54-0.63 indicates only moderate agreement); Selection bias in journals chosen (only 11 major ophthalmology journals); Selection bias: Limited to 11 major ophthalmology journals only, which may not represent all ophthalmology research; Reviewer bias: Human assessment may be subject to individual interpretation differences despite inter-rater reliability checks; AI platform differences: Discrepancy between ChatGPT-4.0 and Gemini Pro 2.5 suggests potential platform-specific biases; Temporal bias: Comparison with 2017 study may reflect changes in field practices rather than checklist effectiveness

Limitations

  • The authors note that "persistent deficiencies remain, particularly in the reporting of bias, sensitivity analyses, and research registration" and that "further refinement is required to ensure consistency and accuracy" of AI assessments, as "The application of AI models offers promising potential for enhancing the efficiency and effectiveness of reporting quality assessments
  • however, further refinement is required to ensure consistency and accuracy."

Open questions raised

  • Persistent deficiencies in risk of bias assessment (item 11)
  • Inadequate reporting of sensitivity analysis (items 13f and 20c)
  • Poor compliance with research registration requirements (items 24a–24c)
  • Need for further refinement of AI models for consistency and accuracy in compliance assessment
  • Explicit guidance needed in future PRISMA iterations regarding the role of AI in research evaluation
  • Future iterations of the PRISMA guidelines should consider explicitly addressing the role of AI in research evaluation. The study identifies the need for improved reporting in risk of bias assessment, sensitivity analysis, and research registration. Further refinement of AI models is required to ensure consistency and accuracy in compliance assessment.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 63%

Explore related topics

Related papers