Large language models in systematic review and meta-analysis of surgical treatments for vaginal vault prolapse
Yunjeong Park, Hyun-Soo Zhang, Sang Wook Bai · npj Digital Medicine · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41746-026-02431-w
Methodology & findings
Study design
Prospectively registered systematic review and meta-analysis of randomized controlled trials (RCTs) with parallel AI-augmented and human expert review across multiple workflow stages: title/abstract screening, full-text review, snowballing, data extraction, and risk-of-bias assessment.
Sample
N = 1668, 7 groups
Primary method
Random-effects meta-analysis models for pooling data. For binary outcomes (objective success, reoperation), pooled odds ratios (ORs) with 95% CIs calculated using random-effects models. For continuous outcomes (POP-Q point C), pooled weighted mean differences (WMDs) with 95% CIs calculated. Between-study heterogeneity assessed using I² statistic and Cochran's Q test (thresholds: 25% low, 50% moderate, 75% high heterogeneity). Zero-event cells handled with 0.5 continuity correction. Analyses performed using R (version 4.4.3). Performance metrics (accuracy, precision, recall, specificity, F1-score, Cohen's κ) calculated with Wilson score confidence intervals for proportions and bootstrap resampling (1,000 iterations) for F1-score confidence intervals.
Main result
In this systematic review and meta-analysis of 18 randomized trials, "SC provided durable anatomical support, with comparable outcomes between abdominal and laparoscopic approaches. Compared with SSF, SC showed a non-significant trend toward fewer reoperations." Additionally, "Transvaginal mesh was associated with higher objective success than SSF, but at the cost of increased mesh-related complications," with "reoperations for mesh-related complications varied by procedure, occurring in 5-16% of women after TVM and 2-4% after SC."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist (quantitative evidence synthesis with AI augmentation)
Author conclusions
"In conclusion, this meta-analysis demonstrates that SC provides durable anatomical support, with similar outcomes between abdominal and laparoscopic approaches, and a possible reduction in reoperation risk compared with SSF that did not reach statistical significance. TVM was associated with higher objective success but also greater mesh-related complications." The authors further conclude: "Beyond these clinical insights, this review also highlights the potential of LLMs to transform evidence synthesis. With appropriate human oversight, AI can enhance efficiency, transparency, and reproducibility, offering a promising paradigm for systematic reviews and meta-analyses not only in urogynecology but across digital medicine and healthcare research more broadly."
Risk of bias
Selection bias: Limited to RCTs; non-RCT evidence excluded; Publication bias: Not formally assessed due to insufficient number of studies per comparison; Attrition bias: Heterogeneity across studies in follow-up duration (1-9 years); Outcome measurement bias: Heterogeneity in patient-reported outcome measures across studies; AI hallucination risk: Mitigated through grounded prompts, audit trails, and human verification, but prior studies show GPT-4 falsely generated non-existent citations in ~28.6% of cases; Recall limitation in AI screening: ChatGPT achieved only 69.8% recall in title/abstract screening, missing ~30% of eligible trials; Technology bias: Performance metrics obtained under optimized prompting and intensive verification; real-world performance may differ; Selection bias: heterogeneous outcome definitions across studies; Reporting bias: some trials reported only figures for outcomes that ChatGPT could not extract; AI hallucination risk: mitigated through grounded prompting and audit trails; AI omission bias: ChatGPT missed 30.2% of eligible trials in title/abstract screening (recall 69.8%); Publication bias: not formally assessed due to insufficient number of trials per comparison; Imbalanced risk of bias judgment distributions affecting Cohen's κ interpretation; Study design heterogeneity (limited long-term follow-up data); Small number of trials per comparison (≤3 trials for most contrasts); High between-study heterogeneity (I²=92.6% for TVM vs SSF at POP-Q point C); Publication bias not formally assessed due to insufficient trials per comparison; Potential systematic biases in LLM-assisted workflow (hallucinations, omissions with 69.8% recall in title/abstract screening); Risk of bias domains identified as challenging even for human reviewers (missing outcome data, outcome measurement, selective reporting)
Limitations
- The authors state: "Generalisability is also limited
- Our performance metrics were obtained in a single, RCT-focused urogynecology review using optimized prompting and intensive human verification
- Performance may differ in other clinical domains, non-RCT evidence bases, non-English literature, and reviews with more complex outcome hierarchies." Additionally, "the model could not extract outcomes presented only as figures," "it struggled with very long reports," and "LLMs could not process multiple full-text PDFs in parallel." Furthermore, "this was a validation-focused study rather than a time-and-motion study
- We did not systematically record person-time at each step of the review process."
Open questions raised
- Need for more long-term follow-up studies to better evaluate durability and late complications after vaginal vault prolapse surgery (e.g., the SALTO trial reported mesh-related complications 5.6-10.2 years after surgery)
- Further adequately powered RCTs with long-term follow-up are needed
- Performance validation of LLMs in other clinical domains, non-RCT evidence bases, non-English literature, and reviews with complex outcome hierarchies
- Development of more sophisticated image and figure recognition capabilities for LLMs to extract outcomes presented graphically
- Improvement in LLM handling of hierarchical tables, long documents (>200 pages), and parallel processing of multiple PDFs
- Standardization of heterogeneous patient-reported outcome measures to enable better quantitative synthesis
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations