12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Large language models in systematic review and meta-analysis of surgical treatments for vaginal vault prolapse

Yunjeong Park, Hyun-Soo Zhang, Sang Wook Bai · npj Digital Medicine · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41746-026-02431-w

Methodology & findings

Study design

Prospectively registered systematic review and meta-analysis of randomized controlled trials (RCTs) with parallel AI-augmented and human expert review across multiple workflow stages: title/abstract screening, full-text review, snowballing, data extraction, and risk-of-bias assessment.

Sample

N = 1668, 7 groups

Primary method

Random-effects meta-analysis models for pooling data. For binary outcomes (objective success, reoperation), pooled odds ratios (ORs) with 95% CIs calculated using random-effects models. For continuous outcomes (POP-Q point C), pooled weighted mean differences (WMDs) with 95% CIs calculated. Between-study heterogeneity assessed using I² statistic and Cochran's Q test (thresholds: 25% low, 50% moderate, 75% high heterogeneity). Zero-event cells handled with 0.5 continuity correction. Analyses performed using R (version 4.4.3). Performance metrics (accuracy, precision, recall, specificity, F1-score, Cohen's κ) calculated with Wilson score confidence intervals for proportions and bootstrap resampling (1,000 iterations) for F1-score confidence intervals.

Main result

In this systematic review and meta-analysis of 18 randomized trials, "SC provided durable anatomical support, with comparable outcomes between abdominal and laparoscopic approaches. Compared with SSF, SC showed a non-significant trend toward fewer reoperations." Additionally, "Transvaginal mesh was associated with higher objective success than SSF, but at the cost of increased mesh-related complications," with "reoperations for mesh-related complications varied by procedure, occurring in 5-16% of women after TVM and 2-4% after SC."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist (quantitative evidence synthesis with AI augmentation)

Author conclusions

"In conclusion, this meta-analysis demonstrates that SC provides durable anatomical support, with similar outcomes between abdominal and laparoscopic approaches, and a possible reduction in reoperation risk compared with SSF that did not reach statistical significance. TVM was associated with higher objective success but also greater mesh-related complications." The authors further conclude: "Beyond these clinical insights, this review also highlights the potential of LLMs to transform evidence synthesis. With appropriate human oversight, AI can enhance efficiency, transparency, and reproducibility, offering a promising paradigm for systematic reviews and meta-analyses not only in urogynecology but across digital medicine and healthcare research more broadly."

Risk of bias

Selection bias: Limited to RCTs; non-RCT evidence excluded; Publication bias: Not formally assessed due to insufficient number of studies per comparison; Attrition bias: Heterogeneity across studies in follow-up duration (1-9 years); Outcome measurement bias: Heterogeneity in patient-reported outcome measures across studies; AI hallucination risk: Mitigated through grounded prompts, audit trails, and human verification, but prior studies show GPT-4 falsely generated non-existent citations in ~28.6% of cases; Recall limitation in AI screening: ChatGPT achieved only 69.8% recall in title/abstract screening, missing ~30% of eligible trials; Technology bias: Performance metrics obtained under optimized prompting and intensive verification; real-world performance may differ; Selection bias: heterogeneous outcome definitions across studies; Reporting bias: some trials reported only figures for outcomes that ChatGPT could not extract; AI hallucination risk: mitigated through grounded prompting and audit trails; AI omission bias: ChatGPT missed 30.2% of eligible trials in title/abstract screening (recall 69.8%); Publication bias: not formally assessed due to insufficient number of trials per comparison; Imbalanced risk of bias judgment distributions affecting Cohen's κ interpretation; Study design heterogeneity (limited long-term follow-up data); Small number of trials per comparison (≤3 trials for most contrasts); High between-study heterogeneity (I²=92.6% for TVM vs SSF at POP-Q point C); Publication bias not formally assessed due to insufficient trials per comparison; Potential systematic biases in LLM-assisted workflow (hallucinations, omissions with 69.8% recall in title/abstract screening); Risk of bias domains identified as challenging even for human reviewers (missing outcome data, outcome measurement, selective reporting)

Limitations

  • The authors state: "Generalisability is also limited
  • Our performance metrics were obtained in a single, RCT-focused urogynecology review using optimized prompting and intensive human verification
  • Performance may differ in other clinical domains, non-RCT evidence bases, non-English literature, and reviews with more complex outcome hierarchies." Additionally, "the model could not extract outcomes presented only as figures," "it struggled with very long reports," and "LLMs could not process multiple full-text PDFs in parallel." Furthermore, "this was a validation-focused study rather than a time-and-motion study
  • We did not systematically record person-time at each step of the review process."

Open questions raised

  • Need for more long-term follow-up studies to better evaluate durability and late complications after vaginal vault prolapse surgery (e.g., the SALTO trial reported mesh-related complications 5.6-10.2 years after surgery)
  • Further adequately powered RCTs with long-term follow-up are needed
  • Performance validation of LLMs in other clinical domains, non-RCT evidence bases, non-English literature, and reviews with complex outcome hierarchies
  • Development of more sophisticated image and figure recognition capabilities for LLMs to extract outcomes presented graphically
  • Improvement in LLM handling of hierarchical tables, long documents (>200 pages), and parallel processing of multiple PDFs
  • Standardization of heterogeneous patient-reported outcome measures to enable better quantitative synthesis
Extracted from: pdfAgreement 59%

Explore related topics

Related papers