AutoForest: Automatically Generating Forest Plots from Biomedical Studies with End-to-End Evidence Extraction and Synthesis
Massimiliano Pronesti, Angelo Miculescu, Mohsin Kapdi, Paul Flanagan, Oisín Redmond, Joao Bettencourt-Silva et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods evaluation combining automatic evaluations on 32 forest plots from 18 Cochrane systematic reviews (56 studies total) with a controlled within-subjects user study involving 4 clinical domain experts and 4 graduate students.
Sample
N = 8, 2 groups
Primary method
Wilcoxon signed-rank tests (non-parametric alternative to paired t-tests); R meta package for statistical synthesis (Balduzzi et al., 2019); quantitative metrics for table detection, structure parsing, and ICO suggestion evaluation; precision, recall, and hallucination rate calculations; accuracy and edit rate calculations.
Main result
The study found that AUTOFOREST significantly outperforms manual workflows across all metrics. "For RQ1, the tool nearly halved the time required to complete a forest plot for both groups (p < 0.001)." Additionally, "the fully automated version ('AUTOFOREST only') achieved over 80% accuracy in data extraction, a substantial improvement over the manual baselines. Human-in-the-loop edits further refined these results, reaching a peak accuracy for experts of 90.2% for data extraction (p = 0.013) and 79.2% for RoB (p = 0.050)."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical (system development and evaluation)
Author conclusions
The authors conclude: "AUTOFOREST addresses this bottleneck by integrating document parsing, automated ICO suggestions, evidence extraction, risk of bias assessment and reasoning to substantially reduce the time required to perform meta-analysis." They further state: "Our user study suggests that AUTOFOREST particularly when used in a human-in-the-loop capacity, can accelerate evidence synthesis substantially in comparison to the standard manual workflow." Importantly, "our results indicate that AI assistance may help bridge the gap between novice and expert performance; students using AUTOFOREST outperformed the manual baseline of domain experts and approached the accuracy of experts using the tool."
Risk of bias
Small sample size (N=8 participants) may not be representative; Limited diversity in participant expertise levels; Selection bias in choice of 18 Cochrane systematic reviews for evaluation; Limited to specific study designs (parallel-group trials); Model-dependent performance (uses Claude Sonnet 4.5; generalizability to other LLMs unclear); Potential bias from participants knowing they are evaluating an automated system; Small sample size (N=8) limiting generalizability; Limited participant diversity in expertise levels; Evaluation restricted to standard parallel-group trial designs; Model selection bias (Claude Sonnet 4.5 selected for strong numerical reasoning); Potential user training/familiarity bias favoring AUTOFOREST in human-in-the-loop condition; Small sample size (N=8) limits generalizability; Limited trial design diversity (only parallel-group trials tested); Potential selection bias in participant recruitment (convenience sample); Model-specific evaluation (only Claude Sonnet 4.5 tested); Hawthorne effect possible in user study setting
Limitations
- The authors acknowledge: "While the user study provides valuable insights, we only conducted the experiment with a sample of 8 participants
- A larger, more diverse cohort, with varied levels of expertise conducting systematic reviews, would be necessary to establish broader usability." Additionally, "the current evaluation focuses on standard parallel-group trial designs
- complex designs such as multi-arm or crossover trials have not yet been explicitly tested in this paper." Furthermore, "AUTOFOREST is designed to accelerate a specific segment of the evidence synthesis pipeline and does not address the full spectrum of tasks in a systematic review, such as literature searching, screening or study selection." Finally, "a current limitation is the absence of a visual feature to directly highlight this evidence in the source document."
Open questions raised
- Need for evaluation with larger, more diverse participant cohorts with varied levels of systematic review expertise
- Extension to complex trial designs (multi-arm, crossover trials)
- Implementation of visual feature to directly highlight evidence in source documents for easier verification
- Extension beyond the meta-analysis segment to address full systematic review pipeline (literature searching, screening, study selection)
- Need for larger, more diverse cohorts with varied levels of systematic review expertise
- Implementation of visual highlighting features to directly mark evidence in source documents
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations
- Co-designing AI Education Curriculum with Cross-Disciplinary High School TeachersBenjamin Xie · 2024 · 28 citations
- GAIDeT (Generative AI Delegation Taxonomy): A taxonomy for humans to delegate tasks to generative artificial intelligence in scientific research and publishingYana Suchikova · 2025 · 24 citations