12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AutoForest: Automatically Generating Forest Plots from Biomedical Studies with End-to-End Evidence Extraction and Synthesis

Massimiliano Pronesti, Angelo Miculescu, Mohsin Kapdi, Paul Flanagan, Oisín Redmond, Joao Bettencourt-Silva et al. · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Mixed-methods evaluation combining automatic evaluations on 32 forest plots from 18 Cochrane systematic reviews (56 studies total) with a controlled within-subjects user study involving 4 clinical domain experts and 4 graduate students.

Sample

N = 8, 2 groups

Primary method

Wilcoxon signed-rank tests (non-parametric alternative to paired t-tests); R meta package for statistical synthesis (Balduzzi et al., 2019); quantitative metrics for table detection, structure parsing, and ICO suggestion evaluation; precision, recall, and hallucination rate calculations; accuracy and edit rate calculations.

Main result

The study found that AUTOFOREST significantly outperforms manual workflows across all metrics. "For RQ1, the tool nearly halved the time required to complete a forest plot for both groups (p < 0.001)." Additionally, "the fully automated version ('AUTOFOREST only') achieved over 80% accuracy in data extraction, a substantial improvement over the manual baselines. Human-in-the-loop edits further refined these results, reaching a peak accuracy for experts of 90.2% for data extraction (p = 0.013) and 79.2% for RoB (p = 0.050)."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical (system development and evaluation)

Author conclusions

The authors conclude: "AUTOFOREST addresses this bottleneck by integrating document parsing, automated ICO suggestions, evidence extraction, risk of bias assessment and reasoning to substantially reduce the time required to perform meta-analysis." They further state: "Our user study suggests that AUTOFOREST particularly when used in a human-in-the-loop capacity, can accelerate evidence synthesis substantially in comparison to the standard manual workflow." Importantly, "our results indicate that AI assistance may help bridge the gap between novice and expert performance; students using AUTOFOREST outperformed the manual baseline of domain experts and approached the accuracy of experts using the tool."

Risk of bias

Small sample size (N=8 participants) may not be representative; Limited diversity in participant expertise levels; Selection bias in choice of 18 Cochrane systematic reviews for evaluation; Limited to specific study designs (parallel-group trials); Model-dependent performance (uses Claude Sonnet 4.5; generalizability to other LLMs unclear); Potential bias from participants knowing they are evaluating an automated system; Small sample size (N=8) limiting generalizability; Limited participant diversity in expertise levels; Evaluation restricted to standard parallel-group trial designs; Model selection bias (Claude Sonnet 4.5 selected for strong numerical reasoning); Potential user training/familiarity bias favoring AUTOFOREST in human-in-the-loop condition; Small sample size (N=8) limits generalizability; Limited trial design diversity (only parallel-group trials tested); Potential selection bias in participant recruitment (convenience sample); Model-specific evaluation (only Claude Sonnet 4.5 tested); Hawthorne effect possible in user study setting

Limitations

  • The authors acknowledge: "While the user study provides valuable insights, we only conducted the experiment with a sample of 8 participants
  • A larger, more diverse cohort, with varied levels of expertise conducting systematic reviews, would be necessary to establish broader usability." Additionally, "the current evaluation focuses on standard parallel-group trial designs
  • complex designs such as multi-arm or crossover trials have not yet been explicitly tested in this paper." Furthermore, "AUTOFOREST is designed to accelerate a specific segment of the evidence synthesis pipeline and does not address the full spectrum of tasks in a systematic review, such as literature searching, screening or study selection." Finally, "a current limitation is the absence of a visual feature to directly highlight this evidence in the source document."

Open questions raised

  • Need for evaluation with larger, more diverse participant cohorts with varied levels of systematic review expertise
  • Extension to complex trial designs (multi-arm, crossover trials)
  • Implementation of visual feature to directly highlight evidence in source documents for easier verification
  • Extension beyond the meta-analysis segment to address full systematic review pipeline (literature searching, screening, study selection)
  • Need for larger, more diverse cohorts with varied levels of systematic review expertise
  • Implementation of visual highlighting features to directly mark evidence in source documents
Data: 32 forest plots from 18 Cochrane systematic reviews with 56 included studies (used in evaluation, availability not explicitly stated); 206 tables from the 56 studies (used for document-to-structure conversion evaluation, availability not explicitly stated); 32 forest plots from 18 Cochrane systematic reviews (56 included studies) - availability not explicitly stated; 32 forest plots from 18 Cochrane systematic reviews (56 included studies) - used for evaluationCode: https://itu-nlp.github.io/projects/autoforest (project website); https://www.youtube.com/watch?v=R6ei97fOyXQ (system demonstration video); https://itu-nlp.github.io/projects/autoforest; https://www.youtube.com/watch?v=R6ei97fOyXQ (demo video); https://www.youtube.com/watch?v=R6ei97fOyXQExtracted from: pdfAgreement 53%

Explore related topics

Related papers