Transforming evidence synthesis: A systematic review of the evolution of automated meta-analysis in the age of AI
Lingbo Li, Anuradha Mathrani, Teo Sušnjak · Research Synthesis Methods · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1017/rsm.2025.10065
Methodology & findings
Study design
PRISMA-compliant systematic review with structured search across five databases (PubMed, Scopus, Google Scholar, IEEE Xplore, ACM Digital Library) from 2006–2024, followed by bidirectional citation chaining.
Main result
This systematic review of 61 studies reveals that automated meta-analysis (AMA) has exhibited "a predominant focus on automating data processing (52.5%), such as extraction and statistical modeling, while only 16.4% address advanced synthesis stages. Just one study (approximately 2%) explored preliminary full-process automation, highlighting a critical gap that limits AMA's capacity for comprehensive synthesis." The study found that "AMA has exhibited distinct implementation patterns and varying degrees of effectiveness in actually improving efficiency, scalability, and reproducibility" across medical (67.2%) and non-medical (32.8%) applications.
Research paradigm
Critical realism with technology adoption and task-technology fit frameworks
Author conclusions
The authors conclude that "As AI systems advance in reasoning and contextual understanding, addressing these gaps is now imperative. Future efforts must focus on bridging automation across all MA stages, refining interpretability, and ensuring methodological robustness to fully realize AMA's potential for scalable, domain-agnostic synthesis." They emphasize that "achieving seamless, end-to-end automation remains an open challenge" and that a "comprehensive review of AMA progress across domains is urgently needed to harness AI's full potential and address persistent limitations in evidence synthesis automation."
Risk of bias
Selection bias: Database-centric search strategy, mitigated by snowball/citation chaining methods; Language bias: Only English-language publications included; Publication bias: Restriction to peer-reviewed articles may exclude gray literature and negative findings; Reviewer bias: One reviewer (L.L.) conducted initial screening, though two additional reviewers independently verified results; Time period bias: 2006–2024 timeframe may exclude earlier foundational work or recent developments post-August 2024; Domain representation bias: Medical applications comprise 67.2% of dataset, potentially overrepresenting medical-specific automation challenges; Language bias (English-only publications); Publication bias (peer-reviewed articles only; gray literature excluded except via snowballing); Database selection bias (five databases used; potential for platform-specific indexing differences); Title/abstract screening bias (single reviewer initial screening, though confirmed by two additional reviewers); Inclusion criteria bias (studies required ≥4 pages with technical detail; shorter works excluded despite potential relevance); Language bias: English-language publications only; Publication bias: Peer-reviewed journals, conference papers, and preprints prioritized; gray literature and opinions excluded; Selection bias: Minimum four-page length requirement may exclude brief but methodologically rigorous studies; Reviewer bias: Initial screening by single reviewer (L.L.) with verification by two additional reviewers, though consensus-based resolution reduces but does not eliminate risk; Database coverage bias: Search strategy optimized for 'meta-analysis' terms, potentially missing relevant automation papers using alternative terminology
Limitations
- The review acknowledges several methodological limitations: "Just one study (approximately 2%) explored preliminary full-process automation, highlighting a critical gap that limits AMA's capacity for comprehensive synthesis." Additionally, "Despite recent breakthroughs in large language models and advanced AI, their integration into statistical modeling and higher-order synthesis, such as heterogeneity assessment and bias evaluation, remains underdeveloped." The authors note that "Integration challenges remain, including workflow fragmentation, analytical limitations, and interoperability barriers, hindering full automation." Furthermore, the review's scope is limited by the exclusion of records with fewer than four pages and the restriction to English-language publications from 2006–2024.
Open questions raised
- Lack of comprehensive end-to-end automation: Only one study (~2%) explored full-process automation, highlighting the need for integrated frameworks across all meta-analytic stages
- Underdeveloped higher-order synthesis automation: Integration of AI into heterogeneity assessment, bias evaluation, and sensitivity analysis remains limited (only 16.4% of studies address advanced synthesis stages)
- Lack of large language model integration in statistical modeling: Despite breakthroughs in LLMs, their application to complex statistical tasks remains underdeveloped
- Workflow fragmentation and interoperability barriers: Integration challenges persist across different automation tools and stages
- Limited cross-domain generalizability: Automation tools often lack transferability across medical and non-medical domains
- Interpretability and transparency gaps: Automated systems require improved explainability for clinical and research contexts
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations
- A Systematic Literature Review of Retrieval-Augmented Generation: Techniques, Metrics, and Challenges2025 · 13 citations
- The Interplay of Learning Analytics and Artificial IntelligenceJelena Jovanović · 2024 · 4 citations