Closing the screening gap but not the writing gap: a two-topic evaluation of LLMs for systematic reviews and meta-analyses in hepatology
Yuntao Zou, Iris Kim, Nan Gao, Michelle Li, Mi‐Ok Kim, Jin Ge · npj gut and liver. · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s44355-026-00068-w
Methodology & findings
Study design
The study employed a computational simulation approach using Large Language Models (LLMs) deployed on the UCSF 'Versa' platform running OpenAI's GPT-4o for screening and OpenAI-o1 for manuscript drafting.
Main result
The study found that LLMs achieved high sensitivity (82-93%) and specificity (95-99%) for article screening in both topics. Specifically, "the LLM achieved a sensitivity of 93% (95% confidence interval [CI]: 66-100%) and a specificity of 96% (95% CI: 94-97%)" for carvedilol in compensated cirrhosis, and "The LLM achieved a sensitivity of 86% (95% CI: 65-97%) and a specificity of 99% (95% CI: 98-100%)" for anticoagulation in PVT. However, "the 'Results' sections were prone to inaccuracies in extracting numerical data or generating spurious (e.g., 'hallucinated') findings." LLM-assisted screening reduced human time from 62 hours to 3 hours (carvedilol) and 30 hours to 2 hours (anticoagulation).
Research paradigm
Empiricist/Pragmatist
Author conclusions
The authors conclude: "our findings suggest that an LLM-assisted approach to systematic reviews in hepatology can yield major efficiency gains in literature screening, with performance metrics that approach those of expert human screeners. LLM-assisted drafting of 'synthetic' manuscripts, however, remains challenging due to data misrepresentation and hallucination errors." They further note that "refining automated pipelines that reduce laborious tasks while preserving scientific rigor may help accelerate evidence synthesis across the continuum of hepatology and other medical field research."
Risk of bias
Selection bias in the choice of two specific hepatology topics may limit generalizability; Institutional platform bias (UCSF Versa) - results may not generalize to other LLM deployments or open-source models; Performance ceiling effects - high sensitivity/specificity in screening may not translate to other clinical domains; Data extraction bias - LLM's difficulty with quantitative data extraction introduces systematic errors in results sections; Human validation bias - all borderline cases still required human adjudication, potentially introducing reviewer variability; Selection bias: Only two hepatology topics studied, limiting generalizability; Model bias: Experiments only on institution-hosted UCSF Versa platform with OpenAI GPT-4o and OpenAI-o1 models; performance may vary on other LLMs; Evaluation bias: Gold standard for screening was manual human review; consistency of human reviewers not reported; Data extraction bias: LLM hallucinations in numerical data extraction from PDFs and tables/figures; Circular evaluation risk: Using LLM-as-judge introduces potential circularity when LLM evaluates LLM-generated outputs; Selection bias in topic choice (only two hepatology topics evaluated); Institutional platform bias (UCSF Versa environment with specific institutional LLM deployment); Model selection bias (OpenAI models only; GPT-4o for screening, o1 for drafting); Evaluation bias in LLM-as-Judge framework (using same LLM family for evaluation as generation); Reference standard bias (human review as gold standard may itself contain errors)
Limitations
- The authors state: "Despite large efficiency gains, screening was not fully automated: all model-flagged inclusions and ambiguous records still required human adjudication, and full-text eligibility decisions remained manual." Additionally, "all experiments were run in the UCSF 'Versa' platform, which is a secure, institution-hosted environment chosen for copyright and privacy considerations
- We did not evaluate frontier public LLMs
- therefore, performances with larger, nondistilled models may differ." Furthermore, "our evaluation exclusively focused on two hepatology topics (carvedilol in compensated cirrhosis and anticoagulation in PVT)
- generalizability to other topics and domains may be limited." Most critically, "the 'synthetic' automated drafting pipeline struggled with numerical extraction from tables/figures and produced spurious ('hallucinated') results
- The 'synthetic' manuscripts, therefore, necessitated rigorous manual verification and were not immediately usable."
Open questions raised
- Refining data extraction pipelines, potentially via specialized modules that parse quantitative elements from figures or tables, to reduce hallucinations and enhance reproducibility
- Improving prompt engineering techniques, including topic-specific phrases and definitions, to boost screening performance and reduce false positives
- Design of integrated human-machine workflows wherein domain experts vet borderline machine screening decisions and verify numeric data in LLM-drafted results
- Benchmarking studies to evaluate cross-platform reproducibility across different LLM models
- Generalization of the approach to other clinical topics and domains beyond hepatology
- Refining data extraction pipelines with specialized modules for quantitative elements from figures/tables
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations