12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Prompt engineering of large language models for paper screening in medical meta-analyses and systematic reviews: A prospective comparative study

Salma AS Abosabie, Max Dittmer, Elise Wolf, Sara A. Abosabie, Clara Behnke, Felix Baier et al. · Research Synthesis Methods · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
2/4
Quality (LMQS)
E
Evidence
4
Citations
45.02
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1017/rsm.2026.10093

Methodology & findings

Study design

Prospective comparative experimental study with systematically varied prompt engineering across three large language models (LLMs) and four gold-standard meta-analyses/systematic reviews.

Sample

N = 12360, 10 groups

Primary method

Linear mixed-effects models (lme4 package version 1.1-37 in R version 4.4.3) with random intercepts for MA/SRs as grouping variable and fixed factors for LLM type, screening section (title vs. abstract), and prompt characteristics. Univariate analyses conducted separately for F1 (primary outcome) and recall/sensitivity (secondary outcome). Analysis of variance (ANOVA) using aov() function in base R to compare differences between individual prompts and across MA/SRs. p-values calculated via Satterthwaite approximation using lmerTest package (version 3.1-3). Dummy coding applied for categorical factors.

Main result

Over 12,360 pipeline runs across four MA/SRs and three LLMs, "the 515 prompts achieved averages of recall/sensitivity = 83.6 ± 17.0%, precision = 18.5 ± 15.6%, F1 = 27.6 ± 17.2%, and accuracy = 61.1 ± 11.0%". The study found that "F1 increased with reference to the methods of the screened papers in the prompt (β = 0.78%; 95% confidence interval (CI) = 0.38%-1.18%; p < 0.001)" and "F1 also increased with each additional level of information about the MA/SR provided, with most improvement when including the MA/SR's inclusion and exclusion criteria (β = 8.05%; CI = 7.68%-8.41%; p < 0.001)".

Reports effect sizes and confidence intervals.

Research paradigm

positivist/empirical

Author conclusions

The authors conclude that "Human sensitivity in abstract screening is matched by that of LLMs (86.6%) but with higher specificity in human screening", and "In the future, LLMs could function as independent reviewers for paper screening, especially for early-stage title and abstract screening, reducing second-reviewer workload or serving as independent third reviewers." They also state that "We identified prompt engineering principles for LLM-based paper screening, highlighting a previously unaddressed recall/sensitivity-precision/specificity trade-off, with prompts including detailed selection criteria achieving higher F1, precision, and specificity at the expense of recall/sensitivity." The authors note that "Progress in developing user-friendly tools that can access and screen data from multiple literature databases, as well as implementation studies on performance are still needed until LLMs can be used reliably in paper screening for peer-reviewed journal-published MA/SRs."

Risk of bias

Modeling of human screening as gold standard may introduce measurement error given documented limitations in human selection reliability; Selection of MA/SRs from Q1-ranked journals may introduce publication bias; Testing only open-source models may limit generalizability to commercial models (though authors note prior research suggests comparable performance); Limited to PubMed database only; Parallel rather than sequential title/abstract screening workflow differs from conventional human practice; Gold-standard comparison bias: Human screening used as reference, but prior research shows limitations in human selection decisions (false positives and false negatives reflect disagreement with human screening rather than objective ground truth); Selection bias in MA/SR selection: Purposive rather than exhaustive sampling of MA/SRs; only articles from Q1-ranked journals and published after July 1, 2024 were selected; Database limitation: Records drawn only from PubMed, potentially missing studies indexed in other databases; Model selection bias: Only open-source LLMs tested; commercial models not explicitly tested for prompt engineering effects; Generalizability concerns: Only four MA/SRs tested across four medical fields; results may not generalize to other fields or study types; Information availability bias: Study selection based on freely available full-text and complete data on search queries, included studies, and estimates; Temporal bias: MA/SRs selected only from after LLM training cut-offs to avoid training data overlap, but this may not apply to all training corpora; Selection bias: MA/SRs were selected from Q1-ranked journals with specific methodological criteria, which may not represent all MA/SRs; Gold standard bias: reliance on human reviewer decisions as truth standard, which may contain errors; Model selection bias: only open-source models tested, not commercial models like GPT-4; Database limitation: only PubMed searched; other biomedical databases not included; Publication date filter: only MA/SRs published after July 1, 2024 (after LLM training cutoffs) included; Language bias: English language only; Sample size constraint: studies limited to 10-50 included studies per MA/SR

Limitations

  • The authors identified six major limitations: First, "we modeled human inclusion/exclusion decisions as the gold-standard comparison, but prior research has shown limitations in human selection"
  • Second, "due to financial constraints, we tested only open-source models, although models with API call costs can also be used with the MA-LLM pipeline"
  • Third, "we limited screening to titles and abstracts
  • Full-text articles are often not publicly available and current LLMs frequently show 'lost-in-the-middle' behavior"
  • Fourth, "we drew records only from PubMed
  • PubMed is one of the largest biomedical databases (>38 million records in 2025), supporting scale and reproducibility, but future work should extend to additional databases"

Open questions raised

  • Lack of prospective examination of specific prompt characteristics on paper selection by LLMs compared to gold-standard MA/SRs prior to this work
  • Need for explicit examination of prompt design impact on commercial LLM models
  • Future work should extend beyond PubMed to additional literature databases
  • Research should examine LLM screening performance against more MA/SRs within and beyond medicine
  • Full-text screening and gradual replacement of traditional stepwise screening sequence may become feasible with increasing LLM capabilities
  • Progress needed in developing user-friendly tools that can access and screen data from multiple literature databases
Data: MA-LLM pipeline with all 515 prompts; MA-LLM pipeline and all prompts; PubMed database; MA-LLM pipeline and all 515 prompts available at https://zenodo.org/records/19097181; All data extracted from PubMed/NIH database via BioC APICode: MA-LLM pipeline; Zenodo: https://zenodo.org/records/19097181; MA-LLM Python-based pipeline available at https://zenodo.org/records/19097181Extracted from: pdfAgreement 42%

Explore related topics

Related papers