REFLECTIVE-TIAB: cost-effective prompt optimization for large language model-based title and abstract screening in literature reviews
Ákos Józwiák, Attila Imre, Judit Hagymásy, Judit Tittmann, Ãgnes Nagy, Sandor Kovacs et al. · Figshare · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.6084/m9.figshare.32756487.v1
Methodology & findings
Study design
Controlled empirical evaluation study.
Sample
N = 8520, 5 groups
Primary method
DSPy/GEPA reflective prompt optimizer with asymmetric loss function penalizing false negatives. Evaluation metrics included recall, accuracy, and F1 score. Specific statistical testing methods not detailed in abstract.
Main result
The study found that "Optimization improved recall across all LLMs (+3.7% to +37.1%)." Additionally, "Gemini 3 Flash Preview achieved the highest performance (91% accuracy, F1 81.6%) while costing 25-fold less per abstract than GPT-5.2, which ranked among the lowest-performing models." The research demonstrates that "A prompt optimized on a single open-source model is generalized to all nine without retraining" with "Total optimization cost was $6.36."
Reports effect sizes.
Research paradigm
empiricist/positivist
Author conclusions
The authors conclude that "REFLECTIVE-TIAB provides automated, model-transferable prompt optimization for literature screening at negligible cost." They further note that "Model price did not predict screening performance" and "The framework could substantially reduce screening workload while preserving comprehensiveness."
Risk of bias
Gold standard construction limited to 100 abstracts from inter-model disagreements only, potentially introducing selection bias; Optimization performed on single model (Llama 3.3 70B) before generalization testing; No mention of blinding in expert labeling process; Performance metrics varied substantially across models, suggesting potential model-specific bias; Gold standard construction from inter-model disagreements may introduce selection bias; Single model optimization (Llama 3.3 70B) may not generalize equally across all models; Limited to COPD exacerbation predictor search context—generalizability to other literature screening domains unclear
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations