Automating in High-Expertise, Low-Label Environments: Evidence-Based Medicine by Expert-Augmented Few-Shot Learning
Rong Liu, Jingjing Li, Marko Zivkovic, Ahmed Abbasi · MIS Quarterly · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.25300/misq/2024/18573
Methodology & findings
Study design
Computational design science research combining theory-guided design (Compositionality Theory), artifact development, benchmarking analysis against multiple baseline models (traditional classifiers, LLMs, semi-supervised learning, few-shot learning methods), ablation studies, and mixed-methods evaluation (quantitative metrics on three datasets + qualitative expert interviews + downstream case study integration with SR company workflow)..
Primary method
Theory-guided design science (grounded in Compositionality Theory from cognitive science). Systematic requirements analysis based on task complexity identification, artifact instantiation and evaluation following computational design genre (Padmanabhan et al., 2022; Rai, 2017), with socio-technical evaluation encompassing benchmarking, ablation analysis, downstream case study integration, and qualitative expert consultation.
Main result
FastSR achieves an average F1@3 of 70% for sentence classification, exceeding SVM by 40%, GPT-4 by 33%, CNN by 11%, fine-tuned BioBERT by 8%, and ProtoNet by 7%. For sequence tagging, "FastSR obtained the highest performance, with an F1word@3 (F1BERT@3) of 61% (67%), about 15% higher than LSTM+CRF and 4% than ProtoNER." The study demonstrates that "FastSR significantly reduces time and cost due to its efficient automation and superior performance compared to benchmarking models. With single screening, FastSR achieves a 65% of time reduction relative to a double-screening manual SR, leading to a cost saving of $73,500 for a project with 550 articles."
Research paradigm
Design Science Research / Computational Design
Author conclusions
"In conclusion, we propose a customized FSL solution with several novel extensions to integrate human annotation logic into a computational framework that requires minimal labeling effort. Our comprehensive multi-faceted evaluation showcases FastSR's advantages in data extraction and SR automation compared to baseline models. Our results have important implications for designing computational artifacts in high-expertise, low-label environments." The authors further conclude that "FastSR significantly reduces time and cost due to its efficient automation and superior performance compared to benchmarking models. With single screening, FastSR achieves a 65% of time reduction relative to a double-screening manual SR, leading to a cost saving of $73,500 for a project with 550 articles."
Risk of bias
Selection bias: Limited to medical SR domain; generalizability to other domains not tested; Data contamination risk: ALMs (GPT-4) potentially contaminated by overlapping pre-training data with EBM-NLP dataset; Labeling bias: PICO element definitions disease-specific and evolving; depends on expert consensus; Class imbalance: Multiple PICO elements with varying prevalence; may bias model learning; Sample selection bias: WD dataset consists of only 222 labeled sentences from a single pharmaceutical company's 1,500-hour project; limited generalizability to other disease domains; Limited labeled data: FSL approach with K=5 support samples may not capture full complexity of PICO element variation; Annotation quality: Kappa scores (0.85 for WD, 0.87 for COVID, 0.67 for EBM-NLP) indicate moderate inter-rater agreement, particularly lower for public dataset; Data contamination risk: Authors acknowledge GPT-4 performance gap reduced significantly on public EBM-NLP dataset, suggesting potential pre-training data overlap; Class imbalance: Countered with one-vs-rest approach but may still affect minority class performance; Evaluation context bias: Authors had access to SR company domain experts for WD/COVID but public dataset (EBM-NLP) was crowd-labeled with lower agreement; Potential domain overfitting to medical literature; Limited to predefined PICO classes; Data contamination risk in LLM benchmarking (GPT-4); Inter-annotator agreement variability across datasets (Kappa: 0.85-0.87 for WD/COVID, 0.67 for EBM-NLP)
Limitations
- "First, we have chosen to concentrate on demonstrating generalizability within the medical SR context, leaving the exploration of its application to other domains for future work
- Second, FastSR assumes predefined PICO classes established by medical professionals, but future work could focus on developing methods to automatically identify new, disease-specific PICO classes
- Third, PICO classes often have a hierarchical structure..
- Future work could explore prototype designs that might better reflect this hierarchy to accommodate small or rare subclasses
- Fourth, while ALMs underperform in data extraction tasks compared to BERT-based models, they might be useful for downstream tasks like evidence synthesis."
Open questions raised
- Generalizability beyond medical SR domain to other high-expertise, low-label contexts
- Automatic identification of new, disease-specific PICO classes
- Support for hierarchical PICO class structures
- Integration of LLMs for downstream evidence synthesis tasks
- Replacement of individual ML components as new technologies emerge
- Application beyond medical SR context to other high-expertise, low-label domains
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations