A Pilot Project Leveraging Large Language Models for Automated Screening and Variable Extraction in Observational Studies
Scott A. Malec, Manjil Pradhan, Rajesh Upadhayaya, Vincent T. Metzger · medRxiv · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.06.13.26355589
Methodology & findings
Study design
Mixed-methods design incorporating: (1) automated LLM-based screening of MEDLINE database for observational studies; (2) retrieval-augmented generation (RAG) with domain-specific prompt engineering for variable extraction; (3) semantic chunking and cross-encoder re-ranking; (4) multi-stage extraction and classification pipeline; (5) manual annotation of reference standard by human experts; (6) semantic matching evaluation using cosine similarity and Hungarian algorithm; (7) performance evaluation with precision, recall, F1, and classification accuracy metrics across primary and secondary validation datasets..
Primary method
Design science research with iterative refinement; prototyping approach for both software artifact (SYNTHESISdbGPT) and knowledge artifact (SYNTHESISdb database)
Main result
The LLM-based screening process effectively filtered relevant papers from extensive databases, and "the pipeline performed best for categories with more frequent and clearly stated covariates, including Health Behaviors, Sociodemographics, Medication Use/Drug Exposure, and Physiological Measurement." Overall, the pipeline achieved 0.871 F1 score for extraction and 0.901 classification accuracy on the primary BPV→ADRD reference standard, and "the pipeline saves 85.7% or approximately 52 minutes of time per reference" compared to manual review.
Research paradigm
Positivist/empiricist with computational methods
Author conclusions
The authors conclude: "SYNTHESISdb represents a significant advancement in ADRD research, providing a comprehensive resource for researchers to identify and select appropriate confounding variables in observational studies. The database produces more reliable and comparable findings across studies by addressing inconsistencies in confounder selection and in patterns of omitted-variable bias." They further state that "the research presented in this paper has profound implications for epidemiology. Leveraging advanced computational approaches and standardizing confounder control improves the quality, efficiency, and consistency of research on ADRD."
Risk of bias
Selection bias from restricting to open-access papers in PubMed Central and Unpaywall; LLM hallucination risk in variable extraction; Inconsistencies in LLM outputs and decision-making; Semantic matching threshold (0.60 cosine similarity) may miss valid variable matches; Classification bias toward frequent, clearly-stated covariates over sparse categories; Selection bias: Restriction to open-access articles in PubMed Central and Unpaywall excludes proprietary/paywalled studies; Classification bias: LLM hallucinations producing variables not present in source text; Domain bias: Performance varies substantially by covariate category (sparse categories like Psychiatric/Medical History and Misc/Other show 0% precision/recall); Semantic similarity threshold bias: Cosine similarity threshold of 0.60 for variable matching may under- or over-count true matches; Coverage bias: Inconsistent reporting of confounders across studies may lead to underrepresentation of less frequently reported variables; Selection bias: restricted to open-access papers in PubMed Central and Unpaywall, potentially missing non-open-access literature; Language bias: likely restricted to English-language papers in MEDLINE; LLM hallucination bias: acknowledged risk of extracting variables not present in source text; Reference standard bias: manually annotated reference standard may reflect human rater biases and inter-rater disagreement; Temporal bias: corpus limited to papers from 2010 onwards (date specified as 'xxx (need to consult with Dr. Scott)'); Semantic matching threshold bias: cosine similarity threshold of 0.60 acknowledged as arbitrary and requiring justification; Category imbalance bias: unequal distribution of covariates across semantic categories affecting classification performance
Limitations
- "One limitation of our project was the restricted scope of our analysis, which limited our data to open-access papers in PubMed Central and Unpaywall
- This constraint limited the breadth and depth of our dataset." Additionally, "SYNTHESISdbGPT is not designed to distinguish the underlying analytic design of observational study" and "hallucination still remains an inevitable of LLMs," therefore "this pipeline should be used as an assistant to human expert rather than replacement for expert judgment." The authors also note "inconsistencies in the outputs produced and decision made by large language models
- The inconsistencies includes, but are not limited to, extracting variables that are not present in the input papers, failing to adhere to the instructions provided."
Open questions raised
- Need to distinguish underlying analytic designs of observational studies (e.g., predictive vs. causal models)
- Identification and incorporation of underreported variables such as colliders and confounders affected by prior treatment (CAPTs)
- Expansion to broader array of risk factors beyond BPV for ADRD
- Integration of causal feature selection methods into systematic review process
- Benchmarking SYNTHESISdb against other causal feature selection outputs using knowledge graphs and SemMedDB
- Development of intuitive user interface with search features and data visualization tools
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations