12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Pilot Project Leveraging Large Language Models for Automated Screening and Variable Extraction in Observational Studies

Scott A. Malec, Manjil Pradhan, Rajesh Upadhayaya, Vincent T. Metzger · medRxiv · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.64898/2026.06.13.26355589

Methodology & findings

Study design

Mixed-methods design incorporating: (1) automated LLM-based screening of MEDLINE database for observational studies; (2) retrieval-augmented generation (RAG) with domain-specific prompt engineering for variable extraction; (3) semantic chunking and cross-encoder re-ranking; (4) multi-stage extraction and classification pipeline; (5) manual annotation of reference standard by human experts; (6) semantic matching evaluation using cosine similarity and Hungarian algorithm; (7) performance evaluation with precision, recall, F1, and classification accuracy metrics across primary and secondary validation datasets..

Primary method

Design science research with iterative refinement; prototyping approach for both software artifact (SYNTHESISdbGPT) and knowledge artifact (SYNTHESISdb database)

Main result

The LLM-based screening process effectively filtered relevant papers from extensive databases, and "the pipeline performed best for categories with more frequent and clearly stated covariates, including Health Behaviors, Sociodemographics, Medication Use/Drug Exposure, and Physiological Measurement." Overall, the pipeline achieved 0.871 F1 score for extraction and 0.901 classification accuracy on the primary BPV→ADRD reference standard, and "the pipeline saves 85.7% or approximately 52 minutes of time per reference" compared to manual review.

Research paradigm

Positivist/empiricist with computational methods

Author conclusions

The authors conclude: "SYNTHESISdb represents a significant advancement in ADRD research, providing a comprehensive resource for researchers to identify and select appropriate confounding variables in observational studies. The database produces more reliable and comparable findings across studies by addressing inconsistencies in confounder selection and in patterns of omitted-variable bias." They further state that "the research presented in this paper has profound implications for epidemiology. Leveraging advanced computational approaches and standardizing confounder control improves the quality, efficiency, and consistency of research on ADRD."

Risk of bias

Selection bias from restricting to open-access papers in PubMed Central and Unpaywall; LLM hallucination risk in variable extraction; Inconsistencies in LLM outputs and decision-making; Semantic matching threshold (0.60 cosine similarity) may miss valid variable matches; Classification bias toward frequent, clearly-stated covariates over sparse categories; Selection bias: Restriction to open-access articles in PubMed Central and Unpaywall excludes proprietary/paywalled studies; Classification bias: LLM hallucinations producing variables not present in source text; Domain bias: Performance varies substantially by covariate category (sparse categories like Psychiatric/Medical History and Misc/Other show 0% precision/recall); Semantic similarity threshold bias: Cosine similarity threshold of 0.60 for variable matching may under- or over-count true matches; Coverage bias: Inconsistent reporting of confounders across studies may lead to underrepresentation of less frequently reported variables; Selection bias: restricted to open-access papers in PubMed Central and Unpaywall, potentially missing non-open-access literature; Language bias: likely restricted to English-language papers in MEDLINE; LLM hallucination bias: acknowledged risk of extracting variables not present in source text; Reference standard bias: manually annotated reference standard may reflect human rater biases and inter-rater disagreement; Temporal bias: corpus limited to papers from 2010 onwards (date specified as 'xxx (need to consult with Dr. Scott)'); Semantic matching threshold bias: cosine similarity threshold of 0.60 acknowledged as arbitrary and requiring justification; Category imbalance bias: unequal distribution of covariates across semantic categories affecting classification performance

Limitations

  • "One limitation of our project was the restricted scope of our analysis, which limited our data to open-access papers in PubMed Central and Unpaywall
  • This constraint limited the breadth and depth of our dataset." Additionally, "SYNTHESISdbGPT is not designed to distinguish the underlying analytic design of observational study" and "hallucination still remains an inevitable of LLMs," therefore "this pipeline should be used as an assistant to human expert rather than replacement for expert judgment." The authors also note "inconsistencies in the outputs produced and decision made by large language models
  • The inconsistencies includes, but are not limited to, extracting variables that are not present in the input papers, failing to adhere to the instructions provided."

Open questions raised

  • Need to distinguish underlying analytic designs of observational studies (e.g., predictive vs. causal models)
  • Identification and incorporation of underreported variables such as colliders and confounders affected by prior treatment (CAPTs)
  • Expansion to broader array of risk factors beyond BPV for ADRD
  • Integration of causal feature selection methods into systematic review process
  • Benchmarking SYNTHESISdb against other causal feature selection outputs using knowledge graphs and SemMedDB
  • Development of intuitive user interface with search features and data visualization tools
Data: SYNTHESISdb (Confounders for Modifiable Risk Factors of Alzheimer's Disease Database) - prototype confounder database; Accompanying Zenodo database (URL not fully specified in text); BPV→ADRD reference standard dataset (10 studies with manually annotated covariates); Secondary validation datasets including Education→Dementia corpus and Synergy validation datasets; SYNTHESISdb (Confounders for Modifiable Risk Factors of Alzheimer's Disease Database) - prototype database populated with extracted covariates; Reference-standard datasets with human-annotated variables (BPV→ADRD and secondary corpora including Education→Dementia, Bos et al., Brouwer et al., van der Valk et al., Wolters et al.); Zenodo database accompanying the paper (URL not explicitly provided in text); Reference standard datasets (manually annotated): BPV→ADRD corpus (10 studies, 109 reference covariates), Education→Dementia corpus (10 studies, 99 reference covariates), and Synergy validation datasets (Bos et al. 2018, Brouwer et al. 2019, van der Valk et al. 2021, Wolters et al. 2018); Zenodo database: https://zenodo.org/ (mentioned as accompanying database but specific URL not provided in text)Code: https://github.com/unmtransinfo/SYNTHESISdbGPT (Python program for SYNTHESISdbGPT pipeline); https://github.com/unmtransinfo/SYNTHESISdbGPT (Python program and pipeline code); GitHub: https://github.com/unmtransinfo/SYNTHESISdbGPTExtracted from: pdfAgreement 56%

Explore related topics

Related papers