Collaborating with large language models in literature screening for a systematic review of college students’ GenAI literacy
Wonchan Choi, Joyce Lee, Besiki Stvilia, Yan (1978 - ) Zhang, Hyerin Bak · Information Research an international electronic journal · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.47989/ir31iconf64265
Methodology & findings
Study design
This is a computational/simulation study evaluating large language models (LLMs) for literature screening in a systematic review.
Main result
The study found that "GPT-4o-mini Fewshot achieved the highest accuracy (.993), significantly outperforming its Zeroshot and Fewshot-Inc-Only" configuration, and "both GPT-5 Zeroshot and GPT-4o-mini Fewshot configurations performed effectively in the binary classification task of screening titles and abstracts, aligning with the goal of maximising sensitivity for inclusion without compromising balanced accuracy in the early phase of SLRs." Additionally, "LLMs helped correct human miscoding, highlighting their potential role as quality assurance tools rather than benchmarks against humans."
Research paradigm
pragmatist/mixed-methods (empirical evaluation with computational experimentation)
Author conclusions
"Our study demonstrates that LLM-assisted screening can reduce researchers' manual screening efforts while preserving high screening quality for SLRs. The error analysis identified methodological guidance for future LLM-assisted screening in SLRs. Lastly, LLMs helped correct human miscoding, highlighting their potential role as quality assurance tools rather than benchmarks against humans. Thus, the study offers methodological contributions." The study "provided a reusable template for recall-oriented, governed screening of a large, high-noise candidate pool."
Risk of bias
Small gold standard dataset (n=162, ~10% of corpus) may not fully represent all edge cases; Domain-specific corpus (GenAI literacy in higher education) limits generalizability; Gold standard developed by only two researchers, though intercoder reliability was high (Cohen's κ = .87); Model performance may be influenced by specific inclusion/exclusion criteria design; Potential selection bias in the initial literature search across five databases; Selection bias: Only 10% of corpus (162 articles) used for gold standard development; may not be representative; Domain specificity bias: Study focused on GenAI literacy in higher education only; findings may not generalize to other domains; Gold standard bias: Initial gold standard contained errors (2 human errors identified and corrected during analysis), potentially affecting model evaluation; Publication bias: Limited to peer-reviewed articles and conference proceedings; gray literature excluded; Temporal bias: Publication years limited to 2022-2025; excludes earlier foundational work; Geographic bias: U.S. college students only; international contexts excluded; Selection bias: Only 10% of the corpus (162 articles) used for gold standard development; limited to 2022-2025 publication years; Small sample size for evaluation (81 articles in evaluation set); Domain specificity: Corpus limited to GenAI literacy in U.S. higher education contexts, potentially limiting generalisability; Temporal bias: Time-limited corpus (first LLM release in 2022); Human coder bias: Although intercoder reliability was high (κ = .87), two human coders made errors that were corrected by LLM decisions (2 of 10 mismatches); Model selection bias: Only OpenAI GPT models tested; other LLMs (Bard, Claude, Llama) mentioned in literature but not evaluated
Limitations
- The authors state that "since the reported performance is based on a relatively small, domain-specific corpus, future work should evaluate these guidelines on larger datasets and in additional domains
- This will help assess generalisability and refine best practices." Additionally, the study was limited to "original research articles in journals and conference proceedings," with publication years restricted to "2022-2025" and data collection "outside the United States" as an exclusion criterion.
Open questions raised
- Future work should evaluate the guidelines on larger datasets and in additional domains to assess generalizability and refine best practices. The authors note that "practical guidelines for integrating [LLMs] into SLRs as collaborative partners remain underexplored" and more work is needed to show how LLM outputs can be governed, audited, and fed into analysis protocols.
- Future work should evaluate screening guidelines on larger datasets and in additional domains beyond GenAI literacy in higher education to assess generalisability and refine best practices. The authors identify need for refinement of eligibility rules to account for borderline cases, clearer differentiation instructions between empirical and secondary source materials, and investigation of confidence score thresholds and adjustments to reduce false negatives in smaller models.
- Need to evaluate guidelines on larger datasets and in additional domains beyond the current domain-specific corpus
- Assessment of generalisability and refinement of best practices for LLM-assisted screening in other SLR contexts
- Further investigation of methodological guidance for interdisciplinary topics, where concerns about reproducibility remain
- Exploration of how LLM outputs can be governed, audited, and fed into the analysis protocol as collaborative partners
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations