Exploring the Use of ChatGPT for a Systematic Literature Review: a Design-Based Research
Qian Huang, Qiyun Wang · arXiv (Cornell University) · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2409.17426
Methodology & findings
Study design
Design-based research with iterative development.
Primary method
Design-based research with iterative refinement across three rounds
Main result
The study confirms that "ChatGPT can be a helpful tool for doing the SLR" and demonstrates that through iterative prompt refinement and careful specification of analytical frameworks, ChatGPT could generate similar outcomes to those presented in the original literature review. However, the findings also reveal significant limitations: "ChatGPT captured related information from the entire document when it was reading the paper in the PDF document. They included the challenges and strategies mentioned in the Literature section. It was supposed to be from the Results section only."
Research paradigm
Pragmatist/Design-based research
Author conclusions
The authors conclude: "ChatGPT can serve as a Research Assistant(s) in helping with reading, analyzing and summarizing existing studies and the generated results can be used as a reference to allow researchers to have a quick understanding of the research in a certain area. The design principles summarized from this study can guide researchers in generating reliable review results." They emphasize that while ChatGPT has limitations, it can enhance efficiency and accuracy when used with appropriate human oversight and structured guidance.
Risk of bias
No independent validation of ChatGPT outputs by blinded reviewers; Single AI tool tested (ChatGPT-4.0 only) - generalizability unclear; Comparison limited to one prior study (Wang & Huang 2023); No inter-rater reliability assessment between human and AI; Selection of papers predetermined from original review - no assessment of ChatGPT's screening capability; Limited sample size (33 papers) in specific domain (Blended Synchronous Learning); Selection bias: only 33 papers from a specific domain (blended synchronous learning) tested; Single AI model tested (ChatGPT-4.0 only); Comparison against single original literature review without independent verification; Researcher bias in prompt engineering and interpretation of results; Limited scope to one research topic area; Selection bias: papers were pre-screened by human researchers rather than systematically retrieved by ChatGPT; Accuracy bias: ChatGPT conflated information from literature review sections with findings sections; Identification bias: ChatGPT failed to accurately identify research methods (design-based research) and theoretical frameworks; Comparability bias: comparison based on single original review (OLR) without multiple independent reviews for triangulation
Open questions raised
- The authors note that "To the best of our knowledge, no empirical study has been conducted to explore how to use GAI like ChatGPT to do an SLR. Some existing conceptual papers have presented the idea of using ChatGPT to do literature reviews but no one has explored and published papers on how to do it in practice." Future research could explore ChatGPT's capability for autonomous paper screening and comparison across multiple AI tools.
- The authors note that 'To the best of our knowledge, no empirical study has been conducted to explore how to use GAI like ChatGPT to do an SLR. Some existing conceptual papers have presented the idea of using ChatGPT to do literature reviews but no one has explored and published papers on how to do it in practice.' The study addresses this gap by providing empirical evidence of ChatGPT's capabilities and limitations in systematic literature review processes.
- The authors note that "no empirical study has been conducted to explore how to use GAI like ChatGPT to do an SLR" and that "some existing conceptual papers have presented the idea of using ChatGPT to do literature reviews but no one has explored and published papers on how to do it in practice."
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations