The Design and Evaluation of the Collaboration between Researchers and Generative AI for Systematic Literature Reviews
V. -K. Pham · Proceedings of the ... Annual Hawaii International Conference on System Sciences/Proceedings of the Annual Hawaii International Conference on System Sciences · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.24251/hicss.2025.564
Methodology & findings
Study design
Qualitative observational case study employing the design science paradigm.
Primary method
Design science paradigm with phenomenological analysis method (Moustakas 1994 approach)
Main result
The total time spent decreased dramatically from 570 minutes working without AI to 230 minutes in HCAI collaboration, "representing a 60% reduction in overall time consumption." The study found that "In HCAI collaboration mode, time allocation was changed when Copilot handled data analysis and synthesis processing in just 52 minutes. 76.02% of the time was utilized for the interactive and collaborative validation and triangulation discussion." Furthermore, "the collaboration also aided literature comprehension and facilitated robust triangulation," suggesting that "HCAI collaboration accelerates the research process and deepens and broadens the analysis and synthesis capabilities."
Research paradigm
design science
Author conclusions
"The observational data revealed author-triangulation as a prominent feature of the researcher's collaboration with Copilot during the literature review process, particularly in the synthesis task." The authors conclude that "the insights obtained from the collaborative process between a researcher and Copilot could serve as the preliminary results of user research for the interaction design of HCAI collaboration apps for literature review." They further note that "the benefits of HCAI collaboration in performing SLR activity were beyond efficiency and productivity. The findings show that partnership collaboration could enhance the overall quality of SLR results."
Risk of bias
Selection bias: Only one researcher conducted the SLR without GAI; single-author analysis may introduce subjective interpretation; Small sample size: Only 11 papers in the testing subset limits generalizability; Narrow domain focus: Qualitative systematic review only; findings may not apply to other SLR types; Researcher experience: First author had prior experience in literature analysis, which may not represent typical researcher capability; Prompt bias: Researcher-designed prompts may reflect researcher's own biases, which could be amplified in AI responses; Memory effects: AI responses influenced by previous exchanges within a conversation, affecting reproducibility; Lack of independent validation: No external reviewers independently validated results; Researcher bias in prompt design and follow-up questions potentially amplified in GAI responses; Limited independence of Copilot—trained to adjust answers according to input rather than critically challenge assumptions; Memory effects within conversation sessions leading to inconsistent reproducibility across sessions; Context-dependent interpretation challenges in qualitative data extraction; Small sample size (11 papers) limiting generalizability; Single domain focus (qualitative SLR) with potential discipline-specific biases; Single researcher conducting both individual and collaborative modes, potential researcher bias in comparison; Small sample size (11 papers from one subset); Limited to one type of SLR (qualitative systematic review); Potential influence of researcher's follow-up questions on GAI outputs; GAI model training effects on responses based on previous conversation history
Limitations
- "One limitation of this study was that merely examined the qualitative systematic review
- Because different types of SLR may vary in their focus, goals, coverage and other aspects, this narrow focus may limit the generalizability of our findings." Additionally, "To generalize the insights, it should increase the size of the literature across different academic domains to develop generalized literature review apps to facilitate HCAI collaboration." The authors also note that "due to Copilot's functional constraints, sustaining continuous collaboration from the analysis to the final synthesis procedures was impossible, which may lead to fragmented interactions and disjointed outputs."
Open questions raised
- Future studies should validate HCAI collaboration on a large set of literature across different academic domains
- Investigation needed into HCAI collaboration in different SLR types (beyond qualitative SLR)
- Need to test whether different SLR types affect HCAI collaboration performance
- Future research should increase sample size of literature across different academic domains to develop generalized literature review apps
- Enhanced technical integration needed: 'it can be enhanced by integrating the functions enabled by GAI in building the SLR app based on the experiences gained from this study'
- Need to develop SLR apps that sustain continuous collaboration from analysis to final synthesis procedures rather than fragmented interactions
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations