Artificial intelligence for literature reviews: opportunities and challenges
F. J. Bolaños, Angelo A. Salatino, Francesco Osborne, Enrico Motta · Artificial Intelligence Review · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s10462-024-10902-3
Methodology & findings
Study design
Systematic literature review using PRISMA methodology.
Primary method
Literature review methodology combined with systematic software tool analysis using a comprehensive feature framework
Main result
This comprehensive review examined 21 leading SLR tools using a framework combining 23 traditional features with 11 AI features, and analysed 11 recent tools leveraging large language models. The study found that "the majority of SLR tools still depend on possibly outdated methodologies. This includes the use of basic classifiers, which are no longer considered state-of-the-art for text and document classification." The analysis revealed that only eight tools implement at least 70% of designated features, with DistillerSR, Nested Knowledge, Dextr, and ExaCT leading with 82% feature coverage.
Research paradigm
Positivist/empiricist
Author conclusions
The authors conclude that "this survey seeks to offer scholars a thorough insight into the application of Artificial Intelligence in this field, while also highlighting potential avenues for future research." They specifically emphasize that "we highlight three primary research challenges: integrating advanced AI solutions, such as large language models and knowledge graphs, improving usability, and developing a standardised evaluation framework. We also propose best practices to ensure more robust evaluations in terms of performance, usability, and transparency."
Risk of bias
Selection bias: Tools requiring user interfaces were prioritized, potentially excluding prototype systems; Temporal bias: Tools not updated in past 10 years were excluded, possibly missing still-functional tools; Repository bias: Reliance on Scopus, SLR Toolbox, and CRAN may miss tools not indexed in these sources; Subjective decision bias: Despite collaborative review, researcher judgment applied inclusion/exclusion criteria; Documentation bias: Tools evaluated based on available documentation; underdocumented tools may be mischaracterized; SLR Toolbox offline status: Repository went offline in March 2024, affecting reproducibility; Selection bias: Tools without user interfaces excluded; tools not updated in 10 years excluded; Language/terminology bias: Standardized vocabulary lacking in field; reliance on specific search terms; Information availability bias: Missing or incomplete information for some tools addressed through developer interviews; Researcher subjectivity: Manual evaluation of tools by researchers despite consensus procedures; Selection bias from exclusion criteria (tools without user interfaces, tools under maintenance, tools not updated in 10 years); Subjective decisions in applying inclusion/exclusion criteria despite collaborative review; Potential for missing tools not described using selected keywords or absent from targeted repositories; Dynamic nature of software development - tools evolve faster than documentation; Single institution affiliation of researchers may introduce perspective bias; Selection bias from tool identification strategy; Researcher subjectivity in applying inclusion/exclusion criteria; Exclusion of tools without user interfaces may bias toward more developed tools; Temporal bias: snapshot of rapidly evolving field; Selection bias: Tools excluded if not updated in 10 years or without user interfaces; potential for missing tools not described using selected keywords; Subjective assessment bias: Researchers made subjective decisions applying inclusion/exclusion criteria, though mitigated through collaborative review; Temporal bias: SLR Toolbox went offline March 2024; tools and their features are dynamic and may change post-publication; Geographic/language bias: Emphasis on English-language tools in major repositories; Reporting bias: Reliance on published documentation which may not reflect actual tool capabilities; Developer responsiveness bias: Information gaps filled through developer interviews, which may reflect developer availability and communication skills; Selection bias from exclusion criteria (tools requiring advanced technical expertise, unmaintained tools); Observer bias in tool feature assessment despite collaborative author review; Incomplete tool coverage due to keyword limitations in search strategy; Temporal bias as tools evolve rapidly during software development
Limitations
- The study acknowledges that "the selection of search engines and the formulation of search strings might have impacted the completeness of the tool identification
- It is possible that some tools were missed because they were not described using the selected keywords or were absent from the targeted repositories and previous surveys." Additionally, "the exclusion of tools that were either under maintenance and unavailable for evaluation or had not been updated in the past ten years" may have "restricted the generalisability of our findings." The authors also note that "AI remains a rapidly evolving field, and our feature set might not encapsulate all current and emerging dimensions." Furthermore, "the dynamic nature of software development" poses a persistent threat, as "software tools frequently evolve, acquiring new functionalities that may not be documented in the published literature."
Open questions raised
- Integration of advanced AI solutions: LLMs and knowledge graphs remain underutilized in current SLR tools; hallucination issues and lack of explainability present challenges
- Improved usability: User interaction features and human-in-the-loop design need enhancement
- Standardized evaluation framework: No consensus on evaluation metrics for AI-enhanced SLR tools; need for robust performance, usability, and transparency assessment
- Incorporation of recent NLP advances: Most tools still rely on outdated classifiers and Bag-of-Words methods rather than state-of-the-art embeddings
- Living reviews: Only 1 of 21 tools supports automated updating of literature with new relevant papers
- Snowballing automation: No tools provide automated snowballing functionality despite its importance in systematic reviews
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations