AI ‐Augmented Search for Systematic Reviews: A Comparative Analysis
Valerie Vera, Vedant Khandelwal, Kaushik Roy, Harshul Surana · Proceedings of the Association for Information Science and Technology · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/pra2.1290
Methodology & findings
Study design
Comparative empirical analysis study with two parts: Study 1 involved 12 academic librarians using NeuroLit Navigator with real-time research questions from patrons in the public health domain, with subsequent evaluation of the same queries on commercial systems (Scite, Consensus, Perplexity).
Main result
NeuroLit Navigator achieved a mean relevance rate of 65% compared to 47% for Consensus, 40% for Perplexity, and 38% for Scite in the computer science domain. "Pairwise comparisons using Cohen's d indicated large effect sizes between NeuroLit Navigator and each system (d = 1.20 for Consensus, 1.45 for Perplexity, and 1.55 for Scite)." In the public health domain, while Consensus achieved the highest raw relevance percentage (38%), "NeuroLit Navigator, while slightly lower in relevance (36%), outperformed all other systems in areas essential for systematic review queries: reproducibility, interpretability, and controlled vocabulary usage."
Research paradigm
Empirical/pragmatist (human-centered AI design with iterative evaluation)
Author conclusions
The authors conclude: "Our study supports the growing call for human-centered AI systems in systematic review workflows. Findings demonstrate that while AI can enhance searches, librarian oversight remains essential to ensure transparency, reproducibility, and alignment with domain-specific knowledge. As AI technologies continue to evolve, we argue for the design of AI systems that prioritize augmentation over human substitution, ensuring that AI serves to enhance, rather than replace, the expertise of librarians."
Risk of bias
Small sample size (12 librarians) limiting generalizability; Selection bias: librarians self-selected from professional organization listservs; Limited to top-ranked articles assessed, not full result set; Evaluation limited to two domains (public health, computer science); Inter-rater reliability, though acceptable (κ = 0.76), indicates some subjectivity; Licensing restrictions prevented integration with other databases, limiting scope; Only three commercial systems compared, not exhaustive of available AI tools; Small sample size (n=12 librarians) limits generalizability; Evaluation limited to top 3-5 results, not comprehensive recall assessment; Selection bias: librarians recruited via professional organization listservs, potentially self-selected; Limited to health sciences and computer science domains; Exemplar articles as 'gold standard' may bias NeuroLit Navigator retrieval; Potential bias toward neurosymbolic approach given NeuroLit Navigator was developed by study authors; Small sample size (12 librarians) may limit generalizability; Selection bias from regional recruitment via professional organization listservs; Limited evaluation scope (top 3-5 results only) may not capture full system performance; Evaluator bias possible despite inter-coder reliability calculation; Domain-specific evaluation may not generalize beyond public health and computer science; Potential funding/affiliation bias as NeuroLit Navigator was developed at authors' institution
Limitations
- "First, our study included a small sample size of 12 librarians, which may not make our findings generalizable." Additionally, "our evaluation was also limited by the small number of top-ranked articles assessed, which provided only a narrow view of the system's overall retrieval effectiveness and thus did not fully assess recall." The authors also note that "we focused on literature in two domains" and "we initially developed NeuroLit Navigator for use with PubMed, which primarily includes coverage of health sciences literature," and "While integrating the system with other databases could broaden its applicability in other domains, we encountered licensing restrictions that precluded direct integration and use." Finally, "while our study compared three widely used commercial LLM-based systems (i.e., Scite, Consensus, and Perplexity), these systems do not encompass the full range of AI tools used in literature searching."
Open questions raised
- Future studies should: (1) expand sample size to include users of varying experience levels; (2) assess broader sets of retrieved articles and incorporate recall-oriented metrics (mean average precision); (3) explore integration with other databases beyond PubMed to improve interdisciplinary applicability; (4) conduct comparative evaluation of additional commercial AI systems beyond Scite, Consensus, and Perplexity; (5) conduct usability testing with researchers and students to ensure accessibility for users with limited search expertise
- Future work should: (1) expand sample size to include users of varying experience levels; (2) assess broader sets of retrieved articles and incorporate recall-oriented metrics (mean average precision); (3) address licensing restrictions to integrate with other databases beyond PubMed; (4) assess system adaptability in interdisciplinary fields with less standardized terminology; (5) conduct comparative evaluation of additional commercial AI systems beyond Scite, Consensus, and Perplexity; (6) conduct usability testing with researchers and students beyond librarians.
- Generalizability beyond small librarian sample (12 participants); need expanded sample including users of varying experience levels
- Limited evaluation scope (top 3-5 results) that does not fully assess recall; future research should assess broader set of retrieved articles and incorporate recall-oriented metrics (e.g., mean average precision)
- Domain limitation: Study focused on only two domains (public health and computer science); future work should address integration challenges with other databases beyond PubMed to assess adaptability in interdisciplinary fields
- Incomplete comparison: Study only compared three commercial LLM-based systems; comparative evaluation of additional commercial AI systems could offer further insights
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations