12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI ‐Augmented Search for Systematic Reviews: A Comparative Analysis

Valerie Vera, Vedant Khandelwal, Kaushik Roy, Harshul Surana · Proceedings of the Association for Information Science and Technology · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
C
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/pra2.1290

Methodology & findings

Study design

Comparative empirical analysis study with two parts: Study 1 involved 12 academic librarians using NeuroLit Navigator with real-time research questions from patrons in the public health domain, with subsequent evaluation of the same queries on commercial systems (Scite, Consensus, Perplexity).

Main result

NeuroLit Navigator achieved a mean relevance rate of 65% compared to 47% for Consensus, 40% for Perplexity, and 38% for Scite in the computer science domain. "Pairwise comparisons using Cohen's d indicated large effect sizes between NeuroLit Navigator and each system (d = 1.20 for Consensus, 1.45 for Perplexity, and 1.55 for Scite)." In the public health domain, while Consensus achieved the highest raw relevance percentage (38%), "NeuroLit Navigator, while slightly lower in relevance (36%), outperformed all other systems in areas essential for systematic review queries: reproducibility, interpretability, and controlled vocabulary usage."

Research paradigm

Empirical/pragmatist (human-centered AI design with iterative evaluation)

Author conclusions

The authors conclude: "Our study supports the growing call for human-centered AI systems in systematic review workflows. Findings demonstrate that while AI can enhance searches, librarian oversight remains essential to ensure transparency, reproducibility, and alignment with domain-specific knowledge. As AI technologies continue to evolve, we argue for the design of AI systems that prioritize augmentation over human substitution, ensuring that AI serves to enhance, rather than replace, the expertise of librarians."

Risk of bias

Small sample size (12 librarians) limiting generalizability; Selection bias: librarians self-selected from professional organization listservs; Limited to top-ranked articles assessed, not full result set; Evaluation limited to two domains (public health, computer science); Inter-rater reliability, though acceptable (κ = 0.76), indicates some subjectivity; Licensing restrictions prevented integration with other databases, limiting scope; Only three commercial systems compared, not exhaustive of available AI tools; Small sample size (n=12 librarians) limits generalizability; Evaluation limited to top 3-5 results, not comprehensive recall assessment; Selection bias: librarians recruited via professional organization listservs, potentially self-selected; Limited to health sciences and computer science domains; Exemplar articles as 'gold standard' may bias NeuroLit Navigator retrieval; Potential bias toward neurosymbolic approach given NeuroLit Navigator was developed by study authors; Small sample size (12 librarians) may limit generalizability; Selection bias from regional recruitment via professional organization listservs; Limited evaluation scope (top 3-5 results only) may not capture full system performance; Evaluator bias possible despite inter-coder reliability calculation; Domain-specific evaluation may not generalize beyond public health and computer science; Potential funding/affiliation bias as NeuroLit Navigator was developed at authors' institution

Limitations

  • "First, our study included a small sample size of 12 librarians, which may not make our findings generalizable." Additionally, "our evaluation was also limited by the small number of top-ranked articles assessed, which provided only a narrow view of the system's overall retrieval effectiveness and thus did not fully assess recall." The authors also note that "we focused on literature in two domains" and "we initially developed NeuroLit Navigator for use with PubMed, which primarily includes coverage of health sciences literature," and "While integrating the system with other databases could broaden its applicability in other domains, we encountered licensing restrictions that precluded direct integration and use." Finally, "while our study compared three widely used commercial LLM-based systems (i.e., Scite, Consensus, and Perplexity), these systems do not encompass the full range of AI tools used in literature searching."

Open questions raised

  • Future studies should: (1) expand sample size to include users of varying experience levels; (2) assess broader sets of retrieved articles and incorporate recall-oriented metrics (mean average precision); (3) explore integration with other databases beyond PubMed to improve interdisciplinary applicability; (4) conduct comparative evaluation of additional commercial AI systems beyond Scite, Consensus, and Perplexity; (5) conduct usability testing with researchers and students to ensure accessibility for users with limited search expertise
  • Future work should: (1) expand sample size to include users of varying experience levels; (2) assess broader sets of retrieved articles and incorporate recall-oriented metrics (mean average precision); (3) address licensing restrictions to integrate with other databases beyond PubMed; (4) assess system adaptability in interdisciplinary fields with less standardized terminology; (5) conduct comparative evaluation of additional commercial AI systems beyond Scite, Consensus, and Perplexity; (6) conduct usability testing with researchers and students beyond librarians.
  • Generalizability beyond small librarian sample (12 participants); need expanded sample including users of varying experience levels
  • Limited evaluation scope (top 3-5 results) that does not fully assess recall; future research should assess broader set of retrieved articles and incorporate recall-oriented metrics (e.g., mean average precision)
  • Domain limitation: Study focused on only two domains (public health and computer science); future work should address integration challenges with other databases beyond PubMed to assess adaptability in interdisciplinary fields
  • Incomplete comparison: Study only compared three commercial LLM-based systems; comparative evaluation of additional commercial AI systems could offer further insights
Data: No publicly available datasets explicitly mentioned. The paper references using research questions from patrons and citations of exemplar articles provided by research teams, but does not state these are publicly available.; No datasets explicitly mentioned as publicly available. Study used research questions and exemplar articles from patron requests.; No external datasets explicitly mentioned as publicly available. The study used research questions and citations provided by librarians and institutional patrons, which appear to be proprietary institutional data.Code: NeuroLit Navigator is described as "an open-source, non-commercial solution" but no specific GitHub or repository URL is provided in the paper. The paper references Roy et al. (2024) for detailed system development documentation.; NeuroLit Navigator is described as "designed as an open-source, non-commercial solution" but no specific repository link provided in paper.; NeuroLit Navigator is described as "an open-source, non-commercial solution" but no specific GitHub or repository URL is provided in the paper. Referenced work by Roy et al. (2024) contains more detailed system description.Extracted from: pdfAgreement 39%

Explore related topics

Related papers