Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resources
Michael Gusenbauer, Neal Haddaway · Research Synthesis Methods · 2019
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/jrsm.1378
Methodology & findings
Study design
Systematic evaluation of 28 academic search systems using 27 test criteria measuring coverage, query functionality, filtering capabilities, citation searching, and reproducibility.
Main result
The study found that "only 14 of the 28 academic search systems examined are well-suited to evidence synthesis in the form of systematic reviews in that they met all necessary performance requirements." Substantially, "These 14 can be used as principal search systems: ACM Digital Library, BASE, ClinicalTrials.gov, Cochrane Library, EbscoHost (tested for ERIC, Medline, EconLit, CINHAL Plus, SporitsDiscus), OVID (tested for Embase, Embase Classic, PsychINFO), ProQuest (tested for Nursing & Allied Health Database, Public Health Database), PubMed, ScienceDirect, Scopus, TRID, Virtual Health Library, Web of Science (tested for Web of Science Core Collection, Medline), and Wiley Online Library." The study revealed "that half the search systems we examined have at least some issues with Boolean queries" and "substantially differences in the performance of search systems, meaning that their usability in systematic searches varies."
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude: "Selection of suitable search systems is essential for the outcome of evidence-synthesis research" and "Reviewers must consider the different functionalities offered, or not offered, when interacting with a given search system." They further state: "We hope our study helps to create awareness of the importance of search literacy. This study shows the limitations of such convenience. This research encourages responsible and knowledgeable researchers to be aware of search system qualities so they can then use the appropriate tool for the task at hand." The authors advocate that "search system operators—Open Access or not—review the capabilities and improve performance criteria where necessary" and recommend that "reviewers should always consult information specialists or librarians and enlist their support in designing systematic review search strategies."
Risk of bias
Selection bias: Search systems selected based on mentions in highly-cited recent papers may not represent the full landscape of available systems; Temporal bias: Tests performed in February-March 2019; systems update frequently and results may be outdated; Test design bias: Choice of specific threshold values (e.g., ≥25 terms, ≥5 field codes) could influence which systems pass/fail necessary criteria; Query bias: Specific test queries chosen (research, define, paper, Asterix, table, analysis) may not be representative of real systematic review searches; Observer bias: Manual assessment of interface features and documentation by authors; Temporal bias: Testing performed at specific time point (Feb-Mar 2019); systems update frequently and results may not reflect current state; Test design bias: Selection of 27 specific criteria may not capture all relevant performance dimensions; Threshold bias: Numeric thresholds (e.g., 25-term minimum) determined by authors based on sample review of 10 studies, may not be universally appropriate; Coverage bias: Sample of 28 search systems selected from 'hot papers' in Web of Science; may not represent all available systems; Language bias: Tests focused primarily on English language search capabilities; Selection bias in choice of search systems (included only those mentioned in ≥2 of 63 'hot papers'); Temporal bias (tests performed February-March 2019; systems update frequently); Threshold setting bias (authors acknowledge this but attempt to mitigate with systematic review guidance and best practices review); Limited generalizability (did not test all hundreds of bibliographic databases that exist); Threshold selection bias: While thresholds were based on evidence synthesis guidance and systematic review best practices, the choice of specific numeric thresholds (e.g., minimum 25-term Boolean strings, 1000 accessible hits) involves subjective judgment; Temporal bias: Testing conducted in February-March 2019 only; databases update frequently, potentially rendering results outdated; Testing design bias: Study tested user interaction with search systems rather than actual precision/recall performance or database completeness, potentially missing important performance dimensions; Selection bias: Included only 28 search systems based on frequency in highly-cited systematic reviews; other important systems may have been omitted; Cross-sectional analysis limitation: Cannot capture temporal performance variations or intermittent failures; Query generalization bias: Used specific test queries that may not represent all systematic review search requirements across disciplines; Selection bias in choice of 28 search systems (limited to those mentioned in top 0.1% cited systematic reviews and meta-analyses); Temporal bias: tests conducted February-March 2019; search system functionalities change over time; Threshold selection bias: numeric thresholds (e.g., 25-term minimum, 1000 records) may not be universally appropriate across disciplines; Testing environment bias: VPN and institutional access differences may not capture all real-world scenarios; Language bias: primarily tested English language functionality; Selection bias in database inclusion - systems were chosen based on citation in highly-cited systematic reviews, which may not represent all commonly-used systems; Temporal specificity - tests were conducted in February-March 2019; system updates may have changed functionality; Threshold selection bias - while authors base thresholds on guidance and best practices, the choice of specific numeric thresholds is somewhat subjective; Potential for incomplete testing - authors acknowledge other tests might exist that could reveal different performance profiles
Limitations
- The authors acknowledge that "while we took the greatest care to include a large evidence-based selection of meaningful methods to test the capacity of search systems, there may be other tests unknown to us that could be performed." Additionally, "from a theoretical standpoint, our study can only provide evidence that search systems behave incorrectly in failing to comply with certain test criteria
- It is impossible to be absolutely certain that a system that has proved successful in our specific tests would not fail in slightly different tests or under different circumstances." They also note that "as most of the search systems update not only their database, but also their search functionalities, the performance results tested in this study might change over time."
Open questions raised
- Authors identify the need for improved Open Access search systems suitable for systematic reviews, noting "it seems there is currently almost no getting around proprietary search systems if one attempts a rigorous systematic review." They call for database owners to implement improvements aligned with evidence synthesis requirements. Future work suggested includes updating assessments as systems evolve, and extending evaluation to previously unexamined search systems. Authors advocate for journal guidelines to mandate reporting of specific databases searched (including indices and access dates) to improve reproducibility and replicability.
- Comprehensive, systematic empirical assessment of search system performance for systematic reviews was lacking prior to this study
- Need for ongoing reassessment of search system advice in evidence-synthesis guidance (e.g., Campbell Collaboration recommendations regarding Boolean searching)
- Need for improved transparency in reporting exact databases and subscriptions accessed in systematic reviews
- Open Access search systems lack necessary functionalities to serve as principal resources for systematic searches across most disciplines
- Semantic search engines (Google Scholar, Microsoft Academic, Semantic Scholar) designed for exploratory rather than systematic search require different approaches
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations