BibliZap: An exploratory evaluation of an automated multi-level citation searching tool for systematic and rapid reviews
Research Synthesis Methods · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1017/rsm.2026.10079
Methodology & findings
Study design
Retrospective evaluation study using a gold-standard corpus derived from 66 published systematic reviews (2012-2021) from high-impact journals.
Primary method
Design science research with artifact development and evaluation using retrospective gold-standard analysis
Main result
Supplementing standard PubMed searches with multi-level citation searching substantially improves recall, as "supplementing standard PubMed searches with multi-level citation searching substantially improves recall, increasing average sensitivity from 75% to 97%." BibliZap proved particularly effective, "retrieving approximately 73% of these missed records within the first 2,000 ranked results," demonstrating that "more than half of the relevant studies missed by PubMed were retrieved within the top 500 BibliZap-ranked articles."
Research paradigm
Positivist/empiricist - quantitative evaluation of tool performance against gold standard
Author conclusions
The authors conclude: "BibliZap is a transparent, configurable, and automated citation searching tool that improves sensitivity in SRs conducted on restricted database searches (here, PubMed). Its feasibility is best demonstrated in rapid or resource-limited review contexts, particularly when screening is limited to the top-ranked outputs. BibliZap operationalizes key recommendations from the TAR-CiS statement and offers practical value for both exhaustive and exploratory evidence retrieval tasks."
Risk of bias
Selection bias: Reviews selected from high-impact journals may not represent typical systematic review practices; Gold standard incompleteness: Inclusion sets may be incomplete or imprecise, particularly with suboptimal reporting; Database restriction: Evaluation limited to PubMed-indexed articles only, not capturing multi-database strategies; Temporal limitation: Retrospective evaluation using published reviews rather than real-time screening; Citation bias: BibliZap may underrepresent very recent or sparsely cited studies; Selection bias: Reviews were selected from high-impact journals only, potentially limiting generalizability to lower-tier or specialty journals; Retrospective study design: No prospective validation in real-time screening workflows; Gold standard definition: Inclusion sets from published SRs may be incomplete or imprecise with suboptimal reporting; Database restriction: Evaluation limited to PubMed-indexed records only; does not reflect multi-database strategies used in most SRs; Time window bias: Time filters applied to restrict results to SR publication window; Publication bias: Citation-based approaches underrepresent very recent or sparsely cited studies; 4 SRs excluded: Original PubMed queries retrieved no results when re-executed; Selection bias: Reviews from high-impact journals may not represent broader SR practices; Incomplete gold standard: Inclusion sets may be incomplete or imprecise due to suboptimal reporting; Database restriction: Evaluation limited to PubMed-indexed records, not representative of multi-database searches used in most SRs; Exclusion of studies without PMID: Approximately 5% of included studies excluded from analysis; Circular reference bias: Mitigated by excluding the SR itself from citation graph; Temporal bias: Citation-based approaches may underrepresent very recent or sparsely cited studies
Limitations
- The authors identified several limitations: "First, it was conducted retrospectively using published SRs
- Although these reviews spanned diverse medical topics and were selected from high-impact journals, prospective validation within real-time screening workflows would provide a more robust assessment of usability in practice." Additionally, "our evaluation was limited to the PubMed component of each review, meaning that 'missed' refer specifically to records not retrieved by PubMed
- While this does not capture the multi-database strategies used in most SRs, fewer than 5% of included studies in our corpus lacked a PubMed ID." Furthermore, "BibliZap currently performs direct citation searching and, at depth-2 expansion, inherently retrieves co-cited and co-citing records
- However, it does not yet implement dedicated co-citation or bibliographic-coupling algorithms as standalone search strategies, which could represent a future enhancement of its functionality."
Open questions raised
- Prospective validation within real-time screening workflows needed
- Assessment of impact on overall conclusions or strength of evidence in resulting SRs
- Direct comparison of depth-1 versus depth-2 configurations for marginal benefit analysis
- Implementation of dedicated co-citation and bibliographic-coupling algorithms
- Development of user-defined weighting schemes within ranking algorithm
- Integration with review management platforms such as Covidence or Rayyan
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations