TiAb Review Plugin: A Browser-Based Tool for AI-Assisted Title and Abstract Screening
Yuki Kataoka, Masahiro Banno, Michihito Kyo, Shuri Nakao, Tomoo Sato, Shunsuke Taito et al. · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-method evaluation comprising: (1) ML equivalence verification using 10-fold cross-validation on six public datasets comparing TypeScript implementation against original Python ASReview implementation; (2) systematic LLM parameter tuning on depression benchmark dataset testing 16 parameter configurations across two model families; (3) cross-dataset LLM validation on five publicly available datasets with retrospective ground-truth labels..
Primary method
Design science research with iterative implementation and empirical validation
Main result
The study found that "The TypeScript classifier produced top-100 rankings 100 percent identical to the original ASReview across all six datasets" and for LLM screening, "recall was 94 to 100 percent with precision of 2 to 15 percent, and Work Saved over Sampling at 95 percent recall (WSS@95) ranged from 48.7 to 87.3 percent." The tool successfully integrates LLM screening and ML active learning into a browser-based, no-code environment.
Research paradigm
Design science / artifact-driven
Author conclusions
"We developed a functional browser extension that integrates LLM screening and ML active learning into a no-code, serverless environment, ready for practical use in systematic review screening." The authors also state that the tool "provides LLM-assisted T&A screening directly within the researcher's existing web workflow" and "operates entirely within the browser, requiring no software installation, server infrastructure, or programming skills."
Risk of bias
Retrospective evaluation only with known ground-truth labels; Single LLM model family tested (Google Gemini); Benchmark datasets concentrated in critical care medicine; Implementation equivalence verified only for top-100 records; Retrospective evaluation only with known ground-truth labels; no prospective validation; Single LLM vendor (Google Gemini) tested; generalizability to other models unknown; Evaluation limited to critical care medicine datasets with well-defined eligibility criteria; ML equivalence tested only on top-100 ranked records, not full ranking lists; No formal ablation study on sensitivity-prioritization prompt instruction; Funding from JSPS and pharmaceutical company speaker honoraria (potential bias on LLM recommendations); Retrospective evaluation only—no prospective validation in live systematic reviews; Limited to single LLM provider (Google Gemini); generalizability to other LLMs unknown; ML evaluation limited to top-100 ranked records; lower-ranked divergence not assessed; Benchmark datasets concentrated in critical care medicine; generalizability to other domains uncertain; Fixed inclusion threshold (0.5) applied across all datasets; threshold optimization per dataset not explored
Limitations
- "Regarding ML evaluation, we verified implementation equivalence only for the top-100 ranked records
- divergence in lower-ranked records, while unlikely given identical algorithms, was not assessed
- Regarding LLM evaluation, we tested only one model family (Google Gemini)
- performance may differ with other LLMs." Additionally, "All evaluations in this study are retrospective, using datasets with known ground-truth labels
- No prospective study has yet been conducted to measure the tool's impact on screening efficiency, reviewer time, or error rates in a live systematic review project."
Open questions raised
- Prospective validation study needed to measure tool's impact on screening efficiency, time savings, and error rates in live systematic review projects
- Support for additional LLM providers (e.g., OpenAI, Anthropic, Alibaba) to increase flexibility and reduce dependence on single API
- Formal usability testing using standardized instruments such as the System Usability Scale (SUS) needed to provide evidence on tool's ease of use across different user populations
- Prospective validation study needed to measure tool impact on screening efficiency, time savings, and error rates in live systematic review projects
- Support for additional LLM providers (OpenAI, Anthropic, Alibaba) to increase flexibility and reduce API vendor lock-in
- Formal usability testing using standardized instruments such as the System Usability Scale (SUS) to assess ease of use across different user populations
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations