12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Rayyan—a web and mobile app for systematic reviews

Mourad Ouzzani, Hossam M. Hammady, Zbys Fedorowicz, Ahmed K. Elmagarmid · Systematic Reviews · 2016

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
D
Evidence
24,664
Citations
166.33
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s13643-016-0384-4

Methodology & findings

Study design

Mixed-methods case study combining pilot testing on two published Cochrane reviews (273 and 1030 records), algorithm evaluation on 15 systematic reviews using two-fold cross-validation, and user survey (66 respondents).

Sample

> 1000, 11 groups

Primary method

Support Vector Machine (SVM) classifier for prediction algorithm. Two-fold cross-validation with 50% training/50% testing, repeated 10 times with results averaged. ROC (Receiver Operating Characteristic) curve analysis. Two metrics: AUC (Area Under the Curve) and WSS@95 (Work Saved over random Sampling at 95% recall).

Main result

The study found that "our users reported a 40% average time savings when using Rayyan compared to others tools, with 37% of the respondents reporting more than 50% time savings." Additionally, "experiments on a set of 15 reviews showed that the prediction embedded in Rayyan can reduce the time for screening articles," with prediction algorithm results showing "AUC = 0.87 ± 0.09 and WSS@95 = 0.49 ± 0.18."

Reports effect sizes and confidence intervals.

Research paradigm

Pragmatist/Design Science

Author conclusions

"Rayyan has been shown to be a very useful app with significant potential to lighten the load of systematic review authors by speeding up the tedious part of the process of selection of studies for inclusion within the review. Experiments on a set of 15 reviews showed that the prediction embedded in Rayyan can reduce the time for screening articles. In addition, our survey showed that our users reported time savings in the order of 40% on average compared to other tools they have been using in the past."

Risk of bias

Selection bias: pilot testing used only two Cochrane reviews, not representative of all systematic review types; Non-blinded testing: tester was familiar with expected results from published reviews; Funding bias: tool fully funded by Qatar Foundation with authors as technical leads; Survey bias: self-selected respondents (66 respondents); potential self-selection bias from users already using the tool; Attrition: no information on survey response rate or non-response bias; Selection bias: Pilot testing used only two Cochrane reviews authored by one of the developers (ZF), not representative of broader user population; Performance bias: Tester was not blinded to expected results; experiment acknowledged as not technically blinded; Lack of comparator validation: No head-to-head comparison with other systematic review tools at time of publication; Self-selection bias: Survey respondents (66) likely represents motivated users who found Rayyan beneficial; Funding bias: Tool fully funded by Qatar Foundation; authors are developers and likely invested in positive outcomes; Tester familiarity bias: Although the experiment was not technically 'blinded,' the tester already knew the results of the selection process; Selection bias: Testing conducted on only two Cochrane reviews from one experienced author; Sampling bias: Survey respondents (66 total) represent self-selected users rather than representative sample; Developer bias: Authors include developers of Rayyan tool being evaluated

Limitations

  • The paper notes that "A comprehensive comparison of Rayyan with other systems would require additional studies to be conducted, more especially those which build on several previous reports." Additionally, the pilot testing was "not technically 'blinded'" as the tester was familiar with the expected results
  • The evaluation was limited to testing on two Cochrane reviews and did not include comprehensive comparison with competing systems.

Open questions raised

  • Better detection and handling of duplicates
  • Assessment of risk of bias with automatic identification and extraction of supporting sentences
  • Automatic extraction of PICO and other data elements from full-text articles
  • Extending Rayyan API for integration with other software platforms
  • Seamless integration with Review Manager (RevMan), the Cochrane software
  • Better detection of duplicates and user-guided handling processes
Data: 15 systematic review test collections from Oregon EPC, Southern California EPC, and RTI/UNC EPC (referenced in [16], not directly provided); 15 systematic reviews from Oregon EPC, Southern California EPC, and RTI/UNC EPC (referenced but not explicitly made available); Rayyan platform at http://rayyan.qcri.org (tool/app itself, not research data); 15 systematic review test collections from Oregon EPC, Southern California EPC, and RTI/UNC EPC (referenced in Table 1); Two Cochrane reviews used for pilot testing (Fedorowicz et al. 2008 and 2012)Extracted from: pdfAgreement 50%

Explore related topics

Related papers