A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews
Pierre Achkar, Tim Gollub, Arno Simons, Harrisen Scells, Martin Potthast · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Large-scale corpus construction with three-stage enrichment pipeline: (i) data collection and metadata filtering from OpenAlex database with title-based heuristics; (ii) full-text acquisition and document parsing using PaddleOCR-VL; (iii) structured information extraction using LLM-based verify-then-repair pipeline with verification against OCR text and semantic validation.
Sample
N = 301871, 13 groups
Primary method
Approximate string matching for OCR text verification; document deduplication using DOI, OpenAlex ID, and normalized title strings; metadata-based inclusion filtering; evidence-centric structured extraction with verify-then-repair pipeline; LLM-based extraction with primary pass and repair pass; two-stage verification (alignment checking and semantic validation); precision, recall, and F-score metrics for retrieval evaluation.
Main result
The study presents Webis-SR4ALL-26, a large-scale corpus comprising "301,871 systematic reviews spanning 27 scientific domains, derived from reviews indexed in OpenAlex." Key retrieval demonstration results showed that "normalized Boolean strategies achieved higher recall than keyword-based queries (0.245 vs. 0.180) at similarly low precision (0.015)."
Reports effect sizes.
Research paradigm
Positivist/empiricist
Author conclusions
"Webis-SR4ALL-26 is a large-scale, cross-disciplinary corpus of systematic reviews comprising 301,871 reviews across 27 scientific disciplines. It integrates resolved OpenAlex reference lists, structured methodological information extracted through a verification-based pipeline, and executable approximations of reported search strategies within a unified open infrastructure." The authors further conclude: "By enabling reproducible research under consistent cross-disciplinary conditions, Webis-SR4ALL-26 extends existing resources that are typically limited in scale and disciplinary scope, and supports work in information retrieval, evidence synthesis, and meta-research on scientific practice."
Risk of bias
Selection bias: Title-based identification excludes reviews not explicitly declaring 'systematic review' terminology; Attrition bias: Only 72,678 of 301,871 reviews (24.1%) had successfully retrieved full texts; Extraction bias: Strict verification protocol prioritizes precision over recall, potentially missing implicit information; Coverage bias: Medicine accounts for ~64% of reviews, creating domain imbalance; Reference set bias: Resolved reference lists include background literature, not just included studies; Selection bias from title-based systematic review identification excludes non-self-identified reviews; Incomplete full-text availability (24.1% of 301,871 reviews) introduces coverage bias; OCR parsing errors may affect extraction accuracy despite use of PaddleOCR-VL; Strict verification protocol filters out implicitly stated information, biasing toward explicitly documented methods; Query normalization abstractions may not capture original database-specific syntax and refinements; Selection bias from title-based identification favoring explicit "systematic review" declarations; Coverage bias due to incomplete full-text availability (72,678 of 301,871 reviews); Extraction precision-recall tradeoff leading to underreporting of distributed information; Citation context bias—reference sets include background and contextual literature beyond included studies
Limitations
- "The identification of systematic reviews relies on explicit title-based declarations, which favors precision but excludes reviews that do not self-identify using this terminology." Additionally, "Full-text availability is incomplete, which constrains methodological extraction to a subset of records and contributes to uneven field coverage
- In addition, the strict verification protocol applied during extraction prioritizes precision over recall, so methodological elements that are implicitly stated or distributed across multiple passages may remain unstructured." Furthermore, "the resolved reference sets used for evaluation do not correspond directly to included-study gold standards, as reviews cite background and contextual literature in addition to eligible studies."
Open questions raised
- Extension of systematic review identification beyond title-based heuristics by incorporating additional metadata signals and full-text cues
- Expanding full-text coverage to increase proportion of reviews with structured methodological information
- Direct approximation or reconstruction of included-study sets to enable evaluation settings closer to screening-level gold standards
- Improvements in extraction methods that increase recall without compromising evidence grounding
- Systematic review identification could be extended beyond title-based heuristics by incorporating additional metadata signals and full-text cues
- Expanding full-text coverage would increase the proportion of reviews for which methodological information can be structured and analyzed
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations