12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Large-Scale, Cross-Disciplinary Corpus of Systematic Reviews

Pierre Achkar, Tim Gollub, Arno Simons, Harrisen Scells, Martin Potthast · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Large-scale corpus construction with three-stage enrichment pipeline: (i) data collection and metadata filtering from OpenAlex database with title-based heuristics; (ii) full-text acquisition and document parsing using PaddleOCR-VL; (iii) structured information extraction using LLM-based verify-then-repair pipeline with verification against OCR text and semantic validation.

Sample

N = 301871, 13 groups

Primary method

Approximate string matching for OCR text verification; document deduplication using DOI, OpenAlex ID, and normalized title strings; metadata-based inclusion filtering; evidence-centric structured extraction with verify-then-repair pipeline; LLM-based extraction with primary pass and repair pass; two-stage verification (alignment checking and semantic validation); precision, recall, and F-score metrics for retrieval evaluation.

Main result

The study presents Webis-SR4ALL-26, a large-scale corpus comprising "301,871 systematic reviews spanning 27 scientific domains, derived from reviews indexed in OpenAlex." Key retrieval demonstration results showed that "normalized Boolean strategies achieved higher recall than keyword-based queries (0.245 vs. 0.180) at similarly low precision (0.015)."

Reports effect sizes.

Research paradigm

Positivist/empiricist

Author conclusions

"Webis-SR4ALL-26 is a large-scale, cross-disciplinary corpus of systematic reviews comprising 301,871 reviews across 27 scientific disciplines. It integrates resolved OpenAlex reference lists, structured methodological information extracted through a verification-based pipeline, and executable approximations of reported search strategies within a unified open infrastructure." The authors further conclude: "By enabling reproducible research under consistent cross-disciplinary conditions, Webis-SR4ALL-26 extends existing resources that are typically limited in scale and disciplinary scope, and supports work in information retrieval, evidence synthesis, and meta-research on scientific practice."

Risk of bias

Selection bias: Title-based identification excludes reviews not explicitly declaring 'systematic review' terminology; Attrition bias: Only 72,678 of 301,871 reviews (24.1%) had successfully retrieved full texts; Extraction bias: Strict verification protocol prioritizes precision over recall, potentially missing implicit information; Coverage bias: Medicine accounts for ~64% of reviews, creating domain imbalance; Reference set bias: Resolved reference lists include background literature, not just included studies; Selection bias from title-based systematic review identification excludes non-self-identified reviews; Incomplete full-text availability (24.1% of 301,871 reviews) introduces coverage bias; OCR parsing errors may affect extraction accuracy despite use of PaddleOCR-VL; Strict verification protocol filters out implicitly stated information, biasing toward explicitly documented methods; Query normalization abstractions may not capture original database-specific syntax and refinements; Selection bias from title-based identification favoring explicit "systematic review" declarations; Coverage bias due to incomplete full-text availability (72,678 of 301,871 reviews); Extraction precision-recall tradeoff leading to underreporting of distributed information; Citation context bias—reference sets include background and contextual literature beyond included studies

Limitations

  • "The identification of systematic reviews relies on explicit title-based declarations, which favors precision but excludes reviews that do not self-identify using this terminology." Additionally, "Full-text availability is incomplete, which constrains methodological extraction to a subset of records and contributes to uneven field coverage
  • In addition, the strict verification protocol applied during extraction prioritizes precision over recall, so methodological elements that are implicitly stated or distributed across multiple passages may remain unstructured." Furthermore, "the resolved reference sets used for evaluation do not correspond directly to included-study gold standards, as reviews cite background and contextual literature in addition to eligible studies."

Open questions raised

  • Extension of systematic review identification beyond title-based heuristics by incorporating additional metadata signals and full-text cues
  • Expanding full-text coverage to increase proportion of reviews with structured methodological information
  • Direct approximation or reconstruction of included-study sets to enable evaluation settings closer to screening-level gold standards
  • Improvements in extraction methods that increase recall without compromising evidence grounding
  • Systematic review identification could be extended beyond title-based heuristics by incorporating additional metadata signals and full-text cues
  • Expanding full-text coverage would increase the proportion of reviews for which methodological information can be structured and analyzed
Data: Webis-SR4ALL-26; OpenAlex; Incorporated benchmark datasets; Webis-SR4ALL-26 (primary dataset): 301,871 systematic reviews with metadata, structured extraction fields, and normalized queries; OpenAlex database (data source): cited as basis for review identification and reference resolutionCode: PaddleOCR; OLM-OCR; OLMoCR; PaddleOCR repository: https://github.com/PaddlePaddle/PaddleOCR; olmOCR repository: https://github.com/allenai/olmocr/tree/main/olmocr/benchExtracted from: pdfAgreement 61%

Explore related topics

Related papers