12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing

Krzysztof Żurawicki, Julia Farganus, Arkadiusz Gaweł, Mateusz Bystroński, Tomasz Kajdanowicz · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Systematic analysis of 135,520 peer reviews from ICLR (2013-2025) and NeurIPS (2021-2025) using psycholinguistic metrics (Token Type Ratio, Flesch Reading Ease, Flesch-Kincaid Grade Level), mathematical engagement detection via LLM classifier (Qwen3.5-8B with 96.6% accuracy on ground truth), citation verification against external databases (Crossref, OpenAlex, DBLP), and inter-rater agreement assessment using Krippendorff's α.

Main result

The study found that "LLM-generated reviews diverge sharply from human prose in surface-level complexity. All general-purpose models produce text with Flesch Reading Ease scores between 13.7 (DeepSeek) and 21.3 (GPT-5), far below the human baseline of 39.1." Furthermore, "GPT-5 achieves the highest agreement with human panels (α r = 0.36/0.26 on ICLR/NeurIPS), matching or slightly exceeding the human-only baseline (α r = 0.34/0.24)." The analysis also revealed that "in the post-ChatGPT era (beginning with ICLR 2024 and NeurIPS 2023), reviews have become measurably more difficult to read. Metrics evaluating educational grade level and linguistic complexity display a sharp, continuous upward trajectory."

Research paradigm

Positivist/Empiricist

Author conclusions

The authors conclude: "Given these stylistic and factual gaps, LLMs are currently unsuitable as wholesale replacements for human expertise. Instead, we propose a synergistic pipeline: leveraging frontier models for mathematical analysis and fine-tuned models for stylistic drafting, while retaining human reviewers for final qualitative judgment and citation verification." They further state that "the metrics and benchmark introduced here provide a principled, extensible foundation for measuring progress in automated peer review—one that we believe will prove essential as LLM adoption in the review process continues to accelerate." Additionally, "These patterns suggest that frontier LLMs can effectively complement human reviewers by consistently engaging with mathematical content and, in some cases, identifying formal structures that humans may overlook."

Risk of bias

Selection bias: Dataset limited to OpenReview (primarily ICLR, NeurIPS); not representative of all peer review systems; Temporal confounding: Pre/post-ChatGPT cutoff may conflate multiple factors (reviewer pool changes, submission quality increases, conference policy changes); Data contamination: LLMs evaluated may have been trained on OpenReview data, creating potential knowledge leakage in the 2024-2025 subset; Measurement bias: Citation verification relies on publicly available APIs that may not index all resources; missing data treated as unverified rather than unfound; Classifier dependency: Oracle classifier used for success ratio may encode biases present in training data; Confounding variables acknowledged by authors: 'exponential increase in submissions or a demographic shift within the reviewer pool' could explain stylistic changes; Mathematical engagement detection: LLM-based classification (Qwen3.5-8B) may introduce systematic biases despite 96.6% validation accuracy on small sample (n=30); Sample imbalance: Only 5% rejected papers in NeurIPS, addressed partially through balanced sampling but ICLR balance not fully specified; Selection bias: Dataset limited to ICLR and NeurIPS with imbalanced NeurIPS representation (only accepted papers fully available); Confounding variables: Paper acknowledges that increased review complexity could reflect exponential growth in submissions or demographic shifts in reviewer pools; LLM contamination: Models may have been trained on peer review data, affecting fair assessment of alignment with human reviews; Evaluation circularity: Oracle classifier used for 'success ratio' may encode dataset biases rather than true discriminative features; Temporal confounding: Post-2023 trend analysis may reflect simultaneous changes in submission volume, paper complexity, or reviewer pool composition; Selection bias: Only publicly available OpenReview data; NeurIPS contains primarily accepted papers (5% rejected), addressed via balanced sampling; Measurement bias: LLM-based mathematical engagement detection relies on single classifier (Qwen3.5-8B) despite 96.6% validation accuracy; Confounding variables: Trends in stylistic complexity may result from exponential increase in submissions, demographic shifts in reviewer pool, or paper complexity changes—though authors argue these are unlikely primary causes; Temporal confounding: Pre-post design without true control; authors note confounding variables such as increasing submission volume; Data leakage risk: Authors acknowledge potential contamination within human review pool from LLM-edited text, though mitigated for 2025 subset; Oracle classifier bias: Success Ratio evaluation risks conflating true discriminative cues with spurious features; Prompt engineering bias: Model outputs highly sensitive to prompt phrasing; near-zero/negative α between prompts for some models indicates scores driven by prompt rather than content

Limitations

  • The authors state: "As original submissions may be edited and the review may not reflect the corresponding paper content, locating original submissions on pre-print repositories like arXiv is logistically unfeasible
  • Nevertheless, high coverage ratios (above 0.9 for GPT-5) suggest that revised versions remain reliable proxies." Additionally, "Our framework is limited by the zero-shot setting, LLM-based mathematical engagement classification, and venue coverage (ICLR, NeurIPS)." The paper further notes that "Our method accounts only for citations structured according to well known formats but unfortunately LLMs as well as reviewers do not always follow standardized citations format because unawareness or mistake
  • To verify citation, we use publicly available APIs, that methods may not contain all possible domains and resources." The authors also acknowledge that "We also observe that LLMs do not always follow the given format of the review, inventing their own sections and scores
  • This leads to a situation where, for around 1% of reviews, we are unable to parse the reviews to extract scores and decide to drop them from the analysis."

Open questions raised

  • Future verification of citation utility beyond existence checks: Authors note 'future research could concentrate on confirming the usefulness of these resources, evaluating whether it is reasonable to cite them and expanding analysis to larger spectrum of possible resources'
  • Expanded venue coverage: Current framework limited to ICLR and NeurIPS; generalizability to other conferences and disciplines unknown
  • Qualitative evaluation of content dimensions: Authors state 'the qualitative evaluation of these individual dimensions is left for future work'
  • Fine-tuning strategies: 'model sensitivity to prompt engineering -where specialized instructions yield denser, mathematically rigorous, and cross-referenced outputs -identifies a key target for future fine-tuning'
  • Citation format standardization and external source integration: 'the struggle to provide valid citations could stem from zero-shot setting and unavailability of external sources'
  • Expanded resource domains for citation verification: Need for broader coverage of academic resources beyond current publicly available APIs
Data: OpenReview ICLR (2013-2025): 35,985 papers, 135,520 reviews; OpenReview NeurIPS (2021-2025): 18,762 papers, 74,472 reviews; Black Hole simulation dataset (MAD/SANE): access restricted, not publicly available; 2025 subset for atomic strength/weakness extraction: 804 reviews from 100 papers; OpenReview peer reviews: ICLR (2013-2025, 35,985 papers, 135,520 reviews) and NeurIPS (2021-2025, 18,762 papers, 74,472 reviews); 2025 subset used for atomic strengths/weaknesses analysis (804 reviews); OpenReview dataset (ICLR 2013-2025, NeurIPS 2021-2025): publicly available via OpenReview.net; Black Hole simulation dataset (MAD/SANE): access noted as restricted per authors; 2025 subset used for atomic strength/weakness extraction and information coverage evaluation: subset of OpenReview data; Validation set for mathematical engagement classifier: 30 papers manually annotated by authorsCode: Anonymous repository mentioned but URL not provided in text; Anonymous repository (mentioned but URL not provided in accessible form); Anonymous repository for source code (mentioned as available but URL not provided in text due to anonymity)Extracted from: pdfAgreement 59%

Explore related topics

Related papers