PRAIB: Peer Review AI Benchmark of Behaviour of LLM-Assisted Reviewing
Krzysztof Żurawicki, Julia Farganus, Arkadiusz Gaweł, Mateusz Bystroński, Tomasz Kajdanowicz · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Systematic analysis of 135,520 peer reviews from ICLR (2013-2025) and NeurIPS (2021-2025) using psycholinguistic metrics (Token Type Ratio, Flesch Reading Ease, Flesch-Kincaid Grade Level), mathematical engagement detection via LLM classifier (Qwen3.5-8B with 96.6% accuracy on ground truth), citation verification against external databases (Crossref, OpenAlex, DBLP), and inter-rater agreement assessment using Krippendorff's α.
Main result
The study found that "LLM-generated reviews diverge sharply from human prose in surface-level complexity. All general-purpose models produce text with Flesch Reading Ease scores between 13.7 (DeepSeek) and 21.3 (GPT-5), far below the human baseline of 39.1." Furthermore, "GPT-5 achieves the highest agreement with human panels (α r = 0.36/0.26 on ICLR/NeurIPS), matching or slightly exceeding the human-only baseline (α r = 0.34/0.24)." The analysis also revealed that "in the post-ChatGPT era (beginning with ICLR 2024 and NeurIPS 2023), reviews have become measurably more difficult to read. Metrics evaluating educational grade level and linguistic complexity display a sharp, continuous upward trajectory."
Research paradigm
Positivist/Empiricist
Author conclusions
The authors conclude: "Given these stylistic and factual gaps, LLMs are currently unsuitable as wholesale replacements for human expertise. Instead, we propose a synergistic pipeline: leveraging frontier models for mathematical analysis and fine-tuned models for stylistic drafting, while retaining human reviewers for final qualitative judgment and citation verification." They further state that "the metrics and benchmark introduced here provide a principled, extensible foundation for measuring progress in automated peer review—one that we believe will prove essential as LLM adoption in the review process continues to accelerate." Additionally, "These patterns suggest that frontier LLMs can effectively complement human reviewers by consistently engaging with mathematical content and, in some cases, identifying formal structures that humans may overlook."
Risk of bias
Selection bias: Dataset limited to OpenReview (primarily ICLR, NeurIPS); not representative of all peer review systems; Temporal confounding: Pre/post-ChatGPT cutoff may conflate multiple factors (reviewer pool changes, submission quality increases, conference policy changes); Data contamination: LLMs evaluated may have been trained on OpenReview data, creating potential knowledge leakage in the 2024-2025 subset; Measurement bias: Citation verification relies on publicly available APIs that may not index all resources; missing data treated as unverified rather than unfound; Classifier dependency: Oracle classifier used for success ratio may encode biases present in training data; Confounding variables acknowledged by authors: 'exponential increase in submissions or a demographic shift within the reviewer pool' could explain stylistic changes; Mathematical engagement detection: LLM-based classification (Qwen3.5-8B) may introduce systematic biases despite 96.6% validation accuracy on small sample (n=30); Sample imbalance: Only 5% rejected papers in NeurIPS, addressed partially through balanced sampling but ICLR balance not fully specified; Selection bias: Dataset limited to ICLR and NeurIPS with imbalanced NeurIPS representation (only accepted papers fully available); Confounding variables: Paper acknowledges that increased review complexity could reflect exponential growth in submissions or demographic shifts in reviewer pools; LLM contamination: Models may have been trained on peer review data, affecting fair assessment of alignment with human reviews; Evaluation circularity: Oracle classifier used for 'success ratio' may encode dataset biases rather than true discriminative features; Temporal confounding: Post-2023 trend analysis may reflect simultaneous changes in submission volume, paper complexity, or reviewer pool composition; Selection bias: Only publicly available OpenReview data; NeurIPS contains primarily accepted papers (5% rejected), addressed via balanced sampling; Measurement bias: LLM-based mathematical engagement detection relies on single classifier (Qwen3.5-8B) despite 96.6% validation accuracy; Confounding variables: Trends in stylistic complexity may result from exponential increase in submissions, demographic shifts in reviewer pool, or paper complexity changes—though authors argue these are unlikely primary causes; Temporal confounding: Pre-post design without true control; authors note confounding variables such as increasing submission volume; Data leakage risk: Authors acknowledge potential contamination within human review pool from LLM-edited text, though mitigated for 2025 subset; Oracle classifier bias: Success Ratio evaluation risks conflating true discriminative cues with spurious features; Prompt engineering bias: Model outputs highly sensitive to prompt phrasing; near-zero/negative α between prompts for some models indicates scores driven by prompt rather than content
Limitations
- The authors state: "As original submissions may be edited and the review may not reflect the corresponding paper content, locating original submissions on pre-print repositories like arXiv is logistically unfeasible
- Nevertheless, high coverage ratios (above 0.9 for GPT-5) suggest that revised versions remain reliable proxies." Additionally, "Our framework is limited by the zero-shot setting, LLM-based mathematical engagement classification, and venue coverage (ICLR, NeurIPS)." The paper further notes that "Our method accounts only for citations structured according to well known formats but unfortunately LLMs as well as reviewers do not always follow standardized citations format because unawareness or mistake
- To verify citation, we use publicly available APIs, that methods may not contain all possible domains and resources." The authors also acknowledge that "We also observe that LLMs do not always follow the given format of the review, inventing their own sections and scores
- This leads to a situation where, for around 1% of reviews, we are unable to parse the reviews to extract scores and decide to drop them from the analysis."
Open questions raised
- Future verification of citation utility beyond existence checks: Authors note 'future research could concentrate on confirming the usefulness of these resources, evaluating whether it is reasonable to cite them and expanding analysis to larger spectrum of possible resources'
- Expanded venue coverage: Current framework limited to ICLR and NeurIPS; generalizability to other conferences and disciplines unknown
- Qualitative evaluation of content dimensions: Authors state 'the qualitative evaluation of these individual dimensions is left for future work'
- Fine-tuning strategies: 'model sensitivity to prompt engineering -where specialized instructions yield denser, mathematically rigorous, and cross-referenced outputs -identifies a key target for future fine-tuning'
- Citation format standardization and external source integration: 'the struggle to provide valid citations could stem from zero-shot setting and unavailability of external sources'
- Expanded resource domains for citation verification: Need for broader coverage of academic resources beyond current publicly available APIs
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations