12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

The State of Peer Review in Empirical Software Engineering: A Community Survey on Review Load, Quality, and GenAI Use

Justus Bogner, Roberto Verdecchia · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Cross-sectional questionnaire survey with 22 questions (20 single-choice or multiple-choice, 2 open free-text questions).

Sample

N = 120, 7 groups

Primary method

Descriptive statistics (frequencies, percentages, medians, means); thematic analysis for open-ended questions; ordinal scale ratings on 5-point Likert scale

Main result

The study found that "a two-thirds majority perceived their review load as high (51) or very high (29), with a median rating of 4 and a mean of 3.88." Additionally, "more than 90% of respondents (110) perceived high workload and too many high-priority tasks as the main obstacle to providing higher-quality reviews." The research also revealed that "the most frequently reported issue was that reviews provided by others are shallow or very short (82)," and "the issue of AI-generated reviews was already reported by a worrisome one-sixth of participants (20)."

Reports effect sizes.

Research paradigm

empiricist/positivist

Author conclusions

The authors conclude that "peer reviewers in the ESE community are under high review load and that something needs to be done: the current state of affairs is clearly not sustainable anymore and likely has not been for a while." They further state: "While some participants seem carefully optimistic about making good use of LLMs to improve peer review, many other responses, especially the many free-text comments, also tell a story of disappointment, anger, frustration, and uncertainty." Regarding solutions, they emphasize: "Whatever we decide as a community, it will be important to accompany any introduced changes with trustworthy evaluations about their effects, both positive and negative, so that we can adjust course if necessary. However, in addition to these solutions to fight the symptoms, we should not forget major root causes, like our 'publish or perish' culture, which also needs fixing."

Risk of bias

Selection bias: Survey advertised via personal contacts and social media; primarily reached experienced academics; Self-selection bias: Only 120 responses from invitation to major ESE venues; likely disproportionate response from those interested in peer review issues; Social-desirability bias: Participants reported higher quality of own reviews vs. received reviews; bias acknowledged for questions about review ratios and LLM use; Geographic bias: 77/120 from Europe, 23 from North America; only 1 from Australia/Oceania, none from Africa; Seniority bias: 81/120 were tenured academics; only 4 PhD students; primarily 11-20+ years ESE experience; Demographic homogeneity: Predominantly senior academic positions from Western institutions; Attrition/non-response: Final question (suggestions for improvement) answered by only 82/120 participants; Selection bias: Participants self-selected through personal contacts and social media; predominantly European/North American researchers; Social-desirability bias: Participants likely overreported quality of their own reviews and underreported unethical LLM use; Survivorship bias: Sample consists of experienced ESE researchers who remained active in peer review; Geographic bias: No participants from Africa; limited representation from Asia, South America, and Oceania; Seniority bias: Majority (81/120) were tenured academics; only 4 PhD students participated; Selection bias: Only 120 participants, predominantly from Europe (77) and North America (23), with no African participants; Self-selection bias: Participants self-selected into survey by responding, likely representing those more engaged with peer review; Social desirability bias: Respondents may over-report quality of their own reviews and adherence to review ratios; Geographic bias: Disproportionate representation from Europe and North America; Seniority bias: 81 of 120 participants held tenured positions, underrepresenting junior researchers; Survivorship bias: Only active reviewers in past 12 months were included

Limitations

  • The authors note that "most survey respondents were seasoned ESE researchers from European or North American institutions, which needs to be considered for interpreting the results." Additionally, they acknowledge that "social-desirability bias is likely" in responses about review ratios and that "participants deemed the reviews they write of higher quality than those they receive" which suggests "either an above-average survey sample of diligent, high-quality reviewers, i.e., a self-selection bias of people truly interested in peer review, or a certain degree of self-assessment bias and/or social-desirability bias." The authors also note that "current usage is both more frequent and more questionable than reported by our selective sample of ESE researchers."

Open questions raised

  • Need for empirical evidence on what effective and efficient peer review should look like
  • Lack of systematic characterization of LLM use extent and impact in ESE peer review
  • Limited data on how ESE research is distributed across workshops, conferences, and journals
  • Need for policy clarification on what constitutes 'unethical use' of LLMs
  • Investigation of which venues successfully maintain high review quality standards
  • Need for trustworthy evaluations of proposed peer review improvements
Data: Survey data published on Zenodo at https://doi.org/10.5281/zenodo.20495019; Survey data; Survey data published online at https://doi.org/10.5281/zenodo.20495019 (as referenced in footnote 1)Extracted from: pdfAgreement 66%

Explore related topics

Related papers