Beyond Detection: Governing GenAI in Academic Peer Review as a Sociotechnical Challenge
Tatiana Chakravorti, Pranav Narayanan Venkit, Sourojit Ghosh, Sarah Rajtmajer · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Convergent parallel mixed methods design combining: (1) qualitative discourse analysis of 448 social media posts from LinkedIn, Reddit, and Twitter/X (September 2025-August 2026) with keywords including 'AI + peer review', 'ChatGPT + peer review', 'LLM + peer review'; (2) 14 semi-structured interviews (30-60 minutes, conducted virtually September-November 2025) with area chairs and program chairs from AI/HCI conferences, using thematic analysis.
Sample
N = 462, 3 groups
Primary method
Qualitative discourse analysis using iterative coding process combining theory-informed deductive categories (from AI governance and academic evaluation literature) and inductive refinement based on recurring patterns. Thematic analysis of interview transcripts using: multiple readings, open coding by individual authors, grouping codes into preliminary themes, refinement through author discussions, verification against transcripts. Convergent parallel mixed methods design with independent analysis of each dataset before triangulation. No quantitative statistical tests reported.
Main result
The study found that "GenAI usage in peer review is already prevalent; Liang et al. [51] estimates that between 6.5% and 17% of reviews submitted to prominent NLP conferences during the 2023-2024 cycles contained AI-generated text, even where conference policies explicitly prohibited such use." Across both interviews and social media discourse, participants broadly agreed that "GenAI may be acceptable for limited supportive tasks, such as improving clarity or organizing feedback, but that key evaluative decisions like judging novelty, contribution, or whether a paper should be accepted should remain the responsibility of humans." The research reveals that "policy ambiguity functions as a governance strategy" with participants explaining "that this ambiguity was motivated by several concerns. First, they noted that GenAI technologies are evolving rapidly, making it difficult to define stable rules that would remain relevant over time."
Reports effect sizes.
Research paradigm
Qualitative/interpretive; mixed methods (qualitative discourse analysis + semi-structured interviews)
Author conclusions
"Across interviews and social media discussions, we find broad agreement that GenAI may be useful for limited support tasks, but that key evaluative decisions such as judging novelty, contribution, and acceptance should remain human responsibilities. Participants also raised concerns about inaccurate feedback, unclear accountability, and new risks introduced by AI use. Our findings show that these concerns are shaped by existing pressures in peer review, including reviewer overload and unclear policies, which place greater burdens on individuals, especially early-career scholars. Together, these perspectives highlight that AI in peer review is not just a technical issue, but a governance challenge. We argue that responsible use of GenAI requires clear boundaries, meaningful human oversight, and attention to the labor and care involved in scholarly evaluation."
Risk of bias
Selection bias: Convenience sample of social media posts retrieved via keyword search shaped by platform algorithms; Selection bias: Interview participants recruited via targeted emails and snowball sampling, potentially skewed toward those engaged with AI governance issues; Sampling bias: English-language only; limited to Western/USA-based academic contexts; Temporal bias: Data collected during period of rapid GenAI evolution (Sept 2025-Nov 2025), limiting generalizability; Disciplinary bias: Sample limited primarily to AI and HCI conference contexts; Platform bias: Social media sample may not represent offline discourse or institutions not active on LinkedIn, Reddit, Twitter/X; Selection bias in social media sample: keyword-based search retrieval shaped by platform ranking algorithms; Convenience sampling in interviews with recruitment via email/social media and snowball sampling; Potential response bias: participants self-selected into study about GenAI in peer review; Interview sample limited to English-speaking researchers in primarily US-based AI/HCI venues; Exclusion of non-English discourse and non-Western academic contexts; Selection bias: Social media sample is a search-retrieved convenience sample shaped by keyword queries and platform ranking algorithms, not representative of all platform discourse; Selection bias: Interview recruitment used snowball sampling and targeted LinkedIn/Twitter, limiting diversity of recruitment channels; Selection bias: Interview participants (n=14) are from AI/HCI conferences only; findings may not generalize to other disciplines or journal contexts; Temporal specificity: Data collected during period of rapid GenAI change (2025-2026); norms may evolve; Language bias: English-language content only; Western/USA-centric perspective
Limitations
- "First, all data are English language and largely reflect Western, primarily USA based academic contexts, which may limit applicability to other scholarly systems
- Second, the social media dataset is not statistically representative
- it is a keyword based, search retrieved convenience sample shaped by platform algorithms, and theme frequencies should not be interpreted as prevalence estimates
- Third, findings are grounded mainly in AI and HCI conference contexts, and their transferability to journal review or disciplines beyond AI or HCI remains uncertain
- Finally, data were collected during a specific period of rapid GenAI change, and norms and governance practices may evolve over time."
Open questions raised
- Limited direct perspectives of area/program chairs and reviewers in existing literature (authors note: "A majority of research around the impact of GenAI on peer reviewing does not consider the direct perspectives of the area/program chairs and reviewers")
- Gap between governance policies and actual practice in peer review systems
- Limited understanding of structural strain as a driver of AI adoption in peer review
- Transferability to journal review contexts and disciplines beyond AI/HCI
- Evolution of norms and governance practices as GenAI capabilities continue to advance
- Comparative analysis across non-English-speaking and non-Western academic systems
Explore related topics
Related papers
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Exploring Students’ Perceptions of ChatGPT: Thematic Analysis and Follow-Up SurveyAbdulhadi Shoufan · 2023 · 464 citations
- AI-generated feedback on writing: insights into efficacy and ENL student preferenceJuan Escalante · 2023 · 461 citations
- Is it harmful or helpful? Examining the causes and consequences of generative AI usage among university studentsMuhammad Abbas · 2024 · 372 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations