Are Non-English Papers Reviewed Fairly? Language-of-Study Bias in NLP Peer Reviews
Ehsan Barkhordar, Abdulfattah Safa, Verena Blaschke, Erika Lombart, Marie-Catherine de Marneffe, Gözde Gül Sahin · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods study combining human annotation with computational modeling.
Main result
The study found that "non-English papers face substantially higher bias rates than English-only ones, with negative bias consistently outweighing positive bias." Specifically, "Reviews of English-focused papers exhibit a bias rate of 0.37%, while reviews of single non-English papers show a collective rate of 14.79%, roughly 40 times higher." Additionally, "the most dominant form" of bias is that reviewers demand "unjustified cross-lingual generalization," which accounted for 62.16% of negative bias instances in the manually annotated dataset.
Research paradigm
Mixed methods (quantitative computational analysis + qualitative interpretation)
Author conclusions
The authors conclude: "In this work, we provide the first systematic evidence that language-of-study bias in NLP peer review is not an isolated artifact but a structural pattern. Our large-scale analysis shows that non-English papers face bias rates roughly 40 times higher than English-only ones, consistently across all six venues we examined. Importantly, the problem compounds: while English papers encounter primarily one bias pattern (generalizability demands), non-English and multilingual papers face a wider repertoire, including English-as-gold-standard framing, language choice interrogation, and impact dismissal. These findings point to concrete interventions: reviewer guidelines could explicitly require evaluation against a paper's stated scope, and the strong performance of our LLM-based detector (87.37 Macro F1) suggests automated screening is a realistic complement to human oversight."
Risk of bias
Selection bias in sampling strategy (LLM-assisted stage designed to surface likely bias cases, introducing non-random sample); Class imbalance in annotated dataset (439 no-bias vs. 73 negative vs. 17 positive segments); Conservative annotation bias (acknowledged undercount of subtle bias instances); Model limitation bias (false negatives dominate over false positives in Gemini classifier, defaulting to NO BIAS); Venue/corpus bias (limited to six NLP venues; rejected papers largely unavailable); Confounding variables (non-English papers may cluster in contribution types that are themselves more bias-prone); Selection bias: Sampling strategy biased toward likely bias cases in Stage 1, may undercount true negatives; Annotator bias: Conservative labeling in ambiguous cases may undercount subtle biases; Exclusion bias: 5 segments (0.9%) with 'Needs Context' excluded from analysis; Corpus composition bias: Primarily accepted papers (94%); rejected paper reviews largely unavailable; Model bias: Stage 1 LLM triage flagged 299 negative but only 38 positive candidates (7.9:1 ratio), skewing annotated pool; Confounding: Non-English papers may cluster in contribution types already prone to bias; independent effects not fully separable; Generalization limitation: Data from 6 NLP venues only; may not represent other peer review systems; Selection bias: Sampling strategy emphasized bias-containing segments, skewing the annotated dataset toward 299 negative versus 38 positive candidates (7.9:1 ratio); Annotator disagreement: Fleiss κ = 0.68 (substantial but not perfect agreement); five segments required exclusion due to lack of consensus; Model limitations: LLM classifier shows conservative bias toward NO BIAS label (false negatives dominate false positives in confusion matrix); Class imbalance in training data: LOBSTER contains 439 NO BIAS vs. 73 NEGATIVE BIAS vs. 17 POSITIVE BIAS segments; Venue composition bias: Majority of corpus from EMNLP 2023 (375 segments); limited representation from other venues; Potential confounding: Non-English papers may cluster in contribution types (Data & Benchmarking, Linguistic Analysis) that are inherently more bias-prone
Limitations
- The authors state several limitations: "although annotating full reviews provides complete context, we adopt conservative labeling in ambiguous or borderline cases, which may ultimately undercount instances of subtle bias." Additionally, "five segments (0.9%) labeled No Majority or Unclear / Needs Context are excluded from evaluation." They note that "our data comes from a limited set of venues and cycles, so prevalence estimates may not generalize to other review systems." Furthermore, "the majority of our corpus consists of accepted papers
- however, a small fraction (≈6%) includes non-accepted submissions from the ARR 2024 (Apr–Jun) cycle
- Reviews of rejected papers from other venues are not publicly available, so bias patterns in reviews leading to rejection remain largely unexplored." The authors also acknowledge that "distinguishing language-related bias from valid language-scoped critique is inherently difficult, and borderline cases can yield annotator disagreement."
Open questions raised
- Limited understanding of how bias patterns manifest in reviews of rejected papers (currently unavailable for most venues)
- Need for controlled analysis separating effects of language scope, contribution type, venue, and year (these factors are likely correlated)
- Underexplored positive bias patterns (insufficient cases for detailed qualitative analysis)
- Broader application of bias detection methods to other review systems beyond NLP
- The paper identifies the need for fairer reviewing practices and concrete interventions: "reviewer guidelines could explicitly require evaluation against a paper's stated scope, and the strong performance of our LLM-based detector (87.37 Macro F1) suggests automated screening is a realistic complement to human oversight." Future work should include controlled analysis separating effects of correlated dimensions and investigation of bias patterns in reviews leading to rejection.
- Bias patterns in reviews leading to rejection: "Reviews of rejected papers from other venues are not publicly available, so bias patterns in reviews leading to rejection remain largely unexplored."
Explore related topics
Related papers
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Publishers’ and journals’ instructions to authors on use of generative artificial intelligence in academic and scientific publishing: bibliometric analysisConner Ganjavi · 2024 · 199 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations