Does language bias GenAI academic evaluation in humanities and social sciences? A mixed‐methods study based on Chinese‐language HSS papers
Yu Zhu, Yujie Jia, Yumeng Zhu, Jiyuan Ye · Journal of the Association for Information Science and Technology · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/asi.70079
Methodology & findings
Study design
Within-subjects experimental design with 1150 expert-selected papers from 23 disciplines evaluated by two GenAI models (GPT-4o and DeepSeek-V3) in both Chinese and English languages.
Sample
N = 1150, 6 groups
Primary method
Within-subjects design methodology; Cohen's d effect size calculation; thematic analysis for qualitative evaluation of reasoning patterns. Specific software and statistical tests not detailed in abstract.
Main result
The study reveals that "GPT‐4o favors English (Cohen's d = 1.10), while DeepSeek‐V3 favors Chinese (Cohen's d = −0.87), persisting across all disciplines." Additionally, "both models generate more critical comments for English papers, yet arrive at opposite scores through different rhetorical strategies—GPT‐4o tends to moderate its positive assessments of Chinese papers while DeepSeek‐V3 amplifies them."
Reports effect sizes and confidence intervals.
Research paradigm
Mixed-methods empirical research combining quantitative measurement and qualitative thematic analysis
Author conclusions
The authors conclude that "language bias in GenAI evaluation is bidirectional and model‐dependent, with scores not directly reflecting evaluative justifications." They further state that "This study provides controlled evidence that language bias in GenAI evaluation is bidirectional and model‐dependent" and that "The findings have implications for designing fairer multilingual academic evaluation systems and for understanding the limitations of GenAI as scholarly evaluation infrastructure."
Risk of bias
Language bias (primary focus of study); Model-dependent bias (bidirectional depending on model type); Selection bias potential: papers were 'expert-selected' which may not represent all HSS scholarship; Potential confounding: discipline-specific language patterns not fully controlled; Paper selection mechanism not detailed in abstract; Language bias as primary research question (inherent to study design); Model selection bias (only two specific GenAI models tested); Paper selection bias (expert-selected papers may not represent full distribution of HSS scholarship); Potential confounding by discipline-specific evaluation norms; Language-induced bias (primary focus of study); Model-dependent bias patterns; Potential selection bias in expert-selected paper sample; Possible confounding variables between disciplines
Open questions raised
- The paper identifies the need for understanding how GenAI systems can be designed to provide fairer multilingual academic evaluation, and highlights the broader need to understand GenAI's limitations as scholarly evaluation infrastructure.
- The study identifies the need for research on designing fairer multilingual academic evaluation systems and for better understanding the limitations of GenAI as scholarly evaluation infrastructure.
- The study identifies the need for research on: (1) designing fairer multilingual academic evaluation systems, and (2) understanding the limitations of GenAI as scholarly evaluation infrastructure in cross-language contexts.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations