AI in peer review: can artificial intelligence be an ally in reducing gender and geographical gaps in peer review? A randomized trial
André L. Teixeira · Research Integrity and Peer Review · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s41073-025-00182-y
Methodology & findings
Study design
Cross-sectional randomized controlled trial design.
Sample
N = 1979, 2 groups
Primary method
Chi-square tests for group differences in proportions of scientists by gender, geographical location, and country income level. Independent sample t-tests for comparing academic metrics (publications, citations, h-index) between groups. Komolgorov-Smirnof test to confirm normality of variables. For repeatability analysis, chi-square tests comparing proportions of identical vs. different names across trials, and 2 x 2 repeated measures ANOVA with group (with vs. without DEI prompt) and condition (Trial 1 vs. Trial 2) as main factors. IBM SPSS 26 used for analysis. Statistical significance set at P < 0.05.
Main result
When not prompted to consider DEI, GPT-4o primarily identified male scientists (68% vs. 32% female) and those affiliated with high-income countries, especially North America (60%) and Europe (28%). In contrast, "when gender and geographical DEI was explicitly prompted, the list of scientists was gender balanced (51% females and 49% males) and geographically diverse. Specifically, the proportion of scientists affiliated with high-income countries decreased from 95.3% to 42.3%, while scientists from upper-middle (3.2% to 26.2%), lower-middle (1.2% to 26.1%), and low-income (0.2% to 5.4%) countries significantly increased." Additionally, "there were no significant differences between groups without and with a DEI prompt in the total number of published works (284 ± 237 vs. 281 ± 245, P = 0.77), Google Scholar-derived total number of citations (48,445 ± 60,270 vs. 53,792 ± 71,903, P = 0.13), and h-index (79 ± 43 vs. 76 ± 43, P = 0.15)."
Reports effect sizes.
Research paradigm
positivist/empiricist
Author conclusions
"In closing, the present study demonstrated that without a DEI prompt, GPT-4o can successfully identify expert scientists in different medical fields, but predominantly males and those affiliated with high income countries. However, a gender-balanced and geographically diverse pool of scientists can be achieved by incorporating DEI into the prompt." The authors conclude that "AI can be an ally in reducing gender and geographical gaps in peer review, though DEI should be explicitly highlighted in the prompt. On the other hand, AI could perpetuate existing biases if not carefully managed."
Risk of bias
Training data bias: LLMs trained on historical publication data reflecting existing gender and geographical inequalities in academia; Single researcher conducting procedures unblinded, though article selection and name generation were automated; Potential misidentification of deceased individuals (14 names, 11 from non-DEI group, 3 from DEI group); Five names could not be identified as scientists; Gender determination reliance on pronouns, names, and external databases (Gender API) with 80% accuracy threshold; Incompleteness of Google Scholar profiles (only 1,427 of 1,979 scientists had profiles for citation/h-index analysis); Potential historical bias in underlying datasets perpetuating existing disparities even with DEI prompting; Historical bias in training data reflecting existing gender and geographical inequalities in academic publishing; Potential selection bias in article choice, though minimized through randomization software; Single researcher conducting unblinded procedures, though article selection and AI generation were automated; Residual bias in AI model even with DEI prompt (42% still from high-income countries); Limited data completeness: 5 names unidentifiable, 14 scientists deceased, only 72% of identified scientists had Google Scholar profiles; Historical bias in training data reflecting gender and geographical inequalities in publishing; Residual bias toward high-income countries even with DEI prompt (42% remained from high-income countries); Selection bias: only 50 articles from 5 journals examined; Unblinded design: single researcher conducted all procedures; Exclusion of deceased individuals and non-identifiable scientists may introduce selection bias; Limited generalizability to non-medical fields
Limitations
- Several limitations were acknowledged by the authors
- First, "GPT-4o, like all LLMs, relies on the datasets it was trained on, which may reflect historical biases in publication and authorship
- For example, ~90% of handling editors of the BMJ publishing group (comprising 21 biomedical journals) and > 60% of the senior authors of manuscripts submitted between 2018 and 2021 are geographically located in Europe and North America." Second, the authors note "there is currently no universally accepted threshold or specific metric that defines a 'qualified' reviewer
- In the present study, the reported academic metrics (i.e., peer-reviewed publications, citations and h-index) provide insight into the academic productivity and potential impact of the suggested scientists
- However, ranking or selecting scientists solely based on these metrics presents significant caveats." Third, "while this study demonstrated that AI could generate a more diverse and inclusive pool of potential reviewers, it remains unclear whether these individuals would agree to participate in the peer review." Additionally, "the present study did not examine other types of bias, such as ethnic or institutional bias, which should be addressed in future research."
Open questions raised
- Whether the identified reviewers would actually agree to participate in peer review
- Impact of diversity changes on the quality of peer review and broader scientific discourse
- Examination of other types of bias (ethnic, institutional) beyond gender and geography
- Universally accepted threshold or specific metric defining a 'qualified' reviewer
- Effectiveness of DEI strategies in journal peer review (original studies still lacking)
- Optimal strategies for prompting AI to identify experts from lower-income countries
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations