12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI in peer review: can artificial intelligence be an ally in reducing gender and geographical gaps in peer review? A randomized trial

André L. Teixeira · Research Integrity and Peer Review · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

10/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s41073-025-00182-y

Methodology & findings

Study design

Cross-sectional randomized controlled trial design.

Sample

N = 1979, 2 groups

Primary method

Chi-square tests for group differences in proportions of scientists by gender, geographical location, and country income level. Independent sample t-tests for comparing academic metrics (publications, citations, h-index) between groups. Komolgorov-Smirnof test to confirm normality of variables. For repeatability analysis, chi-square tests comparing proportions of identical vs. different names across trials, and 2 x 2 repeated measures ANOVA with group (with vs. without DEI prompt) and condition (Trial 1 vs. Trial 2) as main factors. IBM SPSS 26 used for analysis. Statistical significance set at P < 0.05.

Main result

When not prompted to consider DEI, GPT-4o primarily identified male scientists (68% vs. 32% female) and those affiliated with high-income countries, especially North America (60%) and Europe (28%). In contrast, "when gender and geographical DEI was explicitly prompted, the list of scientists was gender balanced (51% females and 49% males) and geographically diverse. Specifically, the proportion of scientists affiliated with high-income countries decreased from 95.3% to 42.3%, while scientists from upper-middle (3.2% to 26.2%), lower-middle (1.2% to 26.1%), and low-income (0.2% to 5.4%) countries significantly increased." Additionally, "there were no significant differences between groups without and with a DEI prompt in the total number of published works (284 ± 237 vs. 281 ± 245, P = 0.77), Google Scholar-derived total number of citations (48,445 ± 60,270 vs. 53,792 ± 71,903, P = 0.13), and h-index (79 ± 43 vs. 76 ± 43, P = 0.15)."

Reports effect sizes.

Research paradigm

positivist/empiricist

Author conclusions

"In closing, the present study demonstrated that without a DEI prompt, GPT-4o can successfully identify expert scientists in different medical fields, but predominantly males and those affiliated with high income countries. However, a gender-balanced and geographically diverse pool of scientists can be achieved by incorporating DEI into the prompt." The authors conclude that "AI can be an ally in reducing gender and geographical gaps in peer review, though DEI should be explicitly highlighted in the prompt. On the other hand, AI could perpetuate existing biases if not carefully managed."

Risk of bias

Training data bias: LLMs trained on historical publication data reflecting existing gender and geographical inequalities in academia; Single researcher conducting procedures unblinded, though article selection and name generation were automated; Potential misidentification of deceased individuals (14 names, 11 from non-DEI group, 3 from DEI group); Five names could not be identified as scientists; Gender determination reliance on pronouns, names, and external databases (Gender API) with 80% accuracy threshold; Incompleteness of Google Scholar profiles (only 1,427 of 1,979 scientists had profiles for citation/h-index analysis); Potential historical bias in underlying datasets perpetuating existing disparities even with DEI prompting; Historical bias in training data reflecting existing gender and geographical inequalities in academic publishing; Potential selection bias in article choice, though minimized through randomization software; Single researcher conducting unblinded procedures, though article selection and AI generation were automated; Residual bias in AI model even with DEI prompt (42% still from high-income countries); Limited data completeness: 5 names unidentifiable, 14 scientists deceased, only 72% of identified scientists had Google Scholar profiles; Historical bias in training data reflecting gender and geographical inequalities in publishing; Residual bias toward high-income countries even with DEI prompt (42% remained from high-income countries); Selection bias: only 50 articles from 5 journals examined; Unblinded design: single researcher conducted all procedures; Exclusion of deceased individuals and non-identifiable scientists may introduce selection bias; Limited generalizability to non-medical fields

Limitations

  • Several limitations were acknowledged by the authors
  • First, "GPT-4o, like all LLMs, relies on the datasets it was trained on, which may reflect historical biases in publication and authorship
  • For example, ~90% of handling editors of the BMJ publishing group (comprising 21 biomedical journals) and > 60% of the senior authors of manuscripts submitted between 2018 and 2021 are geographically located in Europe and North America." Second, the authors note "there is currently no universally accepted threshold or specific metric that defines a 'qualified' reviewer
  • In the present study, the reported academic metrics (i.e., peer-reviewed publications, citations and h-index) provide insight into the academic productivity and potential impact of the suggested scientists
  • However, ranking or selecting scientists solely based on these metrics presents significant caveats." Third, "while this study demonstrated that AI could generate a more diverse and inclusive pool of potential reviewers, it remains unclear whether these individuals would agree to participate in the peer review." Additionally, "the present study did not examine other types of bias, such as ethnic or institutional bias, which should be addressed in future research."

Open questions raised

  • Whether the identified reviewers would actually agree to participate in peer review
  • Impact of diversity changes on the quality of peer review and broader scientific discourse
  • Examination of other types of bias (ethnic, institutional) beyond gender and geography
  • Universally accepted threshold or specific metric defining a 'qualified' reviewer
  • Effectiveness of DEI strategies in journal peer review (original studies still lacking)
  • Optimal strategies for prompting AI to identify experts from lower-income countries
Data: Not explicitly stated; supplementary data referenced (Supplementary Table 1) but full dataset availability not mentionedCode: Not mentionedExtracted from: pdfAgreement 53%

Explore related topics

Related papers