12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer review

Mohammad Hosseini, Serge P. J. M. Horbach · Research Integrity and Peer Review · 2023

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
I
Evidence
209
Citations
7.40
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s41073-023-00133-5

Methodology & findings

Study design

Qualitative analysis using five core thematic frameworks (reviewers' role, editors' role, functions and quality of peer reviews, reproducibility, and social/epistemic functions) applied to ChatGPT.

Main result

The study found that "LLMs have the potential to substantially alter the role of both peer reviewers and editors. Through supporting both actors in efficiently writing constructive reports or decision letters, LLMs can facilitate higher quality review and address issues of review shortage." However, the authors also identified that "the fundamental opacity of LLMs' training data, inner workings, data handling, and development processes raise concerns about potential biases, confidentiality and the reproducibility of review reports."

Research paradigm

Interpretive/critical analysis of technology adoption in scholarly communication

Author conclusions

The authors conclude that "LLMs are likely to have a profound impact on academia and scholarly communication. While potentially beneficial to the scholarly communication system, many uncertainties remain and their use is not without risks." They specifically recommend that "if LLMs are used to write scholarly reviews and decision letters, reviewers and editors should disclose their use and accept full responsibility for data security and confidentiality, and their reports' accuracy, tone, reasoning and originality." Furthermore, they emphasize that "the question is therefore not whether these systems find their way to our daily practices of producing and reviewing scientific content, but how to use them responsibly."

Risk of bias

Selection bias in choice of ChatGPT as primary example system; Temporal bias: LLMs develop rapidly during study period, making conclusions potentially outdated; LLMs may amplify existing biases in training data related to geography, race, class, and research geography; Opacity of LLM training prevents assessment of underlying biases; Authors' selection of specific ChatGPT interactions as illustrative examples may not be fully representative; Rapid evolution of LLM systems during manuscript review period; Framework from Tennant and Ross-Hellauer may not capture all relevant dimensions of peer review impact; Focus primarily on ChatGPT rather than comprehensive LLM landscape; Selection bias: Only one primary LLM (ChatGPT) examined in depth; examples from other systems (BARD) limited; Scope bias: Analysis limited to journal article peer review; grants and other review types excluded; Temporal bias: Rapid LLM development during manuscript review period; findings may be outdated; Funding bias: Research supported by NIH/NCATS, which may have interest in health informatics applications

Limitations

  • The authors state that "this short essay has specific limitations (we only discussed review of journal articles and not other object types like grants, we used examples from ChatGPT, and were constrained by limitations of the used framework)." Additionally, they note uncertainty because "Even in the relatively short time that this manuscript was under review, several new developments challenged some of the manuscript's assertions
  • Among others, this includes the launch of GPT-4 as a successor of GPT-3.5 used in the examples in our manuscript."

Open questions raised

  • Unclear how LLMs will develop and why they perform in specific ways due to opacity of training processes
  • Lack of understanding regarding confidentiality and data handling by LLM providers
  • Need for continuous monitoring of LLM capabilities as they develop rapidly
  • Unclear how skill development for peer review will occur if LLMs become routinely used
  • Limited investigation of reproducibility issues when systems change frequently
  • Need for research on social and epistemic impacts of LLM-assisted peer review on academic communities
Extracted from: pdfAgreement 67%

Explore related topics

Related papers