Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer review
Mohammad Hosseini, Serge P. J. M. Horbach · Research Integrity and Peer Review · 2023
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s41073-023-00133-5
Methodology & findings
Study design
Qualitative analysis using five core thematic frameworks (reviewers' role, editors' role, functions and quality of peer reviews, reproducibility, and social/epistemic functions) applied to ChatGPT.
Main result
The study found that "LLMs have the potential to substantially alter the role of both peer reviewers and editors. Through supporting both actors in efficiently writing constructive reports or decision letters, LLMs can facilitate higher quality review and address issues of review shortage." However, the authors also identified that "the fundamental opacity of LLMs' training data, inner workings, data handling, and development processes raise concerns about potential biases, confidentiality and the reproducibility of review reports."
Research paradigm
Interpretive/critical analysis of technology adoption in scholarly communication
Author conclusions
The authors conclude that "LLMs are likely to have a profound impact on academia and scholarly communication. While potentially beneficial to the scholarly communication system, many uncertainties remain and their use is not without risks." They specifically recommend that "if LLMs are used to write scholarly reviews and decision letters, reviewers and editors should disclose their use and accept full responsibility for data security and confidentiality, and their reports' accuracy, tone, reasoning and originality." Furthermore, they emphasize that "the question is therefore not whether these systems find their way to our daily practices of producing and reviewing scientific content, but how to use them responsibly."
Risk of bias
Selection bias in choice of ChatGPT as primary example system; Temporal bias: LLMs develop rapidly during study period, making conclusions potentially outdated; LLMs may amplify existing biases in training data related to geography, race, class, and research geography; Opacity of LLM training prevents assessment of underlying biases; Authors' selection of specific ChatGPT interactions as illustrative examples may not be fully representative; Rapid evolution of LLM systems during manuscript review period; Framework from Tennant and Ross-Hellauer may not capture all relevant dimensions of peer review impact; Focus primarily on ChatGPT rather than comprehensive LLM landscape; Selection bias: Only one primary LLM (ChatGPT) examined in depth; examples from other systems (BARD) limited; Scope bias: Analysis limited to journal article peer review; grants and other review types excluded; Temporal bias: Rapid LLM development during manuscript review period; findings may be outdated; Funding bias: Research supported by NIH/NCATS, which may have interest in health informatics applications
Limitations
- The authors state that "this short essay has specific limitations (we only discussed review of journal articles and not other object types like grants, we used examples from ChatGPT, and were constrained by limitations of the used framework)." Additionally, they note uncertainty because "Even in the relatively short time that this manuscript was under review, several new developments challenged some of the manuscript's assertions
- Among others, this includes the launch of GPT-4 as a successor of GPT-3.5 used in the examples in our manuscript."
Open questions raised
- Unclear how LLMs will develop and why they perform in specific ways due to opacity of training processes
- Lack of understanding regarding confidentiality and data handling by LLM providers
- Need for continuous monitoring of LLM capabilities as they develop rapidly
- Unclear how skill development for peer review will occur if LLMs become routinely used
- Limited investigation of reproducibility issues when systems change frequently
- Need for research on social and epistemic impacts of LLM-assisted peer review on academic communities
Explore related topics
Related papers
- ChatGPT in education: Strategies for responsible implementationMohanad Halaweh · 2023 · 576 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Using AI to write scholarly publicationsMohammad Hosseini · 2023 · 264 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- The ethics of disclosing the use of artificial intelligence tools in writing scholarly manuscriptsMohammad Hosseini · 2023 · 202 citations