Artificial Intelligence in Peer Review: Enhancing Efficiency While Preserving Integrity
Bohdana Doskaliuk, Olena Zimba, Marlen Yessirkepov, І. П. Кліщ, Roman Yatsyshyn · Journal of Korean Medical Science · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3346/jkms.2025.40.e92
Methodology & findings
Study design
Systematic literature search conducted on November 21, 2024, using Scopus, MEDLINE/PubMed, and DOAJ databases, focusing on English-language articles.
Main result
The paper finds that "AI can significantly reduce the burden on reviewers by automatically flagging grammatical errors, spelling mistakes, and awkward phrasing" and that "AI is becoming an important tool for uncovering potential ethical issues, such as plagiarism or data manipulation." However, the authors emphasize that "AI tools are limited in assessing these factors" regarding novelty and significance, and "they may not fully acknowledge the importance of groundbreaking findings or the value of new theoretical approaches."
Research paradigm
Interpretivist/Qualitative
Author conclusions
The authors conclude that "the integration of AI into academic publishing brings notable advantages, such as streamlining repetitive tasks and improving the efficiency of peer review. Nonetheless, its shortcomings in specialized knowledge, contextual interpretation, and ethical decision-making emphasize the need for human involvement." They further state that "rather than replacing human expertise, AI should serve as a supportive tool, guided by comprehensive policies and adequate training. By combining AI's capabilities with ethical standards, the academic community can refine the peer review process while preserving its fundamental values."
Risk of bias
AI training data bias: if datasets contain inherent biases (gender, race, geographic region, publication trends), AI tools may perpetuate those biases; Geographic bias: models disproportionately trained on research from Western institutions may overlook valuable contributions from underrepresented regions; Selection bias in literature search (English-language only, specific database choices); Potential for over-reliance on AI reducing quality of human expert judgment; Selection bias in literature search (English-language articles only); Potential bias in included studies based on publication trends; AI training data biases (gender, race, geographic region, institutional representation); Incomplete coverage of gray literature (exclusion of conference papers and book chapters); Dependence on indexing comprehensiveness of three databases (Scopus, MEDLINE/PubMed, DOAJ); Publication bias in literature selection (non-systematic review); Potential dataset bias in AI training (acknowledged by authors); Geographic bias in training data (Western institution overrepresentation noted); Selection bias in article inclusion (language restricted to English); Potential confirmation bias in narrative synthesis without meta-analysis
Limitations
- The review acknowledges that "While ChatGPT and other AI tools are highly effective at language processing and handling general tasks, they lack the in-depth subject-matter expertise required to understand or critically evaluate complex scientific content fully." Additionally, "AI's inability to recognize the potential long-term impact of research is a significant limitation in academic publishing." The authors also note that "AI models are trained on vast datasets, and the quality of their outputs is directly influenced by the data they have been exposed to
- If these datasets contain inherent biases (whether related to gender, race, geographic region, or publication trends) AI tools may inadvertently perpetuate those biases."
Open questions raised
- Future studies could explore: (1) development of AI tools specifically designed to assess the quality of scientific evidence within manuscripts using advanced natural language processing techniques; (2) comparative studies assessing effectiveness of AI-driven versus traditional peer review processes in detecting errors and biases.
- The authors identify that "Future studies could explore the development of AI tools specifically designed to assess the quality of scientific evidence within manuscripts, potentially using advanced natural language processing techniques. Additionally, comparative studies assessing the effectiveness of AI-driven versus traditional peer review processes in detecting errors and biases could provide valuable insights."
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations