Are AI Tools Tougher Reviewers Than Toxicologists? Artificial Intelligence as Editor and Peer Reviewer of Scientific Manuscripts in Toxicology
Jose L. Domingo · Qeios · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.32388/lrqpsd
Methodology & findings
Study design
Proof-of-concept exploratory analysis.
Sample
N = 8, 2 groups
Primary method
No formal statistical analysis performed. Author explicitly states: "Given the exploratory nature of the study, no formal sample size calculation or statistical power analysis was performed." Descriptive analysis of AI tool recommendations (counts and frequencies) presented in tables.
Main result
The study found that AI tools demonstrated divergent recommendations when evaluating published toxicology manuscripts as if they were submitted for peer review. Specifically, "Not a single one of the eight evaluated papers, treated as submitted manuscripts by the AI tools, received a recommendation of 'Accept in its present form' from any of them", whereas all eight papers had been accepted for publication by human reviewers. The AI tools showed greater stringency, with "the majority of the five AI tools, for most of the eight manuscripts" recommending major revisions. Among the AI tools used as reviewers, "ChatGPT and Copilot were most frequently identified by the meta-reviewer tools as producing comparatively useful reports", while Grok was identified as least reliable.
Reports effect sizes.
Research paradigm
empirical-exploratory
Author conclusions
The authors conclude: "Despite the very significant limitations of this preliminary study, given that none of the evaluated manuscripts received a recommendation of 'Accept in its present form', the findings suggest that AI tools may tend to produce more conservative recommendations than those reflected in final editorial decisions." They further state: "These observations should not be interpreted as evidence that AI tools outperform human reviewers, but rather as indicative of their potential role as complementary decision-support systems." Importantly, they emphasize that "the use of AI tools in peer review raises important ethical questions regarding transparency, accountability, confidentiality, and potential bias that will need to be addressed through clear policy frameworks before any large-scale implementation is contemplated."
Risk of bias
Selection bias: Only 8 papers included, all from April 2026 issues; papers were already published (not blinded to publication status); Selection bias in AI tool choice: Based solely on author's personal experience; Evaluation bias: Meta-review conducted by AI tools rather than human experts; Confounding: Reviewers evaluated already-published and revised papers, not original submissions; Publication bias: Only successful papers evaluated (no rejected manuscripts included); Selection bias: Non-random selection of journals (though author claims random selection, then states it was 'entirely random' with subjective quality judgments); Small sample size (n=8 papers) with no formal sample size calculation or power analysis; Use of already-published papers rather than original submitted versions, confounding comparison with original editorial decisions; Single-instance AI queries without iterative refinement, limiting generalizability; AI-based meta-review rather than human expert evaluation introduces cascading bias; Author's extensive personal experience as Editor-in-Chief may introduce confirmation bias in AI tool selection; Limited generalizability: restricted to toxicology and specific AI tools based on author's personal experience; Selection bias: Author selected papers based on personal experience with AI tools, not random/systematic selection of tools; Selection bias: Only first experimental article from each journal's April 2026 issue (or May for one journal) was selected; Outcome bias: All eight selected papers were already published and accepted, so cannot assess AI ability to identify unsuitable manuscripts; Measurement bias: Evaluation of review quality conducted by AI tools rather than human experts; Confounding: Already-published, revised manuscripts reviewed rather than original submissions; impossible to separate AI stringency from manuscript improvement during human review process
Limitations
- The author states: "This is a highly preliminary study, which the author has restricted exclusively to his area of specialization, toxicology, and to a limited number of papers, journals, and AI tools
- Given the exploratory nature of the study, no formal sample size calculation or statistical power analysis was performed." Additionally, "the AI tools evaluated here lacked access to the original submitted versions of the manuscripts and therefore reviewed already-revised, accepted, and published papers
- This is a factor that must be considered when interpreting the discrepancy between AI recommendations (Major revisions in most cases) and the decisions of the original human reviewers (acceptance for publication)." The author also notes that "evaluation of reviewer report quality was conducted using AI tools rather than human experts, which may introduce additional layers of bias and reduces the availability of an external reference standard."
Open questions raised
- The author identifies the need for "a large-scale study evaluating a substantial number of already-published articles (access to the original submitted versions would be ideal), with multiple AI tools acting as Editors for desk decisions, and as reviewers to formulate comments and suggestions, and to recommend corresponding final decisions." The author also notes unresolved questions including "effectiveness, fairness, and efficiency" of AI in peer review, and the need for "clear policy frameworks" addressing "transparency, accountability, confidentiality, and potential bias."
- Need for large-scale study evaluating substantial number of already-published articles with access to original submitted versions
- Need for multiple AI tools acting as Editors for desk decisions with extended evaluation
- Further investigation of AI tools' capacity to identify manuscripts unsuitable for peer review
- Research on ethical frameworks for AI use in peer review including transparency, accountability, confidentiality, and potential bias
- Need for studies in disciplines beyond toxicology
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Leveraging ChatGPT for Enhancing Critical Thinking SkillsYing Guo · 2023 · 223 citations