How to Report the Use of Artificial Intelligence in Scientific Articles: A Scoping Review and Taxonomy of Editorial Policies
Juan Aníbal González-Rivera · Revista Caribeña de Psicología · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.37226/rcp.v10i1.17451
Methodology & findings
Study design
Scoping review following PRISMA-ScR guidelines.
Sample
N = 19, 3 groups
Primary method
Qualitative thematic analysis. No inferential statistics reported. Search strategy optimized using PRESS 2015 checklist. Data management in spreadsheet format under version control (Git). Analyses and figure preparation conducted in R and Python. Coverage analysis: 12 policies (63%) request version and date of use; 8 policies (42%) request prompts/parameters when AI affects results/figures/code.
Main result
The review found that "the field converges on general principles but still diverges on how to report them in practice—where to place the disclosure, which minimum elements to include (tool and version, assisted tasks, prompts, human oversight, source verification, and treatment of images/code), and how to harmonize expectations across disciplines with different publishing cultures." Additionally, "the corpus converges on three principles: (a) AI does not meet authorship criteria; (b) the use of AI in manuscript preparation must be explicitly disclosed; and (c) responsibility for the content rests exclusively with human authors, with safeguards for confidentiality during peer review."
Reports effect sizes.
Research paradigm
Qualitative/Descriptive (policy analysis and thematic synthesis)
Author conclusions
"Taken together, these findings support a pragmatic roadmap: institutionalize a standardized AI-use disclosure, aligned with AI-Use-12; prohibit AI authorship and protect confidentiality in peer review; and discourage detector-only decisions. Implemented consistently, these measures can improve transparency, enable sharper peer review, and bolster trust in the scientific record in the era of generative AI."
Risk of bias
Single reviewer conducting screening and extraction without inter-rater reliability verification; Grey literature bias (19 documents are institutional policies and editorials, not peer-reviewed empirical studies); Terminology interpretation variability across publishers requiring subjective judgment; Lack of formal external validation for the derived AI-Use-12 instrument; Policy landscape changing rapidly; snapshot at specific date (October 21, 2025); Single-reviewer design (no inter-rater reliability metrics such as Cohen's κ); No double extraction performed due to single-reviewer constraint; Interpretive judgment required for terminology variation across publishers, which may introduce misclassification; AI-Use-12 instrument lacks formal external validation; Single reviewer conducting screening and extraction (no inter-rater reliability metrics; single-reviewer bias mitigated by detailed decision logging and second-pass verification); Grey literature quality assessment using AACODS rather than formal risk-of-bias tool; Interpretive judgment required for terminology variation across publishers (potential misclassification); Personal calibration on 10% of records may not fully standardize inclusion parameters
Limitations
- "First, as a scoping review, it maps and describes rather than quantifies effects
- Second, policies change quickly
- despite recording access dates, maintaining currency requires periodic updates
- Third, much of the corpus is grey literature
- I verified official status and described quality with AACODS, but this is not a risk-of-bias assessment
- Fourth, terminology varies across publishers, which required interpretive judgment and may introduce misclassification
Open questions raised
- First, field-wide cross-sectional audits of 'Instructions for Authors' and submission forms across large, representative journal samples could estimate adoption and maturity of AI policies. Second, formal validation of AI-Use-12 (Delphi, CVI, usability testing with editors and reviewers) would strengthen its acceptance. Third, research on editorial workflows could balance disclosure, traceability, and verification (e.g., reference audits) without imposing undue burden on authors or staff. Fourth, empirical studies could assess how disclosure affects editorial timelines, corrections/retractions, and reader/reviewer trust.
- Field-wide cross-sectional audits of Instructions for Authors and submission forms across large, representative journal samples to estimate adoption and maturity of AI policies
- Formal validation of AI-Use-12 (Delphi, content validity index, usability testing with editors and reviewers)
- Research on editorial workflows to balance disclosure, traceability, and verification without imposing undue burden
- Empirical studies assessing how disclosure affects editorial timelines, corrections/retractions, and reader/reviewer trust
- Detection research pivoting from general classifiers toward provenance, metadata, and integrity checks for images and code
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations