12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A call for clarity: a unified checklist for reporting use of large language models in writing scientific manuscripts

Neil Mehta, Ryan Marshall Felder, Eric Kodish, Peter B. Imrey, Richard M. Frankel · Research Integrity and Peer Review · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s41073-026-00212-3

Methodology & findings

Study design

Literature synthesis with participatory design validation.

Primary method

Framework design based on synthesis of existing editorial statements and guidelines, with iterative grouping and expert team review

Main result

The paper proposes that "frameworks integrating ethical boundaries with reporting expectations are needed" and that "a practical checklist for AI use and reporting in scholarly writing" can "operationalize those principles for consistent application by authors, reviewers, and journal editors and editorial staff." The authors identify that existing guidance is fragmented, with "journals, publishers, and professional associations have responded quickly to perceived challenges LLMs present to the scientific publication process, yet their approaches have been fragmented and inconsistent."

Research paradigm

Normative/prescriptive (establishing standards and guidelines)

Author conclusions

The authors propose that "this checklist can serve both as a starting point for ongoing dialogue and as a practical mechanism for embedding transparency into the fabric of scientific communication." They further conclude that by making checklists publicly accessible, "the process of LLM utilization disclosure would not only be harmonized but also elevated to a new level of openness-fostering trust, accountability, and confidence in the integrity of the scholarly record." They also recommend treating this as "a living document, regularly updated to reflect technological advances, ethical insights, and community consensus."

Risk of bias

Social desirability bias in voluntary disclosures; Potential undisclosed LLM use if enforcement mechanisms are lacking; Disciplinary variation in acceptable uses not fully addressed; Social desirability bias in voluntary disclosure; Potential incomplete disclosures of non-acceptable LLM uses; Variable access to literature across institutions affecting reproducibility; Social desirability bias in voluntary checklist adoption; Potential incomplete disclosures of generally not acceptable uses; Variation in access to literature due to licensing and paywall restrictions

Limitations

  • The authors acknowledge that "a meaningful limitation of any voluntary checklist is social desirability bias: authors may be reluctant to acknowledge 'generally not acceptable' uses, potentially resulting in incomplete disclosures." Additionally, they note that "LLM-enabled literature retrieval may be constrained by limitations like licensing restrictions, paywalls, and institutional access policies, potentially differing from what a human researcher could access, which may affect reproducibility and independent verification."

Open questions raised

  • Need for consensus-based framework to guide authors and journal editors in everyday practice
  • Disciplinary differences in acceptable LLM uses (e.g., differences between science/medicine vs. English/philosophy journals)
  • Integration with existing systems like EQUATOR Network and manuscript submission platforms (ScholarOne, Editorial Manager)
  • Development of audit trail mechanisms with logging prompts and version history
  • The authors identify that further work is needed to determine discipline-specific guidelines, noting that "acceptable uses of AI in science and medicine, for instance, may not be acceptable in writing submitted to English or philosophy journals." They also call for consideration of the tradeoffs between productivity gains and skill atrophy in AI use.
  • Further work would be necessary to identify discipline-specific differences in acceptable uses of LLMs
Code: https://mehta nb1.github.io/AIDis closureForm/; https://mehta nb1.github.io/AIDisclosureForm/Extracted from: pdfAgreement 68%

Explore related topics

Related papers