12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers

Catherine A. Gao, Frederick M. Howard, Nikolay S. Markov, Emma Dyer, Siddhi Ramesh, Yuan Luo et al. · npj Digital Medicine · 2023

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
657
Citations
23.16
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41746-023-00819-6

Methodology & findings

Study design

Experimental comparison study with human evaluation and detector tool assessment.

Sample

N = 500, 3 groups

Primary method

Descriptive statistics including medians and interquartile ranges; AUROC calculation for detector performance; plagiarism detection via automated tools (iThenticate and plagiarism detector website); blinded human review with accuracy assessment.

Main result

The study found that most generated abstracts were detected using an AI output detector with "median [interquartile range] of 99.98% 'fake' [12.73%, 99.98%]" compared to "median 0.02% [IQR 0.02%, 0.09%] for the original abstracts." Additionally, "when given a mixture of original and general abstracts, blinded human reviewers correctly identified 68% of generated abstracts as being generated by ChatGPT, but incorrectly identified 14% of original abstracts as being generated."

Reports effect sizes and confidence intervals.

Research paradigm

Empiricist/positivist

Author conclusions

The authors conclude that "ChatGPT writes believable scientific abstracts, though with completely generated data." They further state that "Depending on publisher-specific guidelines, AI output detectors may serve as an editorial tool to help maintain scientific standards."

Risk of bias

Selection bias: Only abstracts from five high-impact factor medical journals were included, which may not represent all scientific domains or journals with different impact factors; Potential detector bias: Reliance on GPT-2 Output Detector, which may have inherent biases in classification; Human reviewer bias: Blinded reviewers' judgments may be influenced by writing style patterns specific to ChatGPT that differ from journal to journal; Limited sample size for human review: Not explicitly stated in abstract how many reviewers participated or how many abstracts were evaluated per reviewer; Selection bias: Abstracts gathered from only five high-impact medical journals may not represent broader journal landscape; Blinding: While human reviewers were blinded, the selection of journals and abstracts could introduce bias; Potential confounding: Journal prestige and field-specific writing conventions not controlled; Selection bias: abstracts gathered from five high-impact journals may not represent full spectrum of scientific writing; Potential human reviewer bias in blinded assessment despite blinding procedures; Limited specification of journal selection criteria; Detector performance may vary across different types of scientific abstracts

Open questions raised

  • The authors identify that "The boundaries of ethical and acceptable use of large language models to help scientific writing are still being discussed, and different journals and conferences are adopting varying policies," indicating need for further research on policy development and standardization across publishing venues.
  • The authors identify that "The boundaries of ethical and acceptable use of large language models to help scientific writing are still being discussed, and different journals and conferences are adopting varying policies," suggesting need for clearer guidelines and continued investigation.
  • The authors identify that "The boundaries of ethical and acceptable use of large language models to help scientific writing are still being discussed, and different journals and conferences are adopting varying policies," indicating need for further work on policy development and guidelines.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 64%

Explore related topics

Related papers