12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Fabrication and errors in the bibliographic citations generated by ChatGPT

William H. Walters, Esther Isabelle Wilder · Scientific Reports · 2023

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
I
Evidence
352
Citations
12.36
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1038/s41598-023-41032-5

Methodology & findings

Study design

Empirical evaluation study.

Main result

The study found that "55% of the GPT-3.5 citations but just 18% of the GPT-4 citations are fabricated." Additionally, "43% of the real (non-fabricated) GPT-3.5 citations but just 24% of the real GPT-4 citations include substantive citation errors." The analysis demonstrates that fabricated citations remain a significant problem in ChatGPT-generated academic text, though GPT-4 shows substantial improvement over GPT-3.5.

Research paradigm

Empirical/positivist (quantitative evaluation of ChatGPT outputs)

Author conclusions

The authors conclude that "in terms of both fabricated citations and citation errors, GPT-4 is a major improvement over GPT-3.5." However, they emphasize that "users of ChatGPT are cautioned to check the citations it generates—and, of course, to evaluate the quality of the cited works themselves." They further note that "even with the latest version of ChatGPT, misinformation can be found throughout the generated texts—not just in the reference lists" and that "these errors are potentially dangerous, and they are exacerbated by the fact that ChatGPT often stands by its incorrect statements when asked to verify them."

Risk of bias

Selection bias: Topics chosen represent first-year composition course assignments, not specialized scientific domains; Temporal limitation: Data collected in first week of April 2023; software behavior may differ at other times; Prompt design bias: Specific prompt structure used may not represent typical user behavior; Verification bias: Reliance on multiple databases and websites for verification may miss obscure but real publications; Selection bias: Topics chosen are predominantly Wikipedia-style overviews of social, political, and environmental topics, not specialized scientific subjects; Limited to first-week April 2023 data collection period; Paper length constraints imposed by ChatGPT response field limits may affect citation behavior; Researcher bias in determining what constitutes a 'match or near-match' for real versus fabricated citations; No blinding of reviewers conducting the manual verification process; Selection bias: Topics selected are broad, Wikipedia-style overviews typical of undergraduate coursework, potentially not representative of specialized scientific or technical writing.; Limited temporal scope: Data generated in first week of April 2023; ChatGPT capabilities may have changed.; Researcher judgment bias: Determination of 'near-match' for real vs. fabricated works relied on researcher assessment rather than fully automated methods.; Database coverage bias: Searching multiple but finite set of databases may miss some real citations available elsewhere.; No blinding: Evaluators knew they were examining ChatGPT-generated citations.

Limitations

  • The study has notable limitations: First, "because detailed information on the use of ChatGPT is not available, we cannot know what proportion of users are taking advantage of the enhanced performance of GPT-4." Second, the study focused only on literature review papers of approximately 2000 words on multidisciplinary topics typical of first-year composition courses, which may not generalize to other document types or specialized domains
  • Third, the evaluation of hyperlinks and other formatting elements was limited.

Open questions raised

  • Need for continued investigation as ChatGPT technology improves
  • Lack of systematic study of fabricated citations in longer, specialized academic works
  • Limited understanding of why ChatGPT generates fabricated citations despite apparent attempt to recognize bibliographic data
  • Need for improvement in AI detection tools to account for fabricated citations
  • Further research needed on misinformation beyond reference lists in ChatGPT-generated texts
  • Whether detected fabricated citations can serve as a distinctive characteristic for AI-generated text detection tools
Data: Supplementary Appendix 1: List of 42 paper topics; Supplementary Appendix 2: Full texts of 84 papers generated by GPT-3.5 and GPT-4; Supplementary Appendix 3: Complete data compilation for all analyses; Supplementary Appendix 1: The 42 paper topics; Supplementary Appendix 2: The complete 84 texts generated by GPT-3.5 and GPT-4; Supplementary Appendix 3: The resulting data file with compiled bibliographic information; Supplementary Appendix 1: 42 paper topics; Supplementary Appendix 2: 84 texts generated by GPT-3.5 and GPT-4; Supplementary Appendix 3: Compiled data file for analyses; Available at: https://doi.org/10.1038/s41598-023-41032-5Extracted from: pdfAgreement 67%

Explore related topics

Related papers