Playing Games with Ais: The Limits of GPT-3 and Similar Large Language Models
Adam Sobieszek, Tadeusz Price · Minds and Machines · 2022
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11023-022-09602-0
Methodology & findings
Study design
Theoretical analysis combining psychometric frameworks (Item Response Theory), formal logic analysis of the Turing Test, critical examination of GPT-3's architecture and conditional probability learning mechanisms, and empirical examples of GPT-3 outputs.
Main result
The authors argue that "GPT-3 is very good at generating plausible text it is a bad truth-teller." They demonstrate that GPT-3 does not fundamentally lack semantic ability, but rather cannot be forced into producing only true continuations. Instead, "to maximise their objective function they strategize to be plausible instead of truthful." The study identifies three fundamental limits: the regularity limit (questions must present as regularities in language), the priming limit (the ability to specify desired outputs through prompts), and the modal limit (GPT cannot reliably produce outputs faithful to the actual world). The authors show that "informativity is not a characteristic of any specific group of questions," including binary, mathematical, and trivia questions, contrary to prior claims.
Research paradigm
Philosophical and theoretical analysis of computational systems; interpretivist/conceptual analysis
Author conclusions
The authors conclude that "any statistical language generator will not be able to display consistent fidelity to the real world, and that while GPT-3 is very good at generating plausible text it is a bad truth-teller." They further conclude: "we highlighted some potential social issues that might arise if language models become widespread tools for writing, namely that the prevalence of these generators of plausible mistruths will permanently pollute our information ecosystems and the training sets of future language models." They propose the "modal drift hypothesis—that because humans are not psychologically equipped to effectively differentiate truth from plausible falsehoods among texts generated by language models, a mass production of plausible texts may deteriorate the quality of our informational ecosystem."
Risk of bias
The analysis is primarily theoretical without systematic empirical validation across diverse question types and contexts; Limited discussion of potential confirmation bias in selecting examples that support the theoretical position; The authors' pre-existing theoretical frameworks (from psychometrics and linguistics) may influence interpretation of GPT-3 capabilities; Selection bias in choice of test questions and examples used to evaluate GPT-3; Limited empirical validation - primarily theoretical argument rather than systematic experimental testing; Confirmation bias in cherry-picking examples that support the plausibility thesis (John Prescott example); Reliance on informal/anecdotal testing rather than controlled experimental conditions; No systematic comparison of multiple LLMs beyond occasional reference to GPT-2; The analysis relies on qualitative interpretation of examples rather than systematic empirical testing; Selection of examples (John Prescott) may not be representative of all GPT-3 failure modes; No quantitative evaluation of actual error rates or frequency of plausible falsehoods; Limited empirical data on human performance in distinguishing AI-generated text
Limitations
- The authors acknowledge that "it is not easy with GPT-3, as it has no knowledge of the context in which it is being used, and the only thing that we can truly control is the prompt we provide it with." They also note limitations in specifying tasks: "it is harder to precisely locate a procedure from its description
- That's why few-shot learning works: giving GPT some examples of correct completions works to locate in the space of tasks the one we wanted GPT-3 to engage in." Additionally, they state that "non-fiction writing is still filled with utterances either beyond the scope of propositional logic, or that stripped of their context can appear non-actual."
Open questions raised
- Future work on how to engineer handicaps or safeguards to prevent language models from generating plausible falsehoods
- Research on methods to make language models explicitly grounded in actual world knowledge
- Studies on how to detect AI-generated text and protect information ecosystems from degradation
- Investigation of how synthetic data training (models trained on previous model outputs) compounds truth-telling problems
- Need for systematic empirical testing of GPT-3's abilities on carefully standardized question sets
- Understanding how to better prime language models for specific tasks
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations