12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyond

Mike Perkins · Journal of University Teaching and Learning Practice · 2023

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

8/10
Relevance
0/4
Quality (LMQS)
I
Evidence
668
Citations
23.66
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.53761/1.20.02.07

Methodology & findings

Study design

Narrative literature review with critical analysis of Large Language Models (LLMs) such as ChatGPT and GPT-3.

Main result

The study found that "LLMs have already progressed to the point that neither trained academic staff or technological tools can consistently determine whether text is generated by an LLM or by a human." The research demonstrates that "the current generation of LLMs are already fluent in their output, and emerging research has suggested that existing LLMs can produce output which humans struggle to identify as being machine created" (Abd-Elaal et al., 2022; Clark et al., 2021; Gunser et al., 2021; Köbis & Mossink, 2021; Wahle et al., 2021). Additionally, the paper shows that "at the same time, given that the text that is produced by LLMs is uniquely created based on the inputs provided, current research suggests that use of the created text by students is unlikely to be spotted by existing text-matching software tools used by HEIs" (Wahle, Ruas, Kirstein, et al., 2022).

Research paradigm

Interpretivist/Critical analysis of academic integrity policy

Author conclusions

The authors conclude that "it is not the use of the tools themselves that defines whether plagiarism or a breach of academic integrity has occurred, but whether any such use is made clear." They state that "Deciding whether any particular use of LLMs by students may be defined as academic misconduct will be determined by the future policies of any given HEI, and this highlights the importance of creating clear academic integrity policies and educating students in any acceptable use cases of LLMs." Furthermore, they argue that "a blanket ban of these tools at an institutional level is neither feasible, nor enforceable" and that "HEIs must consider the implications of this in future policy development." The authors ultimately believe that "the future integration of LLMs and other AI supported digital tools into the classroom environment is highly likely, and therefore HEIs must consider the implications of this in future policy development."

Risk of bias

Selection bias in reviewed studies - not all studies on LLM detection capability included; Publication bias - studies finding successful detection may be more likely to be published; Methodological heterogeneity across reviewed detection studies; Limited sample sizes in some detection studies reviewed (e.g., Clark et al., 2021 used n=780 but was one of larger studies); Participant awareness bias in detection experiments - participants knew LLM text was present; Selection bias in reviewed studies: Research on LLM detection relies on small, non-representative samples; Temporal bias: Tools and capabilities evolving rapidly; findings may be outdated quickly; Publication bias: May preferentially review published studies over grey literature; Confirmation bias: Authors may selectively highlight studies supporting their narrative on AI detection difficulty; Lack of direct empirical data from the authors themselves on detection rates; Selection bias in reviewed studies - experimental participants aware of AI-generated text presence; Temporal bias - rapid evolution of LLM capabilities makes findings quickly outdated; Publication bias - tendency to publish studies showing detection difficulties; Potential discrimination bias in detection tools against EFL students

Limitations

  • The authors acknowledge that "an unavoidable methodological issue with these experimental studies is that in order to determine whether participants can accurately identify text as machine or human generated, participants need to be aware that some of the text they are about to encounter may be machine generated before participating in an experiment." They further note that "given the novelty of these tools, it is likely that even experienced academic staff are simply not aware of the capabilities that these tools have, as demonstrated with the participants in Clark et al.'s study (2021)
  • This may result in their ability to identify any LLM produced output 'in situ' when evaluating work being even lower than demonstrated in an experimental design." Additionally, they state that "given that both academic staff, as well as technological methods of detection are unable to accurately detect machine generated text and therefore student uses of LLM based tools, this presents a clear threat to academic integrity for HEIs, requiring a range of adjustments to be made in both practice and policy."

Open questions raised

  • Further empirical research needed on how academic staff can be trained or supported to detect AI tool use in student work
  • Need for studies assessing detection accuracy of newer tools (GPTZero, Crossplag AI detect) in academic settings
  • Research required on the long-term impacts of LLM integration into Digital Writing Assistants
  • Studies needed on how different HEIs define and enforce academic integrity policies regarding LLM use
  • Exploration of human-AI co-creation dynamics and control issues raised by writers using current generation LLMs
  • Investigation of copyright and fairness issues in human-AI co-creation across creative and academic domains
Extracted from: pdfAgreement 66%

Explore related topics

Related papers