Can ChatGPT enhance business student creativity? Evidence from a randomised controlled trial
Rosemary Fisher, Taylor Gogan, James S. Williams, Richard Laferriere, Gordon Campbell, Asanka Gunasekara et al. · Studies in Higher Education · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1080/03075079.2025.2515512
Methodology & findings
Study design
Randomised controlled trial (RCT) with stratified random sampling across three academic disciplines (Entrepreneurship, Marketing, Management).
Sample
N = 1190, 17 groups
Primary method
Two Multivariate Analysis of Variance (MANOVA) tests were conducted: (1) 2 (Discipline: Entrepreneurship vs Marketing) × 2 (Group: Intervention vs Control) MANOVA with creativity components as dependent variables; (2) separate MANOVA for Management discipline due to different task. Pillai's trace reported for first MANOVA due to Box's M violation; Wilks Lambda for second MANOVA. Follow-up univariate F-tests for each creativity component. Pearson correlations used to assess inter-rater reliability. Pairwise comparisons for discipline-specific effects. Data pre-processing included removal of outliers based on Mahalanobis distances, Q-Q plot inspection, and multicollinearity assessment.
Main result
The study found that "while participants who used ChatGPT-3.5 generally scored higher in creativity compared to those in the control condition, this effect varied depending on the students' discipline and the specific component of creativity being measured." Specifically, "ChatGPT-3.5 use had little to no impact on creativity among Entrepreneurship students but tended to enhance creativity among Marketing students and, to a lesser extent, Management students." Additionally, "the intervention group scored significantly higher than the control group across all outcomes" in Marketing, while for Management students, "the intervention group scoring significantly higher than the control group" in fluency and elaboration but not originality.
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/quantitative empirical
Author conclusions
"This research highlights the potential and challenges of genAI, specifically ChatGPT-3.5, for university student creativity. While the use of ChatGPT-3.5 significantly enhances fluency and elaboration in some disciplines, its impact on originality was inconsistent and seemingly varied by task and student proclivities. The findings underscore the need for educators to tailor pedagogical approaches to balance AI-enhanced outputs with activities that foster self-efficacy and intrinsic creative motivation." The authors also conclude that "By fostering both creative and AI self-efficacy in students, educators can equip students to critically engage with genAI as a collaborative tool that serves as a catalyst for growth rather than a substitute for their abilities or crutch."
Risk of bias
Selection bias: Students self-selected into elective Entrepreneurship subject vs. core Marketing and Management subjects, potentially confounding discipline effects with motivation; Attrition: Duplications removed during data cleansing (students attending two subjects); Confounders: Prior experience with AI and prompt engineering not measured; prompt quality not standardized; task type differences (paperclip task vs. return-to-workplace task); Measurement bias: Low inter-rater reliability for originality (rs = .32-.49), comparatively lower agreement for elaboration (rs = .60-.66); Assumption violations: Box's M test significant for first MANOVA (unequal variance-covariance matrices); multicollinearity between fluency and flexibility; non-normal residuals requiring outlier removal; Assessment bias: Raters may have perceived AI-assisted ideas as less unique based on quantity of responses; Selection bias: Students self-selected into elective Entrepreneurship versus core Marketing and Management subjects, potentially confounding discipline effects with motivation/ability differences; Measurement bias: Low inter-rater reliability for originality (rs = .32-.49) compared to fluency (rs = .83-.95); Confounding: Different task types used for Management (return-to-workplace) versus Entrepreneurship/Marketing (paperclip), though authors note findings did not appear to vary as a function of task type; Subjective assessment: Rater perception of AI-assisted responses may have been influenced by quantity of ideas generated; Lack of blinding: Raters aware of group assignment based on response quantity and AI assistance indicators; Selection bias: Students self-selected into elective Entrepreneurship subject while Marketing and Management were core requirements, affecting intrinsic motivation and sample composition; Measurement bias: Low inter-rater reliability for originality (rs = .32-.49) and elaboration (rs = .60-.66) measures; Confounding: Different tasks used across disciplines (paperclip vs. return-to-office) may confound disciplinary effects; Attrition/data cleaning: Duplications removed during data cleansing; multivariate outliers identified and removed (n=2); Violation of MANOVA assumptions: Box's M test was significant for first MANOVA indicating unequal variance-covariance matrices; residuals significantly non-normal before outlier removal; Potential experimenter bias: Five different researchers administered the study across 26 classes; Uncontrolled prompt quality: "a lack of standardisation with respect to prompt quality" across participants in the ChatGPT intervention group
Limitations
- "The study is limited by a lack of generalisability
- Additionally, individual differences in prior experience with AI and prompt engineering, a lack of standardisation with respect to prompt quality, the nature of the tasks themselves, and the subjective assessment of creativity potentially affected the outcomes." The authors also note that "intrinsic motivation and individual baseline creativity were not explicitly measured in this study, this interpretation remains speculative and a limitation" and acknowledge uncertainty about "whether the results tended towards conventional ideas that did not sufficiently push the boundaries of originality."
Open questions raised
- Direct measurement of intrinsic motivation and baseline creativity using validated instruments
- Examination of how pedagogical practices can be tailored to influence genAI's capacity to support student creativity
- Investigation of the interplay between task type and whether a subject is core or elective
- Study of AI's impact on originality across different contexts
- Research on how students actively engage with generative AI for creative tasks in educational settings (identified as scarce in literature)
- Domain-specific investigations of creative processes as they vary across different fields of study
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Do AI chatbots improve students learning outcomes? Evidence from a meta‐analysisRong Wu · 2023 · 469 citations