Investigating the Potential of Large Language Models for Automated Writing Scoring
Shan Wang · Atlantis Highlights in Computer Sciences/Atlantis highlights in computer sciences · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2991/978-94-6463-502-7_116
Methodology & findings
Study design
Mixed-methods approach combining quantitative evaluation of agreement metrics (confusion matrix, Quadratic Weighted Kappa) with qualitative analysis of feedback quality
Primary method
Confusion matrix analysis, Quadratic Weighted Kappa metric for agreement measurement, qualitative analysis for feedback evaluation
Main result
The study found that "the results demonstrate a high level of agreement between GPT-4 scores and human raters, as evidenced by the confusion matrix and Quadratic Weighted Kappa metric." Additionally, "Qualitative analysis of GPT-4 feedback suggests its ability to provide constructive and comprehensive suggestions for improving student writing."
Reports effect sizes.
Research paradigm
Mixed-methods (quantitative and qualitative)
Author conclusions
The authors conclude that "this study proposes the use of LLM-based systems as formative assessment tools to complement human judgment," indicating they view GPT-4 as supplementary rather than a replacement for human assessment.
Risk of bias
Single model tested (GPT-4 only); Potential rater bias in human scorer selection not described; Limited information on essay sample diversity; Potential rater bias in human scoring comparison; Unknown sample representativeness (essay types, student populations, discipline domains); Unclear blinding procedures for human raters; Limited information on inter-rater reliability among human raters; Not explicitly stated in the abstract
Limitations
- The abstract states that "there are still limitations surrounding LLM-based automated scoring and feedbacks." However, specific limitations are not detailed in the provided abstract text.
Open questions raised
- The study identifies the need for further investigation into the limitations of LLM-based automated scoring and feedback systems, and recommends positioning such systems as complementary tools rather than standalone assessment solutions.
- The study identifies the need for further refinement of LLM-based automated scoring systems and suggests that these systems should be used as complementary tools rather than standalone assessment solutions.
- Not explicitly stated in the abstract
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations