12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Investigating the Potential of Large Language Models for Automated Writing Scoring

Shan Wang · Atlantis Highlights in Computer Sciences/Atlantis highlights in computer sciences · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
0/4
Quality (LMQS)
E
Evidence
1
Citations
0.80
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2991/978-94-6463-502-7_116

Methodology & findings

Study design

Mixed-methods approach combining quantitative evaluation of agreement metrics (confusion matrix, Quadratic Weighted Kappa) with qualitative analysis of feedback quality

Primary method

Confusion matrix analysis, Quadratic Weighted Kappa metric for agreement measurement, qualitative analysis for feedback evaluation

Main result

The study found that "the results demonstrate a high level of agreement between GPT-4 scores and human raters, as evidenced by the confusion matrix and Quadratic Weighted Kappa metric." Additionally, "Qualitative analysis of GPT-4 feedback suggests its ability to provide constructive and comprehensive suggestions for improving student writing."

Reports effect sizes.

Research paradigm

Mixed-methods (quantitative and qualitative)

Author conclusions

The authors conclude that "this study proposes the use of LLM-based systems as formative assessment tools to complement human judgment," indicating they view GPT-4 as supplementary rather than a replacement for human assessment.

Risk of bias

Single model tested (GPT-4 only); Potential rater bias in human scorer selection not described; Limited information on essay sample diversity; Potential rater bias in human scoring comparison; Unknown sample representativeness (essay types, student populations, discipline domains); Unclear blinding procedures for human raters; Limited information on inter-rater reliability among human raters; Not explicitly stated in the abstract

Limitations

  • The abstract states that "there are still limitations surrounding LLM-based automated scoring and feedbacks." However, specific limitations are not detailed in the provided abstract text.

Open questions raised

  • The study identifies the need for further investigation into the limitations of LLM-based automated scoring and feedback systems, and recommends positioning such systems as complementary tools rather than standalone assessment solutions.
  • The study identifies the need for further refinement of LLM-based automated scoring systems and suggests that these systems should be used as complementary tools rather than standalone assessment solutions.
  • Not explicitly stated in the abstract
Data: not_statedCode: not_statedExtracted from: pdfAgreement 69%

Explore related topics

Related papers