12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

LLMs Do Not Grade Essays Like Humans

Jerin George Mathew, Sumayya Taher, Anindita Kundu, Denilson Barbosa · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
2/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Comparative empirical evaluation study.

Sample

N = 5269, 3 groups

Primary method

Quadratic Weighted Kappa (QWK) for measuring agreement between raters; Pearson correlation for measuring linear relationships between predicted and human scores; Mean Absolute Error (MAE) for measuring average absolute difference between scores; SHAP values for feature importance analysis; Aspect-Based Sentiment Analysis (ABSA) using DeBERTa-v3 model with confidence threshold of 0.9

Main result

The study found that "agreement between LLM and human scores remains relatively weak and varies with essay characteristics." In particular, "LLMs tend to assign higher scores to short or underdeveloped essays, while assigning lower scores to longer essays that contain minor grammatical or spelling errors." Additionally, "the scores generated by LLMs are generally consistent with the feedback they generate: essays receiving more praise tend to receive higher scores, while essays receiving more criticism tend to receive lower scores."

Reports effect sizes.

Research paradigm

Empirical-positivist

Author conclusions

The authors conclude: "Across several GPT and Llama models, we found that agreement between LLM-generated scores and human annotations is weak. In particular, disagreement varies systematically with essay characteristics: LLMs tend to assign higher scores to short or underdeveloped essays, while assigning lower scores to longer essays that contain minor grammatical or spelling errors." They further state that "the consistency between scores and feedback suggests that LLMs could assist in identifying strengths and weaknesses in student essays and provide rubric-aligned preliminary assessments. For students, the feedback may be useful for identifying language-level issues such as grammatical or spelling errors. However, the systematic differences observed between LLM and human grading indicate that such systems should currently be used as support tools rather than fully automated graders."

Risk of bias

Selection bias: Use of only two essay datasets (ASAP and DREsS); generalizability to other essay types or domains unclear; Model-specific bias: Different LLMs exhibit distinct score distributions; GPT-3.5 notably shows different behavior than other models; Annotation bias: Human raters may themselves have biases; only one rater used for DREsS dataset; Feature extraction bias: ABSA model may introduce classification errors; manually defined trait-related vocabularies may be incomplete; Prompt template bias: Single prompt template used across all models; acknowledged by authors as potential limitation; Anonymization effect: ASAP dataset uses NER placeholders that may affect model behavior in unknown ways; Selection bias in manually defined trait-related vocabularies that may miss linguistic expressions; Classification errors from ABSA model potentially introducing systematic bias in sentiment analysis; Single prompt template may favor certain model architectures; Use of NER placeholders in ASAP dataset anonymization may affect model perception; Potential overfitting of proxy model used for SHAP analysis to derived features rather than actual LLM decision process; Selection bias in essay datasets (ASAP and DREsS may not represent all essay types); Potential bias in human grader annotations used as ground truth; ABSA model classification errors may introduce systematic bias in feedback sentiment analysis; Single prompt template may introduce prompt design bias; Proxy model assumptions in SHAP analysis may not fully capture LLM scoring behavior; Manually defined trait vocabularies may miss linguistic expressions used by models

Limitations

  • The authors acknowledge several limitations: "First, we relied on a single prompt template for all models
  • Prompt design can influence LLM performance in automated essay scoring, as discussed by Mansour et al
  • However, identifying a prompt that consistently improves performance across different models and essay types remains challenging." Additionally, "our analysis of feedback relied on an ABSA model to identify praise and criticism mentions in LLM-generated feedback
  • Although this approach enables large-scale analysis, the model may introduce classification errors." Furthermore, "the extraction of trait-related terms relied on manually defined vocabularies derived from an initial inspection of feedback samples
  • These term lists may not capture all linguistic expressions used by the models to refer to rubric traits, which could lead to missed mentions in the analysis." Finally, "the SHAP analysis relies on a proxy model trained to predict LLM-assigned scores using derived features (i.e., counts of positive and negative trait-specific mentions)
  • This approach assumes that the proxy model reasonably approximates the scoring behavior of the LLM and that the selected features capture the most relevant signals used during scoring."

Open questions raised

  • Need to evaluate analysis on more essay datasets beyond ASAP and DREsS
  • Further study of how prompt variations influence grading behavior
  • Examination of whether similar patterns hold across a wider range of LLM architectures
  • Investigation of how instructors perceive and use LLM-generated evaluations in real grading settings
  • Study of whether teachers agree with model-assigned scores and feedback
  • Evaluation of whether feedback generated by LLMs leads to measurable improvements in student writing
Data: ASAP (Automated Student Assessment Prize) - Kaggle competition dataset; DREsS (Dataset for Rubric-based Essay Scoring) - includes DREsS_New subset with 1,917 essays; ASAP (Automated Student Assessment Prize): https://kaggle.com/competitions/asap-aes; DREsS (Dataset for Rubric-based Essay Scoring): referenced in paper [45]; ASAP++ dataset: referenced as providing trait-level grades [28]; DREsS (Dataset for Rubric-based Essay Scoring): Available but no URL provided in paper; ASAP++ dataset with trait-level gradesExtracted from: pdfAgreement 62%

Explore related topics

Related papers