LLMs Do Not Grade Essays Like Humans
Jerin George Mathew, Sumayya Taher, Anindita Kundu, Denilson Barbosa · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Comparative empirical evaluation study.
Sample
N = 5269, 3 groups
Primary method
Quadratic Weighted Kappa (QWK) for measuring agreement between raters; Pearson correlation for measuring linear relationships between predicted and human scores; Mean Absolute Error (MAE) for measuring average absolute difference between scores; SHAP values for feature importance analysis; Aspect-Based Sentiment Analysis (ABSA) using DeBERTa-v3 model with confidence threshold of 0.9
Main result
The study found that "agreement between LLM and human scores remains relatively weak and varies with essay characteristics." In particular, "LLMs tend to assign higher scores to short or underdeveloped essays, while assigning lower scores to longer essays that contain minor grammatical or spelling errors." Additionally, "the scores generated by LLMs are generally consistent with the feedback they generate: essays receiving more praise tend to receive higher scores, while essays receiving more criticism tend to receive lower scores."
Reports effect sizes.
Research paradigm
Empirical-positivist
Author conclusions
The authors conclude: "Across several GPT and Llama models, we found that agreement between LLM-generated scores and human annotations is weak. In particular, disagreement varies systematically with essay characteristics: LLMs tend to assign higher scores to short or underdeveloped essays, while assigning lower scores to longer essays that contain minor grammatical or spelling errors." They further state that "the consistency between scores and feedback suggests that LLMs could assist in identifying strengths and weaknesses in student essays and provide rubric-aligned preliminary assessments. For students, the feedback may be useful for identifying language-level issues such as grammatical or spelling errors. However, the systematic differences observed between LLM and human grading indicate that such systems should currently be used as support tools rather than fully automated graders."
Risk of bias
Selection bias: Use of only two essay datasets (ASAP and DREsS); generalizability to other essay types or domains unclear; Model-specific bias: Different LLMs exhibit distinct score distributions; GPT-3.5 notably shows different behavior than other models; Annotation bias: Human raters may themselves have biases; only one rater used for DREsS dataset; Feature extraction bias: ABSA model may introduce classification errors; manually defined trait-related vocabularies may be incomplete; Prompt template bias: Single prompt template used across all models; acknowledged by authors as potential limitation; Anonymization effect: ASAP dataset uses NER placeholders that may affect model behavior in unknown ways; Selection bias in manually defined trait-related vocabularies that may miss linguistic expressions; Classification errors from ABSA model potentially introducing systematic bias in sentiment analysis; Single prompt template may favor certain model architectures; Use of NER placeholders in ASAP dataset anonymization may affect model perception; Potential overfitting of proxy model used for SHAP analysis to derived features rather than actual LLM decision process; Selection bias in essay datasets (ASAP and DREsS may not represent all essay types); Potential bias in human grader annotations used as ground truth; ABSA model classification errors may introduce systematic bias in feedback sentiment analysis; Single prompt template may introduce prompt design bias; Proxy model assumptions in SHAP analysis may not fully capture LLM scoring behavior; Manually defined trait vocabularies may miss linguistic expressions used by models
Limitations
- The authors acknowledge several limitations: "First, we relied on a single prompt template for all models
- Prompt design can influence LLM performance in automated essay scoring, as discussed by Mansour et al
- However, identifying a prompt that consistently improves performance across different models and essay types remains challenging." Additionally, "our analysis of feedback relied on an ABSA model to identify praise and criticism mentions in LLM-generated feedback
- Although this approach enables large-scale analysis, the model may introduce classification errors." Furthermore, "the extraction of trait-related terms relied on manually defined vocabularies derived from an initial inspection of feedback samples
- These term lists may not capture all linguistic expressions used by the models to refer to rubric traits, which could lead to missed mentions in the analysis." Finally, "the SHAP analysis relies on a proxy model trained to predict LLM-assigned scores using derived features (i.e., counts of positive and negative trait-specific mentions)
- This approach assumes that the proxy model reasonably approximates the scoring behavior of the LLM and that the selected features capture the most relevant signals used during scoring."
Open questions raised
- Need to evaluate analysis on more essay datasets beyond ASAP and DREsS
- Further study of how prompt variations influence grading behavior
- Examination of whether similar patterns hold across a wider range of LLM architectures
- Investigation of how instructors perceive and use LLM-generated evaluations in real grading settings
- Study of whether teachers agree with model-assigned scores and feedback
- Evaluation of whether feedback generated by LLMs leads to measurable improvements in student writing
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations