AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition
Y. P. Wang, Zhaojun Ding, Xuansheng Wu, Siyue Sun, Ninghao Liu, Xiaoming Zhai · Proceedings of the AAAI Conference on Artificial Intelligence · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v40i48.42123
Methodology & findings
Study design
Comparative empirical evaluation using benchmark datasets.
Primary method
Metrics include QWK (quadratic weighted kappa for inter-rater agreement), Pearson/Spearman correlations, MAE (mean absolute error), and RMSE (root mean squared error). Specific statistical software or detailed analytical procedures are not described in the abstract.
Main result
AutoSCORE demonstrates that "across diverse tasks and rubrics, AutoSCORE predominantly improves scoring accuracy, human-machine agreement (QWK, correlations), and reduces error metrics (MAE, RMSE) compared to single-agent baselines, with particularly strong benefits on complex, multidimensional rubrics, and especially large relative gains on smaller LLMs."
Reports effect sizes.
Research paradigm
Positivist/empiricist
Author conclusions
The authors conclude that "structured component recognition combined with multi-agent design offers a scalable, reliable, and interpretable solution for automated scoring."
Risk of bias
Model selection bias: Only specific LLMs tested (GPT-4o, LLaMA-3.1-8B, LLaMA-3.1-70B); generalizability to other models unclear; Dataset bias: Evaluation limited to ASAP benchmark datasets; may not represent all scoring domains; Baseline comparison bias: Comparison only against 'single-agent baselines'; unclear if other structured approaches were compared; Potential selection bias in choice of benchmark datasets; potential model-dependent bias related to proprietary vs. open-source LLM performance differences; no mention of blinding or randomization in evaluation protocols.; Potential selection bias in dataset choice (ASAP benchmark only); model variation effects across different LLMs; prompt sensitivity mentioned as a challenge in existing approaches; potential confounding from model size differences.
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations