12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AutoSCORE: Enhancing Automated Scoring with Multi-Agent Large Language Models via Structured Component Recognition

Y. P. Wang, Zhaojun Ding, Xuansheng Wu, Siyue Sun, Ninghao Liu, Xiaoming Zhai · Proceedings of the AAAI Conference on Artificial Intelligence · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

5/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v40i48.42123

Methodology & findings

Study design

Comparative empirical evaluation using benchmark datasets.

Primary method

Metrics include QWK (quadratic weighted kappa for inter-rater agreement), Pearson/Spearman correlations, MAE (mean absolute error), and RMSE (root mean squared error). Specific statistical software or detailed analytical procedures are not described in the abstract.

Main result

AutoSCORE demonstrates that "across diverse tasks and rubrics, AutoSCORE predominantly improves scoring accuracy, human-machine agreement (QWK, correlations), and reduces error metrics (MAE, RMSE) compared to single-agent baselines, with particularly strong benefits on complex, multidimensional rubrics, and especially large relative gains on smaller LLMs."

Reports effect sizes.

Research paradigm

Positivist/empiricist

Author conclusions

The authors conclude that "structured component recognition combined with multi-agent design offers a scalable, reliable, and interpretable solution for automated scoring."

Risk of bias

Model selection bias: Only specific LLMs tested (GPT-4o, LLaMA-3.1-8B, LLaMA-3.1-70B); generalizability to other models unclear; Dataset bias: Evaluation limited to ASAP benchmark datasets; may not represent all scoring domains; Baseline comparison bias: Comparison only against 'single-agent baselines'; unclear if other structured approaches were compared; Potential selection bias in choice of benchmark datasets; potential model-dependent bias related to proprietary vs. open-source LLM performance differences; no mention of blinding or randomization in evaluation protocols.; Potential selection bias in dataset choice (ASAP benchmark only); model variation effects across different LLMs; prompt sensitivity mentioned as a challenge in existing approaches; potential confounding from model size differences.

Data: not_statedCode: not_statedExtracted from: pdfAgreement 71%

Explore related topics

Related papers