Graph RAG for Automated Short Answer Grading with Feedback: Bridging Pedagogical Needs and Technical Capabilities
Guoliang Xu, James E. Corter · Proceedings of the AAAI Conference on Artificial Intelligence · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v40i48.42125
Methodology & findings
Study design
Empirical evaluation using benchmark datasets with multiple experimental conditions: unseen-question and unseen-answer splits on the Short Answer Feedback (SAF) dataset with 31 topics.
Primary method
Comparative evaluation across model conditions; performance scaling analysis. Specific statistical tests and software not detailed in abstract.
Main result
The study found that "GraphRAG achieves grading accuracy comparable to vector-based RAG and generally superior to a fine-tuned LLM baseline model while providing more transparent source attribution." Additional key findings include that "Instructing the LLM to discretize continuous scores to match pedagogical rubrics, such as the 0.25 increments common in SAF, improves grading accuracy" and that "prompt-based length control substantially enhances feedback quality and its stability, achieving optimal balance of instructional richness and conciseness."
Reports effect sizes.
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude that "GraphRAG offers a robust, explainable, pedagogy-aligned, and cost-effective solution for large-scale educational applications, enabling transparent automated grading with effective pedagogical feedback and practical deployment costs."
Risk of bias
Not explicitly discussed in the abstract; Dataset bias: Evaluation limited to SAF dataset with 31 topics; Model selection bias: Comparison may not include all relevant LLM baselines; Domain specificity: Results may not generalize to other educational contexts; Dataset bias: Evaluation limited to Short Answer Feedback dataset with 31 topics - generalizability to other domains unclear; Model selection bias: Comparison limited to specific LLM variants (GPT-4o-mini, Claude-Opus-4); Potential instructor bias in ground truth annotations used for the SAF dataset; No mention of blinding or independent verification of grading accuracy
Open questions raised
- The paper identifies challenges in "transparency, pedagogical alignment, and cost-effectiveness that limit their real-world deployment" of current ASAG-F systems, which GraphRAG addresses.
- The paper addresses gaps in "transparency, pedagogical alignment, and cost-effectiveness that limit their real-world deployment" of automated short answer grading systems.
- The abstract indicates that automated short answer grading systems "currently face challenges in transparency, pedagogical alignment, and cost-effectiveness that limit their real-world deployment," suggesting these as areas where the research contributes.
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations