12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Graph RAG for Automated Short Answer Grading with Feedback: Bridging Pedagogical Needs and Technical Capabilities

Guoliang Xu, James E. Corter · Proceedings of the AAAI Conference on Artificial Intelligence · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

6/10
Relevance
0/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v40i48.42125

Methodology & findings

Study design

Empirical evaluation using benchmark datasets with multiple experimental conditions: unseen-question and unseen-answer splits on the Short Answer Feedback (SAF) dataset with 31 topics.

Primary method

Comparative evaluation across model conditions; performance scaling analysis. Specific statistical tests and software not detailed in abstract.

Main result

The study found that "GraphRAG achieves grading accuracy comparable to vector-based RAG and generally superior to a fine-tuned LLM baseline model while providing more transparent source attribution." Additional key findings include that "Instructing the LLM to discretize continuous scores to match pedagogical rubrics, such as the 0.25 increments common in SAF, improves grading accuracy" and that "prompt-based length control substantially enhances feedback quality and its stability, achieving optimal balance of instructional richness and conciseness."

Reports effect sizes.

Research paradigm

Empiricist/Positivist

Author conclusions

The authors conclude that "GraphRAG offers a robust, explainable, pedagogy-aligned, and cost-effective solution for large-scale educational applications, enabling transparent automated grading with effective pedagogical feedback and practical deployment costs."

Risk of bias

Not explicitly discussed in the abstract; Dataset bias: Evaluation limited to SAF dataset with 31 topics; Model selection bias: Comparison may not include all relevant LLM baselines; Domain specificity: Results may not generalize to other educational contexts; Dataset bias: Evaluation limited to Short Answer Feedback dataset with 31 topics - generalizability to other domains unclear; Model selection bias: Comparison limited to specific LLM variants (GPT-4o-mini, Claude-Opus-4); Potential instructor bias in ground truth annotations used for the SAF dataset; No mention of blinding or independent verification of grading accuracy

Open questions raised

  • The paper identifies challenges in "transparency, pedagogical alignment, and cost-effectiveness that limit their real-world deployment" of current ASAG-F systems, which GraphRAG addresses.
  • The paper addresses gaps in "transparency, pedagogical alignment, and cost-effectiveness that limit their real-world deployment" of automated short answer grading systems.
  • The abstract indicates that automated short answer grading systems "currently face challenges in transparency, pedagogical alignment, and cost-effectiveness that limit their real-world deployment," suggesting these as areas where the research contributes.
Data: not_statedCode: not_statedExtracted from: pdfAgreement 67%

Explore related topics

Related papers