12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in Science

Clayton Cohn, Nicole Hutchins, Tuan Anh Le, Gautam Biswas · Proceedings of the AAAI Conference on Artificial Intelligence · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

6/10
Relevance
2/4
Quality (LMQS)
E
Evidence
38
Citations
51.41
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v38i21.30364

Methodology & findings

Study design

Empirical case study with iterative human-in-the-loop refinement.

Sample

N = 270, 6 groups

Primary method

Cohen's kappa (κ) for inter-rater reliability during consensus-building (target κ > 0.7); Cohen's Quadratic Weighted Kappa (QWK) for model-human agreement (accounts for degree of disagreement in ordinal data); Macro F1-Score for overall model performance (chosen for imbalanced dataset); Accuracy (reported for reference but not used for primary comparisons); Inductive coding (Charmaz 2006) for qualitative analysis of model-human disagreements; Thematic analysis using researcher memos (Hatch 2002) to identify patterns

Main result

The study found that "Across all questions, the model's scoring mostly aligned with the human scorers. Of the 11 subscores and total scores, 9 of them saw 'strong' agreement or better (QWK >= 0.8) at some point in the process (i.e., across the 4 implementations: 3 baselines and our Chain-of-Thought Prompting + Active Learning approach). 4 subscores achieved 'almost perfect' (QWK > 0.9) agreement." The Chain-of-Thought Prompting + Active Learning approach successfully scored and provided meaningful explanations for formative assessment responses in middle school Earth Science.

Reports effect sizes and confidence intervals.

Research paradigm

Pragmatist/Mixed-methods (combining empirical evaluation with qualitative analysis)

Author conclusions

"In this paper, we employed a Chain-of-Thought Prompting + Active Learning approach for scoring and explaining formative assessment question responses in a middle school Earth Science curriculum. Our results show that GPT-4, CoT reasoning, and active learning can be effectively leveraged toward accurate grading of science formative assessments. In several cases, the model achieved 'almost perfect' alignment with humans. The model generated relevant evidence linked to the rubric to help explain its scoring, which could benefit students and teachers."

Risk of bias

Selection bias: Data came from two studies at a single Southeastern U.S. public middle school; limited geographic/demographic diversity; Data impoverishment: Small, imbalanced dataset with non-canonical syntax and semantics; Overfitting risk: Active learning and CoT prompting showed tendency to overfit, particularly on simpler conceptual items; Scorer bias: Only 20% of data independently scored by two raters; consensus-building process may have introduced confirmation bias; Validation set bias: Validation-to-training ratio of 43:1 creates potential for spurious pattern detection during active learning; Model-centric bias: Only one LLM model (GPT-4) tested; no comparison with other LLMs; Selection bias: Dataset from single school in Southeastern U.S.; Data imbalance: Small, imbalanced dataset across subscores; Labeling bias: Human raters conducted initial scoring; potential for systematic human errors; Overfitting risk: Limited training instances led to overfitting in active learning iterations; Question design bias: Q2 had multiple correct answers leading to ambiguity even for human scorers; Selection bias: data from two specific SPICE studies at one southeastern U.S. middle school; Data impoverishment: small, imbalanced, non-canonical dataset; Inter-rater reliability dependency: model alignment dependent on human consensus scoring; Overfitting risk: particularly with CoT and active learning on simpler subscores; Limited generalizability: study focused only on Earth Science water runoff curriculum

Limitations

  • "With LLM approaches, ethical concerns arise with regard to privacy, bias, and hallucinations, and these concerns are amplified when they are deployed in high-stakes environments (e.g., classrooms with children)
  • In addition, while CoT has been shown to improve model performance over traditional ICL, the degree to which the reasoning chains guide the model's decision-making (if at all) is still an open question." Additionally, "our results also show that CoT and active learning can lead to overfitting, in particular, with simpler, easier-to-define subproblems."

Open questions raised

  • Need to explore whether the level of human IRR agreement can provide quantitative expectations for model performance
  • Investigation of 1-shot active learning to mitigate overfitting
  • Exploration of rubric refinement during active learning (currently excluded)
  • Investigation of how to provide feedback for positive performances without unexplained predictions
  • Extension of partnership with classroom teachers to determine how LLM output best fits their pedagogical needs
  • Investigation of how to use the method to evaluate students' learning performance and improve student learning
Data: Data and code available on GitHub repository (https://github.com/oele-isis-vanderbilt/EAAI24) with test code and sample data, though full dataset availability not explicitly stated.; Test code and sample data available in GitHub repository: https://github.com/oele-isis-vanderbilt/EAAI24; Assessment data from two Vanderbilt University-approved SPICE studies; data available via GitHub repository: https://github.com/oele-isis-vanderbilt/EAAI24 (test code and sample data mentioned)Code: https://github.com/oele-isis-vanderbilt/EAAI24Extracted from: pdfAgreement 56%

Explore related topics

Related papers