A Chain-of-Thought Prompting Approach with LLMs for Evaluating Students’ Formative Assessment Responses in Science
Clayton Cohn, Nicole Hutchins, Tuan Anh Le, Gautam Biswas · Proceedings of the AAAI Conference on Artificial Intelligence · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v38i21.30364
Methodology & findings
Study design
Empirical case study with iterative human-in-the-loop refinement.
Sample
N = 270, 6 groups
Primary method
Cohen's kappa (κ) for inter-rater reliability during consensus-building (target κ > 0.7); Cohen's Quadratic Weighted Kappa (QWK) for model-human agreement (accounts for degree of disagreement in ordinal data); Macro F1-Score for overall model performance (chosen for imbalanced dataset); Accuracy (reported for reference but not used for primary comparisons); Inductive coding (Charmaz 2006) for qualitative analysis of model-human disagreements; Thematic analysis using researcher memos (Hatch 2002) to identify patterns
Main result
The study found that "Across all questions, the model's scoring mostly aligned with the human scorers. Of the 11 subscores and total scores, 9 of them saw 'strong' agreement or better (QWK >= 0.8) at some point in the process (i.e., across the 4 implementations: 3 baselines and our Chain-of-Thought Prompting + Active Learning approach). 4 subscores achieved 'almost perfect' (QWK > 0.9) agreement." The Chain-of-Thought Prompting + Active Learning approach successfully scored and provided meaningful explanations for formative assessment responses in middle school Earth Science.
Reports effect sizes and confidence intervals.
Research paradigm
Pragmatist/Mixed-methods (combining empirical evaluation with qualitative analysis)
Author conclusions
"In this paper, we employed a Chain-of-Thought Prompting + Active Learning approach for scoring and explaining formative assessment question responses in a middle school Earth Science curriculum. Our results show that GPT-4, CoT reasoning, and active learning can be effectively leveraged toward accurate grading of science formative assessments. In several cases, the model achieved 'almost perfect' alignment with humans. The model generated relevant evidence linked to the rubric to help explain its scoring, which could benefit students and teachers."
Risk of bias
Selection bias: Data came from two studies at a single Southeastern U.S. public middle school; limited geographic/demographic diversity; Data impoverishment: Small, imbalanced dataset with non-canonical syntax and semantics; Overfitting risk: Active learning and CoT prompting showed tendency to overfit, particularly on simpler conceptual items; Scorer bias: Only 20% of data independently scored by two raters; consensus-building process may have introduced confirmation bias; Validation set bias: Validation-to-training ratio of 43:1 creates potential for spurious pattern detection during active learning; Model-centric bias: Only one LLM model (GPT-4) tested; no comparison with other LLMs; Selection bias: Dataset from single school in Southeastern U.S.; Data imbalance: Small, imbalanced dataset across subscores; Labeling bias: Human raters conducted initial scoring; potential for systematic human errors; Overfitting risk: Limited training instances led to overfitting in active learning iterations; Question design bias: Q2 had multiple correct answers leading to ambiguity even for human scorers; Selection bias: data from two specific SPICE studies at one southeastern U.S. middle school; Data impoverishment: small, imbalanced, non-canonical dataset; Inter-rater reliability dependency: model alignment dependent on human consensus scoring; Overfitting risk: particularly with CoT and active learning on simpler subscores; Limited generalizability: study focused only on Earth Science water runoff curriculum
Limitations
- "With LLM approaches, ethical concerns arise with regard to privacy, bias, and hallucinations, and these concerns are amplified when they are deployed in high-stakes environments (e.g., classrooms with children)
- In addition, while CoT has been shown to improve model performance over traditional ICL, the degree to which the reasoning chains guide the model's decision-making (if at all) is still an open question." Additionally, "our results also show that CoT and active learning can lead to overfitting, in particular, with simpler, easier-to-define subproblems."
Open questions raised
- Need to explore whether the level of human IRR agreement can provide quantitative expectations for model performance
- Investigation of 1-shot active learning to mitigate overfitting
- Exploration of rubric refinement during active learning (currently excluded)
- Investigation of how to provide feedback for positive performances without unexplained predictions
- Extension of partnership with classroom teachers to determine how LLM output best fits their pedagogical needs
- Investigation of how to use the method to evaluate students' learning performance and improve student learning
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations