Automatic Short Answer Grading for Finnish with ChatGPT
Li-Hsin Chang, Filip Ginter · Proceedings of the AAAI Conference on Artificial Intelligence · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v38i21.30363
Methodology & findings
Study design
Empirical experiment using two LLM-based chatbots (ChatGPT built on GPT-3.5 and GPT-4) to grade 2,000 Finnish short-answer student responses from ten undergraduate courses under zero-shot and one-shot settings.
Sample
N = 2377, 5 groups
Primary method
Quadratic-Weighted Kappa (QWK) for inter-rater agreement; Tolerance-Adjusted Accuracy (TAA) for percentage of accurately scored answers; Relative Merit Consensus (RMC) for preservation of relative answer merit; Pearson correlation coefficient for relationship between QWK and standard deviation; Survival curves showing percentage of questions meeting threshold values; Stratified sampling by grade for test answer selection
Main result
The study found that "GPT-4 achieves a good QWK score (0.6+) in 44% of one-shot settings, clearly outperforming GPT-3.5 at 21%." Additionally, "we observe a negative association between student answer length and model performance, as well as a correlation between a smaller standard deviation among a set of predictions and lower performance."
Reports effect sizes.
Research paradigm
empiricist/positivist
Author conclusions
"This study examined the feasibility of directly using LLM-based chatbots for the assessment of short answers. While immediate deployment presents certain challenges, the performance exhibited by one-shot GPT-4 justifies a more indepth exploration across multiple dimensions. These avenues encompass investigating the impact of employing additional shots, enhancing question clarity, and providing the model with comprehensive information, including grading criteria and reference answers. Lastly, investing more computational resources to include explanation generation alongside grading and exploring alternative scoring methods are all examples of eligible avenues for future research in this field."
Risk of bias
Language bias: ChatGPT performs less optimally on non-English data (Finnish); GPT tokenization optimized for English; Selection bias: Questions selected based on answer pool size and grade distribution balance, not random sampling; Evaluation bias: Possible that different evaluators graded answers to same question; Context loss: Data extracted from original educational context may not reflect true learning outcome measurement; Grading scale sensitivity: Metrics highly sensitive to number of possible score categories; Selection bias: Answer selection prioritized questions with greater number of answers and balanced grade distribution; Potential evaluator bias: Different evaluators may have graded answers to the same question; Language bias: Non-English dataset (Finnish) may negatively impact model performance; Context loss: Extraction of data from original educational context may not reflect true learning outcome measurement; Model tokenization bias: GPT models optimized for English tokenization may perform suboptimally on Finnish; Selection bias: Answers selected from larger pools with assumed balanced grade distribution may not represent typical answer distributions; Evaluator variability: Multiple evaluators may have graded answers to the same question, introducing inconsistency in gold standards; Language bias: Non-English dataset (Finnish) shown to negatively impact ChatGPT performance; tokenization optimized for English; Contextual loss: Data extracted from original educational context, reducing ecological validity; Lack of grading criteria: Absence of reference answers and explicit grading rubrics limits prompt quality and model guidance; Question selection bias: Prioritization of questions with balanced grades and specific answer lengths may not represent typical exam questions
Limitations
- "This exploratory study does have certain limitations
- The three metrics used are sensitive to the number of possible scores, and the questions have varying grading scales
- There are no results available from a few-shot baseline using smaller language models like BERT for comparison
- The absence of grading criteria and reference answer in the dataset constrains the content of the prompt
- Due to the data collection method, it is possible that answers to the same question were graded by different evaluators
- The extraction of data from its original educational context also poses challenges in assessing whether the model's accuracy truly reflects the accuracy of measuring underlying learning outcomes
Open questions raised
- Effects of increasing shots (more examples) in prompting
- Integration of grading criteria in prompts
- Chain-of-thought prompting for eliciting explanations before grade prediction
- Using models as second grader with human intervention detection
- Ranking or comparison of answers versus absolute grading
- Keyword extraction approaches
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- Artificial intelligence in higher education: the state of the fieldHelen Crompton · 2023 · 1,378 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- Ethics of AI in Education: Towards a Community-Wide FrameworkW. Holmes · 2021 · 1,056 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations