How Teachers Can Use Large Language Models and Bloom’s Taxonomy to Create Educational Quizzes
Sabina Elkins, Ekaterina Kochmar, Jackie Chi Kit Cheung, Iulian Vlad Serban · Proceedings of the AAAI Conference on Artificial Intelligence · 2024
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v38i21.30353
Methodology & findings
Study design
Controlled experiment with real teachers (n=24) writing three types of quizzes (handwritten, simple generation, controlled generation) with Bloom's taxonomy-aligned questions.
Sample
N = 32, 11 groups
Primary method
Inter-annotator agreement measured using Cohen's κ for binary metrics (question-level) and Kendall's τ for ordinal metrics (quiz-level). Significance testing at α=0.05 level. Pairwise comparisons between quiz types. Coverage metric calculated as ratio of mapped text length to total passage length (pyramid method). No explicit mention of correction for multiple comparisons or specific statistical software used.
Main result
The study found that "teachers strongly prefer writing quizzes with the help of controlled generations" and "the quizzes with both controlled and simple generations are of comparable quality. Some metrics even point towards their superior quality, when compared to handwritten quizzes." The results demonstrate that "teachers prefer to write quizzes with automatically generated questions, and that such quizzes have no loss in quality compared to handwritten versions."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical (mixed-methods: experimental with qualitative feedback)
Author conclusions
"This paper aims to show that LLMs are capable of generating different types of questions from a given context that teachers find useful to create a quiz that is of comparable quality to a handwritten version... The results demonstrate that teachers strongly prefer writing quizzes with the help of controlled generations. They also directly copy more of the controlled generations than the simple generations, indicating that these questions are of higher quality or better suited to a teacher's goals. This confirms our hypothesis that teachers find automatically generated pedagogical questions useful for quiz writing."
Risk of bias
Selection bias: Teachers recruited through Upwork (BIO domain) and word-of-mouth (ML domain) may not be representative of all teachers; Unblinded outcome: Annotators were blinded to quiz type, reducing assessment bias; Single annotator for coverage metric: Only first author annotated coverage, introducing potential bias without agreement measurement; Contrived setting: Controlled lab environment may not reflect real teaching practices; Small sample: Only 24 teachers (12 per domain) limits generalizability; LLM selection bias: Only GPT-3.5 tested; results may not generalize to other LLMs; Selection bias: Teachers recruited through Upwork (BIO) and word-of-mouth (ML) may not be representative; Hawthorne effect: Teachers knew they were being video-recorded and evaluated; Experimenter bias: First author conducted coverage metric annotation with no inter-rater reliability measurement; Domain specificity: Only two domains tested (biology and machine learning); Language bias: English-language setting only; Annotator bias: Annotators blinded to quiz type, but training and agreement varied (κ=0.3-0.6 for question-level, τ>0.5 for quiz-level); Limited LLM: Only GPT-3.5 tested, not generalizable to other models; Selection bias: Teachers recruited via different mechanisms (Upwork for biology, word-of-mouth for ML) with different experience levels (high school vs. university); Ordering bias: Mitigated by randomizing quiz type order; Annotator bias: Blinded to quiz generation method, but only 4 annotators per domain evaluated majority of quizzes; Inter-rater variability: Fair to moderate agreement on question-level metrics (Cohen's κ=0.3-0.6), with acknowledgment that unbalanced dataset may not represent agreement on problem cases; Experimenter bias: Coverage metric annotated only by first author with no agreement measurement; Hawthorne effect: Controlled experimental setting may not reflect realistic quiz writing
Limitations
- The authors acknowledge that "the quiz writing setting is to some extent contrived
- In reality, teachers' quiz writing experiences will be subjective: they might use additional resources or existing knowledge, write a draft and return to it later to edit, not constrain the number of questions, have different learning goals in mind, and more." Additionally, "this work only considers one LLM, two domains, the English-language setting, and a limited number of teachers" and "a missing consideration in this work is the other half of the educational setting: students."
Open questions raised
- Need to remove controlled setting constraints to assess realistic teacher use of generated questions
- Expand beyond single LLM (GPT-3.5) to assess generalizability across models
- Expand beyond two domains (biology and machine learning) and English-language setting
- Include student goals, opinions, and performance to comprehensively understand implications
- Establish inter-annotator agreement for coverage metric (currently only annotated by first author)
- Need to remove experimental constraints to better assess how teachers realistically use generated questions
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations