12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

How Teachers Can Use Large Language Models and Bloom’s Taxonomy to Create Educational Quizzes

Sabina Elkins, Ekaterina Kochmar, Jackie Chi Kit Cheung, Iulian Vlad Serban · Proceedings of the AAAI Conference on Artificial Intelligence · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
3/4
Quality (LMQS)
E
Evidence
26
Citations
9.73
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1609/aaai.v38i21.30353

Methodology & findings

Study design

Controlled experiment with real teachers (n=24) writing three types of quizzes (handwritten, simple generation, controlled generation) with Bloom's taxonomy-aligned questions.

Sample

N = 32, 11 groups

Primary method

Inter-annotator agreement measured using Cohen's κ for binary metrics (question-level) and Kendall's τ for ordinal metrics (quiz-level). Significance testing at α=0.05 level. Pairwise comparisons between quiz types. Coverage metric calculated as ratio of mapped text length to total passage length (pyramid method). No explicit mention of correction for multiple comparisons or specific statistical software used.

Main result

The study found that "teachers strongly prefer writing quizzes with the help of controlled generations" and "the quizzes with both controlled and simple generations are of comparable quality. Some metrics even point towards their superior quality, when compared to handwritten quizzes." The results demonstrate that "teachers prefer to write quizzes with automatically generated questions, and that such quizzes have no loss in quality compared to handwritten versions."

Reports effect sizes and confidence intervals.

Research paradigm

Empirical (mixed-methods: experimental with qualitative feedback)

Author conclusions

"This paper aims to show that LLMs are capable of generating different types of questions from a given context that teachers find useful to create a quiz that is of comparable quality to a handwritten version... The results demonstrate that teachers strongly prefer writing quizzes with the help of controlled generations. They also directly copy more of the controlled generations than the simple generations, indicating that these questions are of higher quality or better suited to a teacher's goals. This confirms our hypothesis that teachers find automatically generated pedagogical questions useful for quiz writing."

Risk of bias

Selection bias: Teachers recruited through Upwork (BIO domain) and word-of-mouth (ML domain) may not be representative of all teachers; Unblinded outcome: Annotators were blinded to quiz type, reducing assessment bias; Single annotator for coverage metric: Only first author annotated coverage, introducing potential bias without agreement measurement; Contrived setting: Controlled lab environment may not reflect real teaching practices; Small sample: Only 24 teachers (12 per domain) limits generalizability; LLM selection bias: Only GPT-3.5 tested; results may not generalize to other LLMs; Selection bias: Teachers recruited through Upwork (BIO) and word-of-mouth (ML) may not be representative; Hawthorne effect: Teachers knew they were being video-recorded and evaluated; Experimenter bias: First author conducted coverage metric annotation with no inter-rater reliability measurement; Domain specificity: Only two domains tested (biology and machine learning); Language bias: English-language setting only; Annotator bias: Annotators blinded to quiz type, but training and agreement varied (κ=0.3-0.6 for question-level, τ>0.5 for quiz-level); Limited LLM: Only GPT-3.5 tested, not generalizable to other models; Selection bias: Teachers recruited via different mechanisms (Upwork for biology, word-of-mouth for ML) with different experience levels (high school vs. university); Ordering bias: Mitigated by randomizing quiz type order; Annotator bias: Blinded to quiz generation method, but only 4 annotators per domain evaluated majority of quizzes; Inter-rater variability: Fair to moderate agreement on question-level metrics (Cohen's κ=0.3-0.6), with acknowledgment that unbalanced dataset may not represent agreement on problem cases; Experimenter bias: Coverage metric annotated only by first author with no agreement measurement; Hawthorne effect: Controlled experimental setting may not reflect realistic quiz writing

Limitations

  • The authors acknowledge that "the quiz writing setting is to some extent contrived
  • In reality, teachers' quiz writing experiences will be subjective: they might use additional resources or existing knowledge, write a draft and return to it later to edit, not constrain the number of questions, have different learning goals in mind, and more." Additionally, "this work only considers one LLM, two domains, the English-language setting, and a limited number of teachers" and "a missing consideration in this work is the other half of the educational setting: students."

Open questions raised

  • Need to remove controlled setting constraints to assess realistic teacher use of generated questions
  • Expand beyond single LLM (GPT-3.5) to assess generalizability across models
  • Expand beyond two domains (biology and machine learning) and English-language setting
  • Include student goals, opinions, and performance to comprehensively understand implications
  • Establish inter-annotator agreement for coverage metric (currently only annotated by first author)
  • Need to remove experimental constraints to better assess how teachers realistically use generated questions
Data: "The input contexts, more details about Bloom's taxonomy, the human authored few-shot examples, all of the generated candidates and quizzes, the annotator demographics, and more can be found at https://anonymous.4open.science/r/EQG in practice2752/README.md"; Input contexts, generated candidates, quizzes, and annotator demographics available at https://anonymous.4open.science/r/EQG in practice2752/README.md; "The input contexts, more details about Bloom's taxonomy, the human authored few-shot examples, all of the generated candidates and quizzes, the annotator demographics, and more can be found at https://anonymous.4open.science/r/EQG in practice2752/README.md."Code: Not explicitly mentioned in the paper; Not mentioned; No code repository explicitly mentioned. Data repository provided but code availability not stated.Extracted from: pdfAgreement 37%

Explore related topics

Related papers