Undergraduate Pacific Studies Exam Generation and Answering Using Retrieval Augmented Generation and Large Language Models
E. P. T. Tyndall, Colleen Gayheart, Alexandre Some, Joseph Genz, Brent Langhals, Torrey Wagner · Proceedings of the ... Annual Hawaii International Conference on System Sciences/Proceedings of the Annual Hawaii International Conference on System Sciences · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.24251/hicss.2025.193
Methodology & findings
Study design
Experimental design with comparative assessment.
Sample
N = 56, 3 groups
Primary method
Text-similarity metrics including ROUGE-1, cosine similarity, and word embeddings were used for performance evaluation. Specific statistical software or additional analytical methods are not detailed in the abstract.
Main result
The study found that "RAG-assisted models outperformed those without access to the textbook, and that ChatGPT-4-Turbo was more accurate than ChatGPT-3.5-Turbo on nearly all exams." The findings demonstrate comparative performance differences between models with and without retrieval-augmented generation access to source material.
Reports effect sizes.
Research paradigm
Empiricist/Positivist
Author conclusions
The abstract indicates "The findings demonstrate the potential of generative artificial intelligence tools in academic assessments and provide insights into comparative performance of these models." However, the full set of author conclusions with complete verbatim quotes cannot be extracted from the abstract alone.
Risk of bias
Selection bias in choice of textbook material; Potential systematic differences in question generation across models; No blinding or independent verification of exam creation process; Limited scope to single subject domain (Pacific Studies); Single textbook source - limited generalizability to other course materials; Single course subject (Pacific Studies) - unclear if findings apply to other disciplines; No human grading comparison - reliance on automated text-similarity metrics may not capture pedagogical validity; Temporal limitation - ChatGPT models may have been updated between exam creation and evaluation; Selection bias: Limited to a single undergraduate textbook and subject matter; Potential confounding: Different model training data and versions may introduce systematic differences; Measurement bias: Reliance on automated text-similarity metrics (ROUGE-1, cosine similarity) may not capture semantic understanding equivalently across response types
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- ChatGPT: Bullshit spewer or the end of traditional assessments in higher education?Jürgen Rudolph · 2023 · 1,674 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations