12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Undergraduate Pacific Studies Exam Generation and Answering Using Retrieval Augmented Generation and Large Language Models

E. P. T. Tyndall, Colleen Gayheart, Alexandre Some, Joseph Genz, Brent Langhals, Torrey Wagner · Proceedings of the ... Annual Hawaii International Conference on System Sciences/Proceedings of the Annual Hawaii International Conference on System Sciences · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

7/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.24251/hicss.2025.193

Methodology & findings

Study design

Experimental design with comparative assessment.

Sample

N = 56, 3 groups

Primary method

Text-similarity metrics including ROUGE-1, cosine similarity, and word embeddings were used for performance evaluation. Specific statistical software or additional analytical methods are not detailed in the abstract.

Main result

The study found that "RAG-assisted models outperformed those without access to the textbook, and that ChatGPT-4-Turbo was more accurate than ChatGPT-3.5-Turbo on nearly all exams." The findings demonstrate comparative performance differences between models with and without retrieval-augmented generation access to source material.

Reports effect sizes.

Research paradigm

Empiricist/Positivist

Author conclusions

The abstract indicates "The findings demonstrate the potential of generative artificial intelligence tools in academic assessments and provide insights into comparative performance of these models." However, the full set of author conclusions with complete verbatim quotes cannot be extracted from the abstract alone.

Risk of bias

Selection bias in choice of textbook material; Potential systematic differences in question generation across models; No blinding or independent verification of exam creation process; Limited scope to single subject domain (Pacific Studies); Single textbook source - limited generalizability to other course materials; Single course subject (Pacific Studies) - unclear if findings apply to other disciplines; No human grading comparison - reliance on automated text-similarity metrics may not capture pedagogical validity; Temporal limitation - ChatGPT models may have been updated between exam creation and evaluation; Selection bias: Limited to a single undergraduate textbook and subject matter; Potential confounding: Different model training data and versions may introduce systematic differences; Measurement bias: Reliance on automated text-similarity metrics (ROUGE-1, cosine similarity) may not capture semantic understanding equivalently across response types

Data: not_statedCode: not_statedExtracted from: pdfAgreement 60%

Explore related topics

Related papers