Mapping the Challenges of HCI: An Application and Evaluation of ChatGPT for Mining Insights at Scale
Jonas Oppenlaender, Joonas Hämäläinen · International Journal of Human-Computer Interaction · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1080/10447318.2026.2647543
Methodology & findings
Study design
Two-step LLM-based text mining approach: Step 1 used ChatGPT (gpt-3.5-turbo-0301) to extract candidate research challenges from 879 CHI 2023 papers; Step 2 used GPT-4 (gpt-4-0314) to filter results to top 5 challenges per paper.
Sample
N = 879, 6 groups
Primary method
Cohen's kappa (κ) for inter-rater agreement; Cosine similarity/distance metrics for semantic analysis; Silhouette score (0.27) for cluster consistency evaluation; Descriptive statistics (means, standard deviations, quartiles, min/max)
Main result
We identify 4,392 HCI research challenges in 113 highly diverse topics from the CHI 2023 proceedings. The study found that "the combination of ChatGPT and GPT-4 makes an excellent cost-efficient means for analyzing a text corpus at scale." ChatGPT extracted 34,638 candidate challenges (M = 39.3 per paper, SD = 21.5), which GPT-4 filtered to 4,392 final challenges. The qualitative evaluation showed "very high inter-rater agreement (κ = 0.86) between the raters on the question of how well the challenges align with human-identified challenges," with "about two thirds of the sampled papers (65.9% and 63.6%) the two raters thought the human-identified list of research challenges matched perfectly with the GPT-4-extracted list."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical - applied evaluation of computational tools on real-world task
Author conclusions
"We are entering an era where foundation-scale models are becoming increasingly influential." The authors conclude that "LLMs represent a significant addition to researchers' toolkit, potentially indicating an evolving shift in some areas of knowledge work." They argue that "Cost-efficiency is key for flexibly prototyping research ideas and analyzing text corpora from different perspectives, with implications for applying LLMs for mining insights in academia and practice." The study demonstrates that "the LLMs' ability for mining insights has wide-ranging and transformative implications for research in academia."
Risk of bias
Single human annotator may introduce subjective bias in evaluation; Human annotator had expertise gaps in certain HCI subdomains (e.g., embodied interaction); Fatigue effects noted in human annotator performance; Prompt design sensitivity - outputs sensitive to specific prompt wording; CHI 2023 proceedings may not represent entire HCI field; Potential bias in LLM training data (acknowledged for ChatGPT and GPT-4); Single human annotator for qualitative evaluation may introduce rater bias, though inter-rater agreement with authors was assessed; Prompt design sensitivity: different prompts could yield different results; Limited domain expertise across all 113 HCI subdomains; Manual labeling of 113 topics by first author only (no inter-rater agreement reported for topic labeling); Human annotator fatigue and task complexity noted as affecting annotation quality; LLM hallucination risks, though authors report finding no clear hallucinations in sampled set; Selection bias: CHI 2023 proceedings represent only a subset of HCI research, excluding work published in other venues; Annotator bias: Single postdoctoral researcher conducted human annotation; fatigue effects noted; Model bias: ChatGPT and GPT-4 trained on data with unknown composition; potential for hallucinations or biases in training data; Prompt design bias: Iterative prompt development focused on achieving 'consistent results' rather than objective ground truth; prompts explicitly instruct models to 'forget previous instructions'; Topic labeling bias: Final 113 topics manually labeled by first author only, without inter-rater reliability assessment; Temporal bias: Text corpus published after LLM training cutoffs, but specific training data composition for both models unknown
Limitations
- The study acknowledged several limitations: "extracting research challenges from text and filtering this list to a concise set of five statements is a complex task." The human annotator "faced several difficulties, particularly in domains outside their direct expertise, such as embodied interaction." Additionally, "the human annotator was subject to fatigue, and later tried to complete the task by scanning headlines and copy-pasting verbatim statements rather than analyzing and rephrasing the extracted statements into research challenges." The evaluation relied on a single human annotator and manual topic labeling by one author rather than multiple independent coders.
Open questions raised
- Need for work that supports scholars in understanding the overall direction of the HCI field
- Alignment between current HCI research challenges and HCI's grand challenges requires reassessment
- Limited research on real-world applications of LLMs versus benchmark-based evaluation
- Gap between academia's slow adoption and industry's rapid development of generative AI applications
- Missing opportunities in applying LLMs for mining insights from text corpora in academia
- Need for evaluation of LLMs on real-world tasks rather than solely academic benchmarks
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations