Leveraging large language models for thematic analysis: a case study in the charity sector
Paul Clough · AI & Society · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00146-025-02487-4
Methodology & findings
Study design
Case study with empirical experiments.
Sample
N = 14.5, 9 groups
Primary method
Cohen's kappa (κ) for inter-rater reliability on multi-label classification; instance-based kappa pooling decisions across instances and codes; precision and recall calculations for excerpt extraction; cosine similarity using embedding vectors (text-embedding-ada-002); weighted scoring system for Agreement Rate on corporate priority categorization; descriptive statistics including mean, standard deviation (SD), and ranges; hallucination rate calculated as proportion of runs not following instructions; wall clock time (execution time in minutes); 5-point Likert scale analysis for qualitative assessment. Python code used for metric computation.
Main result
In deductive coding, "GPT-4o achieved a substantial agreement in multi-label thematic categorization (κ = 0.61–0.65), while GPT-4o-mini showed a moderate agreement (κ = 0.41–0.58). Both models excelled in sentiment analysis (κ = 0.91–0.95), but struggled with evaluating evidence of impact due to contextual complexity (κ ≤ 0.01)." For inductive coding, "GPT-4o demonstrated a strong semantic alignment with human-generated themes (cosine similarity = 0.76–0.79) though its tendency toward broad themes required human refinement."
Reports effect sizes and confidence intervals.
Research paradigm
Pragmatist/Mixed-methods (combining positivist measurement with interpretivist qualitative assessment)
Author conclusions
"Results show that GPT-4o models can effectively serve as both an initial coding tool and a validation mechanism for inductive and deductive coding processes within an LLM-human collaborative framework." Furthermore, "our findings underline the indispensable role of human expertise—from prompt engineering and managing hallucinations to final verification—to ensure accurate and trustworthy AI-assisted analyses. While LLMs can enhance qualitative analysis, their full potential is only realized under skilled human guidance."
Risk of bias
Single expert coder (Tearfund staff) as gold standard—no inter-rater reliability check against multiple human coders; Selection bias: datasets from single organization (Tearfund) with specific evaluation workflows; Task complexity varies across datasets; not all datasets fully coded for all coding types; Prompt engineering bias: prompts iteratively refined based on outputs, potential circularity; Limited diversity in codebooks tested (Light Wheel and SDGs only); Temperature settings fixed at 0 and 0.5—no systematic exploration of parameter sensitivity; Selection bias: Single case study organization (Tearfund) may not represent broader charity sector contexts; Evaluator bias: Expert coder annotations serve as gold standard but generated by single individual at Tearfund; Model bias: GPT-4o trained on data with potential societal biases that may manifest in qualitative coding; Prompt bias: Iterative prompt refinement by researchers may introduce unintentional bias toward model outputs; Language bias: One report in French tests multilingual capabilities but limited non-English representation; Task complexity bias: Some coding tasks (evidence of impact) inherently harder; may reflect task difficulty rather than model capability; Single case study organization (Tearfund) limits generalizability; Single expert coder as gold standard for annotations (potential expert bias); Model selection bias—only GPT-4o and GPT-4o-mini tested; no comparison with other state-of-the-art LLMs; Potential overfitting during prompt iterative refinement; Small dataset size (Dataset 1: 5 reports, 113 excerpts; Dataset 2: 5 reports, 49 excerpts; Dataset 3: 5 reports, no coded data); Unrepresentative sampling—datasets from charity sector only
Limitations
- The authors state: "First, few-shot learning results in negligible improvement or worse performance in theme classification compared to zero-shot learning
- Despite selecting diverse examples, factors, such as overfitting, task complexity, unintentional bias, and differences in how LLMs process different codebooks, may affect the learning process
- Further experimentation is needed to explore this issue in more depth
- Second, prompt design remains a key challenge
- While iterative improvements were made, further refinement is essential, particularly by incorporating domain expertise to allow real-time adjustments
- Lastly, the reliance on a single case study, Tearfund's evaluation meta-synthesis, limits the generalizability of the findings."
Open questions raised
- Limited research combining both deductive and inductive coding approaches within a single comprehensive framework
- Few studies exploring dual roles of LLMs (initial coder and validator) in real-life workflows
- Need for testing framework across diverse sectors and contexts beyond single case study
- Further investigation of few-shot learning optimization and why it underperforms in theme classification
- Potential for using GPT models as evaluators (as suggested by Chiang & Lee 2023)
- Multi-agent AI and multimodal models for thematic analysis
Explore related topics
Related papers
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- Exploring Students’ Perceptions of ChatGPT: Thematic Analysis and Follow-Up SurveyAbdulhadi Shoufan · 2023 · 464 citations
- AI-generated feedback on writing: insights into efficacy and ENL student preferenceJuan Escalante · 2023 · 461 citations
- Is it harmful or helpful? Examining the causes and consequences of generative AI usage among university studentsMuhammad Abbas · 2024 · 372 citations
- The Perception by University Students of the Use of ChatGPT in EducationThi Thuy An Ngo · 2023 · 257 citations