12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Leveraging large language models for thematic analysis: a case study in the charity sector

Paul Clough · AI & Society · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
4/4
Quality (LMQS)
E
Evidence
6
Citations
23.75
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s00146-025-02487-4

Methodology & findings

Study design

Case study with empirical experiments.

Sample

N = 14.5, 9 groups

Primary method

Cohen's kappa (κ) for inter-rater reliability on multi-label classification; instance-based kappa pooling decisions across instances and codes; precision and recall calculations for excerpt extraction; cosine similarity using embedding vectors (text-embedding-ada-002); weighted scoring system for Agreement Rate on corporate priority categorization; descriptive statistics including mean, standard deviation (SD), and ranges; hallucination rate calculated as proportion of runs not following instructions; wall clock time (execution time in minutes); 5-point Likert scale analysis for qualitative assessment. Python code used for metric computation.

Main result

In deductive coding, "GPT-4o achieved a substantial agreement in multi-label thematic categorization (κ = 0.61–0.65), while GPT-4o-mini showed a moderate agreement (κ = 0.41–0.58). Both models excelled in sentiment analysis (κ = 0.91–0.95), but struggled with evaluating evidence of impact due to contextual complexity (κ ≤ 0.01)." For inductive coding, "GPT-4o demonstrated a strong semantic alignment with human-generated themes (cosine similarity = 0.76–0.79) though its tendency toward broad themes required human refinement."

Reports effect sizes and confidence intervals.

Research paradigm

Pragmatist/Mixed-methods (combining positivist measurement with interpretivist qualitative assessment)

Author conclusions

"Results show that GPT-4o models can effectively serve as both an initial coding tool and a validation mechanism for inductive and deductive coding processes within an LLM-human collaborative framework." Furthermore, "our findings underline the indispensable role of human expertise—from prompt engineering and managing hallucinations to final verification—to ensure accurate and trustworthy AI-assisted analyses. While LLMs can enhance qualitative analysis, their full potential is only realized under skilled human guidance."

Risk of bias

Single expert coder (Tearfund staff) as gold standard—no inter-rater reliability check against multiple human coders; Selection bias: datasets from single organization (Tearfund) with specific evaluation workflows; Task complexity varies across datasets; not all datasets fully coded for all coding types; Prompt engineering bias: prompts iteratively refined based on outputs, potential circularity; Limited diversity in codebooks tested (Light Wheel and SDGs only); Temperature settings fixed at 0 and 0.5—no systematic exploration of parameter sensitivity; Selection bias: Single case study organization (Tearfund) may not represent broader charity sector contexts; Evaluator bias: Expert coder annotations serve as gold standard but generated by single individual at Tearfund; Model bias: GPT-4o trained on data with potential societal biases that may manifest in qualitative coding; Prompt bias: Iterative prompt refinement by researchers may introduce unintentional bias toward model outputs; Language bias: One report in French tests multilingual capabilities but limited non-English representation; Task complexity bias: Some coding tasks (evidence of impact) inherently harder; may reflect task difficulty rather than model capability; Single case study organization (Tearfund) limits generalizability; Single expert coder as gold standard for annotations (potential expert bias); Model selection bias—only GPT-4o and GPT-4o-mini tested; no comparison with other state-of-the-art LLMs; Potential overfitting during prompt iterative refinement; Small dataset size (Dataset 1: 5 reports, 113 excerpts; Dataset 2: 5 reports, 49 excerpts; Dataset 3: 5 reports, no coded data); Unrepresentative sampling—datasets from charity sector only

Limitations

  • The authors state: "First, few-shot learning results in negligible improvement or worse performance in theme classification compared to zero-shot learning
  • Despite selecting diverse examples, factors, such as overfitting, task complexity, unintentional bias, and differences in how LLMs process different codebooks, may affect the learning process
  • Further experimentation is needed to explore this issue in more depth
  • Second, prompt design remains a key challenge
  • While iterative improvements were made, further refinement is essential, particularly by incorporating domain expertise to allow real-time adjustments
  • Lastly, the reliance on a single case study, Tearfund's evaluation meta-synthesis, limits the generalizability of the findings."

Open questions raised

  • Limited research combining both deductive and inductive coding approaches within a single comprehensive framework
  • Few studies exploring dual roles of LLMs (initial coder and validator) in real-life workflows
  • Need for testing framework across diverse sectors and contexts beyond single case study
  • Further investigation of few-shot learning optimization and why it underperforms in theme classification
  • Potential for using GPT models as evaluators (as suggested by Chiang & Lee 2023)
  • Multi-agent AI and multimodal models for thematic analysis
Data: Datasets derived from Tearfund evaluation reports (three datasets with varying degrees of coding completion). Data availability statement: 'No datasets were generated or analyzed during the current study.' This indicates datasets are not publicly available—they are proprietary Tearfund materials.; No datasets were generated or analyzed during the current study. The authors state: "No datasets were generated or analyzed during the current study." Data from Tearfund evaluation reports were used but not made publicly available due to organizational sensitivity.; No new datasets generated. Authors state "No datasets were generated or analyzed during the current study." The study used proprietary Tearfund evaluation reports; these are not publicly available.Code: Reusable prompt templates documented at: https://github.com/pauldclough/tearfund-meta-synthesis (mentioned footnote 8: 'For detailed prompt designs used in the experiments, see (site visited: 21/11/2024): https://github.com/pauldclough/tearfund-metasynthesis'); Detailed prompt designs used in experiments are available at: https://github.com/pauldclough/tearfund-meta-synthesis (site visited 21/11/2024); GitHub repository mentioned for detailed prompt designs: https://github.com/pauldclough/tearfund-meta-synthesis (site visited 21/11/2024)Extracted from: pdfAgreement 49%

Explore related topics

Related papers