Using Large Language Models to Support Thematic Analysis in Empirical Legal Studies
Jakub Drápal, Hannes Westermann, Jaromír Šavelka · Frontiers in artificial intelligence and applications · 2023
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3233/faia230965
Methodology & findings
Study design
Empirical case study using computational framework for thematic analysis.
Sample
N = 785, 4 groups
Primary method
Manual evaluation of code quality using predefined evaluation scheme. Recall@K (R@1, R@3) metrics for theme prediction evaluation. Comparative analysis of LLM-discovered themes (n=8) versus manually-discovered themes (n=14) using qualitative mapping. No inferential statistical tests were performed.
Main result
The study found that "72.6% of the 785 predicted codes were deemed reasonable" initially, improving to "88.8% of the codes were perceived as reasonable (+16.2% improvement)" after expert feedback. For theme prediction, "the overall R@1 of .66 and R@3 of .82 appear to suggest that the proposed approach is promising but clear limitations exist." Additionally, "the LLM-discovered theft in commercial settings maps to data points from manually discovered theft at work and breaking into another object," and "all of the manually identified themes could be mapped to one or more of the potential themes within a higher-level theme."
Reports effect sizes.
Research paradigm
Pragmatist/Mixed-methods (combining qualitative thematic analysis with computational NLP evaluation)
Author conclusions
"We proposed a novel LLM-powered framework supporting thematic analysis, and evaluated its performance on an analysis of criminal courts' opinions focused on the categories of thefts in Czechia. We found that the initial coding of data was performed with reasonable quality (RQ1), and further improved when expert feedback was provided (RQ2). The performance on zero-shot classification of the data (facts descriptions) in terms of themes (categories of theft) was promising (RQ3) but could likely benefit from expert feedback (future work). The evaluation of the end-to-end performance of the pipeline on discovering and predicting themes suggested viability of the proposed framework (RQ4) while highlighting the importance of subject matter expert supervision."
Risk of bias
Evaluator bias: single author evaluated autonomously generated codes (potential confirmation bias or memory effects); Selection bias: 49 cases removed from dataset due to pilot use or errors, with slight over-representation of serious offenses; Limited ground truth validation: manual themes developed by 3 law students (modest sample), with inter-rater reliability resolved through one student's judgment; Model-specific findings: evaluation limited to GPT-4; generalizability to other LLMs unclear; Language bias: analysis conducted on Czech legal texts; may not generalize across jurisdictions or languages; Evaluator bias: single evaluator (subject matter expert/author) assessed initial codes without blinding to feedback condition; Selection bias: slight over-representation of serious offenses in dataset construction; Confirmation bias: potential bias in manual expert assessment of theme mapping; Proprietary model opacity: GPT-4 is a black-box model limiting interpretability and reproducibility; Evaluator knowledge of experimental condition (rater bias); Single evaluator for initial code assessment (RQ1); Small expert panel for manual theme assignment (3 law students); Single case study domain (Czech criminal law theft cases); Proprietary black-box model limits reproducibility
Limitations
- The authors acknowledge "a possible limitation of this experiment in that the author knew in which round the initial code was produced." Additionally, they note that "the black-box nature of the proprietary LLMs is especially problematic" for maintaining researcher agency over the process
- The framework's applicability is limited to phases 2 and 3 of thematic analysis, and "subject matter expert interventions might be desirable at various stages of the processing to improve the quality of the resulting themes and their alignment with the research questions."
Open questions raised
- Extension beyond phases 2 and 3 of thematic analysis to cover full pipeline
- Validation of findings in domains beyond court opinions and criminal law
- Investigation of expert feedback effects at the theme prediction stage (RQ3)
- Exploration of fine-tuning approaches using limited expert-labeled data
- Comparison of performance across multiple LLM models (not just GPT-4)
- Methods to improve alignment of LLM-discovered themes with research questions
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- What if the devil is my guardian angel: ChatGPT as a case study of using chatbots in educationAhmed Tlili · 2023 · 1,587 citations
- Embracing the future of Artificial Intelligence in the classroom: the relevance of AI literacy, prompt engineering, and critical thinking in modern educationYoshija Walter · 2024 · 805 citations
- Leveraging ChatGPT for Enhancing Critical Thinking SkillsYing Guo · 2023 · 223 citations
- Students’ use of large language models in engineering education: A case study on technology acceptance, perceptions, efficacy, and detection chancesMargherita Bernabei · 2023 · 150 citations
- Human-in-the-Loop AI Reviewing: Feasibility, Opportunities, and RisksIddo Drori · 2024 · 38 citations