12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Using Large Language Models to Support Thematic Analysis in Empirical Legal Studies

Jakub Drápal, Hannes Westermann, Jaromír Šavelka · Frontiers in artificial intelligence and applications · 2023

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
16
Citations
118.30
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3233/faia230965

Methodology & findings

Study design

Empirical case study using computational framework for thematic analysis.

Sample

N = 785, 4 groups

Primary method

Manual evaluation of code quality using predefined evaluation scheme. Recall@K (R@1, R@3) metrics for theme prediction evaluation. Comparative analysis of LLM-discovered themes (n=8) versus manually-discovered themes (n=14) using qualitative mapping. No inferential statistical tests were performed.

Main result

The study found that "72.6% of the 785 predicted codes were deemed reasonable" initially, improving to "88.8% of the codes were perceived as reasonable (+16.2% improvement)" after expert feedback. For theme prediction, "the overall R@1 of .66 and R@3 of .82 appear to suggest that the proposed approach is promising but clear limitations exist." Additionally, "the LLM-discovered theft in commercial settings maps to data points from manually discovered theft at work and breaking into another object," and "all of the manually identified themes could be mapped to one or more of the potential themes within a higher-level theme."

Reports effect sizes.

Research paradigm

Pragmatist/Mixed-methods (combining qualitative thematic analysis with computational NLP evaluation)

Author conclusions

"We proposed a novel LLM-powered framework supporting thematic analysis, and evaluated its performance on an analysis of criminal courts' opinions focused on the categories of thefts in Czechia. We found that the initial coding of data was performed with reasonable quality (RQ1), and further improved when expert feedback was provided (RQ2). The performance on zero-shot classification of the data (facts descriptions) in terms of themes (categories of theft) was promising (RQ3) but could likely benefit from expert feedback (future work). The evaluation of the end-to-end performance of the pipeline on discovering and predicting themes suggested viability of the proposed framework (RQ4) while highlighting the importance of subject matter expert supervision."

Risk of bias

Evaluator bias: single author evaluated autonomously generated codes (potential confirmation bias or memory effects); Selection bias: 49 cases removed from dataset due to pilot use or errors, with slight over-representation of serious offenses; Limited ground truth validation: manual themes developed by 3 law students (modest sample), with inter-rater reliability resolved through one student's judgment; Model-specific findings: evaluation limited to GPT-4; generalizability to other LLMs unclear; Language bias: analysis conducted on Czech legal texts; may not generalize across jurisdictions or languages; Evaluator bias: single evaluator (subject matter expert/author) assessed initial codes without blinding to feedback condition; Selection bias: slight over-representation of serious offenses in dataset construction; Confirmation bias: potential bias in manual expert assessment of theme mapping; Proprietary model opacity: GPT-4 is a black-box model limiting interpretability and reproducibility; Evaluator knowledge of experimental condition (rater bias); Single evaluator for initial code assessment (RQ1); Small expert panel for manual theme assignment (3 law students); Single case study domain (Czech criminal law theft cases); Proprietary black-box model limits reproducibility

Limitations

  • The authors acknowledge "a possible limitation of this experiment in that the author knew in which round the initial code was produced." Additionally, they note that "the black-box nature of the proprietary LLMs is especially problematic" for maintaining researcher agency over the process
  • The framework's applicability is limited to phases 2 and 3 of thematic analysis, and "subject matter expert interventions might be desirable at various stages of the processing to improve the quality of the resulting themes and their alignment with the research questions."

Open questions raised

  • Extension beyond phases 2 and 3 of thematic analysis to cover full pipeline
  • Validation of findings in domains beyond court opinions and criminal law
  • Investigation of expert feedback effects at the theme prediction stage (RQ3)
  • Exploration of fine-tuning approaches using limited expert-labeled data
  • Comparison of performance across multiple LLM models (not just GPT-4)
  • Methods to improve alignment of LLM-discovered themes with research questions
Data: Dataset of 785 facts descriptions from Czech criminal court opinions (2017) regarding theft cases. Dataset is not explicitly stated as publicly available; authors received 834 cases from Prosecution Service and removed 49 cases.; Czech criminal court opinions dataset (theft cases)Code: Tiktoken; OpenAI Python LibraryExtracted from: pdfAgreement 57%

Explore related topics

Related papers