12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Leveraging Large Language Models for Automating Inductive Qualitative Coding: A Comparative Study of Prompt Engineering Techniques

Elias Frigui · Inquiry Queen s Undergraduate Research Conference Proceedings · 2024

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
E
Evidence
1
Citations
0.14
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.24908/iqurcp18054

Methodology & findings

Study design

Comparative empirical study testing multiple prompt engineering techniques (Zero-shot, Few-shot, and Chain-of-Thought learning) applied to inductive qualitative coding of interview transcripts

Main result

The study found that "Few-shot learning showed consistent performance with moderate amounts data, while CoT proved most effective in reducing partial hallucinations." Additionally, the authors determined that "LLMs cannot fully replace human coders" but "can aid the process with a human-in-the-loop approach."

Reports effect sizes.

Research paradigm

Pragmatist/Mixed-methods (technology evaluation with human-in-the-loop validation)

Author conclusions

The authors conclude that "tailored LLM with adequate prompting techniques can help assist researchers when performing qualitative analysis." They also determined through their research pivot that a human-in-the-loop approach offers better outcomes than full automation.

Risk of bias

Selection bias in interview transcripts chosen for coding; Evaluator bias in assessing code quality and hallucination reduction; Model-specific bias (GPT family models only); No mention of blinding or inter-rater reliability checks; Selection bias: No information provided about how interview transcripts were selected for coding; Evaluator bias: No mention of inter-rater reliability or blinding in evaluation of LLM coding performance; Model selection bias: Only GPT family models tested; generalizability to other LLM families unclear; Context/token limitation bias: Authors note LLM context and token limits affected performance; Limited sample description in abstract; Single domain focus (interview transcripts only); Potential selection bias in choice of LLM models tested; No mention of inter-rater reliability validation

Limitations

  • The authors note that "challenges of context and token limits in LLMs" constrain full automation
  • Additionally, the study "pivoted to testing prompt strategies after realizing that a human-in-the-loop process would offer better accuracy and flexibility."

Open questions raised

  • The authors identify the need for human-in-the-loop approaches and suggest future research should focus on optimizing prompt strategies for practical qualitative research applications rather than pursuing full automation.
  • Future research directions include: (1) testing different LLM families beyond GPT, (2) exploring optimal data amounts for Few-shot learning, (3) evaluating human-in-the-loop workflow efficiency, (4) addressing context and token limit challenges, and (5) validating findings across different types of qualitative data and research domains.
  • The study suggests future research should explore: (1) full automation feasibility with improved LLM architectures, (2) optimization of prompt engineering techniques for different coding contexts, (3) scaling to larger qualitative datasets, and (4) generalization across social science and software engineering domains
Data: not_statedCode: not_statedExtracted from: pdfAgreement 69%

Explore related topics

Related papers