Using Artificial Intelligence to Code Unstructured Research Data
Alan Dennis, Warren Rosengren, Joseph Steed, Tucker Todd · Journal of the Association for Information Systems · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Two illustrative case studies demonstrating a five-step framework integrating generative AI with human coders and inter-rater reliability assessments on unstructured research data
Sample
N = 2500, 3 groups
Primary method
Inter-rater reliability assessments; ensemble methods with multiple LLMs; Delphi-style iterative revision process. Specific statistical tests or software not mentioned in abstract.
Main result
The study demonstrates that "the first used five LLMs to score more than 2,500 participant-generated ideas on novelty, workability, and relevance, achieving sufficient reliability levels for analysis, comparable to human coding" and "the second applied the method to a different dataset using a different set of three LLMs and again achieved acceptable reliability for analysis."
Reports effect sizes.
Research paradigm
Pragmatist/Mixed-methods (combining computational and human-centered approaches)
Author conclusions
The authors conclude that the proposed framework "integrates generative artificial intelligence (AI) with human coders and inter-rater reliability assessments to deliver faster, transparent, and replicable coding" and offer "guidelines and informed suggestions for prompt design, tool selection, bias checks, and opportunities for large-scale qualitative research."
Risk of bias
Selection bias in choice of LLMs and datasets; Potential algorithmic bias in LLM outputs not explicitly addressed in abstract; Lack of explicit bias checks mentioned beyond the framework's proposal for bias checks; Limited information on inter-rater reliability methodology between human coders; Potential bias in rubric development (human-designed); Selection bias in dataset choice for demonstration; LLM model-specific biases not fully characterized; Limited transparency on specific prompt design choices; No comparison against gold standard or external validation; Potential AI model bias in coding decisions; Limited diversity in datasets used for illustration; Dependence on rubric quality for reproducibility
Open questions raised
- The authors identify opportunities for applying this framework to large-scale qualitative research, suggesting a gap in methodologies for efficiently coding unstructured research data at scale.
- Opportunities for large-scale qualitative research using AI-assisted coding methods; need for further development of guidelines for prompt design, tool selection, and bias checks
Explore related topics
Related papers
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Role of AI chatbots in education: systematic literature reviewLasha Labadze · 2023 · 791 citations