A Human-Centered Workflow for Using Large Language Models in Content Analysis
ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Narrative review synthesizing methodological literature across multiple disciplines (political science, sociology, computer science, psychology, management) with supplementary Python code and prompt library resources for practical implementation..
Primary method
Design science methodology; human-centered workflow design synthesizing methodological literature across multiple disciplines
Main result
The paper establishes that "LLMs demonstrate remarkable versatility through their emergent abilities: capabilities that appear when models are scaled up in size and training data" and that "because LLMs can perform content analysis at scale with measurable accuracy, they can significantly expand the empirical scope of content analysis research." The authors synthesize guidance into a comprehensive workflow where "researchers design, supervise, and validate each stage of the LLM process to ensure rigor and transparency."
Research paradigm
Pragmatist/Design-oriented
Author conclusions
The authors conclude: "The division of labor reflects a pragmatic position on the capabilities and limitations of current LLMs. Models excel at applying well-specified categorical distinctions" and that "the promptbook is the functional analogue of a codebook: it specifies the operationalization of constructs, provides inclusion and exclusion criteria, and establishes a transparent audit trail of decision rules." They state their position that "LLMs need not interpret meaning autonomously. They perform a bounded, well-specified text transformation designed, supervised, and validated by human researchers who remain responsible for construct definition, theoretical interpretation, and the ultimate warrant attached to research findings."
Risk of bias
Data contamination: LLMs may have seen validation data in training, inflating performance metrics; Representational bias: LLMs tend to misinterpret texts from underrepresented demographic groups and reflect majority, Western perspectives; Training data bias: LLMs trained on scraped internet data reflect systemic biases and may misrepresent marginalized groups; Model memorization vs. reasoning: Unclear whether high validity scores result from actual reasoning or data memorization; Data contamination risk: proprietary LLMs may have seen validation data during pre-training, artificially inflating performance metrics; Representational bias: LLMs tend to misportray marginalized groups and may systematically misinterpret texts from underrepresented demographic groups; Training data bias: models trained on large-scale scraped text data reflect majority, often Western perspectives; Model dependency: reliance on external proprietary systems that may change or be discontinued
Limitations
- The authors explicitly state: "We do not empirically compare models or validation techniques
- We do not cover fine-tuning, retrieval-augmented generation (RAG) models and other advanced techniques
- We are focused on providing guidelines and resources for researchers that do not require significant up-front investment or exquisite technical expertise that are beyond typical management researcher's capabilities." Additionally, "LLMs are sensitive to how instructions are phrased, prone to generating plausible but inaccurate outputs, and their internal workings remain largely opaque."
Open questions raised
- The authors identify a methodological gap: "while the pace of LLM adoption in management and organizational research has accelerated dramatically, published guidance has tended to address narrow aspects of the process (e.g., prompt engineering, model selection, or validity assessment) rather than providing an integrated end-to-end account." They aim to provide integrated end-to-end guidance synthesized into "a single, practically grounded workflow that researchers can follow from research design through to interpretation."
- The paper identifies that 'published guidance has tended to address narrow aspects of the process (e.g., prompt engineering, model selection, or validity assessment) rather than providing an integrated end-to-end account' and emphasizes the need for 'a human-centered workflow for using LLMs for content analysis' as a gap in the literature.
- Need for empirical comparisons of different LLM models and validation techniques
- Limited guidance on fine-tuning and retrieval-augmented generation (RAG) models
- Lack of standardized validation procedures for tasks without objective ground truth (e.g., summarization, inductive coding)
- Need for research on LLM performance with non-Western and underrepresented organizational contexts
Explore related topics
Related papers
- Co-designing AI Education Curriculum with Cross-Disciplinary High School TeachersBenjamin Xie · 2024 · 28 citations
- GAIDeT (Generative AI Delegation Taxonomy): A taxonomy for humans to delegate tasks to generative artificial intelligence in scientific research and publishingYana Suchikova · 2025 · 24 citations
- Towards an Integrative Approach for Automated Literature Reviews Using Machine LearningChristoph Tauchert · 2020 · 19 citations
- PEER: Empowering Writing with Large Language ModelsKathrin Seßler · 2023 · 19 citations
- Using Large Language Models to Support Thematic Analysis in Empirical Legal StudiesJakub Drápal · 2023 · 16 citations
- BioRAG: A RAG-LLM Framework for Biological Question ReasoningChengrui Wang · 2024 · 12 citations