12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

A Human-Centered Workflow for Using Large Language Models in Content Analysis

ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Narrative review synthesizing methodological literature across multiple disciplines (political science, sociology, computer science, psychology, management) with supplementary Python code and prompt library resources for practical implementation..

Primary method

Design science methodology; human-centered workflow design synthesizing methodological literature across multiple disciplines

Main result

The paper establishes that "LLMs demonstrate remarkable versatility through their emergent abilities: capabilities that appear when models are scaled up in size and training data" and that "because LLMs can perform content analysis at scale with measurable accuracy, they can significantly expand the empirical scope of content analysis research." The authors synthesize guidance into a comprehensive workflow where "researchers design, supervise, and validate each stage of the LLM process to ensure rigor and transparency."

Research paradigm

Pragmatist/Design-oriented

Author conclusions

The authors conclude: "The division of labor reflects a pragmatic position on the capabilities and limitations of current LLMs. Models excel at applying well-specified categorical distinctions" and that "the promptbook is the functional analogue of a codebook: it specifies the operationalization of constructs, provides inclusion and exclusion criteria, and establishes a transparent audit trail of decision rules." They state their position that "LLMs need not interpret meaning autonomously. They perform a bounded, well-specified text transformation designed, supervised, and validated by human researchers who remain responsible for construct definition, theoretical interpretation, and the ultimate warrant attached to research findings."

Risk of bias

Data contamination: LLMs may have seen validation data in training, inflating performance metrics; Representational bias: LLMs tend to misinterpret texts from underrepresented demographic groups and reflect majority, Western perspectives; Training data bias: LLMs trained on scraped internet data reflect systemic biases and may misrepresent marginalized groups; Model memorization vs. reasoning: Unclear whether high validity scores result from actual reasoning or data memorization; Data contamination risk: proprietary LLMs may have seen validation data during pre-training, artificially inflating performance metrics; Representational bias: LLMs tend to misportray marginalized groups and may systematically misinterpret texts from underrepresented demographic groups; Training data bias: models trained on large-scale scraped text data reflect majority, often Western perspectives; Model dependency: reliance on external proprietary systems that may change or be discontinued

Limitations

  • The authors explicitly state: "We do not empirically compare models or validation techniques
  • We do not cover fine-tuning, retrieval-augmented generation (RAG) models and other advanced techniques
  • We are focused on providing guidelines and resources for researchers that do not require significant up-front investment or exquisite technical expertise that are beyond typical management researcher's capabilities." Additionally, "LLMs are sensitive to how instructions are phrased, prone to generating plausible but inaccurate outputs, and their internal workings remain largely opaque."

Open questions raised

  • The authors identify a methodological gap: "while the pace of LLM adoption in management and organizational research has accelerated dramatically, published guidance has tended to address narrow aspects of the process (e.g., prompt engineering, model selection, or validity assessment) rather than providing an integrated end-to-end account." They aim to provide integrated end-to-end guidance synthesized into "a single, practically grounded workflow that researchers can follow from research design through to interpretation."
  • The paper identifies that 'published guidance has tended to address narrow aspects of the process (e.g., prompt engineering, model selection, or validity assessment) rather than providing an integrated end-to-end account' and emphasizes the need for 'a human-centered workflow for using LLMs for content analysis' as a gap in the literature.
  • Need for empirical comparisons of different LLM models and validation techniques
  • Limited guidance on fine-tuning and retrieval-augmented generation (RAG) models
  • Lack of standardized validation procedures for tasks without objective ground truth (e.g., summarization, inductive coding)
  • Need for research on LLM performance with non-Western and underrepresented organizational contexts
Code: Python code in Jupyter Notebook format provided as supplementary materials (specific repository URL not mentioned in extracted text); The paper mentions "supplementary materials, including a prompt library and Python code in Jupyter Notebook format, accompanied by detailed usage instructions" available where the paper states "Click here for the latest version and supplementary materials (Python code and prompt library)."; Python code in Jupyter Notebook format and prompt library provided as supplementary materials (specific repository URL not stated in document)Extracted from: pdfAgreement 67%

Explore related topics

Related papers