12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

The use and methodological reporting of large language models in qualitative research: a scoping review

Christian Kempny, Julian Frings, Paul Rust, Sven Meister, Leonard Fehring · BMC Medical Research Methodology · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1186/s12874-026-02913-1

Methodology & findings

Study design

Scoping review conducted following Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews (PRISMA-ScR) guidelines and the methodological framework by Arksey and O'Malley refined by the Joanna Briggs Institute.

Main result

The review found that "LLMs are being applied across the full spectrum of qualitative research stages, from developing interview materials and transcribing data to supporting coding and theme development, summarizing datasets, and drafting analytic texts." A critical finding is that "a substantial proportion of studies (75%) did not report specific parameter settings" for the LLMs used, and "while 46 studies provided complete or partial prompts, 10 reported no prompting details at all." Additionally, "the substantial variation in reported agreement between LLM and human coders across these domains (from 36% to 99% agreement with human coders) reflects several interconnected factors," with agreement rates depending on task complexity, prompt engineering quality, and data characteristics.

Research paradigm

Interpretive/constructivist (qualitative research methodology)

Author conclusions

"This scoping review demonstrates that LLMs are being explored and applied across qualitative research workflows. However, the aims of the included studies indicate that many applications remain exploratory or evaluative in nature, rather than reflecting established or routine adoption. Additional technical reporting remains highly inconsistent: critical parameters such as temperature, context length, and exact prompts are reported in a minority of studies, directly undermining reproducibility and comparability. Given these findings, the development and adoption of dedicated reporting guidelines, such as the COREQ + LLM extension, is urgently needed to ensure that LLM-assisted qualitative research meets the standards of transparency, rigor, and interpretive depth that the field demands."

Risk of bias

Language bias: English-language restriction excludes relevant non-English publications; Geographic bias: Concentration in US (35%), UK (12%), limiting global applicability; Publication bias: Peer-reviewed articles only; excludes grey literature and preprints; Selection bias: Restriction to LLM-only tools excludes BERT-based and traditional NLP approaches; No quality appraisal conducted: Studies treated equally regardless of methodological rigor; Technology obsolescence: Rapid LLM evolution means findings may quickly become dated; Language bias: Restriction to English-language publications excludes non-English research traditions and relevant work in other languages; Geographic bias: Studies concentrated in United States (35%), United Kingdom (12%), with limited representation from Africa, South America, and parts of Asia; Publication bias: Restriction to peer-reviewed journal articles excludes conference proceedings, preprints, and grey literature where methodological innovations may first appear; Selection bias: Studies selected based on explicit LLM use; studies mentioning LLMs without substantive integration may be missed or included inconsistently; Technology bias: Focus on LLMs only; exclusion of other AI-based tools (BERT-based classifiers, traditional NLP approaches) limits scope of technological assessment; Language bias: Restriction to English-language publications only; Geographic bias: 35% of studies from United States, limiting global representation; Publication bias: Inclusion limited to peer-reviewed journal articles, excluding grey literature, conference proceedings, and preprints; Selection bias: No quality assessment conducted, treating all studies equally; Technology obsolescence bias: Rapid evolution of LLM technology dates findings quickly

Open questions raised

  • Lack of standardized reporting guidelines tailored to LLM use in qualitative research (addressed by COREQ + LLM)
  • Absence of systematic comparative methodological studies examining different models, prompting strategies, and validation approaches across diverse qualitative tasks
  • Limited investigation of LLM applications in non-English and multilingual contexts
  • Lack of longitudinal research examining effects of sustained LLM use on researcher skill development and analytical practices
  • Critical gap in robust evidence regarding effectiveness of LLMs in qualitative research - few studies provide systematic evaluations comparing LLM-generated outputs with high-quality human qualitative analysis
  • Need for carefully designed empirical studies directly comparing LLM-assisted and human-led analyses with experienced qualitative researchers and blinded evaluation procedures
Extracted from: pdfAgreement 53%

Explore related topics

Related papers