Is Llm-Based Synthetic Data Research In Information Systems Replicable?
Katarina Stanoevska‐Slabeva, Roman Mattoscio · Journal of the Association for Information Systems · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Concept-centric systematic literature review (SLR) with inductive coding analysis.
Sample
N = 56, 4 groups
Primary method
Inductive analysis approach with systematic coding of research pipeline elements; descriptive frequency counts of reporting gaps across coded studies
Main result
The analysis reveals "substantial heterogeneity in both methodological design and reporting practices, alongside notable reporting gaps that may hinder replicability." Among the 56 coded papers, critical omissions were identified: "6 do not specify the exact LLM model and version used, 15 do not clearly report the LLM access mode, 34 omit key configuration settings and controls (e.g., temperature, top-p) and 37 do not explicitly describe the output format of the generated data."
Reports effect sizes.
Research paradigm
Critical realism / Methodological rigor
Author conclusions
The authors conclude: "LLM-based SR show high potential to reduce time, costs, and recruitment bottlenecks for human respondents in empirical IS research. However, synthetic data generation is probabilistic, prompt-sensitive, and subject to drift across model updates and persona modelling approaches." They propose that "the IS community needs to build on SR findings if the conditions under which they were produced remain transparent" and present a preliminary reporting framework to stimulate scholarly discourse on methodological rigor.
Risk of bias
Selection bias in paper inclusion criteria for the SLR not fully detailed; Incomplete coding of the full corpus (only 56 of 111 papers coded in depth); Potential publication bias favoring certain reporting practices; Reviewer bias in assessment of reporting quality across heterogeneous studies; Selective coding of papers (only 56 of 111 papers coded in depth); Potential publication bias in literature review (only published papers included); Subjective assessment of reporting completeness by reviewers
Limitations
- The paper acknowledges that "open questions remain: Which additional items are needed to ensure that SR studies can be meaningfully replicated? To what extent should researchers be required to share exact artefacts such as prompts, code, persona-specifying data and evaluation/validation methods to enable research replicability?" The analysis is limited to identifying current reporting practices and proposing a preliminary framework rather than establishing definitive standards.
Open questions raised
- Need for LLM-specific research reporting standards
- Clarification of which additional items are needed to ensure SR studies can be meaningfully replicated
- Determination of the extent to which researchers should share exact artifacts such as prompts, code, and persona-specifying data
- Development of guidelines for evaluation and validation methods in SR research
- Which additional items are needed to ensure that SR studies can be meaningfully replicated? To what extent should researchers be required to share exact artefacts such as prompts, code, persona-specifying data and evaluation/validation methods to enable research replicability? The authors note that LLM-specific research reporting standards may be required to ensure replicability and that the IS community is invited to critically examine and extend the proposed framework toward reliable and transparent LLM-based SR research.
- Lack of LLM-specific research reporting standards
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations