Estimating the reproducibility of psychological science
Alexander A. Aarts · Science · 2015
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1126/science.aac4716
Methodology & findings
Study design
Large-scale collaborative replication project involving 100 replications of experimental and correlational studies published in 2008 issues of three major psychology journals (Psychological Science, Journal of Personality and Social Psychology, Journal of Experimental Psychology: Learning, Memory, and Cognition).
Sample
N = 185, 8 groups
Primary method
Conversion of effect sizes to correlation coefficients (r) with confidence intervals; Fisher transformation of correlation coefficients for analysis; Two-tailed significance tests with alpha = 0.05; McNemar test for comparing proportions of statistically significant results in original vs. replication studies; Wilcoxon signed-rank test for comparing central tendency of P value and effect size distributions; Paired two-sample t-test for effect size distributions; Goodness-of-fit chi-square test for coverage analysis; Binomial test for proportion of studies with stronger original effect sizes; Spearman's rank-order correlations of reproducibility indicators with study characteristics; Fixed-effect meta-analyses using R package metafor on Fisher-transformed correlations; Odds ratio meta-analyses; R statistical programming language for all analyses and reproduction; Standardization of moderator variables (M = 0, SD = 1) and aggregation into summary indices
Main result
The study found that "Replication effects were half the magnitude of original effects, representing a substantial decline. Ninety-seven percent of original studies had significant results. Thirty-six percent of replications had significant results; 47% of original effect sizes were in the 95% confidence interval of the replication effect size; 39% of effects were subjectively rated to have replicated the original result; and, if no bias in original results is assumed, combining original and replication results left 68% with significant effects."
Reports effect sizes and confidence intervals.
Research paradigm
Empiricist/Positivist
Author conclusions
The authors conclude that "Correlational tests suggest that replication success was better predicted by the strength of original evidence than by characteristics of the original and replication teams." More broadly, they note that replication success depends primarily on the original effect size and P-value: "A negative correlation of replication success with the original study P value indicates that the initial strength of evidence is predictive of reproducibility. For example, 26 of 63 (41%) original studies with P < 0.02 achieved P < 0.05 in the replication, whereas 6 of 23 (26%) that had a P value between 0.02 < P < 0.04 and 2 of 11 (18%) that had a P value > 0.04 did so."
Risk of bias
Selection bias in choice of studies to replicate (limited to three psychology journals); Publication bias in original studies (97% had statistically significant results); Potential differences between replication teams and original teams; Availability of original materials affecting replication feasibility; Selection bias in study eligibility: only 32% of 488 articles became eligible for selection during project period; Selection bias in article matching: feasibility constraints meant specialized samples and resources were underrepresented; Publication bias in original studies: original effect sizes may be inflated due to publication, selection, and reporting biases; Incomplete replication of all studies: 88% completion rate (100 of 113 replications completed); Single effect selection per article: identified key result may not be central to article aims; Potential moderating influences of sample, setting, or replication quality; Selection bias in study sampling: Only 158 of 488 articles (32%) became eligible during project period; 111 articles (70%) were selected by teams, but only 100 (88%) completed replications; Non-random selection: Articles were matched to teams by interests and expertise, potentially selecting for studies more feasible to replicate; Attrition: 13 of 113 attempted replications (12%) were not completed by project deadline; Publication bias in original studies: Authors acknowledge original studies likely have inflated effect sizes due to publication and selection bias; Selection bias in excluded studies: 47 eligible articles not claimed; 41 (87%) were eligible but not claimed, often requiring specialized samples or resources; Measurement bias: Single key effect selected from multi-study articles may not be central to overall aims; Potential bias toward replicable studies: Feasibility constraints may have excluded studies that are harder to replicate; Multiple comparisons: 5 different reproducibility indicators evaluated; aggregation of variables into indices to reduce false positives; Selection bias: Only 158 of 488 articles (32%) became eligible for selection during project period; only 111 (70%) were selected by teams; Attrition: 113 replications were initiated but only 100 (88%) completed by deadline; Non-random sampling: Matching of articles with replication teams by interests and expertise could introduce selection bias; Publication bias in original studies: Original studies may have inflated effect sizes due to publication, selection, and reporting biases; Feasibility constraints: 47 articles from eligible pool not claimed, mostly due to requiring specialized samples or resources; Interdisciplinary variation: Replications completed for only 56-72% of eligible articles depending on journal; Selection bias in study selection (though minimized through quasi-random sampling from 2008 journal issues); Selection bias from studies not replicated due to feasibility constraints; Potential publication bias in original studies (acknowledged by authors); Reporting bias in original studies (acknowledged as likely source of inflated effect sizes); Selective analysis in original studies; Attrition: 13 of 113 replications not completed by deadline; Publication bias in original studies; Selective reporting in original research; Selection bias in article sampling (though quasi-random selection attempted); Potential moderation by sample, setting, or replication quality; Feasibility constraints leading to incomplete coverage (32% of 488 articles became eligible; 70% of eligible articles selected); Team matching based on interest and expertise may introduce selection bias
Open questions raised
- The study identifies the need for better understanding of factors that affect replication success and the reproducibility of psychological science more broadly.
- The authors identify the need for further investigation of causes of reproducibility differences, noting "The resulting open dataset provides an initial estimate of the reproducibility of psychology and correlational data to support development of hypotheses about the causes of reproducibility." They also note the need for future research on why replication success differs between social and cognitive psychology, and between main effects and interaction effects.
- The paper addresses a major gap: 'There is plenty of concern about the rate and predictors of reproducibility, but limited evidence.' The authors identify need for greater transparency and specification of conditions sufficient to obtain results. They note that with only single effect selected per study, broader article aims may not be fully captured. The work is positioned as 'an initial estimate of the reproducibility of psychology' suggesting need for broader and ongoing reproducibility assessment across disciplines.
- The authors identify the gap that "There is plenty of concern about the rate and predictors of reproducibility, but limited evidence." They aimed to address this gap by providing "a large-scale, collaborative effort to obtain an initial estimate of the reproducibility of psychological science." The open dataset provides "correlational data to support development of hypotheses about the causes of reproducibility."
- The authors note they were "inspired to address the gap in direct empirical evidence about reproducibility" and state this research provides "a largescale, collaborative effort to obtain an initial estimate of the reproducibility of psychological science." The paper indicates that reproducibility varies by subdiscipline and research design, suggesting future research should investigate these differences.
- The authors identify that "There is plenty of concern about the rate and predictors of reproducibility, but limited evidence." They note the need for greater transparency about methodology and results: "With no transparency, the reasons for low reproducibility cannot be evaluated." The study was designed to provide "an initial estimate of the reproducibility of psychological science" and to provide "correlational data to support development of hypotheses about the causes of reproducibility."
Explore related topics
Related papers
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations