Dead Science Walking: Publication Bias and the AI Scientist Pipeline
Kargi Chauhan · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Conceptual and theoretical analysis with formal mathematical definition of the null result gap (∆) and amplification index (A∆).
Main result
The paper formalises publication bias distortion as the null result gap ∆, finding that "a standard three-stage pipeline can amplify corpus distortion by a factor of 2.18×" across retrieval, generation, and evaluation stages. Specific domain estimates are provided: "drug discovery ∆≈0.60, psychology ∆≈0.56, cancer biology ∆≈0.35". The authors argue that "AI scientist systems can turn publication bias from a slow epistemic tax into a fast systems failure" by inheriting biased training corpora and amplifying them through automated processes.
Research paradigm
Critical theory / governance analysis
Author conclusions
"The null result gap is not a property of language models. It is a property of the publication system that language models learn from. AI scientist systems can inherit that gap, amplify it through retrieval and evaluation, and return old failures as new discoveries." The authors conclude: "The response should be infrastructural: index null results, evaluate retraction awareness, and disclose training corpora. None of these steps requires slowing scientific AI. They make acceleration more trustworthy." Most critically: "The AI-for-science community builds the systems, designs the benchmarks, and runs the venues. It can set these norms before biased automation becomes scientific common sense; after that point, the field will not be accelerating discovery so much as accelerating the recycling of its own uncorrected errors."
Risk of bias
Publication bias in training corpora (over-representation of positive results); Selection bias in retrieval systems (null results less likely to be indexed); Generation bias in language models (tendency to produce confident narratives from mixed evidence); Evaluation bias in automated review (LLM-as-judge preference for fluent, positive, novel claims); File-drawer effects (null results never published); Retracted paper persistence in literature (Retraction Watch database contains >50,000 entries); Citation manipulation in preprint networks; Hallucination bias in LLM systems generating scientific content; Publication bias in training corpora (documented across multiple domains); Selection bias in retrieval systems favoring positive abstracts; LLM evaluation bias favoring fluency, confidence, and novelty over accuracy; Retraction lag and continued citation of retracted work; Positive framing bias in literature summarization; Publication bias in scientific literature (positive results over-represented); File-drawer effects and selective reporting; Abstract-reporting bias making negative findings less retrievable; Hallucination in LLM-generated citations; LLM-as-judge biases favoring fluent, confident, novel claims; Position bias and verbosity bias in language model evaluations; Retracted papers remaining cited in scientific literature
Limitations
- The authors acknowledge: "The exact value is less important than the order of magnitude" and note that "These values should be read as order-of-magnitude estimates, not precise field constants." They further state that "the multipliers are first-order estimates, not direct measurements of a deployed AI scientist" and that "The multiplicative model assumes independence across stages, but the biases are likely positively correlated: positive retrieval makes confident generation and positive evaluation more likely
- Thus true amplification may exceed A∆." The authors also note they "do not claim that current systems have already produced every failure mode below, nor that all scientific domains have the same publication bias." Additionally: "Replication laundering remains supported structurally rather than directly."
Open questions raised
- The authors identify several directions for future work: (1) Empirical calibration of amplification indices through "benchmark corpora with known null-result coverage, retrieval and generation tests over matched hypotheses, and LLM-judge evaluations of positive versus null or mixed narratives under controlled factual support"; (2) Direct measurement of deployed AI scientist systems rather than first-order estimates; (3) Empirical testing of the four proposed failure modes in real AI scientist deployments; (4) Investigation of how AI-generated outputs recursively enter the corpus as preprints and fine-tuning data; (5) Refinement of the retraction-aware evaluation metric (Ra) to distinguish between inappropriate positive citation, neutral historical citation, and explicit discussion of retraction.
- Empirical calibration of multiplier values through benchmark corpora with known null-result coverage
- Retrieval and generation tests over matched hypotheses with controlled factual support
- LLM-judge evaluations of positive versus null or mixed narratives under controlled conditions
- Direct measurement of deployed AI scientist systems rather than first-order estimates
- Investigation of recursive amplification when AI outputs re-enter training corpora
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- ChatGPT and a new academic reality: Artificial Intelligence‐written research papers and the ethics of the large language models in scholarly publishingBrady Lund · 2023 · 769 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations