12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Dead Science Walking: Publication Bias and the AI Scientist Pipeline

Kargi Chauhan · arXiv (Cornell University) · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
0/4
Quality (LMQS)
I
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Conceptual and theoretical analysis with formal mathematical definition of the null result gap (∆) and amplification index (A∆).

Main result

The paper formalises publication bias distortion as the null result gap ∆, finding that "a standard three-stage pipeline can amplify corpus distortion by a factor of 2.18×" across retrieval, generation, and evaluation stages. Specific domain estimates are provided: "drug discovery ∆≈0.60, psychology ∆≈0.56, cancer biology ∆≈0.35". The authors argue that "AI scientist systems can turn publication bias from a slow epistemic tax into a fast systems failure" by inheriting biased training corpora and amplifying them through automated processes.

Research paradigm

Critical theory / governance analysis

Author conclusions

"The null result gap is not a property of language models. It is a property of the publication system that language models learn from. AI scientist systems can inherit that gap, amplify it through retrieval and evaluation, and return old failures as new discoveries." The authors conclude: "The response should be infrastructural: index null results, evaluate retraction awareness, and disclose training corpora. None of these steps requires slowing scientific AI. They make acceleration more trustworthy." Most critically: "The AI-for-science community builds the systems, designs the benchmarks, and runs the venues. It can set these norms before biased automation becomes scientific common sense; after that point, the field will not be accelerating discovery so much as accelerating the recycling of its own uncorrected errors."

Risk of bias

Publication bias in training corpora (over-representation of positive results); Selection bias in retrieval systems (null results less likely to be indexed); Generation bias in language models (tendency to produce confident narratives from mixed evidence); Evaluation bias in automated review (LLM-as-judge preference for fluent, positive, novel claims); File-drawer effects (null results never published); Retracted paper persistence in literature (Retraction Watch database contains >50,000 entries); Citation manipulation in preprint networks; Hallucination bias in LLM systems generating scientific content; Publication bias in training corpora (documented across multiple domains); Selection bias in retrieval systems favoring positive abstracts; LLM evaluation bias favoring fluency, confidence, and novelty over accuracy; Retraction lag and continued citation of retracted work; Positive framing bias in literature summarization; Publication bias in scientific literature (positive results over-represented); File-drawer effects and selective reporting; Abstract-reporting bias making negative findings less retrievable; Hallucination in LLM-generated citations; LLM-as-judge biases favoring fluent, confident, novel claims; Position bias and verbosity bias in language model evaluations; Retracted papers remaining cited in scientific literature

Limitations

  • The authors acknowledge: "The exact value is less important than the order of magnitude" and note that "These values should be read as order-of-magnitude estimates, not precise field constants." They further state that "the multipliers are first-order estimates, not direct measurements of a deployed AI scientist" and that "The multiplicative model assumes independence across stages, but the biases are likely positively correlated: positive retrieval makes confident generation and positive evaluation more likely
  • Thus true amplification may exceed A∆." The authors also note they "do not claim that current systems have already produced every failure mode below, nor that all scientific domains have the same publication bias." Additionally: "Replication laundering remains supported structurally rather than directly."

Open questions raised

  • The authors identify several directions for future work: (1) Empirical calibration of amplification indices through "benchmark corpora with known null-result coverage, retrieval and generation tests over matched hypotheses, and LLM-judge evaluations of positive versus null or mixed narratives under controlled factual support"; (2) Direct measurement of deployed AI scientist systems rather than first-order estimates; (3) Empirical testing of the four proposed failure modes in real AI scientist deployments; (4) Investigation of how AI-generated outputs recursively enter the corpus as preprints and fine-tuning data; (5) Refinement of the retraction-aware evaluation metric (Ra) to distinguish between inappropriate positive citation, neutral historical citation, and explicit discussion of retraction.
  • Empirical calibration of multiplier values through benchmark corpora with known null-result coverage
  • Retrieval and generation tests over matched hypotheses with controlled factual support
  • LLM-judge evaluations of positive versus null or mixed narratives under controlled conditions
  • Direct measurement of deployed AI scientist systems rather than first-order estimates
  • Investigation of recursive amplification when AI outputs re-enter training corpora
Extracted from: pdfAgreement 75%

Explore related topics

Related papers