PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience
Xinyang Liao, Lingyu Li, Huacan Liu, Tianle Gu, Yang Yao, Tong Zhu et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Benchmarking study with adversarial evaluation.
Sample
N = 200, 8 groups
Primary method
Evaluation metrics computed as averages of Likert-scale (1-5) judge ratings. Dimension-level raw scores computed as mean of subcriteria (Equation 1). Item-level scores computed as mean of three dimensions (Equation 2). Raw scores converted to percentage capability scores via linear transformation: (S-1)/4 × 100 (Equation 3). System-level metrics computed as averages across the 200 items (Equation 4). Resistance defined as inverse of capability score (Equation 5). Refusal rate computed as proportion of items triggering refusal (Equation 7).
Main result
The study found that "all evaluated auto-research systems readily complete the full pseudoscientific projects with near-zero refusal rates in minutes" and "the best resistance score is only 27.4%". Testing seven state-of-the-art agents revealed that "current systems readily produce persuasive reports that align with pseudoscientific premises with near-zero refusal rates and the highest resistance of only 27.4%".
Reports effect sizes and confidence intervals.
Research paradigm
Empirical/Computational - Benchmarking study with mixed methods evaluation
Author conclusions
"In this work, we introduce PseudoBench, a benchmark for evaluating whether agentic auto-research systems can resist pseudoscientific research tasks. PseudoBench contains 200 curated pseudoscientific claim-evidence pairs across five categories and evaluates auto-research systems through an end-to-end pipeline spanning experimental design, analysis, and report writing. Across seven auto-research systems, we find that current systems readily transform pseudoscientific premises into structured, polished, and persuasive paper-style reports. These results highlight the urgent need for scientific alignment in auto-research systems."
Risk of bias
Judge model selection bias: Use of GPT-5.4 as primary judge model; ablation study shows different judges assign different scores (80.0%-89.5% range); Dataset construction bias: Filtering and rewriting procedures may introduce subjective judgments despite multi-stage pipeline; Coverage bias: Limited to 200 items from 1,271 candidates; representation of pseudoscientific domains may be unequal; Prompt bias: The report generation prompt explicitly instructs systems to support the pseudoscientific proposition, potentially amplifying sycophancy; Judge model bias: All evaluation uses LLM-as-judge which may have inherent biases in assessing pseudoscience; Dataset curation bias: Manual selection and rewriting of pseudoscientific claims may introduce author bias; Prompt design bias: The report generation prompt explicitly instructs agents to support the proposition, creating favorable conditions for pseudoscience generation; Judge model bias - evaluation depends on LLM-as-judge which may have its own biases toward certain types of content; Dataset curation bias - selection of 'not even wrong' claims may not represent full spectrum of pseudoscience; Model selection bias - tested models skew toward GPT-5.4 (6 out of 7 systems); Prompt engineering bias - system prompts explicitly instruct agents to 'actively choose material organization, analysis methods, and writing strategies that are more favorable to supporting the proposition'; Evaluation metric bias - persuasiveness dimension may penalize sophisticated legitimate science
Limitations
- The authors state that "PseudoBench deliberately focuses on curated pseudoscientific claim–evidence pairs
- This scope is designed to evaluate the cognitive bottom line of auto-research systems, that is, whether these systems can resist claims that are 'not even wrong'
- PseudoBench provides a foundation on which future work can extend toward a wider spectrum of scientific scenarios and more fine-grained epistemic risks such as borderline scientific controversies, low-quality studies, or domain-specific technical falsehoods." Additionally, "as with any publicly released benchmark, PseudoBench cannot fully prevent data contamination once the dataset is incorporated into future training corpora."
Open questions raised
- Need to extend toward a wider spectrum of scientific scenarios including borderline scientific controversies, low-quality studies, and domain-specific technical falsehoods
- Need for scientific alignment techniques specifically designed for epistemic rigor in research, beyond general harmlessness and task completion optimization
- Requirement for safeguards and refusal mechanisms specifically targeting pseudoscientific and unsupported claims in agentic auto-research systems
- The authors identify several research gaps and future directions: (1) extending evaluation beyond curated pseudoscience to borderline scientific controversies, low-quality studies, and domain-specific technical falsehoods; (2) developing scientific alignment techniques specifically for auto-research systems to identify unsupported or unreasonable claims; (3) preventing data contamination when benchmarks are incorporated into future training corpora; (4) addressing the broader epistemic and systemic risks of agentic auto-research in academic literature and public decision-making.
- Need to extend toward wider spectrum of scientific scenarios beyond 'not even wrong' claims to include borderline scientific controversies, low-quality studies, and domain-specific technical falsehoods
- Need for scientific alignment techniques specifically designed for research integrity and epistemic rigor in auto-research systems
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations