State-of-the-Art Claims Require State-of-the-Art Evidence
YongKyung Oh · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Cross-domain diagnostic analysis applied to ten public leaderboards (HELM MMLU, LiveBench, Open ASR, Open VLM, VBench, TabArena Binary/Multiclass/Regression, TSFM-Bench MAE/MSE).
Sample
N = 1000, 4 groups
Primary method
Wilcoxon signed-rank test (referenced as standard recommendation but not primary analysis method); Cohen's d effect size calculation; Win rate (proportion) analysis; Breakdown point analysis (leave-one-out influence diagnostics); Sensitivity analysis varying individual thresholds; Fragility rate calculation (proportion of pairs failing at least one diagnostic test)
Main result
The study found that "more than half of top-model comparisons lack support for at least one commonly assumed property of superiority" across ten public leaderboards. The authors demonstrate that "a mean score across k tasks establishes only that a model ranks first on average. It does not show that the model wins on most tasks or would retain its rank under different dataset compositions." Across more than 1,000 pairwise comparisons spanning six domains, the fragility rate—the proportion of model pairs where the mean-score winner fails at least one elementary test—ranges from 39.47% to 74.73% depending on the benchmark.
Reports effect sizes.
Research paradigm
Empirical-analytical; quantitative measurement of benchmark fragility across multiple leaderboards
Author conclusions
The authors conclude: "The majority of top-model comparisons lack the statistical support that 'SOTA' implicitly promises. The tools to close this gap already exist. Therefore, the barrier is institutional rather than methodological." They further state: "Matching Claims to Evidence. Our goal is to align claim language with the strength of the supporting evidence. In current practice, aggregate rankings are often interpreted as broad superiority, even when the underlying evidence may be narrow or heterogeneous. Reporting language should reflect these limitations, especially when performance varies across tasks or depends on specific evaluation choices." Additionally: "We do not propose that authors must pass specific statistical tests. Instead, we propose that claim language should reflect what the evidence actually supports. When aggregate gains are narrow or inconsistent, phrases like 'achieves lowest average error' or 'ranks first on this benchmark' are more accurate than 'state-of-the-art,' which serves as a marketing term rather than a precise scientific claim."
Risk of bias
Selection bias: Analysis limited to top 20 models; patterns in lower-ranked models unknown; Snapshot bias: Single-time analysis does not capture temporal evolution of leaderboard fragility; Benchmark design bias: Analysis does not control for inherent benchmark heterogeneity or design choices; Publication bias: Analysis does not distinguish between peer-reviewed and non-reviewed models on leaderboards; Interpretation bias: Threshold choices (τ_w = 0.6, τ_d = 0.2, τ_b = 0.2) may reflect researcher assumptions despite sensitivity analysis; Selection bias: Only top 20 models analyzed, not full leaderboard rankings; Temporal bias: Single snapshot analysis (December 31, 2025) does not capture leaderboard evolution; Metric bias: Analysis focuses on single-metric benchmarks; multi-metric evaluation frameworks not addressed; Task composition bias: Fragility patterns may reflect task heterogeneity rather than model inadequacy; Publication bias: Only peer-reviewed benchmarks from major venues analyzed; Selection bias: Analysis limited to top 20 models per benchmark; patterns in lower-ranked models unknown; Snapshot bias: Temporal analysis limited to single snapshot per leaderboard (as of December 31, 2025 or benchmark-specific dates); fragility patterns over time not tracked; Model representation bias: Top-ranked positions dominated by frontier models from well-resourced organizations, may not reflect diversity of approaches; Threshold arbitrariness: Despite sensitivity analysis, choice of diagnostic thresholds (τ_d=0.2, τ_w=0.6, τ_b=0.2) may influence conclusions; Publication bias in benchmark design: Benchmarks selected were published at top venues; patterns in obscure or archived benchmarks unknown
Limitations
- The authors state: "Model Selection: We analyze only top-ranked models
- SOTA claims concentrate among top models, making this the relevant population
- However, rankings further down may exhibit different fragility patterns." Additionally: "Threshold Choice: We set thresholds for three tests at lenient levels grounded in prior literature
- Our sensitivity analysis shows that conclusions hold across reasonable variations
- However, we do not claim these thresholds are uniquely correct." Furthermore: "Causal Explanation: We document fragility but do not explain its sources
- Why some benchmarks exhibit higher fragility likely depends on task diversity, metric choice, and model similarity." The analysis also addresses "Single-Metric Focus: Our analysis examines benchmarks where a single aggregate metric determines rankings
Open questions raised
- Extension to multi-metric evaluation frameworks rather than single-metric rankings
- Longitudinal tracking of fragility patterns over time as leaderboards evolve
- Analysis of domain-specific patterns and structural properties predicting dominant failure modes
- Investigation of causal sources explaining why certain benchmarks exhibit higher fragility
- Study of how reviewer-author interactions shape SOTA claims in practice
- Quantification of task heterogeneity and its relationship to ranking stability
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- ChatGPT in higher education: Considerations for academic integrity and student learningMiriam Sullivan · 2023 · 740 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations