12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

State-of-the-Art Claims Require State-of-the-Art Evidence

YongKyung Oh · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Cross-domain diagnostic analysis applied to ten public leaderboards (HELM MMLU, LiveBench, Open ASR, Open VLM, VBench, TabArena Binary/Multiclass/Regression, TSFM-Bench MAE/MSE).

Sample

N = 1000, 4 groups

Primary method

Wilcoxon signed-rank test (referenced as standard recommendation but not primary analysis method); Cohen's d effect size calculation; Win rate (proportion) analysis; Breakdown point analysis (leave-one-out influence diagnostics); Sensitivity analysis varying individual thresholds; Fragility rate calculation (proportion of pairs failing at least one diagnostic test)

Main result

The study found that "more than half of top-model comparisons lack support for at least one commonly assumed property of superiority" across ten public leaderboards. The authors demonstrate that "a mean score across k tasks establishes only that a model ranks first on average. It does not show that the model wins on most tasks or would retain its rank under different dataset compositions." Across more than 1,000 pairwise comparisons spanning six domains, the fragility rate—the proportion of model pairs where the mean-score winner fails at least one elementary test—ranges from 39.47% to 74.73% depending on the benchmark.

Reports effect sizes.

Research paradigm

Empirical-analytical; quantitative measurement of benchmark fragility across multiple leaderboards

Author conclusions

The authors conclude: "The majority of top-model comparisons lack the statistical support that 'SOTA' implicitly promises. The tools to close this gap already exist. Therefore, the barrier is institutional rather than methodological." They further state: "Matching Claims to Evidence. Our goal is to align claim language with the strength of the supporting evidence. In current practice, aggregate rankings are often interpreted as broad superiority, even when the underlying evidence may be narrow or heterogeneous. Reporting language should reflect these limitations, especially when performance varies across tasks or depends on specific evaluation choices." Additionally: "We do not propose that authors must pass specific statistical tests. Instead, we propose that claim language should reflect what the evidence actually supports. When aggregate gains are narrow or inconsistent, phrases like 'achieves lowest average error' or 'ranks first on this benchmark' are more accurate than 'state-of-the-art,' which serves as a marketing term rather than a precise scientific claim."

Risk of bias

Selection bias: Analysis limited to top 20 models; patterns in lower-ranked models unknown; Snapshot bias: Single-time analysis does not capture temporal evolution of leaderboard fragility; Benchmark design bias: Analysis does not control for inherent benchmark heterogeneity or design choices; Publication bias: Analysis does not distinguish between peer-reviewed and non-reviewed models on leaderboards; Interpretation bias: Threshold choices (τ_w = 0.6, τ_d = 0.2, τ_b = 0.2) may reflect researcher assumptions despite sensitivity analysis; Selection bias: Only top 20 models analyzed, not full leaderboard rankings; Temporal bias: Single snapshot analysis (December 31, 2025) does not capture leaderboard evolution; Metric bias: Analysis focuses on single-metric benchmarks; multi-metric evaluation frameworks not addressed; Task composition bias: Fragility patterns may reflect task heterogeneity rather than model inadequacy; Publication bias: Only peer-reviewed benchmarks from major venues analyzed; Selection bias: Analysis limited to top 20 models per benchmark; patterns in lower-ranked models unknown; Snapshot bias: Temporal analysis limited to single snapshot per leaderboard (as of December 31, 2025 or benchmark-specific dates); fragility patterns over time not tracked; Model representation bias: Top-ranked positions dominated by frontier models from well-resourced organizations, may not reflect diversity of approaches; Threshold arbitrariness: Despite sensitivity analysis, choice of diagnostic thresholds (τ_d=0.2, τ_w=0.6, τ_b=0.2) may influence conclusions; Publication bias in benchmark design: Benchmarks selected were published at top venues; patterns in obscure or archived benchmarks unknown

Limitations

  • The authors state: "Model Selection: We analyze only top-ranked models
  • SOTA claims concentrate among top models, making this the relevant population
  • However, rankings further down may exhibit different fragility patterns." Additionally: "Threshold Choice: We set thresholds for three tests at lenient levels grounded in prior literature
  • Our sensitivity analysis shows that conclusions hold across reasonable variations
  • However, we do not claim these thresholds are uniquely correct." Furthermore: "Causal Explanation: We document fragility but do not explain its sources
  • Why some benchmarks exhibit higher fragility likely depends on task diversity, metric choice, and model similarity." The analysis also addresses "Single-Metric Focus: Our analysis examines benchmarks where a single aggregate metric determines rankings

Open questions raised

  • Extension to multi-metric evaluation frameworks rather than single-metric rankings
  • Longitudinal tracking of fragility patterns over time as leaderboards evolve
  • Analysis of domain-specific patterns and structural properties predicting dominant failure modes
  • Investigation of causal sources explaining why certain benchmarks exhibit higher fragility
  • Study of how reviewer-author interactions shape SOTA claims in practice
  • Quantification of task heterogeneity and its relationship to ranking stability
Data: HELM MMLU (https://crfm.stanford.edu/helm/mmlu/); LiveBench; Open ASR; Open VLM; VBench; TabArena (https://tabarena.ai/); TSFM-Bench (https://github.com/decisionintelligence/TSFM-Bench); HELM MMLU: https://crfm.stanford.edu/helm/mmlu/; TabArena: https://tabarena.ai/; TSFM-Bench: https://github.com/decisionintelligence/TSFM-Bench; Code and analysis results: https://github.com/yongkyung-oh/SOTACode: https://github.com/yongkyung-oh/SOTA; Analysis code: https://github.com/yongkyung-oh/SOTA; TSFM-Bench: https://github.com/decisionintelligence/TSFM-BenchExtracted from: pdfAgreement 56%

Explore related topics

Related papers