Intelligence Is Not the Bottleneck: Validating an LLM First-Pass Manuscript Score Against Peer-Review Outcomes
Costa Georgantas · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Observational validation study with pre-registered hypotheses and outcome blinding.
Sample
N = 300, 14 groups
Primary method
Bootstrap confidence intervals (4,000-resample BCa/bias-corrected and accelerated): used for AUROC, Spearman correlations, and effect sizes (primary deviation from pre-registration: BCa substituted for registered percentile interval due to superior coverage at sample size; justified in Appendix F). Jonckheere-Terpstra trend test (10,000 Monte Carlo permutations) for monotonic gradient across ordered tiers. Spearman rank correlation for ordinal and continuous associations. Stratified bootstrap for class-conditional statistics. Mann-Whitney U tests for pairwise tier comparisons. Wilcoxon signed-rank test (exact) for paired within-paper SD comparison (p=0.014). Logistic regression with cross-validated AUROC (5-fold stratified) for covariate-control descriptive analysis. Benjamini-Hochberg multiple-comparison correction for per-dimension analysis. Wilson intervals (descriptive) for band reject-rate proportions. Label-shuffle null control at analysis time (permuted labels yield AUROC≈0.5). All analyses deterministic under fixed seed; analysis code released for reproducibility.
Main result
The study found that "the score separates rejected from accepted submissions (AUROC 0.82 (95% CI 0.78-0.87) on cohort M, confirmed at 0.87 (95% CI 0.79-0.93) on cohort H)" and that "the lowest score quintile (the bottom fifth of the 100-paper cohort H; integer-score ties shrink the strict quantile band to n = 15) is rejected at 100.0 % (95% CI [79.6, 100.0]), a 2.00 × lift over the 50.0 % base rate." The key insight is that "Direct (GPT-5.4), a one-paragraph prompt, already separates rejected from accepted submissions on its own (AUROC 0.80 (95% CI 0.70-0.88)), indistinguishably from AIPR (∆AUROC 0.07, p=0.09)," demonstrating that "Intelligence is not the bottleneck; reliability is."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical validation; positivist epistemology grounded in measurement against external ground truth
Author conclusions
"We validated a training-free, prompt-only LLM manuscript score against the public decisions and reviewer ratings of a major machine learning venue, under a pre-registered, reproducible protocol. The score separates rejected from accepted submissions (AUROC 0.82), rises across decision tiers, and tracks reviewer ratings (ρ = 0.52); it is strongest at the low end, where the lowest quintile is rejected at 100.0% on the production model, with strong work absent. Two findings locate the system's value. The validity comes mostly from the model: a one-paragraph prompt on the same model already discriminates, indistinguishably from the full pipeline (∆AUROC 0.07, p=0.09). What the engineering adds is reliability: AIPR's score barely moves across runs (0.7 vs. 2.8 points) where the bare prompt swings, delivered with a grounded, anchored review in one pass. Intelligence is not the bottleneck; reliability is."
Risk of bias
Author conflict of interest: author founded and operates AIPR, the system being evaluated; Single venue studied (ICLR 2026 only); cross-field generalization not claimed; Ground truth contamination risk: mean reviewer ratings may themselves be partly LLM-assisted; Proprietary scorer limits independent verification of manuscript-to-score function; Cheap-to-frontier model bridge is weakest at the low end where the triage claim resides; Paper-text leakage risk mitigated but not completely eliminated (arXiv pre-cutoff papers excluded); No measurement of actual reviewer decision-making behavior; Temporal leakage risk: model's knowledge cutoff (August 2025) precedes outcome release (January 2026), mitigating memorization but residual arXiv-text leakage possible; Sampling bias: balanced sampling across decision tiers in cohort M does not reflect venue's natural acceptance rate (27.4%); mitigated by reporting both sampled and re-weighted metrics; Selection bias: arXiv-twin submissions excluded; pre-registered exclusions (desk-rejects, parse failures, leakage flags) apply uniformly; Author/founder conflict of interest: the grading system (AIPR) is operated by the study's author; Weak bridge at low end: cohort H's cheap-model-to-frontier correlation (ρ=0.81 globally) weakens in bottom quintile (ρ=0.52); Noisy ground truth: reviewer disagreement and potential LLM-assisted review composition; Proprietary scorer: audit limited to outputs; full prompt/configuration not released; Single venue: all data from ICLR 2026; field/venue specificity unresolved; Temporal leakage risk: model's August 2025 knowledge cutoff precedes January 2026 ICLR outcome release, controlling for outcome memorization; arXiv-twin leakage: pre-cutoff papers on arXiv excluded via self-identity step in pipeline; Selection bias from balanced sampling across decision tiers (natural acceptance rate would starve accepted tiers); Ground-truth reliability: NeurIPS consistency experiments show review is noisy; decisions are themselves unreliable; Potential LLM-assistance in ground-truth reviewer ratings (acknowledged as ceiling on validity); Low-end bridge validity: cheap-model (GPT-5.4-mini) to frontier-model (GPT-5.4) correlation weakest at low score quintile (ρ=0.52); Confirmation bias: author founded and operates AIPR system under evaluation; Halo effect: four informative dimensions moderately correlated (mean r=0.49), not fully independent assessments
Limitations
- "The reject class is 'weak relative to the venue bar,' not weak absolutely, and the decision is itself noisy [3, 8]
- we mitigate by also validating against the mean rating, but no analysis exceeds its ground truth, and that ground truth may itself be partly LLM-assisted [22], a ceiling on H4." Additionally, "The scale rests on the cheap-model bridge, which is weakest at the low end, which is why we report the deployable flag on AIPR (GPT-5.4) directly." The study is limited to "one venue in one field, so we do not claim cross-field generalization." The authors also note that "the scorer is proprietary: the released data, labels, and analysis code make the score-to-outcome study fully reproducible, but the manuscript-to-score function is audited through its outputs rather than re-implemented." A "second-venue replication (ICLR 2025, AIPR (GPT-5.4-mini), as a fully pre-cutoff contaminated contrast) is reserved for future work."
Open questions raised
- Cross-field generalization: study limited to one venue in machine learning; applicability to other fields unstudied
- Second-venue replication: ICLR 2025 fully pre-cutoff contaminated contrast deferred to future work
- Prospective cohort: grading future venue before decisions released to eliminate leakage entirely
- Prestige-perturbation experiment: whether model priors on authors/institutions affect scores (separate study mentioned)
- Per-review ratings: full per-review ratings not released; future work awaits their release to compute alternate rating aggregations
- Citation dimension: lightly weighted citation audit uninformative this run for technical reason; grounded search-backed citation signal identified as obvious next step
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations