Are Large Language Models Truly Smarter Than Humans?
Eshwar Reddy M, Sourav Karmakar · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Three mutually reinforcing computational experiments: (1) Experiment 1: lexical contamination scan using web search (Tavily API) to detect benchmark questions in online sources, with dual-condition thresholds (8-gram overlap ≥0.30 AND correct answer present); (2) Experiment 2: behavioral paraphrase diagnostic testing six models on 100 MMLU questions in three surface-form variants (original, paraphrased, indirect-reference); (3) Experiment 3: TS-Guessing protocol probing internal memorization via masked option reconstruction and masked word reconstruction tasks across 513 MMLU questions and six frontier language models..
Main result
The study found that "the overall contamination rate was 13.8%" with "Estimated performance gains from contamination range from +0.030 to +0.054 accuracy points by category." Additionally, "Across all six models, average accuracy dropped from 0.615 on original questions to 0.545 on indirect-reference variants-a drop of 7.0 percentage points," and "The average combined flagging rate across all six models was 72.5%-far above both random baselines-constituting strong behavioral evidence that the vast majority of sampled MMLU questions are internally memorized by frontier LLMs."
Research paradigm
Empirical investigation with computational/behavioral analysis
Author conclusions
"The evidence from three independent experiments points to the same answer: the question cannot be definitively resolved with current benchmark practice, and the available evidence gives substantial reason for skepticism." The authors conclude that "Claims of human-level or superhuman AI performance derived from contaminated public benchmarks deserve the same scrutiny we would apply to any scientific claim where the evaluation procedure is known to be compromised." They further state that "all three experiments agree on the category-level ordering of contamination severity: STEM > Professional > Social Sciences > Humanities. This convergence across three methodologically independent approaches-external web detection, accuracy degradation under surface-form perturbation, and internal reconstruction probing-constitutes the strongest available multi-method evidence that MMLU contamination is a real, structural, and practically consequential phenomenon."
Risk of bias
Selection bias in Experiment 2: 100 questions sampled with deliberate stratification by contamination level and category, potentially biasing toward detecting effects where they exist; Sample size limitation: Only 9 questions per subject in Experiment 1 (513 total) may not be fully representative of each subject's contamination profile; Web search tool limitation: Tavily search API may not comprehensively cover all pretraining data sources; Model selection bias: Six specific frontier models chosen; results may not generalize to other model families; Methodological artifact risk: Paraphrase quality in Experiment 2 depends on GPT-4o's ability to generate semantically equivalent variants; Selection bias in question sampling (random seed 42, but limited to 9 questions per subject); Search engine limitations: Tavily may not comprehensively mirror LLM pretraining corpora; API temperature settings (0) may not reflect deployment conditions; Contamination detection relies on observable web presence, not actual pretraining data contents; Web search API (Tavily) may not capture full pretraining corpora, leading to underestimation of contamination; Conservative two-condition detection rule (8-gram overlap + correct answer presence) may miss paraphrased variants; Selection of specific six MMLU subjects in Experiment 2 may not be representative of all subjects; Temperature set to 0 in model evaluations may not reflect all deployment scenarios; Sampling strategy with fixed seed 42 may introduce systematic bias
Limitations
- The authors state that "our two-condition threshold is conservative: true contamination rates may be higher if paraphrased questions or near-duplicates are counted" and "Tavily does not comprehensively mirror pretraining corpora, so this method is best interpreted as a lower bound on web exposure." Additionally, they note that "web-search-only contamination detection is a significant underestimate for closed-source models whose pretraining corpora extend far beyond what public web indices capture."
Open questions raised
- Need for properly powered private benchmark comparison with 100-200 difficulty-calibrated questions per domain authored after model training cutoffs
- Extension of TS-Guessing protocol to sentence-level reconstructions and multihop compositional probes
- Dedicated investigation of distributed memorization patterns (DeepSeek-R1 anomaly) to determine if it is architectural, corpus-related, or deliberate training choice
- Systematic study of which web-presence features best predict behavioral memorization
- Direct causal study of interaction between contamination and hallucination using controlled synthetic benchmarks
- Need for properly powered private benchmark comparison with 100-200 carefully difficulty-calibrated questions per domain
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations