TAB-AUDIT: Detecting AI-Fabricated Scientific Tables via Multi-View Likelihood Mismatch
Shuo Huang, Yan Pen, Lizhen Qu · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational study combining dataset construction, feature engineering, and machine learning classification.
Main result
The study found that "Random Forest on FUSEDALL gives the strongest completed holdout result" with "0.987 AUROC on the in-domain test set and 0.883 AUROC on the out-of-domain test set". Furthermore, "the core mismatch signal is real. Even the compact PPL4 version of TAB-AUDIT already separates human and fabricated papers substantially better than external zero-shot baselines adapted from prose detection," with TAB-AUDIT reaching "AUROC/AUPRC of 0.884/0.804, 0.902/0.855, and 0.866/0.820" across different observer models, whereas "Binoculars reaches 0.683/0.701 and the DetectGPT-style baseline drops to 0.397/0.430".
Research paradigm
Empirical computational; discriminative machine learning for forensic detection
Author conclusions
The authors conclude that "the strongest evidence comes from combining a provenance-sensitive likelihood contrast with complementary structure and numeric regularities." They emphasize that "the useful evidence is not only whether a table looks fluent or well formatted, but whether its numeric behavior is appropriately matched to its scientific skeleton." The authors further conclude that TAB-AUDIT should be viewed as "an artifact-only, low-FPR screening tool for prioritizing manual review," and note that "if the benchmark is realistic, then a useful detector should be traceable to a concrete distributional gap rather than to opaque classifier behavior."
Risk of bias
Small human evaluation study (n=3 participants) may not generalize to broader population; Benchmark restricted to recent empirical NLP papers, limiting domain generalization; Models trained on GPT-4o-generated papers may overfit to that generator's specific artifacts; Extraction noise and false positives in table extraction pipeline could introduce systematic errors; Evaluation on ACL proceedings conducted without ground-truth labels, limiting validation; Selection bias: Human papers drawn only from recent arXiv NLP papers (within one month), potentially limiting representativeness to other venues and fields; Domain bias: Benchmark restricted to experimental NLP papers, potentially limiting generalization to other scientific domains; Generator bias: Main benchmark uses only GPT-4o for fabrication; out-of-domain evaluation uses GPT-5.2, Qwen2.5, and Llama3.1, but representativeness of these generators for future LLMs unknown; Extraction bias: Table extraction pipeline quality not fully characterized; invalid tables filtered out; Observer model dependency: Detection performance varies substantially by choice of observer model (GPT-2, Qwen, Llama), suggesting the signal may be model-specific; Data leakage risk: Recency constraint employed (papers within one month of collection) to reduce pretraining data contamination, but not fully eliminated; Human study sample bias: Only 3 CS PhD candidates in human evaluation study; not representative of general population or domain experts; Selection bias: Human corpus restricted to recent arXiv papers in empirical NLP only, potentially introducing domain bias; Generator bias: Main benchmark uses GPT-4o; transfer evaluated on GPT-5.2, Qwen2.5, and Llama3.1; Extraction bias: Table extraction and validity filtering may introduce systematic errors; Confounding: Observer model choice affects detection performance differently for in-domain vs out-of-domain data
Limitations
- The authors state: "Our intent/event proxies are lightweight and may be noisy when section headers or reliable topic cues are absent." Additionally, "The relatively low AUC suggests that AI-generated tables can exhibit a high degree of surface-level realism, making human discrimination challenging in the absence of additional textual or methodological context." The authors also note that "This study should be interpreted with caution due to the small number of participants" (three PhD candidates) in the human evaluation study
- Furthermore, "The surrounding prose is included primarily to support realistic presentation of the tables, not to simulate a fully validated scientific contribution," and the synthetic outputs should be interpreted as "literature-grounded but not experiment-grounded."
Open questions raised
- The authors identify that "there has been no systematic study on evaluating and detecting AI-generated fabricated experimental tables. Prior work has primarily focused on textual artifacts, such as hallucinated citations or generated narratives, while largely overlooking structured empirical evidence in tables." They also note the need for cross-generator evaluation and improved transfer learning under domain shift.
- The paper identifies several gaps: (1) prior work has focused on textual artifacts (hallucinated citations, generated narratives) while largely overlooking structured empirical evidence in tables; (2) lack of quantitative analyses and benchmark datasets limits ability to characterize prevalence, patterns, and detectability of fabricated tables; (3) generalization to non-English scientific papers and other domains beyond empirical NLP remains to be validated
- The authors identify that "there has been no systematic study on evaluating and detecting AI-generated fabricated experimental tables" and that "Prior work has primarily focused on textual artifacts, such as hallucinated citations or generated narratives... while largely overlooking structured empirical evidence in tables." They note that "the lack of quantitative analyses and benchmark datasets limits our ability to characterize the prevalence, patterns, and detectability of fabricated tables."
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations