Toward an Engineering of Science: Rebalancing Generation and Verification in the Age of AI
Jiaqi Ma · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Main result
The paper identifies that "AI dramatically lowers generation costs but has not proportionally reduced verification costs" which "creates conditions for epistemic pollution: the contamination of scientific literature by plausible yet unreliable artifacts." The authors conclude that "a central bottleneck of AI-era science is therefore shifting from generating plausible research artifacts toward verifying them to maintain a reliable scientific knowledge base." They demonstrate through a prototype artifact called blueprints that restructuring scientific artifacts from "prose-first narratives into structured, decomposed representations of shorter local arguments" can rebalance verification costs.
Research paradigm
Design theory / Engineering of infrastructure systems
Author conclusions
The authors conclude: "AI has made plausible scientific artifacts cheap to generate while leaving reliable verification comparatively expensive. This is not only a model-capability problem, but an infrastructure problem: scientific systems built around papers, peer review, and citation were calibrated to a world in which plausibility itself was expensive to produce. We have suggested for one artifact-level response: blueprints, structured and decomposed argument graphs that make support relations inspectable before a verifier has to reconstruct them from prose. Papers remain the narrative interface of science; blueprints are proposed as a verification interface. The broader argument is that AI-era scientific artifacts should be designed for verification as deliberately as they are designed for communication, so the infrastructure that admits scientific claims can keep pace with the systems that generate them."
Limitations
- The authors explicitly state: "This paper focuses on staking out a position and proposing a preliminary prototype design
- The claim that recording argument structure upstream can reduce the reconstruction burden paid by downstream verifiers is a hypothesis that has yet to be fully tested." They further note: "We do not claim to address novelty or significance verification
- Those problems are real, but they are not where the artifact-level intervention developed here is most likely to pay off." Additionally, "Blueprints therefore do not claim novelty in representing claims, evidence, provenance, or scholarly relations as structured data."
Open questions raised
- The authors identify three natural starting points for future empirical validation: (1) studies of blueprint authoring overhead and verification time on small case-study projects, (2) comparison of reviewer agreement and error-detection rates between prose papers and corresponding blueprints, and (3) validation of vocabulary extensions in disciplines outside their default. They also note that blueprints may eventually serve as harnesses for AI-assisted science generation, beyond their current verification focus.
- The authors identify three natural starting points for empirical validation: (1) studies of blueprint authoring overhead and verification time on small case-study projects, (2) comparison of reviewer agreement and error-detection rates between prose papers and corresponding blueprints, and (3) validation of vocabulary extensions in disciplines outside their default. They also note that "Blueprints may matter not only for verification, but eventually for generation" and could serve as "harnesses for AI-assisted science" similar to how test suites function in software development.
- The authors identify three natural starting points for empirical validation: "studies of blueprint authoring overhead and verification time on small case-study projects, comparison of reviewer agreement and error-detection rates between prose papers and corresponding blueprints, and validation of vocabulary extensions in disciplines outside our default." They also propose that "Blueprints could play a similar role for scientific arguments" in generation, where "blueprints are a candidate harness for AI-assisted science: they make verification more local now, and may make generation more disciplined later."
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT in education: Strategies for responsible implementationMohanad Halaweh · 2023 · 576 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations