12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Toward an Engineering of Science: Rebalancing Generation and Verification in the Age of AI

Jiaqi Ma · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Main result

The paper identifies that "AI dramatically lowers generation costs but has not proportionally reduced verification costs" which "creates conditions for epistemic pollution: the contamination of scientific literature by plausible yet unreliable artifacts." The authors conclude that "a central bottleneck of AI-era science is therefore shifting from generating plausible research artifacts toward verifying them to maintain a reliable scientific knowledge base." They demonstrate through a prototype artifact called blueprints that restructuring scientific artifacts from "prose-first narratives into structured, decomposed representations of shorter local arguments" can rebalance verification costs.

Research paradigm

Design theory / Engineering of infrastructure systems

Author conclusions

The authors conclude: "AI has made plausible scientific artifacts cheap to generate while leaving reliable verification comparatively expensive. This is not only a model-capability problem, but an infrastructure problem: scientific systems built around papers, peer review, and citation were calibrated to a world in which plausibility itself was expensive to produce. We have suggested for one artifact-level response: blueprints, structured and decomposed argument graphs that make support relations inspectable before a verifier has to reconstruct them from prose. Papers remain the narrative interface of science; blueprints are proposed as a verification interface. The broader argument is that AI-era scientific artifacts should be designed for verification as deliberately as they are designed for communication, so the infrastructure that admits scientific claims can keep pace with the systems that generate them."

Limitations

  • The authors explicitly state: "This paper focuses on staking out a position and proposing a preliminary prototype design
  • The claim that recording argument structure upstream can reduce the reconstruction burden paid by downstream verifiers is a hypothesis that has yet to be fully tested." They further note: "We do not claim to address novelty or significance verification
  • Those problems are real, but they are not where the artifact-level intervention developed here is most likely to pay off." Additionally, "Blueprints therefore do not claim novelty in representing claims, evidence, provenance, or scholarly relations as structured data."

Open questions raised

  • The authors identify three natural starting points for future empirical validation: (1) studies of blueprint authoring overhead and verification time on small case-study projects, (2) comparison of reviewer agreement and error-detection rates between prose papers and corresponding blueprints, and (3) validation of vocabulary extensions in disciplines outside their default. They also note that blueprints may eventually serve as harnesses for AI-assisted science generation, beyond their current verification focus.
  • The authors identify three natural starting points for empirical validation: (1) studies of blueprint authoring overhead and verification time on small case-study projects, (2) comparison of reviewer agreement and error-detection rates between prose papers and corresponding blueprints, and (3) validation of vocabulary extensions in disciplines outside their default. They also note that "Blueprints may matter not only for verification, but eventually for generation" and could serve as "harnesses for AI-assisted science" similar to how test suites function in software development.
  • The authors identify three natural starting points for empirical validation: "studies of blueprint authoring overhead and verification time on small case-study projects, comparison of reviewer agreement and error-detection rates between prose papers and corresponding blueprints, and validation of vocabulary extensions in disciplines outside our default." They also propose that "Blueprints could play a similar role for scientific arguments" in generation, where "blueprints are a candidate harness for AI-assisted science: they make verification more local now, and may make generation more disciplined later."
Code: https://github.com/PatrickMassot/leanblueprint; https://github.com/jiaqima/blueprintExtracted from: pdfAgreement 74%

Explore related topics

Related papers