Deterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics Architecture
Yoojin Nam, Jinhoon Jeong, Namkug Kim · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Mixed-methods evaluation combining three end-to-end demonstrations on public datasets (Wisconsin Breast Cancer for STARD, BCG vaccine trials for PRISMA, NHANES 2017-2018 for STROBE) and a seeded-defect ablation study.
Sample
> 1000, 4 groups
Primary method
Seeded-defect harness: offline injection with single-defect-per-copy design measuring recall (detected/injected) and clean false-positive rates. Content-hash manifest verification using SHA-256. Reproducibility enforced through pinned package versions and per-column value hashing. Ablation: single-prompt LLM review on identical 27 defects using one model, one fixed prompt, one date, provider-default sampling. Demonstration analyses used standard scientific packages (scikit-learn, metafor, R survey package) with pinned versions. No inferential statistical testing performed on toolkit evaluation.
Main result
The study found that "Across the three pipelines every content-hash manifest verified clean and the gates surfaced real defects, including a prognostic claim in a single-time-point study corrected to association-only language. On 27 identical injected defects the deterministic gates detected all 27 with no false positives on the matched clean fixtures, whereas a generic single-prompt LLM reviewer detected 11, its misses concentrated in generated-code, bibliography-internal, and style defects the prose does not expose." This demonstrates that deterministic verification gates achieve perfect detection of seeded defects while avoiding false positives, outperforming generic LLM self-review.
Reports effect sizes.
Research paradigm
Engineering/design science with empirical validation
Author conclusions
The authors conclude: "Determinism-where-possible verification yields an auditable, re-executable trail that exposes the evidence a human needs to check an LLM-assisted manuscript. This is feasibility and reproducibility evidence, not a claim of human-competitive quality, which a separate blinded study addresses." They further state the contribution is "an architecture that produces an auditable trail a human can verify, not a system that replaces human peer review or is shown here to match human writing quality."
Risk of bias
Author-evaluated defects: first author judged whether each flag was genuine defect or false positive against source artifacts, introducing potential confirmation bias; Non-representative defect set: seeded defects deliberately derived from known failure modes, not sampling real-world error distribution; Limited LLM comparison: ablation used only one model, one fixed prompt, one date against identical defects; not universal model-superiority claim; Training data contamination possible: demonstration datasets are public and may lie within model training corpus; Single-author implementation: demonstration analyses executed by first author; no independent replication mentioned; Author conflict of interest: Y.N. is founder and principal developer of Aperivue and MedSci Skills toolkit being evaluated; Selection bias in seeded-defect cases: deliberately chosen from known failure modes rather than random sampling; Model selection bias in ablation study: single model, single fixed prompt, one date, not universal claim; Potential training-data contamination: public datasets used in demonstrations may be in LLM training corpus; Evaluation harness first-author judgment calls for false-positive classification (mitigated by re-executable diagnostics); Missing network-required citation validators in defect injection (2 of 19 distinct injectors not run); Conflict of interest: authors are developers of the toolkit being evaluated; Training data contamination: demonstration datasets are public and may lie within the model's training corpus; Limited defect set: 27 injected defects may not represent all real-world error types; Single model evaluated in ablation: comparison uses one model family, one fixed prompt, one date; Lack of blinded evaluation: first author judged whether flags were genuine defects; Self-review bias: toolkit authors evaluated their own instrument
Limitations
- The authors explicitly state: "We are explicit about what these demonstrations do and do not establish
- They establish that the pipeline runs end to end across three study types, that its outputs are reproducible from archived artifacts, and that its integrity gates surface real defects with re-executable evidence
- They do not establish that the resulting manuscripts are of publishable quality or competitive with human writing." Additionally, "The contribution here is the auditable trail, not a verdict on the prose it accompanies." The seeded-defect set is described as "deliberately family-complete and grounded in recurring failure modes
- we do not claim it is exhaustive, and it is best read as a regression-style challenge suite for known integrity-failure modes rather than a representative sample of real-world manuscript errors." One honest false positive is reported: "an expected baseline signal, such as the offline reference checker marking every entry unverified, is not counted."
Open questions raised
- Quality assessment: authors defer manuscript quality evaluation to a separate blinded study
- Training data contamination: direct contamination assessment deferred to future work
- Extended LLM-review comparison: a bounded self-review convergence loop with MI-CLEAR-LLM logging is specified but not executed for this release
- Detector coverage gaps: 18 of 21 detectors ship with regression tests; some citation-family detectors lack fixtures or regression tests
- Exhaustive defect taxonomy: current defect set acknowledged as not exhaustive; future work to extend representative error sampling
- Blinded evaluation of manuscript quality (separate study acknowledged as deferred work)
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations