12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Deterministic Integrity Gates for LLM-Assisted Clinical Manuscript Preparation: An Auditable Biomedical Informatics Architecture

Yoojin Nam, Jinhoon Jeong, Namkug Kim · ArXiv.org · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
D
Evidence
0
Citations
0.00
FWCI

Methodology & findings

Study design

Mixed-methods evaluation combining three end-to-end demonstrations on public datasets (Wisconsin Breast Cancer for STARD, BCG vaccine trials for PRISMA, NHANES 2017-2018 for STROBE) and a seeded-defect ablation study.

Sample

> 1000, 4 groups

Primary method

Seeded-defect harness: offline injection with single-defect-per-copy design measuring recall (detected/injected) and clean false-positive rates. Content-hash manifest verification using SHA-256. Reproducibility enforced through pinned package versions and per-column value hashing. Ablation: single-prompt LLM review on identical 27 defects using one model, one fixed prompt, one date, provider-default sampling. Demonstration analyses used standard scientific packages (scikit-learn, metafor, R survey package) with pinned versions. No inferential statistical testing performed on toolkit evaluation.

Main result

The study found that "Across the three pipelines every content-hash manifest verified clean and the gates surfaced real defects, including a prognostic claim in a single-time-point study corrected to association-only language. On 27 identical injected defects the deterministic gates detected all 27 with no false positives on the matched clean fixtures, whereas a generic single-prompt LLM reviewer detected 11, its misses concentrated in generated-code, bibliography-internal, and style defects the prose does not expose." This demonstrates that deterministic verification gates achieve perfect detection of seeded defects while avoiding false positives, outperforming generic LLM self-review.

Reports effect sizes.

Research paradigm

Engineering/design science with empirical validation

Author conclusions

The authors conclude: "Determinism-where-possible verification yields an auditable, re-executable trail that exposes the evidence a human needs to check an LLM-assisted manuscript. This is feasibility and reproducibility evidence, not a claim of human-competitive quality, which a separate blinded study addresses." They further state the contribution is "an architecture that produces an auditable trail a human can verify, not a system that replaces human peer review or is shown here to match human writing quality."

Risk of bias

Author-evaluated defects: first author judged whether each flag was genuine defect or false positive against source artifacts, introducing potential confirmation bias; Non-representative defect set: seeded defects deliberately derived from known failure modes, not sampling real-world error distribution; Limited LLM comparison: ablation used only one model, one fixed prompt, one date against identical defects; not universal model-superiority claim; Training data contamination possible: demonstration datasets are public and may lie within model training corpus; Single-author implementation: demonstration analyses executed by first author; no independent replication mentioned; Author conflict of interest: Y.N. is founder and principal developer of Aperivue and MedSci Skills toolkit being evaluated; Selection bias in seeded-defect cases: deliberately chosen from known failure modes rather than random sampling; Model selection bias in ablation study: single model, single fixed prompt, one date, not universal claim; Potential training-data contamination: public datasets used in demonstrations may be in LLM training corpus; Evaluation harness first-author judgment calls for false-positive classification (mitigated by re-executable diagnostics); Missing network-required citation validators in defect injection (2 of 19 distinct injectors not run); Conflict of interest: authors are developers of the toolkit being evaluated; Training data contamination: demonstration datasets are public and may lie within the model's training corpus; Limited defect set: 27 injected defects may not represent all real-world error types; Single model evaluated in ablation: comparison uses one model family, one fixed prompt, one date; Lack of blinded evaluation: first author judged whether flags were genuine defects; Self-review bias: toolkit authors evaluated their own instrument

Limitations

  • The authors explicitly state: "We are explicit about what these demonstrations do and do not establish
  • They establish that the pipeline runs end to end across three study types, that its outputs are reproducible from archived artifacts, and that its integrity gates surface real defects with re-executable evidence
  • They do not establish that the resulting manuscripts are of publishable quality or competitive with human writing." Additionally, "The contribution here is the auditable trail, not a verdict on the prose it accompanies." The seeded-defect set is described as "deliberately family-complete and grounded in recurring failure modes
  • we do not claim it is exhaustive, and it is best read as a regression-style challenge suite for known integrity-failure modes rather than a representative sample of real-world manuscript errors." One honest false positive is reported: "an expected baseline signal, such as the offline reference checker marking every entry unverified, is not counted."

Open questions raised

  • Quality assessment: authors defer manuscript quality evaluation to a separate blinded study
  • Training data contamination: direct contamination assessment deferred to future work
  • Extended LLM-review comparison: a bounded self-review convergence loop with MI-CLEAR-LLM logging is specified but not executed for this release
  • Detector coverage gaps: 18 of 21 detectors ship with regression tests; some citation-family detectors lack fixtures or regression tests
  • Exhaustive defect taxonomy: current defect set acknowledged as not exhaustive; future work to extend representative error sampling
  • Blinded evaluation of manuscript quality (separate study acknowledged as deferred work)
Data: Wisconsin Breast Cancer dataset (scikit-learn); BCG vaccine trials (metafor dat.bcg package); NHANES 2017-2018 (CDC), publicly available; All demonstrations archived with content-hash manifests at Zenodo; Wisconsin Breast Cancer (scikit-learn); BCG vaccine trials (metafor dat.bcg); NHANES 2017-2018 (CDC); All datasets are public and openly available; specific URLs not provided in paper; BCG vaccine trials meta-analysis data (metafor dat.bcg); NHANES 2017–2018 (CDC); All three demonstration datasets are public and redistributableCode: MedSci Skills: https://github.com/Aperivue/medsci-skills; Archived on Zenodo - version 3.8.0: concept DOI 10.5281/zenodo.20155321, version DOI 10.5281/zenodo.20582972, commit 60ce35c; Demonstration version v3.7.0: version DOI 10.5281/zenodo.20577997, commit 5adda7c; MedSci Skills: https://github.com/Aperivue/medsci-skills (MIT license); Zenodo archived: concept DOI 10.5281/zenodo.20155321, version 3.8.0 DOI 10.5281/zenodo.20582972 (commit 60ce35c); Demonstrations generated on v3.7.0 snapshot (DOI 10.5281/zenodo.20577997, commit 5adda7c); https://github.com/Aperivue/medsci-skills (open-source, MIT license); Archived on Zenodo: concept DOI 10.5281/zenodo.20155321 (v3.8.0), version DOI 10.5281/zenodo.20582972 (tag commit 60ce35c); v3.7.0 snapshot: version DOI 10.5281/zenodo.20577997 (commit 5adda7c)Extracted from: pdfAgreement 46%

Explore related topics

Related papers