12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

AI-Driven Scholarly Peer Review via Persistent Workflow Prompting, Meta-Prompting, and Meta-Reasoning

Evgeny Markhasin · ArXiv.org · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
D
Evidence
0
Citations

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2505.03332

Methodology & findings

Study design

Design science methodology involving iterative prompt engineering through meta-prompting, meta-meta-prompting, and meta-reasoning techniques.

Primary method

Design science research with iterative prompt engineering using meta-prompting techniques, meta-reasoning, and collaborative AI-human development. Formalization of tacit expert knowledge through deconstruction and abstraction. PWP (Persistent Workflow Prompting) as core architectural principle.

Main result

The study found that "all tested models, when guided by the PeerReviewPrompt, relatively reliably identified major methodological flaws within the test paper [91] and converged on the conclusion that its central claim (regarding isotopic enrichment) was highly dubious or unsupported by the described methods." Additionally, "the LLM analyses highlighted at least two potentially significant issues not initially noted by the author during manual review," demonstrating that "PWP-guided LLMs not only structure analysis but also augment human review by identifying flaws that might be overlooked due to differing expertise or attention patterns."

Research paradigm

Pragmatist/Engineering-focused

Author conclusions

"Eliciting deep, reliable, domain-specific reasoning from frontier Large Language Models (LLMs) using accessible methods remains a significant challenge, particularly for complex analytical tasks like critical scholarly peer review. This work addressed this challenge by introducing Persistent Workflow Prompting (PWP), a methodology centered on a detailed, hierarchical prompt acting as a persistent workflow library, developed through iterative meta-prompting and meta-reasoning designed to codify expert knowledge." The authors conclude that "sophisticated prompt engineering, informed by meta-reasoning, can translate expert workflows (including tacit knowledge) into structured instructions that actively condition the model for critical evaluation. This provides a feasible 'zero-code' pathway to unlock specialized analytical capabilities within general-purpose LLMs, using only standard chat interfaces."

Risk of bias

Single test case selection bias - only one deliberately flawed paper tested; Author evaluation bias - author's conventional (human-driven) evaluation used as ground truth; Confirmation bias - prompt designed to encourage critical/negative stance may bias LLM toward false positives; Lack of baseline comparison - no zero-shot or simple prompt baseline provided in main results; Limited model diversity - primarily tested on Gemini Advanced 2.5 Pro; Single test case selection bias (deliberately flawed paper); Author expertise bias (test case in author's domain of physical chemistry); Subjective evaluation by author without blinded assessment; No comparison against baseline prompting techniques; Limited model testing (primarily Gemini Advanced 2.5 Pro); Confirmation bias in selecting deliberately flawed test paper; Single test case selection bias - only one known-flawed paper used for development; Author expertise bias - developer is expert in experimental chemistry, limiting generalizability claims; Subjective evaluation - qualitative assessment based on author's conventional evaluation without objective metrics; Model-dependent bias - different LLM models may exhibit different behaviors not systematically characterized; Confirmation bias risk - developer's expectations about methodology flaws may influence prompt engineering; Context window limitations - may affect LLM performance in unpredictable ways; Lack of baseline comparison - no comparison with simpler prompting strategies or human reviews

Limitations

  • "The PeerReviewPrompt was developed and primarily tested using a single publication [91]
  • Although chosen deliberately for its known flaws, this reliance on one test case limits the assessment of the prompt's generalizability to other experimental chemistry papers, particularly those that are methodologically sound or contain different types of errors." Additionally, "the assessment of the prompt's performance presented herein is qualitative and observational
  • No quantitative benchmark was constructed for systematic evaluation against ground truth or objective metrics." The authors also note that "while the PWP aims to guide LLMs towards rigorous analysis, inherent LLM limitations like potential hallucination or inconsistent context recall were observed occasionally but were not systematically characterized or quantified within this study."

Open questions raised

  • Need for evaluation against diverse set of experimental chemistry manuscripts (both methodologically sound and flawed)
  • Expansion of analytical scope beyond core experimental protocol to include product characterization, subsequent experiments, data presentation, and statistical validity
  • Systematic performance evaluation with quantitative benchmarks and comparison to human expert reviews
  • Systematic investigation of multimodal (visual-textual) analysis capabilities
  • Exploration of PWP methodology for other scientific disciplines beyond chemistry
  • Investigation of foundational model development to intrinsically enhance critical evaluation capabilities
Data: Test paper (manuscript + Supporting Information) available via Open Science Framework (OSF): https://osf.io/nq68y/files/osfstorage?view_only=fe29ffe96a8340329f3ebd660faedd43; Test paper (manuscript + Supporting Information) available via OSF repository: https://osf.io/nq68y/files/osfstorage?view_only=fe29ffe96a8340329f3ebd660faedd43; Test paper (manuscript + SI) available via OSF: https://osf.io/nq68y/files/osfstorage?view_only=fe29ffe96a8340329f3ebd660faedd43Code: Prompt files and demonstration materials available via OSF: https://osf.io/nq68y/files/osfstorage?view_only=fe29ffe96a8340329f3ebd660faedd43; OSF repository with prompt files: https://osf.io/nq68y/files/osfstorage?view_only=fe29ffe96a8340329f3ebd660faedd43; Prompt files available via OSF: https://osf.io/nq68y/files/osfstorage?view_only=fe29ffe96a8340329f3ebd660faedd43Extracted from: pdfAgreement 55%

Explore related topics

Related papers