AI-Driven Scholarly Peer Review via Persistent Workflow Prompting, Meta-Prompting, and Meta-Reasoning
Evgeny Markhasin · ArXiv.org · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.48550/arxiv.2505.03332
Methodology & findings
Study design
Design science methodology involving iterative prompt engineering through meta-prompting, meta-meta-prompting, and meta-reasoning techniques.
Primary method
Design science research with iterative prompt engineering using meta-prompting techniques, meta-reasoning, and collaborative AI-human development. Formalization of tacit expert knowledge through deconstruction and abstraction. PWP (Persistent Workflow Prompting) as core architectural principle.
Main result
The study found that "all tested models, when guided by the PeerReviewPrompt, relatively reliably identified major methodological flaws within the test paper [91] and converged on the conclusion that its central claim (regarding isotopic enrichment) was highly dubious or unsupported by the described methods." Additionally, "the LLM analyses highlighted at least two potentially significant issues not initially noted by the author during manual review," demonstrating that "PWP-guided LLMs not only structure analysis but also augment human review by identifying flaws that might be overlooked due to differing expertise or attention patterns."
Research paradigm
Pragmatist/Engineering-focused
Author conclusions
"Eliciting deep, reliable, domain-specific reasoning from frontier Large Language Models (LLMs) using accessible methods remains a significant challenge, particularly for complex analytical tasks like critical scholarly peer review. This work addressed this challenge by introducing Persistent Workflow Prompting (PWP), a methodology centered on a detailed, hierarchical prompt acting as a persistent workflow library, developed through iterative meta-prompting and meta-reasoning designed to codify expert knowledge." The authors conclude that "sophisticated prompt engineering, informed by meta-reasoning, can translate expert workflows (including tacit knowledge) into structured instructions that actively condition the model for critical evaluation. This provides a feasible 'zero-code' pathway to unlock specialized analytical capabilities within general-purpose LLMs, using only standard chat interfaces."
Risk of bias
Single test case selection bias - only one deliberately flawed paper tested; Author evaluation bias - author's conventional (human-driven) evaluation used as ground truth; Confirmation bias - prompt designed to encourage critical/negative stance may bias LLM toward false positives; Lack of baseline comparison - no zero-shot or simple prompt baseline provided in main results; Limited model diversity - primarily tested on Gemini Advanced 2.5 Pro; Single test case selection bias (deliberately flawed paper); Author expertise bias (test case in author's domain of physical chemistry); Subjective evaluation by author without blinded assessment; No comparison against baseline prompting techniques; Limited model testing (primarily Gemini Advanced 2.5 Pro); Confirmation bias in selecting deliberately flawed test paper; Single test case selection bias - only one known-flawed paper used for development; Author expertise bias - developer is expert in experimental chemistry, limiting generalizability claims; Subjective evaluation - qualitative assessment based on author's conventional evaluation without objective metrics; Model-dependent bias - different LLM models may exhibit different behaviors not systematically characterized; Confirmation bias risk - developer's expectations about methodology flaws may influence prompt engineering; Context window limitations - may affect LLM performance in unpredictable ways; Lack of baseline comparison - no comparison with simpler prompting strategies or human reviews
Limitations
- "The PeerReviewPrompt was developed and primarily tested using a single publication [91]
- Although chosen deliberately for its known flaws, this reliance on one test case limits the assessment of the prompt's generalizability to other experimental chemistry papers, particularly those that are methodologically sound or contain different types of errors." Additionally, "the assessment of the prompt's performance presented herein is qualitative and observational
- No quantitative benchmark was constructed for systematic evaluation against ground truth or objective metrics." The authors also note that "while the PWP aims to guide LLMs towards rigorous analysis, inherent LLM limitations like potential hallucination or inconsistent context recall were observed occasionally but were not systematically characterized or quantified within this study."
Open questions raised
- Need for evaluation against diverse set of experimental chemistry manuscripts (both methodologically sound and flawed)
- Expansion of analytical scope beyond core experimental protocol to include product characterization, subsequent experiments, data presentation, and statistical validity
- Systematic performance evaluation with quantitative benchmarks and comparison to human expert reviews
- Systematic investigation of multimodal (visual-textual) analysis capabilities
- Exploration of PWP methodology for other scientific disciplines beyond chemistry
- Investigation of foundational model development to intrinsically enhance critical evaluation capabilities
Explore related topics
Related papers
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Ethical Dilemmas in Using AI for Academic Writing and an Example Framework for Peer Review in Nephrology Academia: A Narrative ReviewJing Miao · 2023 · 93 citations
- Artificial Intelligence in Peer Review: Enhancing Efficiency While Preserving IntegrityBohdana Doskaliuk · 2025 · 59 citations