Are we still able to recognize pearls? Machine-driven peer review and the risk to creativity: An explainable RAG-XAI detection framework with markers extraction
Alin-Gabriel Văduva, Simona-Vasilica Oprea, Adela Bâra · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Multi-stage machine learning study combining: (1) dataset construction from multiple sources (PeerRead corpus for human reviews, GPT-4o and adversarial LLMs for AI reviews), (2) LLM-based feature extraction using 8-marker taxonomy, (3) multi-classifier training (XGBoost, Random Forest, LightGBM, Logistic Regression) with stratified train-test split (70%-30%), (4) SHAP-based explainability analysis, and (5) Retrieval-Augmented Generation (RAG) evaluation using FAISS for semantic similarity search..
Primary method
Design science research with iterative development of detection framework components (marker extraction, classification, explainability, retrieval modules)
Main result
The experimental results demonstrate that "the three ensemble methods, XGBoost, RF and LightGBM, achieve near-identical performance, while the LR baseline performs substantially lower" with "accuracy exceeding 99.6% and AUC-ROC values approaching 1.0, while maintaining extremely low false positive and false negative rates." The framework integrates "a curated dataset of human and AI-generated reviews, retrieval-based similarity analysis (RAG) and a marker-detection layer designed to capture structural and stylistic signals specific to the peer-review genre."
Research paradigm
Positivist/Empiricist
Author conclusions
"Overall, the results indicate that the proposed framework achieves state-of-the-art detection performance and addresses practical requirements such as robustness to adversarial manipulation, model interpretability and human-centered validation, making it well-suited for deployment in real peer review monitoring and editorial decision-support systems." However, the authors recognize that "a more practical and future approach is therefore to shift the focus from authorship detection to evaluation of review quality and relevance, assessing whether the feedback meaningfully engages with the manuscript's content, methodology and contributions."
Risk of bias
Dataset imbalance (74.3% human vs. 25.7% AI) addressed through cost-sensitive learning but may still introduce bias; LLM self-recognition bias during feature extraction (mentioned: early prompts framing LLM as 'expert detector' produced polarized scores); Adversarial reviews generated specifically to evade detection may not reflect real-world obfuscation strategies; Dataset primarily from computer science venues (ICLR, ACL, CoNLL), limiting generalizability to other disciplines; Feature extraction relies on a single LLM (Claude Sonnet 4.6) which may introduce model-specific bias; Dataset composition bias: Human reviews dominated by ICLR 2017 (n=5,458) vs. ACL 2017 (n=275) and CoNLL 2016 (n=39), limiting generalizability across venues; Class imbalance: 74.3% human vs. 25.7% AI-generated reviews; AI generation bias: Adversarial reviews deliberately trained to evade detection, creating artificial difficulty not representative of typical AI usage; Self-recognition bias: LLM feature extractor exhibited bias toward recognizing its own outputs in early experiments; Domain specificity: Framework trained on computer science peer reviews may not generalize to other disciplines; Dataset composition bias: 74.3% human reviews vs. 25.7% AI reviews may not reflect actual prevalence in real peer review systems; Source selection bias: Human reviews drawn primarily from ICLR 2017 (n=5,458), with minimal representation from ACL 2017 (n=275) and CoNLL 2016 (n=39); Self-recognition bias in LLM-based feature extraction: Authors acknowledge that "prompts framing the LLM as 'an expert in detecting AI-generated reviews' produced highly polarized scores" and required mitigation through neutral framing; Adversarial review generation bias: Adversarial prompts explicitly designed to suppress markers may not represent realistic author obfuscation attempts; Marker taxonomy bias: Eight markers are theoretically motivated but may miss other discriminative patterns in peer reviews; Generalization concerns: Framework trained on computer science reviews may not transfer to other domains
Limitations
- The authors acknowledge that "detection remains unreliable, particularly in cases of partial AI assistance/well-edited outputs" and note that "the increasing use of AI-assisted or AI-curated peer reviews presents a challenge that is difficult to regulate or prevent, as such usage is often indistinguishable from human-written content and can be easily disguised through editing or paraphrasing." Additionally, the framework relies on "a curated dataset of human and AI-generated reviews" which may not represent all peer review contexts.
Open questions raised
- Need for domain-specific detection models tailored to peer-review genre rather than generic AI-text detectors
- Lack of clear editorial protocols for identifying and addressing suspected AI-generated reviews
- Limited research on interaction between human reviewers and AI tools
- Need for transparent policies, disclosure requirements and accountability mechanisms in peer review
- Shift from detection to quality-based evaluation frameworks
- Authors identify the need to shift focus from detection to quality evaluation; the inherent limitations of detection at scale given disguise/evasion techniques; the lack of clear editorial protocols for handling suspected AI-generated reviews; and the need for transparent policies and disclosure requirements.
Explore related topics
Related papers
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Ethical Dilemmas in Using AI for Academic Writing and an Example Framework for Peer Review in Nephrology Academia: A Narrative ReviewJing Miao · 2023 · 93 citations
- Artificial Intelligence in Peer Review: Enhancing Efficiency While Preserving IntegrityBohdana Doskaliuk · 2025 · 59 citations