AiraXiv: An AI-Driven Open-Access Platform for Human and AI Scientists
Junshu Pan, Panzhong Lu, Yixuan Weng, Qiyao Sun, Fang Guo, Zijie Yang et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Real-world deployment and case study evaluation.
Primary method
Design science research with real-world deployment validation; iterative platform development integrating user feedback from ICAIS 2025 deployment
Main result
AiraXiv successfully enabled rapid feedback and quality evolution in its deployment as the official infrastructure for ICAIS 2025. The platform "reduced the traditional 9-month conference cycle to 1.5 months, and successfully supported academic exchanges among over 200 participants, including 6 Nobel laureates and dozens of prominent researchers." Furthermore, "accepted and rejected papers by human experts exhibit a clear separation in AI scores, with accepted papers achieving higher median scores (>4.5) and rejected papers clustering around a median of approximately 3.0," and "resubmitted versions of manuscripts generally achieved higher median AI scores compared to initial submissions, suggesting that authors were able to effectively interpret and act upon the automated feedback to refine their work."
Research paradigm
Design science / socio-technical systems research
Author conclusions
"We introduced AiraXiv, an AI-driven, open-access preprint platform for iterative scholarly communication by both human and AI scientists." The authors conclude that "These results suggest that AI-assisted, community-feedback publishing can reduce review bottlenecks, accelerate iteration, and broaden valuable scientific outputs," and they "hope AiraXiv serves as a foundation for an open, fast, and inclusive scientific ecosystem in the AI era, while continuing to evolve through future community participation, system extension, and broader research collaboration."
Risk of bias
Potential bias from limited deployment scope (single conference); AI reviewer bias across domains; vulnerability to adversarial feedback; selection bias toward AI-generated papers in submission landscape.; Domain-specific bias: AI reviewers may perform inconsistently across different research domains; Limited deployment scope: Only one real-world deployment (ICAIS 2025) conducted; generalizability unclear; Reviewer bias: AI reviewers may have systematic biases that are not detected by human experts; Selection bias: Conference participants (200+) may not represent broader research community; Model training bias: AI reviewer models trained on existing data may perpetuate historical biases; Adversarial manipulation: System vulnerable to coordinated or low-quality feedback attacks; AI reviewer bias across domains; Potential model overconfidence and hallucination in reviews; Single-deployment evaluation (ICAIS 2025 only); Potential adversarial feedback in community-driven evaluation; Self-selection bias of authors submitting to novel platform
Limitations
- The authors state that "AI-assisted reviewing signals in our work are imperfect and may be biased or unstable across domains, which can mislead readers if interpreted as final judgments." Additionally, "the end-to-end author-reader feedback loop may be vulnerable to low-quality, adversarial, or coordinated feedback, which requires robust moderation and abuse prevention mechanisms." Finally, "we have only validated AiraXiv in limited real-world deployments, and broader evidence is needed to understand long-term community dynamics, incentive alignment, and the generalization of our work."
Open questions raised
- The authors identify the need for: (1) scaling across diverse fields; (2) strengthening governance and auditability; (3) defending against adversarial content; (4) understanding long-term community dynamics and incentive alignment; (5) improving AI review quality and inference efficiency; and (6) robust moderation and abuse prevention mechanisms for feedback systems.
- Scaling across diverse fields: Unclear how AI reviewing generalizes to non-computer science domains
- Strengthening governance and auditability: Need for robust mechanisms to prevent abuse and ensure transparency
- Defending against adversarial content: Adversarial robustness requires further development
- Long-term community dynamics: Limited evidence on sustainability and incentive alignment over extended periods
- Computational efficiency: Large-scale AI reviewing may introduce substantial computational overhead as submission volumes grow
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Systematic review of research on artificial intelligence applications in higher education – where are the educators?Olaf Zawacki‐Richter · 2019 · 5,282 citations
- State of the art and practice in AI in educationW. Holmes · 2022 · 758 citations
- Co-designing AI Education Curriculum with Cross-Disciplinary High School TeachersBenjamin Xie · 2024 · 28 citations
- GAIDeT (Generative AI Delegation Taxonomy): A taxonomy for humans to delegate tasks to generative artificial intelligence in scientific research and publishingYana Suchikova · 2025 · 24 citations