From Isolated Scoring to Collaborative Ranking: A Comparison-Native Framework for LLM-Based Paper Evaluation
Pujun Zheng, Jiacheng Yao, Jinquan Zheng, Chenyang Gu, Guoxiu He, Jiawei Liu et al. · ArXiv.org · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Empirical computational study using supervised fine-tuning (SFT) and reinforcement learning (RL) to train a 7B parameter language model.
Sample
N = 1634, 14 groups
Primary method
Mann-Whitney U test (non-parametric test for comparing rank distributions across unseen datasets); Maximum Likelihood Estimation (MLE) for Bradley-Terry model; LoRA adaptation for efficient fine-tuning; GRPO (Group Relative Policy Optimization) with modifications for reinforcement learning; 30 independent runs with distinct random seeds for hyperparameter analysis
Main result
The study found that "our framework achieves an average relative improvement of 21.8% over the strong baseline DeepReview-14B" on paper evaluation tasks. The framework "achieves an average relative improvement of 21.8% over the strong baseline DeepReview-14B, while exhibiting robust generalization to five previously unseen datasets." Additionally, "Our model achieves MAP@20 of 0.7076 and NDCG@20 of 0.8153, with MAP@20 exhibiting a 52.6% improvement over the second-best model."
Reports effect sizes and confidence intervals.
Research paradigm
Empirical - computational experiment and comparative evaluation
Author conclusions
The authors conclude: "We propose a comparison-native framework that reformulates paper evaluation as collaborative ranking rather than isolated scoring. By integrating graph-based pair sampling with comparison-aware training, the framework consistently improves both ranking and decision performance and generalizes well to unseen venues. These findings show that our framework enables LLMs to learn more accurate and transferable paper comparison capabilities."
Risk of bias
Positional bias in LLM comparisons (addressed through SFT and RL training); Dataset-specific factors affecting generalization (addressed through comparison-native approach); Potential information leakage from training on 2025 data when evaluating 2024-era models; Limited to computer science domain papers only; Potential biases present in original human review scores used as ground truth; Dataset specificity: Training only on computer science conference papers (2025 timeframe) limits domain generalization; Model size constraint: 7B parameter model may have reduced knowledge breadth compared to larger models; Information limitation: Using only titles and abstracts omits full-text details that could affect quality assessment; Positional bias: LLMs show inherent preference for position order, though authors demonstrate mitigation through SFT/RL; Data leakage risk: Despite precautions, agents and comparison baselines (AIScientist, AgentReview, PairReview) were noted as not ruling out possibility of data leakage; Review score dependency: Ground truth defined by mean human reviewer scores, which may contain review biases; Selection bias: Limited to computer science conference papers only; Dataset bias: Only six top-tier ML/AI conferences in 2025; Positional bias in LLM preferences (though mitigated through training); Training data leakage risk if base model inadvertently learned test data; Potential reinforcement of biases present in training data from top-tier venues
Limitations
- "Our experiments rely exclusively on computer science conference papers, which limits the generalizability of our findings
- The dataset includes only papers from six leading machine learning and artificial intelligence conferences in 2025." Additionally, "resource constraints restricted us to training a 7B model" and "Using only titles and abstracts significantly reduces computational cost and allows wider applicability, but it inevitably constrains information available in full papers, which may cause mild performance degradation." The authors further note "although our results represent notable gains over prior work, they still fall short of human reviewing quality."
Open questions raised
- The authors identify several gaps: (1) Limited application beyond computer science domain; (2) Resource constraints limiting exploration of larger model sizes; (3) Need for more sophisticated designs and extension to additional fields; (4) Gap between automated systems and human reviewing quality; (5) Early-stage research on pairwise and listwise approaches requiring further development; (6) Need for integration of complementary mechanisms for broader assessment of scientific progress.
- Authors identify that existing methods using pairwise or listwise comparisons either do not train model parameters for comparison tasks or still output absolute scores, limiting both performance and generalization. They note that systematically strengthening comparison-native modeling from perspectives of data construction, model learning, and inference remains largely unexplored. Future directions include extending to additional fields beyond computer science, leveraging larger models with more computational resources, and conducting continued bias audits.
- Extension to additional scientific fields beyond computer science
- Leveraging larger models with more computational resources
- More sophisticated pair sampling designs
- Integration of complementary mechanisms for identifying emerging research directions
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Beware of metacognitive laziness: Effects of generative artificial intelligence on learning motivation, processes, and performanceYizhou Fan · 2024 · 419 citations
- Impact of AI assistance on student agencyAli Darvishi · 2023 · 384 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations