Beyond Paper-to-Paper: Structured Profiling and Rubric Scoring for Paper-Reviewer Matching
Yicheng Pan, Zhiyuan Ning, Ludi Wang, Yi Du · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Computational framework design with empirical evaluation on three public benchmarks.
Primary method
Design science research with iterative component validation. The design follows a coarse-to-fine pipeline approach combining multiple computational stages, with ablation studies validating each component's necessity.
Main result
P2R achieves state-of-the-art performance on reviewer matching benchmarks, with "P2R establishes a new state-of-the-art, securing the top rank in 10 out of 13 metrics and ranking second in the remaining three. This superiority is consistent across diverse domains, including machine learning (NeurIPS), information retrieval (SIGIR), and the broad scientific landscape (SciRepEval)." On the NeurIPS dataset specifically, "P2R achieves a substantial gain of +5.14 percentage points in Hard P@5 compared to the strongest baseline."
Research paradigm
computational design science
Author conclusions
"P2R advances reviewer matching by directly modeling paper–reviewer fit with multi-dimensional signals rather than relying solely on paper representations. Beyond improved accuracy, our framework offers a practical recipe for applying LLMs to reviewer matching: use LLMs to build structured expertise profiles, retrieve candidates with high recall, and enforce decision criteria via rubrics during final scoring."
Risk of bias
Model-specific bias: reliance on Qwen3 for profiling may introduce model-dependent biases; Limited dataset representation: evaluation restricted to three conference benchmarks; Annotation bias: ground truth labels derived from Area Chair decisions, which may reflect their subjective preferences; Sparse annotation: limited number of ground-truth reviewers per submission; LLM hallucination risk: despite JSON-only prompting constraints, potential for model-generated inaccuracies in profile extraction; Dataset selection bias: evaluation limited to three specific conference benchmarks; Potential LLM hallucination despite JSON-only prompting constraints; Reliance on expert-curated labels which may reflect annotator biases; No discussion of potential fairness issues in reviewer recommendation; Limited ground truth annotations per submission in NeurIPS dataset; Potential LLM hallucination in profile generation despite schema constraints; Stochasticity in LLM outputs - authors report mean performance over three runs to mitigate; Hyperparameter choices (RRF weights, discretization thresholds) not fully justified or sensitivity-tested; Training-free approach claims may mask implicit biases from pre-trained LLM embeddings
Open questions raised
- The authors identify that prior paper-to-paper matching methods implicitly represent reviewers through publication history and rely on textual similarity alone, which is insufficient for capturing multi-dimensional expertise. They note that structural mismatches (e.g., shared topics but differing methodologies or application domains) are difficult to resolve by aggregating paper similarities, motivating the shift toward explicit multi-dimensional characterization.
- The paper identifies that existing paper-to-paper matching paradigms are insufficient because "effective reviewer matching requires capturing multi-dimensional expertise, and textual similarity to past papers alone is often insufficient." Future work is implied to explore applications of this framework beyond the three tested conference domains.
- The paper identifies that "effective reviewer matching requires capturing multi-dimensional expertise, and textual similarity to past papers alone is often insufficient." The authors note that prior work reduces "expertise to coarse textual or topical similarity" and that "implicit embedding similarity can blur distinct expertise dimensions and hide structural mismatches." Future work could address the scalability of LLM-based profiling for very large conferences and explore task-specific fine-tuning of structured profiling.
Explore related topics
Related papers
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- AI-assisted peer reviewAlessandro Checco · 2021 · 261 citations
- Fighting reviewer fatigue or amplifying bias? Considerations and recommendations for use of ChatGPT and other large language models in scholarly peer reviewMohammad Hosseini · 2023 · 209 citations
- Artificial intelligence to support publishing and peer review: A summary and reviewKayvan Kousha · 2023 · 137 citations
- Ethical Dilemmas in Using AI for Academic Writing and an Example Framework for Peer Review in Nephrology Academia: A Narrative ReviewJing Miao · 2023 · 93 citations
- Artificial Intelligence in Peer Review: Enhancing Efficiency While Preserving IntegrityBohdana Doskaliuk · 2025 · 59 citations