Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications
Zhiyuan Peng, Yi Fang · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18653/v1/2025.aisd-main.5
Methodology & findings
Study design
Empirical benchmark evaluation using a novel dataset (SchNovel) of 15,000 paper pairs from arXiv across six research fields.
Sample
N = 15000, 13 groups
Primary method
McNemar test for comparing paired categorical outcomes (novelty assessment accuracy). Pairwise comparisons used throughout experiments with position bias controls (order swapping). Majority voting employed for Self-Consistency method (10 generated sequences). Accuracy metric reported across different categorical variables (field, start year, year gap).
Main result
RAG-Novelty achieved the highest results overall, significantly outperforming the second-best method, except in the mathematics field. Specifically, "RAG-Novelty achieves the highest results overall, significantly outperforming the second-best method, except in the mathematics field." The study found that "pairwise is consistently much better than pointwise across different year gaps," and that "GPT-4o-mini outperformed all other models, demonstrating a substantial advantage over smaller models."
Reports effect sizes and confidence intervals.
Research paradigm
Positivist/empiricist (computational evaluation)
Author conclusions
The authors conclude that "RAG-Novelty achieves the highest results overall, significantly outperforming the second-best method, except in the mathematics field" and that "Our extensive experiments demonstrate that RAG-Novelty outperforms recent baseline models in assessing novelty in scholarly papers." They also note important findings regarding position bias: "Notably, GPT-4o-mini outperformed all other models, demonstrating a substantial advantage over smaller models like LLaMA 3.1-8b, Mistral 7b, and Gemma 2-9b. Despite such success, even ChatGPT 4o-mini and ChatGPT 3.5 exhibit position bias, where the order of papers in the prompt affects their decision-making instead of content alone."
Risk of bias
Position bias in LLM decision-making (order of papers affects evaluation); Field/category bias (lower performance in mathematics and physics); Affiliation bias (preference for top research universities); Lack of domain knowledge in training data for specialized fields; Temporal bias potential in retrieval-augmented approach; Position bias in LLM outputs (order of papers in prompt affects decision-making); Affiliation bias (models show preference for papers from top research universities); Categorical bias (disparities in performance across research fields); Training data bias (LLMs may lack domain-specific knowledge in certain fields); Position bias in LLMs: smaller models exhibit substantial position bias (Mistral 7b favors last paper; LLaMA 3.1-8b favors middle position); Affiliation bias: LLMs showed preference for papers from top research universities over teaching universities, particularly with larger year gaps; Domain bias: Lower accuracy in mathematics and physics possibly due to lack of specialized content in LLM training data; Data source bias: Exclusive reliance on arXiv may not be representative of all scholarly publications; Temporal bias: Model evaluation uses 2024 as reference point but is evaluated on papers up to 2023; potential knowledge cutoff effects; Selection bias in sample: Only six of eight arXiv fields included; authors cite insufficient data in other fields
Limitations
- The authors state: "Our study evaluates an LLM's ability to assess novelty using a research paper's title, abstract, and metadata
- While the abstract provides a strong indication of a paper's content and key findings, it may not fully capture the novelty of the research compared to the complete text
- Abstracts often summarize the main ideas but may omit important technical details." Additionally, "the exclusive use of arXiv data is limiting
- We selected arXiv as an initial step for its broad, publicly accessible range of publications
- Future work can improve robustness using peer-reviewed publications and sampling papers from more sources."
Open questions raised
- Limited evaluation on peer-reviewed publications beyond arXiv
- Need for improved LLM performance in mathematics and physics domains
- Further investigation needed on how LLMs process affiliation information
- Potential to improve robustness using peer-reviewed publications and sampling from more sources
- Need for better understanding and mitigation of affiliation and categorical biases in LLMs
- Understanding how LLMs process affiliation information when making novelty assessments
Explore related topics
Related papers
- Rayyan—a web and mobile app for systematic reviewsMourad Ouzzani · 2016 · 24,664 citations
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Autonomous chemical research with large language modelsDaniil A. Boiko · 2023 · 809 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations