12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Evaluating and Enhancing Large Language Models for Novelty Assessment in Scholarly Publications

Zhiyuan Peng, Yi Fang · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
3
Citations
1.71
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18653/v1/2025.aisd-main.5

Methodology & findings

Study design

Empirical benchmark evaluation using a novel dataset (SchNovel) of 15,000 paper pairs from arXiv across six research fields.

Sample

N = 15000, 13 groups

Primary method

McNemar test for comparing paired categorical outcomes (novelty assessment accuracy). Pairwise comparisons used throughout experiments with position bias controls (order swapping). Majority voting employed for Self-Consistency method (10 generated sequences). Accuracy metric reported across different categorical variables (field, start year, year gap).

Main result

RAG-Novelty achieved the highest results overall, significantly outperforming the second-best method, except in the mathematics field. Specifically, "RAG-Novelty achieves the highest results overall, significantly outperforming the second-best method, except in the mathematics field." The study found that "pairwise is consistently much better than pointwise across different year gaps," and that "GPT-4o-mini outperformed all other models, demonstrating a substantial advantage over smaller models."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist (computational evaluation)

Author conclusions

The authors conclude that "RAG-Novelty achieves the highest results overall, significantly outperforming the second-best method, except in the mathematics field" and that "Our extensive experiments demonstrate that RAG-Novelty outperforms recent baseline models in assessing novelty in scholarly papers." They also note important findings regarding position bias: "Notably, GPT-4o-mini outperformed all other models, demonstrating a substantial advantage over smaller models like LLaMA 3.1-8b, Mistral 7b, and Gemma 2-9b. Despite such success, even ChatGPT 4o-mini and ChatGPT 3.5 exhibit position bias, where the order of papers in the prompt affects their decision-making instead of content alone."

Risk of bias

Position bias in LLM decision-making (order of papers affects evaluation); Field/category bias (lower performance in mathematics and physics); Affiliation bias (preference for top research universities); Lack of domain knowledge in training data for specialized fields; Temporal bias potential in retrieval-augmented approach; Position bias in LLM outputs (order of papers in prompt affects decision-making); Affiliation bias (models show preference for papers from top research universities); Categorical bias (disparities in performance across research fields); Training data bias (LLMs may lack domain-specific knowledge in certain fields); Position bias in LLMs: smaller models exhibit substantial position bias (Mistral 7b favors last paper; LLaMA 3.1-8b favors middle position); Affiliation bias: LLMs showed preference for papers from top research universities over teaching universities, particularly with larger year gaps; Domain bias: Lower accuracy in mathematics and physics possibly due to lack of specialized content in LLM training data; Data source bias: Exclusive reliance on arXiv may not be representative of all scholarly publications; Temporal bias: Model evaluation uses 2024 as reference point but is evaluated on papers up to 2023; potential knowledge cutoff effects; Selection bias in sample: Only six of eight arXiv fields included; authors cite insufficient data in other fields

Limitations

  • The authors state: "Our study evaluates an LLM's ability to assess novelty using a research paper's title, abstract, and metadata
  • While the abstract provides a strong indication of a paper's content and key findings, it may not fully capture the novelty of the research compared to the complete text
  • Abstracts often summarize the main ideas but may omit important technical details." Additionally, "the exclusive use of arXiv data is limiting
  • We selected arXiv as an initial step for its broad, publicly accessible range of publications
  • Future work can improve robustness using peer-reviewed publications and sampling papers from more sources."

Open questions raised

  • Limited evaluation on peer-reviewed publications beyond arXiv
  • Need for improved LLM performance in mathematics and physics domains
  • Further investigation needed on how LLMs process affiliation information
  • Potential to improve robustness using peer-reviewed publications and sampling from more sources
  • Need for better understanding and mitigation of affiliation and categorical biases in LLMs
  • Understanding how LLMs process affiliation information when making novelty assessments
Data: SchNovel dataset; arXiv datasetCode: RAG-Novelty codeExtracted from: pdfAgreement 62%

Explore related topics

Related papers