12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Can small and reasoning large language models score journal articles for research quality and do averaging and few-shot help?

Mike Thelwall, Ehsan Mohammadi · Scientometrics · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
2/4
Quality (LMQS)
E
Evidence
3
Citations
21.33
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-026-05585-2

Methodology & findings

Study design

Empirical comparative study using multiple large language models to score journal articles (n=2,780 unique articles from 6 health and life sciences fields).

Sample

N = 2780, 6 groups

Primary method

Spearman rank correlation coefficient (primary metric for assessing alignment between LLM and expert scores); Bootstrapped 95% confidence intervals for correlation estimates; Binomial probability tests for comparing strategies across UoAs (e.g., p=0.09 for 5 of 6 successes); Violin plots for visualizing distribution of predicted scores by actual article score; Word Association Thematic Analysis (WATA) with chi-square tests and Benjamini-Hochberg correction for identifying differences in LLM reports; Evolutionary optimization using differential evolution with tenfold cross-validation for weighted sum models

Main result

The study found that "All medium sized LLMs give similar correlations to the cloud-based LLMs and, based on the confidence interval estimates, there is insufficient evidence to conclude that cloud-based LLMs are superior for this task overall" and that "In all 16 contexts tested, averaging five LLM scores gives a higher Spearman correlation than taking a single score."

Reports effect sizes and confidence intervals.

Research paradigm

Positivist/empiricist - quantitative correlation-based evaluation

Author conclusions

"The results show that a range of different open weights LLMs have an ability to score journal articles for research quality, including reasoning models, and that this ability may be present even for relatively small 4 billion parameter models in most UoAs, with 12 billion parameters being a safer minimum that worked for all six UoAs here." Furthermore, "smaller models seem to be nearly as effective as larger ones, making their use practical in situations where high performance computing is not available, and data security is essential. Reasoning is not recommended because it substantially increases computing time/cost without improving the results."

Risk of bias

Proxy-based gold standard (departmental mean scores) may not reflect true individual article quality; Single-rater bias in second gold standard (first author assigned all individual scores); Departmental-level biases in LLM scores (e.g., favoring certain writing styles or methodological orientations); Temporal bias: LLMs trained after 2021 evaluating articles published before 2021, affecting novelty assessments; Selection bias: Articles without DOIs or with very short abstracts were excluded; Gemma 3 results obtained from API calls to Google (not locally run); local validation showed higher correlations with overlapping confidence intervals; Departmental-level bias in LLM scores (e.g., favoring research from certain departments based on writing style); Single-rater bias for the individual article-level gold standard (assigned by first author only); Temporal bias: models trained on 2024-2025 data evaluating articles from 2014-2020; Potential data leakage from cloud-based LLMs using submitted queries; UK-focused dataset may not generalize internationally; First author not expert in all six Units of Assessment (only published in four of six); Gold standard bias: departmental average proxy dampens correlations and may not reflect true article-level quality; Rater bias: first author assigned individual quality scores without blinding to LLM scores; Single rater bias: individual article scores from only one author who is not an expert in all six assessed UoAs; Temporal bias: LLMs evaluated articles published before their training data cutoff, affecting novelty assessment; Selection bias: articles without DOIs or very short abstracts were excluded; Model availability bias: pragmatic selection of models based on computational feasibility, excluding larger open weights models

Limitations

  • "The dataset is UK-focused and both gold standards are imperfect: one is indirect, and the other is subjective to a single individual." Furthermore, "the confidence intervals are relatively wide compared to the differences found, despite the sample sizes of 500, so the differences may be artefacts of the datasets used." The authors also note that "LLMs compared became available years after the REF2021 articles were published and this may influence their judgements of novelty, because a novel contribution in 2014 would not be novel in 2023." Additionally, "only a single performance metric has been reported, Spearman correlation, whereas LLM studies usually include a range to give wider performance information."

Open questions raised

  • Unknown whether reasoning models provide advantage over non-reasoning models for this task
  • Unclear relationship between LLM size and research quality scoring ability
  • Unknown effectiveness of few-shot prompting for this specific task
  • Lack of large-scale multidisciplinary datasets with post-publication quality scores from multiple experts to establish baseline expert agreement rates
  • Need for investigation of alternative few-shot strategies (different number of articles, different presentation methods)
  • Unexplored effectiveness of fine-tuning, agent-based designs, and parameter variations
Data: 2,780 unique journal articles (500 articles randomly selected from each of 6 UK REF2021 Units of Assessment in health and life sciences). Supplementary materials available at https://doi.org/10.6084/m9.figshare.30382651; Supplementary materials available at https://doi.org/10.6084/m9.figshare.30382651 (includes prompts and violin plots); REF2021 submission aggregate scores (publicly released); 2,780 unique health and life sciences journal articles from six UoAs; 2,780 unique journal articles from UK REF2021 Main Panel A (six UoAs with 500 articles each); Article scores and supplementary materials available via Figshare: https://doi.org/10.6084/m9.figshare.30382651Extracted from: pdfAgreement 55%

Explore related topics

Related papers