12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Research quality evaluation by AI in the era of large language models: advantages, disadvantages, and systemic effects – An opinion paper

Mike Thelwall · Scientometrics · 2025

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
3/4
Quality (LMQS)
I
Evidence
15
Citations
13.88
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-025-05361-8

Methodology & findings

Study design

Narrative review synthesizing empirical evidence from multiple published studies on LLM-based research quality evaluation.

Primary method

The paper reviews empirical studies that employed correlation analysis (Pearson's r), comparing ChatGPT scores with expert quality scores, departmental averages, and peer review decisions. Methods include averaging scores across multiple iterations (15-30 iterations reported in cited studies) to improve reliability.

Main result

The paper finds that "LLM-based quality evaluations seem to have at least three clear advantages over bibliometrics for research quality indicators" including greater accuracy, coverage of science across years and fields, and assessment of more quality dimensions. ChatGPT 4o scores "seem to be more accurate than bibliometrics in the sense of higher correlations with human scores in most fields" (Thelwall & Yaghi, 2024b), with "correlations were positive in all UoAs except Clinical Medicine (-0.15) and were weak (0.05 to 0.3) mainly in the arts and humanities, moderate (0.3 to 0.5) mainly in the social sciences and strong (0.5 to 0.8) mainly in the health and physical sciences and engineering."

Research paradigm

Critical realism / epistemological critique

Author conclusions

The authors conclude that "LLM scores have the technical potential to complement or surpass bibliometrics for research quality indicators, but there are too many unknowns yet to use them in any important context. If more information can be gained about their biases, limitations, and potential for gaming then it might be possible to start using them in minor roles to support peer review. In the longer term, successful applications, and a lack of issues from gaming might allow them to be used in less minor roles, taking over from bibliometrics. For example, they might be offered instead of bibliometrics as supporting information for expert reviewers in future versions of the UK REF national research evaluation exercise."

Risk of bias

Unknown biases in LLM training data and algorithms; Potential gender prejudice in LLM models; Favoritism towards work of successful authors or prestigious institutions; Methodological hierarchy biases in LLM training; Document age bias (more recent articles receive higher scores); Field-specific differences in average ChatGPT scores; Abstract length bias (longer abstracts receive higher scores); Potential gaming through author abstract overclaiming; Potential gender prejudice or favoritism towards successful authors; Hierarchy of methodological approaches influencing judgments; Bias towards recent articles and longer abstracts; Field-dependent performance variations; Potential for LLM gaming through inflated abstract claims; Unknown AI biases from training data (gender, author reputation, institutional prestige); Potential document age bias (more recent articles receiving higher scores); Field-specific biases in average ChatGPT scores; Abstract length bias (longer abstracts receiving higher scores); Potential for gaming through overclaimed author abstracts; Limited transparency preventing bias detection; Opacity of training data for commercial LLM variants

Limitations

  • The author notes "research into their value for research quality assessment is limited now (February, 2025) and it would be helpful to have a range of different studies and approaches to confirm or contradict those published so far." Additionally, "LLMs currently work most effectively with article titles and abstracts and therefore are clearly guessing at research quality rather than measuring it." The author also states that "the one large-scale analysis of ChatGPT quality scores so far..
  • relied on public data in terms of the departmental average REF scores
  • Thus, whilst they are consistent with Chat-GPT having a near-universal ability to detect research quality, they do not prove it, because ChatGPT might have leveraged public information about departmental REF quality profiles when scoring individual articles."

Open questions raised

  • Limited research into LLM biases in research quality assessment
  • Need for studies across different use contexts and disciplines
  • Lack of experience about how to use LLMs in practical research evaluations
  • Unknown potential for gaming and manipulation of LLM scores
  • Need for research into conditions under which LLMs make mistakes
  • Limited understanding of LLM effectiveness across all fields (particularly weak in clinical medicine)
Code: Huggingface.co (open source LLM models referenced but not created by authors)Extracted from: pdfAgreement 76%

Explore related topics

Related papers