Research quality evaluation by AI in the era of large language models: advantages, disadvantages, and systemic effects – An opinion paper
Mike Thelwall · Scientometrics · 2025
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-025-05361-8
Methodology & findings
Study design
Narrative review synthesizing empirical evidence from multiple published studies on LLM-based research quality evaluation.
Primary method
The paper reviews empirical studies that employed correlation analysis (Pearson's r), comparing ChatGPT scores with expert quality scores, departmental averages, and peer review decisions. Methods include averaging scores across multiple iterations (15-30 iterations reported in cited studies) to improve reliability.
Main result
The paper finds that "LLM-based quality evaluations seem to have at least three clear advantages over bibliometrics for research quality indicators" including greater accuracy, coverage of science across years and fields, and assessment of more quality dimensions. ChatGPT 4o scores "seem to be more accurate than bibliometrics in the sense of higher correlations with human scores in most fields" (Thelwall & Yaghi, 2024b), with "correlations were positive in all UoAs except Clinical Medicine (-0.15) and were weak (0.05 to 0.3) mainly in the arts and humanities, moderate (0.3 to 0.5) mainly in the social sciences and strong (0.5 to 0.8) mainly in the health and physical sciences and engineering."
Research paradigm
Critical realism / epistemological critique
Author conclusions
The authors conclude that "LLM scores have the technical potential to complement or surpass bibliometrics for research quality indicators, but there are too many unknowns yet to use them in any important context. If more information can be gained about their biases, limitations, and potential for gaming then it might be possible to start using them in minor roles to support peer review. In the longer term, successful applications, and a lack of issues from gaming might allow them to be used in less minor roles, taking over from bibliometrics. For example, they might be offered instead of bibliometrics as supporting information for expert reviewers in future versions of the UK REF national research evaluation exercise."
Risk of bias
Unknown biases in LLM training data and algorithms; Potential gender prejudice in LLM models; Favoritism towards work of successful authors or prestigious institutions; Methodological hierarchy biases in LLM training; Document age bias (more recent articles receive higher scores); Field-specific differences in average ChatGPT scores; Abstract length bias (longer abstracts receive higher scores); Potential gaming through author abstract overclaiming; Potential gender prejudice or favoritism towards successful authors; Hierarchy of methodological approaches influencing judgments; Bias towards recent articles and longer abstracts; Field-dependent performance variations; Potential for LLM gaming through inflated abstract claims; Unknown AI biases from training data (gender, author reputation, institutional prestige); Potential document age bias (more recent articles receiving higher scores); Field-specific biases in average ChatGPT scores; Abstract length bias (longer abstracts receiving higher scores); Potential for gaming through overclaimed author abstracts; Limited transparency preventing bias detection; Opacity of training data for commercial LLM variants
Limitations
- The author notes "research into their value for research quality assessment is limited now (February, 2025) and it would be helpful to have a range of different studies and approaches to confirm or contradict those published so far." Additionally, "LLMs currently work most effectively with article titles and abstracts and therefore are clearly guessing at research quality rather than measuring it." The author also states that "the one large-scale analysis of ChatGPT quality scores so far..
- relied on public data in terms of the departmental average REF scores
- Thus, whilst they are consistent with Chat-GPT having a near-universal ability to detect research quality, they do not prove it, because ChatGPT might have leveraged public information about departmental REF quality profiles when scoring individual articles."
Open questions raised
- Limited research into LLM biases in research quality assessment
- Need for studies across different use contexts and disciplines
- Lack of experience about how to use LLMs in practical research evaluations
- Unknown potential for gaming and manipulation of LLM scores
- Need for research into conditions under which LLMs make mistakes
- Limited understanding of LLM effectiveness across all fields (particularly weak in clinical medicine)
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT in education: Strategies for responsible implementationMohanad Halaweh · 2023 · 576 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations