Large language models and responsible research evaluation: an extension of the Leiden Manifesto
Mike Thelwall · Scientometrics · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-026-05552-x
Methodology & findings
Study design
Narrative review and extension of the Leiden Manifesto framework.
Main result
The article argues that "LLM-based quality scores potentially provide a more useful overall level of quantitative evidence, in the sense of correlating more strongly with expert scores in most fields" and that "Large Language Models (LLMs) can be configured to assess/guess research quality scores through their system instructions and so, in theory, they can cover the concept more completely than can citation-based indicators." The authors also find that "LLM prompts can be explicitly tailored to request higher scores for locally relevant research, and there is some small scale evidence that this can work."
Research paradigm
Critical interpretivist
Author conclusions
The authors conclude that "all the main responsible research evaluation considerations for bibliometrics also apply to LLMs but need to be adapted. For example, transparency in the context of LLMs might entail declaring LLM names and versions as well as the prompts and the score processing algorithms." They further assert that "The extended Leiden Manifesto principles are nevertheless intended to guide future research evaluations, and especially considerations about which approach to use when LLMs are a possible choice" and emphasize that "it seems clear that research evaluations ought to be as responsible as possible, in the sense of minimising inaccuracy and bias and maximising transparency as far as is practical in the context of the goals and available resources."
Risk of bias
Unknown LLM biases: author notes that biases in LLM training data and outputs are largely unexplored; Field-level disparities: ChatGPT-based evaluations show that 'some fields get substantially higher average scores than others'; Language bias: articles with longer abstracts receive higher scores; Systemic biases from training data: LLMs may learn and generate biases from inputs; Copyright-based selection bias: only articles with compliant copyright status may be evaluated; Gender and international biases: established for bibliometrics but extent unknown for LLMs; Insufficient empirical validation of LLM biases across fields; Field-dependent variation in average LLM scores (some fields receive substantially higher scores than others); Potential for systematic biases learned from LLM training data; International bias risks (some countries receive higher scores); Bias risk in article-level characteristics (longer abstracts receive higher scores); Ephemerality of LLM scores across different model versions; Potential for gaming behavior if authors learn LLM scoring preferences; Gender bias in LLM evaluations (unknown extent); International/geographical bias in LLM scores; Field-specific disparities in average LLM scores; Articles with longer abstracts receiving higher LLM scores; Potential publisher/editor gaming of LLM-friendly formatting; Unknown biases learned from LLM training data; Potential bias against locally relevant research
Limitations
- The authors acknowledge that "little is known about [LLM] biases" in contrast to bibliometrics where "some bibliometrics have been shown to have gender biases (e.g., career citations favour males), most have international biases, and there may also be institutional, reputational and interdisciplinary disparities." Additionally, they note that "it is not clear whether these disparities are biases or reflect underlying quality differences." They also state that "LLMs fail to be simple because of their high complexity
- This complexity (e.g., at least 7 billion parameters for most small current LLMs) makes any model parameter information useless for anything except replicability." Furthermore, "it is not clear yet that these would be helpful" regarding LLM-generated reasoned explanations because "LLMs seem to be good at creating plausible explanations even for poor decisions."
Open questions raised
- Limited understanding of LLM biases: 'little is known about their biases' and systematic testing needed
- Unknown disparities in LLM scores: 'This issue is not well understood yet' regarding field variation
- Lack of benchmark datasets: 'it would be very useful (albeit challenging) to create benchmark datasets of articles with expert evaluation scores that LLMs could be tested against'
- Unclear effectiveness of LLM explanations: whether LLM-generated reasoned explanations improve transparency
- Scientific impact assessment: 'whether LLMs can guess scientific impact better than citation-based indicators is likely to remain unknown for a long time'
- LLM expertise gaps: 'currently less LLM expertise for research evaluation than bibliometric expertise'
Explore related topics
Related papers
- Guidance for conducting systematic scoping reviewsMicah D.J. Peters · 2015 · 7,472 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- Artificial intelligence in higher education: the state of the fieldHelen Crompton · 2023 · 1,378 citations
- A SWOT analysis of ChatGPT: Implications for educational practice and researchMohammadreza Farrokhnia · 2023 · 1,171 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- Ethics of AI in Education: Towards a Community-Wide FrameworkW. Holmes · 2021 · 1,056 citations