12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Large language models and responsible research evaluation: an extension of the Leiden Manifesto

Mike Thelwall · Scientometrics · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
I
Evidence
2
Citations
23.31
FWCI
Top 10%
Impact

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1007/s11192-026-05552-x

Methodology & findings

Study design

Narrative review and extension of the Leiden Manifesto framework.

Main result

The article argues that "LLM-based quality scores potentially provide a more useful overall level of quantitative evidence, in the sense of correlating more strongly with expert scores in most fields" and that "Large Language Models (LLMs) can be configured to assess/guess research quality scores through their system instructions and so, in theory, they can cover the concept more completely than can citation-based indicators." The authors also find that "LLM prompts can be explicitly tailored to request higher scores for locally relevant research, and there is some small scale evidence that this can work."

Research paradigm

Critical interpretivist

Author conclusions

The authors conclude that "all the main responsible research evaluation considerations for bibliometrics also apply to LLMs but need to be adapted. For example, transparency in the context of LLMs might entail declaring LLM names and versions as well as the prompts and the score processing algorithms." They further assert that "The extended Leiden Manifesto principles are nevertheless intended to guide future research evaluations, and especially considerations about which approach to use when LLMs are a possible choice" and emphasize that "it seems clear that research evaluations ought to be as responsible as possible, in the sense of minimising inaccuracy and bias and maximising transparency as far as is practical in the context of the goals and available resources."

Risk of bias

Unknown LLM biases: author notes that biases in LLM training data and outputs are largely unexplored; Field-level disparities: ChatGPT-based evaluations show that 'some fields get substantially higher average scores than others'; Language bias: articles with longer abstracts receive higher scores; Systemic biases from training data: LLMs may learn and generate biases from inputs; Copyright-based selection bias: only articles with compliant copyright status may be evaluated; Gender and international biases: established for bibliometrics but extent unknown for LLMs; Insufficient empirical validation of LLM biases across fields; Field-dependent variation in average LLM scores (some fields receive substantially higher scores than others); Potential for systematic biases learned from LLM training data; International bias risks (some countries receive higher scores); Bias risk in article-level characteristics (longer abstracts receive higher scores); Ephemerality of LLM scores across different model versions; Potential for gaming behavior if authors learn LLM scoring preferences; Gender bias in LLM evaluations (unknown extent); International/geographical bias in LLM scores; Field-specific disparities in average LLM scores; Articles with longer abstracts receiving higher LLM scores; Potential publisher/editor gaming of LLM-friendly formatting; Unknown biases learned from LLM training data; Potential bias against locally relevant research

Limitations

  • The authors acknowledge that "little is known about [LLM] biases" in contrast to bibliometrics where "some bibliometrics have been shown to have gender biases (e.g., career citations favour males), most have international biases, and there may also be institutional, reputational and interdisciplinary disparities." Additionally, they note that "it is not clear whether these disparities are biases or reflect underlying quality differences." They also state that "LLMs fail to be simple because of their high complexity
  • This complexity (e.g., at least 7 billion parameters for most small current LLMs) makes any model parameter information useless for anything except replicability." Furthermore, "it is not clear yet that these would be helpful" regarding LLM-generated reasoned explanations because "LLMs seem to be good at creating plausible explanations even for poor decisions."

Open questions raised

  • Limited understanding of LLM biases: 'little is known about their biases' and systematic testing needed
  • Unknown disparities in LLM scores: 'This issue is not well understood yet' regarding field variation
  • Lack of benchmark datasets: 'it would be very useful (albeit challenging) to create benchmark datasets of articles with expert evaluation scores that LLMs could be tested against'
  • Unclear effectiveness of LLM explanations: whether LLM-generated reasoned explanations improve transparency
  • Scientific impact assessment: 'whether LLMs can guess scientific impact better than citation-based indicators is likely to remain unknown for a long time'
  • LLM expertise gaps: 'currently less LLM expertise for research evaluation than bibliometric expertise'
Extracted from: pdfAgreement 77%

Explore related topics

Related papers