Evaluating large language models for abstract evaluation tasks: an empirical study
Yinuo Liu, Emre Sezgin, Eric A. Youngstrom · Frontiers in Research Metrics and Analytics · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.3389/frma.2026.1807672
Methodology & findings
Study design
Comparative empirical study using 160 research abstracts from a local research retreat.
Sample
N = 160, 4 groups
Primary method
Intraclass Correlation Coefficient (ICC) using two approaches: (1) two-way random effects model with absolute agreement definition and rater average unit [ICC (2, k)] for LLM-to-LLM agreement; (2) one-way random effects model [ICC (1, k)] for human-to-LLM agreement. Sensitivity power analyses with fixed sample constraints (n=160, k=3, α=0.05, 1-β=0.80) calculated minimum detectable ICC of 0.119. Visual agreement patterns assessed using Bland-Altman plots with mean difference and 95% limits of agreement (±1.96 SD). Spearman's rank correlation (ρ) and paired t-tests used to assess score-dependent agreement. ANOVA used to test for primacy/recency effects. Boxplots visualized median, interquartile range, and outliers.
Main result
The study found that "ChatGPT-5, Gemini-3-Pro, and Claude-Sonnet-4.5 demonstrated scoring efficiency and consistency, with ChatGPT and Claude achieving moderate agreement with human reviewers on overall quality and objective dimensions, but all models showed limited reliability in assessing subjective criteria." Specifically, inter-rater reliability analysis revealed that the three LLMs achieved "good-to-excellent agreement" with each other (ICC ranging from 0.59-0.87 across criteria), but showed only "moderate agreement with human experts on abstract overall quality and content-specific criteria" with composite ICC scores ranging from 0.38-0.55, while reliability deteriorated markedly on subjective dimensions such as impact, engagement, and applicability (ICC 0.04-0.38).
Reports effect sizes and confidence intervals.
Research paradigm
positivist/empiricist
Author conclusions
The authors conclude that "LLMs could serve as complementary tools to augment, rather than replace, human expertise in research evaluation." They further state: "ChatGPT-5, Gemini-3-Pro, and Claude-Sonnet-4.5 demonstrated scoring efficiency and consistency, with ChatGPT and Claude achieving moderate agreement with human reviewers on overall quality and objective dimensions, but all models showed limited reliability in assessing subjective criteria. Based on these findings, LLMs could serve as complementary tools to augment, rather than replace, human expertise in research evaluation." The authors recommend that "LLMs function best as complementary tools in research evaluation, while human expertise remains essential for assessing the subjective dimensions of research quality."
Risk of bias
Selection bias: Sample limited to pediatric health research abstracts from single institution (Abigail Wexner Research Institute); Expertise mismatch: Human reviewers not always content experts in their randomly assigned abstract domains; Measurement interdependence: Batch processing of 10 abstracts per API call violates independence assumption for ICC; Rater heterogeneity: Variable workloads across 14 human reviewers (1-28 abstracts each); Rubric generality: Single generic rubric used across diverse research topics; Potential order effects: Abstract position within batch could influence scoring despite stated non-significant findings; Selection bias: abstracts predominantly from pediatric health research, not representative of national/international conferences; Expertise-content mismatch: human reviewers were randomly assigned abstracts outside their areas of expertise; Rater heterogeneity: 14 human reviewers with varying numbers of abstracts evaluated (1-28 each), leading to differential rating experience; Score interdependence: batch processing of 10 abstracts per API call violates independence assumption; human reviewers evaluated subsets rather than full pool; Systematic bias: Gemini showed significant positive bias with mean difference of 0.42; ChatGPT showed positive bias of 0.24; Potential criterion contamination: LLMs may apply trained patterns differently from humans; Selection bias: Abstracts from single institution (Nationwide Children's Hospital Research Retreat); Expertise mismatch: Human reviewers not always matched to content area; Reviewer heterogeneity: 14 human reviewers with varying number of assignments (1-28 abstracts each); Score interdependence: Batch processing of 10 abstracts per prompt may introduce comparative effects between abstracts; Model selection bias: Only three contemporary LLMs tested; results may not generalize to other models
Limitations
- The authors state: "Limitations include that sample abstracts predominantly focused on pediatric health research, which does not represent the diversity of topics encountered at national or international conferences
- The feasibility of LLM-assisted abstract review in larger-scale, more diverse conference settings, or their performance with humanities and physical science content, requires further investigation." Additionally, "the human reviewers' expertise did not always align closely with the content of their randomly assigned abstracts, and all reviewers used a single 'one size fits all' rubric." Finally, "the batch strategy of sending 10 abstracts per prompt and the assignment of human reviewers to abstract subsets introduced potential score interdependence, as scores may be influenced by the other abstracts evaluated together
- This violates the strict independence assumption underlying ICC calculations."
Open questions raised
- Feasibility of LLM-assisted abstract review in larger-scale, more diverse conference settings
- LLM performance with humanities and physical science content
- Prompt variations to identify most effective prompts for abstract evaluation tasks
- Score-dependent agreement patterns requiring larger samples and clearer quality stratification across low-, medium-, and high-quality abstracts
- Mechanisms underlying observed misalignment between Gemini and human reviewers on subjective criteria
- Outcome-oriented AI approaches using substantially more detailed rubrics (30-40 items) emphasizing objective and quantitative criteria
Explore related topics
Related papers
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- Fabrication and errors in the bibliographic citations generated by ChatGPTWilliam H. Walters · 2023 · 352 citations
- Generative AI tools and assessment: Guidelines of the world's top-ranking universitiesBenjamin Luke Moorhouse · 2023 · 343 citations
- Human-AI collaboration patterns in AI-assisted academic writingAndy Nguyen · 2024 · 301 citations
- Hallucination Rates and Reference Accuracy of ChatGPT and Bard for Systematic Reviews: Comparative AnalysisMikaël Chelli · 2024 · 295 citations
- AI literacy and its implications for prompt engineering strategiesNils Knoth · 2024 · 277 citations