← Browse all papers
Research theme
Evaluation & Benchmarks
The Evaluation & Benchmarks theme comprises 1,115 papers in this corpus published between 2002 and 2026. Work here is dominated by Benchmarking, Empirical Study, Experimental. 5 open research gaps have been surfaced in this area.
Methodology profile
- Benchmarking534 (48%)
- Empirical Study212 (19%)
- Experimental139 (12%)
- Design Science84 (8%)
- Literature Review51 (5%)
- Conceptual17 (2%)
Research domains
- Data Analysis467 (42%)
- Research Productivity150 (13%)
- Research Integrity104 (9%)
- Scholarly Infrastructure88 (8%)
- Knowledge Synthesis71 (6%)
- AI Governance67 (6%)
Frequent sub-topics
LLM performance in clinical decision support · 2machine learning for mass cytometry data analysis in chronic lymphocytic leukaemia · 1information extraction from epidemiological data using generative AI · 1Explainable machine learning for predictive modeling · 1domain-specific LLM evaluation for supply chain management · 1inter-model agreement in LLM-based cultural annotation of consumer reviews · 1machine learning for food security prediction · 1long document question answering with transformer models and contrastive learning · 1
Open research gaps
- The authors identify that existing benchmarks lack comprehensive coverage and fine-grained metadata for evaluating foundation models, and that current benchmarks may miss performance differences acros
- The abstract does not explicitly identify future research directions or gaps, though it implies questions about why LLMs fail to match theoretical predictions and whether different architectural or tr
- The authors identify that "existing evaluation frameworks do not adequately address multi-agent systems that combine simulation, retrieval, and manufacturing preparation," indicating this gap motivate
- The paper does not explicitly identify future research directions in the provided abstract. The focus is on synthesizing existing pathologies and identifying the shared design principle of measurement
- The authors identify the gap that existing benchmarks lack "a competency-based structure aligned with the real-world clinical responsibilities encountered in general practice", and note that "further
Representative papers
- Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstractsJoost de Winter · 2024 · 29 citations
- Can small and reasoning large language models score journal articles for research quality and do averaging and few-shot help?Mike Thelwall · 2026 · 3 citations
- Assessing the Performance of 8 AI Chatbots in Bibliographic Reference Retrieval: Grok and DeepSeek Outperform ChatGPT, but None are Entirely AccurateÁlvaro Cabezas-Clavijo · 2026 · 1 citations
- Open-source LLMs for text annotation: a practical guide for model setting and fine-tuningMeysam Alizadeh · 2024 · 56 citations
- PM-LLM-Benchmark: Evaluating Large Language Models on Process Mining TasksAlessandro Berti · 2025 · 23 citations
- Can generative AI effectively perform quality evaluation within social sciences? A case study in library and information scienceYu Zhu · 2026 · 3 citations
- SOMD@NSLP2024: Overview and Insights from the Software Mention Detection Shared TaskFrank Krüger · 2024 · 2 citations
- Fine-tuning SciBERT to enable ASJC-based assessments of the disciplinary orientation of research collectionsMichael Gusenbauer · 2025 · 1 citations