← Browse all papers
Research theme
Evaluation & Benchmarks
The Evaluation & Benchmarks theme comprises 1,104 papers in this corpus published between 2002 and 2026. Work here is dominated by Benchmarking, Empirical Study, Experimental. 5 open research gaps have been surfaced in this area.
Methodology profile
- Benchmarking533 (48%)
- Empirical Study211 (19%)
- Experimental137 (12%)
- Design Science84 (8%)
- Literature Review47 (4%)
- Conceptual15 (1%)
Research domains
- Data Analysis461 (42%)
- Research Productivity150 (14%)
- Research Integrity103 (9%)
- Scholarly Infrastructure86 (8%)
- Knowledge Synthesis71 (6%)
- AI Governance67 (6%)
Frequent sub-topics
LLM performance in clinical decision support · 2machine learning for mass cytometry data analysis in chronic lymphocytic leukaemia · 1information extraction from epidemiological data using generative AI · 1Explainable machine learning for predictive modeling · 1domain-specific LLM evaluation for supply chain management · 1inter-model agreement in LLM-based cultural annotation of consumer reviews · 1adversarial robustness of multimodal medical RAG systems · 1Topic modeling for health discourse analysis using SC-BERTopic · 1
Open research gaps
- The authors identify that existing benchmarks are confined to single domains and languages, and fail to evaluate generalization capabilities for real-world industrial applications or reflect coding pr
- The abstract does not explicitly identify future research directions or gaps, though it implies questions about why LLMs fail to match theoretical predictions and whether different architectural or tr
- Not explicitly stated in the provided abstract
- The paper does not explicitly identify future research directions in the provided abstract. However, it implicitly suggests the need for further investigation into how safety-tuning mechanisms affect
- The authors identify that "existing evaluation frameworks do not adequately address multi-agent systems that combine simulation, retrieval, and manufacturing preparation," indicating this gap motivate
Representative papers
- Can ChatGPT be used to predict citation counts, readership, and social media interaction? An exploration among 2222 scientific abstractsJoost de Winter · 2024 · 29 citations
- Can small and reasoning large language models score journal articles for research quality and do averaging and few-shot help?Mike Thelwall · 2026 · 3 citations
- Assessing the Performance of 8 AI Chatbots in Bibliographic Reference Retrieval: Grok and DeepSeek Outperform ChatGPT, but None are Entirely AccurateÁlvaro Cabezas-Clavijo · 2026 · 1 citations
- Open-source LLMs for text annotation: a practical guide for model setting and fine-tuningMeysam Alizadeh · 2024 · 56 citations
- PM-LLM-Benchmark: Evaluating Large Language Models on Process Mining TasksAlessandro Berti · 2025 · 23 citations
- Can generative AI effectively perform quality evaluation within social sciences? A case study in library and information scienceYu Zhu · 2026 · 3 citations
- SOMD@NSLP2024: Overview and Insights from the Software Mention Detection Shared TaskFrank Krüger · 2024 · 2 citations
- Fine-tuning SciBERT to enable ASJC-based assessments of the disciplinary orientation of research collectionsMichael Gusenbauer · 2025 · 1 citations