← Browse all papers
Research method
Benchmarking
The Benchmarking method comprises 1,056 papers in this corpus published between 2000 and 2026.
Themes covered
- Evaluation & Benchmarks533 (50%)
- Hallucination Control149 (14%)
- RAG for Research105 (10%)
- AI-Assisted Discovery76 (7%)
- Responsible AI Use40 (4%)
- Citation Integrity27 (3%)
Research domains
- Research Integrity236 (22%)
- Data Analysis227 (21%)
- Knowledge Synthesis134 (13%)
- Literature Discovery114 (11%)
- Research Productivity83 (8%)
- Scholarly Infrastructure60 (6%)
Frequent sub-topics
information extraction from epidemiological data using generative AI · 1abductive reasoning evaluation in LLMs · 1LLM safety evaluation for moral rationalization facilitation · 1domain-specific LLM evaluation for supply chain management · 1inter-model agreement in LLM-based cultural annotation of consumer reviews · 1adversarial robustness of multimodal medical RAG systems · 1knowledge graph-enhanced retrieval-augmented generation for biomedical QA · 1Topic modeling for health discourse analysis using SC-BERTopic · 1
Representative papers
- Synthesizing scientific literature with retrieval-augmented language modelsAkari Asai · 2026 · 14 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Artificial intelligence adoption in the physical sciences, natural sciences, life sciences, social sciences and the arts and humanities: A bibliometric analysis of research publications from 1960-2021Stefan Hajkowicz · 2023 · 119 citations
- An Evidence-Grounded Research Assistant for Functional Genomics and Drug Target AssessmentKsenia Sokolova · 2025 · 2 citations
- SurveyGen: Quality-Aware Scientific Survey Generation with Large Language ModelsTong Bao · 2025 · 2 citations
- LLM4SCREENLIT: Recommendations on assessing the performance of large language models for screening literature in systematic reviewsLech Madeyski · 2026 · 1 citations
- Validating Large Language Models for Title-Abstract Screening in Low-Prevalence Systematic Reviews: An Environmental Science Case StudyMaximilian Nawrath · 2026 · 1 citations
- Assessing the Performance of 8 AI Chatbots in Bibliographic Reference Retrieval: Grok and DeepSeek Outperform ChatGPT, but None are Entirely AccurateÁlvaro Cabezas-Clavijo · 2026 · 1 citations