Useful for Exploration, Risky for Precision: Evaluating AI Tools in Academic Research
Anthea Dathe, Kiran Hoffmann, Aline Mangold · arXiv (Cornell University) · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
Methodology & findings
Study design
Systematic benchmarking study combining computer-centered and human-centered evaluation metrics.
Primary method
Benchmarking study with human-centered evaluation framework; comparative evaluation of existing commercial tools rather than design/development of new artifact
Main result
The study found that "GenAI tools can improve efficiency in early stages of the research workflow, their outputs still require careful human verification due to limitations in explainability, reproducibility, and source transparency." Specifically, regarding Q&A tools, "while GenAI tools produced reasonably accurate outputs, their xAI accuracy was low. This means, that the highlighted passages in the source document were often irrelevant, or relevant passages were missed entirely." For literature review tools, "the amount of sources that were usable (peer-reviewed, scientific source and matching user intent) was significantly lower" than the total sources retrieved, and "Reproducibility across repeated runs with identical prompts was low."
Research paradigm
Pragmatist/Mixed-methods (combining computational and human-centered evaluation)
Author conclusions
The authors conclude: "In this study, we developed and applied a benchmarking framework combining human-centered and computer-centered metrics to evaluate AI-based Q&A and literature review tools for research use. The results show that Q&A tools can provide useful overviews and generally accurate summaries, but remain unreliable for precise information extraction due to hallucinations and weak explainability. Literature review tools were helpful for exploratory searches, yet their low reproducibility, limited transparency, and inconsistent source quality make them unsuitable for systematic reviews. Our main contribution lies in the human-centered focus of of the evaluation, which assesses whether AI tool outputs are actually usable for researchers in practice rather than evaluating technical performance alone."
Risk of bias
Single rater per tool evaluation - high risk of subjective bias on human-centered metrics; Limited sample size (5 Q&A tools, 5 literature review tools); Domain expertise concentrated in one evaluator's background (psychology) may limit generalizability; Ground-truth definition allowed room for interpretation; Tools rapidly evolving, results may be outdated; Short, non-iterative prompts in literature review evaluation do not reflect real-world usage patterns; Single rater per tool (selection and subjective bias); Rater familiarity with content domain (potential for lenient evaluation); Limited sample of research papers (domain-specific psychology papers may not generalize); Short, non-iterative prompts in literature review task (does not reflect real-world refinement practices); Ground truth definitions subject to interpretation (Likert scale rather than binary); Single rater per tool increases subjective bias for human-centered metrics; Selection bias in tool choice based on Google search rankings rather than usage frequency; Evaluator familiarity with domain (BPD research) may influence judgment on intent match; Limited prompt refinement in literature review benchmarking may not reflect real-world usage patterns; Subjective interpretation of 'ground truth' for consistency evaluation
Limitations
- The authors acknowledge that "due to constraints in resources, each tool was only evaluated by one rater
- This is a significant source of bias, especially for subjective metrics, such as usability measures." Additionally, "the AI-tool landscape is rapidly evolving
- While our results might have been up to date during assessment, tool performance might have enhanced in the meantime." They further note that "the 'ground truth' to measure Consistency external was complex and multifaceted
- This means, that the determination if the output was true or false incorporated some room for interpretation." Finally, "the prompt in the literature review benchmarking was short and due to comparability reasons we did not allow for refinements in further prompts
- In the real world however, users would probably specify their prompt, if they did not receive the desired output."
Open questions raised
- Need for systematic benchmarks from human-centered perspective
- Evaluation dimensions (explainability, transparency, reproducibility) frequently discussed but rarely operationalized as explicit benchmark metrics
- Integrated evaluation approaches combining technical performance with human-centered criteria across realistic research tasks are lacking
- Future research should apply benchmarking approach with multiple raters
- Need for reassessment as LLMs and tools constantly evolve
- Literature review tools should be evaluated on additional quality criteria beyond intent match and peer-review (predatory journals, journal metrics)
Explore related topics
Related papers
- Estimating the reproducibility of psychological scienceAlexander A. Aarts · 2015 · 8,669 citations
- Which academic search systems are suitable for systematic reviews or meta‐analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resourcesMichael Gusenbauer · 2019 · 2,116 citations
- Academic Integrity considerations of AI Large Language Models in the post-pandemic era: ChatGPT and beyondMike Perkins · 2023 · 668 citations
- Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewersCatherine A. Gao · 2023 · 657 citations
- ChatGPT and the rise of generative AI: Threat to academic integrity?Damian Eke · 2023 · 476 citations
- Nonhuman “Authors” and Implications for the Integrity of Scientific Publication and Medical KnowledgeAnnette Flanagin · 2023 · 399 citations