Evaluation of output of AI models ScholarAI and SciSpace: Implications for use of generative artificial intelligence models as research assistants
Meenakshi Aggarwal, Sonia S. Kharay, Priya Bansal, Seema Gupta, Hitant Vohra, Sarit Sharma et al. · The Journal of Community Health Management · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18231/j.jchm.17551.1781068642
Methodology & findings
Study design
Experimental comparative performance study evaluating two custom GPT models (ScholarAI and SciSpace) across four temporal sessions (September 5, November 13-15, 2025).
Sample
N = 4, 2 groups
Primary method
Descriptive statistics including absolute frequencies and percentages to summarize distribution of valid articles, invalid DOIs, and hallucinated references across study sessions. Inter-rater agreement monitoring across evaluation axes. No inferential statistics reported.
Main result
The study found that both AI models demonstrated significant limitations in research assistance. "The results of our study clearly indicate that AI model outputs are not free from misinformation both in interpretation of research results as well as research article citations." Specifically, "Cronbach's alpha (if item deleted) was interpreted correctly in only one session by both AI models, ScholarAI (session four) more detailed explanation than SciSpace (session two)." Additionally, "AI misinformation and bias in literature search continuing to increase in each subsequent session in both AI models, available through advanced ChatGPT5 platform is an important finding of our study."
Reports effect sizes.
Research paradigm
Positivist/empiricist
Author conclusions
"This study demonstrates that while specialized AI models assist in research workflows, they are not a substitute for human expertise. Both models repeatedly failed to properly interpret complex biostatistical diagnostic parameters, such as Cronbach's alpha if item deleted test outputs. The results of our study clearly indicate that AI model outputs are not free from misinformation both in interpretation of research results as well as research article citations. It is imperative that researchers critically review the AI model outputs and validate the inferences and citations generated themselves."
Risk of bias
Evaluator subjectivity in assessment despite standardization training; Selection bias: only two custom GPTs evaluated from research and analysis category; Temporal specificity: data collected in specific months with rapid AI model evolution; Opacity of underlying AI algorithms prevents reproducibility assessment; Small number of sessions (4 per model) limiting generalizability; Temporal specificity (September-November 2025) affecting reproducibility; Limited sample of only 2 AI models from same platform (ChatGPT); Potential evaluator familiarity bias despite inter-rater agreement monitoring; model outputs may vary with updates; Small sample size: Only 4 sessions per AI model; Lack of standardized evaluation framework: Rubrics were modified using CLEAR tool but no validated instrument for assessing AI model output
Limitations
- The authors identified five key limitations: "1
- Researchers are novice in use of GPT AI models
- Evaluation of AI output was subjective
- Opacity of the underlying AI algorithms and lack of reproducibility
- Rapid model evolution and temporal specificity of data
- Contextual and domain-specific limits to generalizability."
Open questions raised
- The authors identify that "investigation of specific generative AI models as research assistants is limited except for ChatGPT." They also note that "it is essential that the robustness of output of these AI models is evaluated and reported in literature to guide students and novice researchers regarding the common pitfalls and limitations."
Explore related topics
Related papers
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations
- Artificial intelligence in higher education: the state of the fieldHelen Crompton · 2023 · 1,378 citations
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- Ethics of AI in Education: Towards a Community-Wide FrameworkW. Holmes · 2021 · 1,056 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- Shaping the Future of Education: Exploring the Potential and Consequences of AI and ChatGPT in Educational SettingsSimone Grassini · 2023 · 921 citations