12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Evaluation of output of AI models ScholarAI and SciSpace: Implications for use of generative artificial intelligence models as research assistants

Meenakshi Aggarwal, Sonia S. Kharay, Priya Bansal, Seema Gupta, Hitant Vohra, Sarit Sharma et al. · The Journal of Community Health Management · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.18231/j.jchm.17551.1781068642

Methodology & findings

Study design

Experimental comparative performance study evaluating two custom GPT models (ScholarAI and SciSpace) across four temporal sessions (September 5, November 13-15, 2025).

Sample

N = 4, 2 groups

Primary method

Descriptive statistics including absolute frequencies and percentages to summarize distribution of valid articles, invalid DOIs, and hallucinated references across study sessions. Inter-rater agreement monitoring across evaluation axes. No inferential statistics reported.

Main result

The study found that both AI models demonstrated significant limitations in research assistance. "The results of our study clearly indicate that AI model outputs are not free from misinformation both in interpretation of research results as well as research article citations." Specifically, "Cronbach's alpha (if item deleted) was interpreted correctly in only one session by both AI models, ScholarAI (session four) more detailed explanation than SciSpace (session two)." Additionally, "AI misinformation and bias in literature search continuing to increase in each subsequent session in both AI models, available through advanced ChatGPT5 platform is an important finding of our study."

Reports effect sizes.

Research paradigm

Positivist/empiricist

Author conclusions

"This study demonstrates that while specialized AI models assist in research workflows, they are not a substitute for human expertise. Both models repeatedly failed to properly interpret complex biostatistical diagnostic parameters, such as Cronbach's alpha if item deleted test outputs. The results of our study clearly indicate that AI model outputs are not free from misinformation both in interpretation of research results as well as research article citations. It is imperative that researchers critically review the AI model outputs and validate the inferences and citations generated themselves."

Risk of bias

Evaluator subjectivity in assessment despite standardization training; Selection bias: only two custom GPTs evaluated from research and analysis category; Temporal specificity: data collected in specific months with rapid AI model evolution; Opacity of underlying AI algorithms prevents reproducibility assessment; Subjectivity in AI output evaluation despite standardization training; Small number of sessions (4 per model) limiting generalizability; Temporal specificity (September-November 2025) affecting reproducibility; Limited sample of only 2 AI models from same platform (ChatGPT); Potential evaluator familiarity bias despite inter-rater agreement monitoring; Selection bias: Only two custom GPTs evaluated (ScholarAI and SciSpace), both among top 10 in research & analysis category on ChatGPT; Evaluator bias: Subjective evaluation of AI output by three researchers, though inter-rater agreement monitoring and consensus-building were employed; Temporal specificity: Study conducted at specific time points (September, November 2025); model outputs may vary with updates; Small sample size: Only 4 sessions per AI model; Lack of standardized evaluation framework: Rubrics were modified using CLEAR tool but no validated instrument for assessing AI model output

Limitations

  • The authors identified five key limitations: "1
  • Researchers are novice in use of GPT AI models
  • Evaluation of AI output was subjective
  • Opacity of the underlying AI algorithms and lack of reproducibility
  • Rapid model evolution and temporal specificity of data
  • Contextual and domain-specific limits to generalizability."

Open questions raised

  • The authors identify that "investigation of specific generative AI models as research assistants is limited except for ChatGPT." They also note that "it is essential that the robustness of output of these AI models is evaluated and reported in literature to guide students and novice researchers regarding the common pitfalls and limitations."
  • The authors note that "investigation of specific generative AI models as research assistants is limited except for ChatGPT." They emphasize that "It is essential that the robustness of output of these AI models is evaluated and reported in literature to guide students and novice researchers regarding the common pitfalls and limitations" and identify that "The reliability and validity of AI models as research assistants is circumspect and needs further research on a larger scale."
  • The authors note that "Investigation of specific generative AI models as research assistants is limited except for ChatGPT." They call for evaluation and reporting of AI model robustness in literature "to guide students and novice researchers regarding the common pitfalls and limitations." They also state that "The reliability and validity of AI models as research assistants is circumspect and needs further research on a larger scale."
Extracted from: pdfAgreement 69%

Explore related topics

Related papers