12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Assessing Large Language Models for Early Article Identification in Otolaryngology—Head and Neck Surgery Systematic Reviews

Ajibola B. Bakare, Young Lee, Jhuree Hong, C. E. Richter, Jonathan P. Kuriakose · Health care science · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

9/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.1002/hcs2.70048

Methodology & findings

Study design

Comparative analysis study.

Sample

N = 3, 3 groups

Primary method

Recall calculation: TP/(TP+FP)×100. Descriptive analysis of outputs including counts of genuine articles, authentic matches, true positives, false positives, and false negatives. Bar graphs constructed to represent variability in publication years and journals. Comparative analysis across three prompts and two LLMs. No inferential statistics reported.

Main result

The study found that "both Bard and ChatGPT could be improved to balance comprehensiveness with adherence to search parameters" and that "ChatGPT-2 identified 13 authentic articles, compared to just 3 by Bard-2, despite Bard generating fewer outputs." Additionally, "Bard demonstrated greater user-friendliness by generating some outputs even with simpler prompts" while "ChatGPT's limitation to five outputs per query remains unclear." The results revealed that "while both models retrieved some authentic articles, many did not align with the selection criteria established by the respective reference studies."

Reports effect sizes.

Research paradigm

Empiricist/positivist (comparative performance evaluation)

Author conclusions

The authors conclude: "Our findings suggest that Bard and ChatGPT are not suitable as primary tools for identifying articles for a systematic review, though LLMs may serve a supplementary role in the systematic review process. This work serves as a cautionary tale of incorporating general knowledge LLMs for medical or scientific usage, particularly for young researchers, researchers inexperienced with AI, and the general public. We advocate relying on standard protocols and established databases for more reliable and valid systematic review results."

Risk of bias

Selection bias: Only three reference reviews selected from multiple eligible criteria (small convenience sample of reviews); Timing/temporal bias: Queries run at different times (July 2023 vs. October 2023) with potential chatbot updates affecting comparisons; Training data bias: Both LLMs exhibit biases in publication year representation and journal selection; Prompt engineering bias: Requires user expertise; different users may formulate prompts differently; Non-reproducibility: LLM outputs vary by timing and context, undermining systematic review replication standards; Hallucination artifact: Both models generate false information with variable frequency; Limited scope: Proof-of-concept with only three reviews; generalizability unclear; Selection bias: Only three PRISMA-compliant systematic reviews selected from multiple options that met criteria; Temporal bias: Queries run at different times (July 8, 2023 and October 14-15, 2023) potentially affecting chatbot behavior; Training data bias: Both LLMs showed biases toward specific publication years and journals; Chatbot evolution bias: Results may differ between different versions of the same LLM; Prompt engineering bias: Outcomes heavily dependent on prompt formulation; Publication bias: Open access requirement in Wu et al. may have assisted article identification; Hallucination/misinformation generation: Both chatbots produced false citations; Selection bias: Only three reference systematic reviews selected from larger pool; Timing bias: Queries executed at different times (July 2023 for Jabbour et al.; October 2023 for Wong et al. and Wu et al.); Evolution bias: LLM models evolved during study period, affecting reproducibility; Training data bias: Both ChatGPT and Bard exhibited biases in training data affecting article identification and matching; Prompt engineering bias: Results heavily dependent on prompt formulation, requiring user expertise; Publication bias: Bard and ChatGPT may exhibit preference for certain journals and publication years; Artificial hallucinations: Both models generated false information, more prevalent in Bard

Limitations

  • The authors state: "Our methodology captures only a limited component of the broader review process, and as such, lacks the depth and comprehensiveness required for a complete review
  • This is an important limitation, particularly in the early stages of literature identification, where direct use of academic search engines or specialized AI tools such as Covidence, DistillerSR, PICO Portal, Rayyan, and RobotReviewer would typically yield a broader and more systematic set of articles." Additionally, "A significant challenge in this study was prompt engineering for the LLMs" and "the potential bias due to the evolution of chatbots, as evidenced by the lack of output by ChatGPT for certain queries." Furthermore, "Current literature review standards often require multiple reviewers to identify articles, which would be undermined if other users cannot reproduce the LLMs' results."

Open questions raised

  • Integration of LLMs into systematic review workflows with improved search accuracy
  • Methods to incorporate human oversight for validation of LLM outputs
  • Prompt tailoring to specific research questions
  • Strategies to minimize artificial hallucinations in LLMs
  • Understanding of ChatGPT's five-output limitation mechanism
  • Investigation of timing effects on LLM output consistency
Extracted from: pdfAgreement 61%

Explore related topics

Related papers