Evaluating and Validating Large Language Models for Health Education on Developmental Dysplasia of the Hip: 2-Phase Study With Expert Ratings and a Pilot Randomized Controlled Trial
Hui Ouyang, Gan Lin, Yiyuan Li, Zhixin Yao, Yating Li, Han Yan et al. · Journal of Medical Internet Research · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2196/73326
Methodology & findings
Study design
This was a 2-phase integrated evaluation study.
Sample
N = 127, 4 groups
Primary method
Kruskal-Wallis tests, ANOVA, post hoc comparisons for phase 1 data; Cohen d effect sizes and linear mixed-effects models used in intention-to-treat manner for phase 2 RCT data
Main result
The study found that "ChatGPT-4 (median 63.67, IQR 63.67-64.67) and DeepSeek-V3 (median 63.33, IQR 63.33-64.67) generate more accurate text than Copilot (median 59.00, IQR 58.67-59.67)" in phase 1 evaluations. In phase 2, "the intervention group showed higher eHealth literacy at T1 (33.62, 95% CI 32.76-34.49; d=0.20, 95% CI 0.13-0.56) and T2 (33.27, 95% CI 32.38-34.17; d=0.36, 95% CI 0.01-0.80), greater DDH knowledge at T1 (7.87, 95% CI 7.48-8.25, d=0.71, 95% CI 0.33-1.11) and T2 (7.12, 95% CI 6.72-7.51; d=0.54, 95% CI 0.17-0.96)".
Reports effect sizes and confidence intervals.
Research paradigm
positivist/empiricist
Author conclusions
"Mainstream LLMs demonstrate varying capacities in generating educational content for DDH. They generated DDH caregiver education materials that were associated with modest improvements in eHealth literacy and knowledge. Although LLMs can address general informational needs, they cannot completely substitute clinical evaluation. Future research should focus on optimizing plain language, refining dialogue design, and enhancing audience personalization to improve the quality of LLMs' materials."
Risk of bias
Assessor blinding only in phase 2 (pilot RCT); phase 1 expert evaluators not blinded; Small pilot sample size (N=127) limits generalizability; Potential selection bias in caregiver recruitment; No mention of allocation concealment details; Intervention group received structured training while control used only web search; Selection bias: Only 127 caregivers recruited (unclear recruitment methods); representativeness unknown; Attrition: Not reported in abstract; unclear dropout rates at T1 and T2; Assessor blinding: Phase 2 is assessor-blinded but Phase 1 expert evaluation blinding status unclear; Small sample size for pilot study may limit generalizability; Expert panel composition: 5 pediatric orthopedic experts may have specific biases; Intervention contamination risk: unclear if control group received no LLM exposure; Short follow-up period: only 2 weeks post-intervention; Assessor-blinded design in phase 2 reduces detection bias; Intention-to-treat analysis reduces attrition bias; Small pilot sample (n=127) may limit generalizability; Short follow-up period (2 weeks)
Open questions raised
- The authors identify the need for future research to "focus on optimizing plain language, refining dialogue design, and enhancing audience personalization to improve the quality of LLMs' materials." Additionally, the study highlights the limitation that LLMs cannot completely substitute clinical evaluation, suggesting gaps in clinical integration and personalization.
- The authors identify the need for "optimizing plain language, refining dialogue design, and enhancing audience personalization to improve the quality of LLMs' materials" and suggest that "LLMs can address general informational needs" but cannot replace clinical evaluation, indicating gaps in LLM development for clinical education contexts
- Future research should focus on "optimizing plain language, refining dialogue design, and enhancing audience personalization to improve the quality of LLMs' materials."
Explore related topics
Related papers
- A comprehensive AI policy education framework for university teaching and learningCecilia Ka Yuk Chan · 2023 · 1,160 citations
- The effects of over-reliance on AI dialogue systems on students' cognitive abilities: a systematic reviewChunpeng Zhai · 2024 · 1,009 citations
- ChatGPT for Education and Research: Opportunities, Threats, and StrategiesMd. Mostafizer Rahman · 2023 · 904 citations
- To use or not to use ChatGPT in higher education? A study of students’ acceptance and use of technologyArtur Strzelecki · 2023 · 691 citations
- Teacher support and student motivation to learn with Artificial Intelligence (AI) based chatbotThomas K. F. Chiu · 2023 · 617 citations
- Do AI chatbots improve students learning outcomes? Evidence from a meta‐analysisRong Wu · 2023 · 469 citations