12,445 papers · continuously updated · last export: 10 Aug 2026livingmeta.ai
← Browse all papers
AI evidence extraction

Evaluating and Validating Large Language Models for Health Education on Developmental Dysplasia of the Hip: 2-Phase Study With Expert Ratings and a Pilot Randomized Controlled Trial

Hui Ouyang, Gan Lin, Yiyuan Li, Zhixin Yao, Yating Li, Han Yan et al. · Journal of Medical Internet Research · 2026

AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.

6/10
Relevance
1/4
Quality (LMQS)
E
Evidence
0
Citations
0.00
FWCI

This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.2196/73326

Methodology & findings

Study design

This was a 2-phase integrated evaluation study.

Sample

N = 127, 4 groups

Primary method

Kruskal-Wallis tests, ANOVA, post hoc comparisons for phase 1 data; Cohen d effect sizes and linear mixed-effects models used in intention-to-treat manner for phase 2 RCT data

Main result

The study found that "ChatGPT-4 (median 63.67, IQR 63.67-64.67) and DeepSeek-V3 (median 63.33, IQR 63.33-64.67) generate more accurate text than Copilot (median 59.00, IQR 58.67-59.67)" in phase 1 evaluations. In phase 2, "the intervention group showed higher eHealth literacy at T1 (33.62, 95% CI 32.76-34.49; d=0.20, 95% CI 0.13-0.56) and T2 (33.27, 95% CI 32.38-34.17; d=0.36, 95% CI 0.01-0.80), greater DDH knowledge at T1 (7.87, 95% CI 7.48-8.25, d=0.71, 95% CI 0.33-1.11) and T2 (7.12, 95% CI 6.72-7.51; d=0.54, 95% CI 0.17-0.96)".

Reports effect sizes and confidence intervals.

Research paradigm

positivist/empiricist

Author conclusions

"Mainstream LLMs demonstrate varying capacities in generating educational content for DDH. They generated DDH caregiver education materials that were associated with modest improvements in eHealth literacy and knowledge. Although LLMs can address general informational needs, they cannot completely substitute clinical evaluation. Future research should focus on optimizing plain language, refining dialogue design, and enhancing audience personalization to improve the quality of LLMs' materials."

Risk of bias

Assessor blinding only in phase 2 (pilot RCT); phase 1 expert evaluators not blinded; Small pilot sample size (N=127) limits generalizability; Potential selection bias in caregiver recruitment; No mention of allocation concealment details; Intervention group received structured training while control used only web search; Selection bias: Only 127 caregivers recruited (unclear recruitment methods); representativeness unknown; Attrition: Not reported in abstract; unclear dropout rates at T1 and T2; Assessor blinding: Phase 2 is assessor-blinded but Phase 1 expert evaluation blinding status unclear; Small sample size for pilot study may limit generalizability; Expert panel composition: 5 pediatric orthopedic experts may have specific biases; Intervention contamination risk: unclear if control group received no LLM exposure; Short follow-up period: only 2 weeks post-intervention; Assessor-blinded design in phase 2 reduces detection bias; Intention-to-treat analysis reduces attrition bias; Small pilot sample (n=127) may limit generalizability; Short follow-up period (2 weeks)

Open questions raised

  • The authors identify the need for future research to "focus on optimizing plain language, refining dialogue design, and enhancing audience personalization to improve the quality of LLMs' materials." Additionally, the study highlights the limitation that LLMs cannot completely substitute clinical evaluation, suggesting gaps in clinical integration and personalization.
  • The authors identify the need for "optimizing plain language, refining dialogue design, and enhancing audience personalization to improve the quality of LLMs' materials" and suggest that "LLMs can address general informational needs" but cannot replace clinical evaluation, indicating gaps in LLM development for clinical education contexts
  • Future research should focus on "optimizing plain language, refining dialogue design, and enhancing audience personalization to improve the quality of LLMs' materials."
Data: not_statedCode: not_statedExtracted from: pdfAgreement 64%

Explore related topics

Related papers