Limitations and mitigation strategies for using generative artificial intelligence in medical writing: a narrative review
Ki-Hyun Jeon · Journal of Korean Medical Association · 2026
AI-generated evidence extraction, verified across multiple analytical personas. Not a substitute for the peer-reviewed original.
This is an AI analysis. Read the peer-reviewed original at the publisher: https://doi.org/10.5124/jkma.25.0163
Methodology & findings
Study design
Narrative review synthesizing literature on LLM limitations and mitigation strategies in medical writing.
Primary method
The paper reviews multiple empirical studies using various statistical approaches: descriptive analysis accuracy comparisons, Fisher's exact test, Levene's test, Mann-Whitney U test, Kaplan-Meier survival analysis, and regression analysis. The reviewed literature employed cross-validation between LLM-generated code and established statistical software (SPSS, R, SAS). Accuracy was measured as percentage of correct vs. fabricated references and statistical test results.
Main result
The paper identifies that "LLM의 환각은 설계상 내재된 한계로" (hallucination in LLM is a design-inherent limitation) and that "2024년 분석에 따르면 생물의학 분야 논문 초록의 최소 13.5%가 LLM을 활용한 흔적을 보였고, 일부 세부 분야에서는 그 비율이 40%에 달하는 등 LLM이 학술 문체와 표현에 큰 영향을 미치고 있다" (according to 2024 analysis, at least 13.5% of biomedical paper abstracts showed traces of LLM use, and in some subfields the ratio reached 40%, indicating LLM has significant impact on academic writing style and expression). The study also found specific error rates: GPT-3.5 generated fake references at 98.1% rate (159 of 162 citations were non-existent), while GPT-4 reduced this to 20.6% (53 of 257 citations), though "여전히 25%가량은 허구의 참고문헌이었다" (approximately 25% were still fabricated references).
Reports effect sizes and confidence intervals.
Research paradigm
Interpretivist/Critical analysis of technology limitations
Author conclusions
The author concludes that "LLM은 의학 연구 전 과정에서 효율성과 접근성을 비약적으로 향상시키는 도구이지만, 동시에 사실 오류, 허위 인용, 통계적 부정확성, 윤리적•법적 위험을 내재한다" (LLMs dramatically improve efficiency and accessibility throughout the medical research process, but simultaneously contain inherent risks of factual errors, false citations, statistical inaccuracy, and ethical-legal hazards). The key recommendation is: "Human-in-the-loop는 연구 과정의 핵심 의사결정 단계 즉, 연구 질문 설정, 분석 설계, 통계 방법 선택, 결과 해석, 결론 도출 등의 각 단계에 인간 연구자가 직접 개입하여 LLM의 산출물을 검토하고 수정 및 승인하는 구조를 의미한다" (Human-in-the-loop means human researchers directly intervene at critical decision-making stages of research including research question formulation, analysis design, statistical method selection, result interpretation, and conclusion drawing to review, modify, and approve LLM outputs).
Risk of bias
As a narrative review, systematic selection bias is possible due to lack of predetermined search strategy. Potential publication bias in cited studies not assessed.; Selection bias in literature reviewed - author may have selectively cited studies demonstrating problems with LLMs; Temporal bias - training data cutoff dates affect model performance; newer models show improved accuracy but review does not comprehensively cover all recent developments; Geographic bias in LLM training data - models trained primarily on Western, English-language internet data, limiting applicability to non-English medical research; Publication bias - positive findings on LLM capabilities may be underrepresented; Author expertise bias - review focuses on limitations; benefits perspective less developed; LLM training data bias toward Western/English-language sources; Structural hallucination bias in probabilistic language models; Temporal data cutoff limitations excluding recent research; Author attribution bias in LLM-generated citations; Selection bias in which studies are reviewed (narrative review format)
Limitations
- The paper states that "LLM의 환각은 설계상 내재된 한계로 인식해야 하며, 이를 완전히 제거하기보다 과소화하고 관리하는 방향으로 접근하는 것이 현실적이다" (hallucinations in current LLMs must be recognized as design-inherent limitations, and realistically should be minimized and managed rather than completely eliminated)
- Additionally, "LLM은 통계적 판단의 적합성이나 연구 맥락을 스스로 평가하지 못하며, 다변량 모델에서의 교란 변수 선택, 결측치 처리 전략, 다중 비교 문제, 과적합 위험 등에 대한 학문적 판단을 수행할 수 없다" (LLMs cannot evaluate statistical assumption appropriateness or research context themselves, and cannot make scholarly judgments about confounding variable selection, missing data handling, multiple comparison issues, or overfitting risks).
Open questions raised
- Future regulatory and legal framework development needed for data analysis and code generation in healthcare contexts. Authors note that "데이터 분석과 코드 생성에 대해서는 아직 법과 제도가 완비되지 않았다. 이에 대해서는 향후 규제 환경의 변화를 주시하여야 할 필요가 있겠다" (legal and institutional frameworks for data analysis and code generation are not yet complete; future changes in regulatory environment must be monitored). Further development of specialized LLM tools for medical research contexts is implied.
- Lack of comprehensive guidelines for responsible LLM use in medical research across all research stages
- Need for standardized protocols for statistical code generation validation
- Insufficient regulatory frameworks for data analysis and code generation in Korean research context
- Limited research on long-term impacts of LLM-assisted writing on research integrity and reproducibility
- Gaps in medical education regarding LLM literacy and critical evaluation skills for trainees
Explore related topics
Related papers
- Preferred reporting items for systematic reviews and meta-analyses: the PRISMA statementDavid Moher · 2009 · 83,271 citations
- PRISMA Extension for Scoping Reviews (PRISMA-ScR): Checklist and ExplanationAndrea C. Tricco · 2018 · 40,391 citations
- Cochrane Handbook for Systematic Reviews of Interventions2019 · 14,420 citations
- PRISMA 2020 explanation and elaboration: updated guidance and exemplars for reporting systematic reviewsMatthew J. Page · 2021 · 10,956 citations
- Updated methodological guidance for the conduct of scoping reviewsMicah D.J. Peters · 2020 · 6,688 citations
- What Is the Impact of ChatGPT on Education? A Rapid Review of the LiteratureChung Kwan Lo · 2023 · 1,725 citations