The Role of large language models in pediatric neurology: A prospective observational study evaluating ChatGPT-3.5 and ChatGPT-4.0

Authors

  • Sevim Turay Department of Pediatric Neurology, Düzce University, Medical Faculty, Düzce, Türkiye
  • Nefise Arıbaş Öz Department of Pediatric Neurology, Düzce University, Medical Faculty, Düzce, Türkiye
  • Elif Meliha Sözbir Department of Pediatric Neurology, Düzce University, Medical Faculty, Düzce, Türkiye
  • Veysel Uludağ Department of Physical Therapy and Rehabilitation, Düzce University, Medical Faculty, Düzce, Türkiye

DOI:

https://doi.org/10.30714/j-ebr.2026.279

Keywords:

Artificial intelligence, ChatGPT, clinical decision support, large language models, pediatric neurology

Abstract

Aim: Artificial intelligence (AI) and large language models (LLMs) are increasingly used in medicine, yet their role in pediatric neurology remains unclear. Given the unique challenges in this field, evaluating the accuracy, reliability, and clinical applicability of AI-generated responses is essential. This study aimed to compare the performance of ChatGPT-3.5 and ChatGPT-4.0 in clinical decision-making in pediatric neurology based on expert evaluations using standardized rating criteria.

Methods: This prospective observational study included 61 board-certified pediatric neurologists who assessed AI-generated responses to ten common pediatric neurology cases. Responses were evaluated on a five-point Likert scale for accuracy, reliability, clinical applicability, and comprehensibility. Since the same raters evaluated both models, comparisons were performed using the non-parametric Wilcoxon signed-rank test. Internal consistency was examined with Cronbach’s alpha, and correlations between experience and model ratings were analyzed using Pearson’s test.

Results: GPT-4.0 achieved higher mean scores than GPT-3.5 across all domains, particularly for accuracy (4.2 ± 0.4 vs 3.5 ± 0.5, p = 0.38) and comprehensibility (4.3 ± 0.4 vs 3.6 ± 0.5, p = 0.62). Although GPT-4.0 performed slightly better overall, none of the differences were statistically significant (p > 0.05). Cronbach’s alpha values indicated low internal consistency (ranging from 0.58 to 0.67), suggesting variability among raters. Expert experience showed no significant correlation with ratings, implying that AI evaluations were largely experienced-independent.

Conclusion: While GPT-4.0 demonstrated modest improvements over GPT-3.5, neither model achieved sufficient reliability for independent clinical use in pediatric neurology. Addressing AI hallucinations, enhancing internal consistency in expert evaluations, and promoting safe clinical integration remain key priorities for future research.

Downloads

Published

2026-07-02

How to Cite

Turay, S., Arıbaş Öz, N., Sözbir, E. M., & Uludağ, V. (2026). The Role of large language models in pediatric neurology: A prospective observational study evaluating ChatGPT-3.5 and ChatGPT-4.0. EXPERIMENTAL BIOMEDICAL RESEARCH, 9(3), 202–211. https://doi.org/10.30714/j-ebr.2026.279