Article detail · 2026
Comparative assessment of AI-generated responses to frequently asked questions by parents on space maintainers in pediatric dentistry
- Year
- 2026
- Type
- article
Data source split
- YÖKSİS YÖKSİS article record
- YÖKSİS venue BMC ORAL HEALTH
- Catalog match (ISSN) BMC Oral Health
- OpenAlex OpenAlex enrichment (abstract, citations, topics)
Abstract
OpenAlex · English
BACKGROUND: Artificial intelligence (AI)–based chatbots are increasingly being used in healthcare services for information access and patient education. However, in pediatric dentistry, evidence regarding the accuracy, reliability, and clinical validity of these systems is limited. This study aimed to evaluate the responses of ChatGPT-3.5, ChatGPT-5, Google Gemini 2.5 Flash, DeepSeek-V3, and Grok 3 about space maintainers in terms of accuracy, reliability, quality, and readability. METHODS: The five chatbots were asked the 20 most frequently asked questions by parents about space maintainers. Responses were evaluated by two pediatric dentistry specialists for accuracy using a 5-point Likert scale, reliability using the Quality Criteria for Consumer Health Information (DISCERN) scale, quality using the Global Quality Scale (GQS), and readability using the Flesch Reading Ease Score (FRES). The Intraclass Correlation Coefficient (ICC) was used to assess agreement between the two specialists. All statistical analyses were performed using SPSS version 26. The Kruskal–Wallis test was applied for Likert, GQS, and DISCERN scores, while one-way ANOVA with post-hoc tests was applied for FRES scores. Statistical significance was set at p < 0.05. RESULTS: In this study, no statistically significant difference was found in Likert scores among ChatGPT-3.5, ChatGPT-5, Google Gemini 2.5 Flash, DeepSeek-V3, and Grok 3 (p > 0.05). In contrast, a statistically significant difference was observed in DISCERN and GQS scores (p < 0.05). In terms of FRES scores, ChatGPT-3.5, ChatGPT-5, and DeepSeek-V3 demonstrated higher readability, whereas Google Gemini 2.5 Flash and Grok 3 obtained lower scores. CONCLUSION: Although AI-based chatbots showed comparable performance in terms of accuracy, they differed significantly in reliability, quality, and readability. These findings suggest that AI chatbots have the potential to provide parents with accurate and understandable information; however, model-specific differences in reliability and readability should be considered.
Topics
Citations
OpenAlex cited_by_count. Not a WoS or Scopus citation count; those sources have no separate column here.
2 citations
OpenAlex cited_by_count (cache / database)
2 publications in the local catalog that cite this work (OpenAlex reference match; not the full global list).
- Performance of artificial intelligence chatbots in providing feeding management and oral health guidance for children with cleft lip and palate: ChatGPT 5.2 vs Gemini 3 Pro 2026
- Evaluating the Accuracy of Large Language Models in Dentistry: A Multi‐Model Study Using Clinical Questions From Turkey's Dental Specialty Exams 2026