İçeriğe geç
akaturk Akademik ölçüm

Makale detayı · 2025

Evaluation of Artificial Intelligence Accuracy in Interpreting Pulmonary Function Tests: A Comparison of ChatGPT 4o, DeepSeek R1, and Claude 3.5 Sonnet

Paediatric respiratory physiology and diagnostic testing

YÖKSİS OpenAlex SJR Q1 JCR Q1 Atıf 0
Yıl
2025
ISSN
1399-3003
Tür
conference-abstract

Veri kaynağı ayrımı

  • YÖKSİS YÖKSİS makale kaydı
  • OpenAlex OpenAlex zenginleştirmesi (özet, atıf, konular)

Özet

İngilizce (OpenAlex)

Large language models (LLMs) offer potential for clinical decision support, yet their role in spirometry interpretation is understudied. This study evaluates the accuracy, clinical utility, and performance of ChatGPT4o, DeepSeekR1, and Claude 3.5. Materials and Methods: Spirometry results from 100 randomly selected patients at Ege University were retrospectively analyzed. Interpretations by two independent pediatric pulmonologists served as the reference standard. ChatGPT, DeepSeek, and Claude were evaluated based on six ATS/ERS criteria. Each LLM was tested using a standardized prompt to assess acceptability and clinical usability, classifying patterns as Normal, Obstructive, Restrictive, or Mixed. Results: ChatGPT had the highest agreement with experts, while DeepSeek and Claude showed lower concordance. Mixed pattern and restrictive disorders were the most misclassified. Logistic regression identified expiratory time as the key factor in LLM decision-making. Conclusion: This study highlights the potential and limitations of LLMs in spirometry interpretation. While they cannot replace expert judgment, LLMs can enhance clinical efficiency. ChatGPT showed the highest accuracy, but further optimization is needed for complex cases. Future research should improve model training, real-time learning, and clinical integration. This study provides one of the first comparisons of AI in spirometry interpretation, contributing valuable insights to the field. erj;66/suppl_69/OA1165/TB1 T1 TB1 Models Usability (ĸ, p-value) SFT pattern (ĸ, p-value) ChatGPT 0.48, 0.82 0.51, <0.05 Claude 0.09, <0.05 0.29, <0.05 DeepSeek 0.00, <0.05 0.29, <0.05

Konular

  • Chronic Obstructive Pulmonary Disease (COPD) Research
  • Artificial Intelligence in Healthcare and Education
  • Inhalation and Respiratory Drug Delivery

Birincil konu Chronic Obstructive Pulmonary Disease (COPD) Research

Yazarlar

  1. KÜBRA ÖZKAYA
  2. GÖKÇEN KARTAL ÖZTÜRK
  3. ECE OCAK
  4. FİGEN GÜLEN EGE ÜNİVERSİTESİ