İçeriğe geç
akaturk Akademik ölçüm

Makale detayı · 2025

Metin temsili ve model seçiminin sınıflandırma performansına etkisi: Covid19-FNIR veri seti üzerinde TF-IDF, BoW ve Transformatör tabanlı yöntemlerin kapsamlı bir karşılaştırması

Niğde Ömer Halisdemir Üniversitesi Mühendislik Bilimleri Dergisi

YÖKSİS OpenAlex Açık erişim · diamond TR Index Atıf 3 Üst %10 Yüzdelik 92.9% FWCI 3.37
Yıl
2025
ISSN
2564-6605
Tür
article

Veri kaynağı ayrımı

  • YÖKSİS YÖKSİS makale kaydı
  • OpenAlex OpenAlex zenginleştirmesi (özet, atıf, konular)

Özet

Türkçe

This study evaluates the performance of various machine learning (ML) models on a dataset split into 80% training and 20% testing using Term Frequency-Inverse Document Frequency (TF-IDF) and Bag of Words (BoW) text vectorization. Transformer-based models like DistilBERT, RoBERTa, and alBERT were integrated with classical ML algorithms and ensemble methods such as Stacking, Hard Voting, and Soft Voting. Stacking achieved the highest performance with both methods—92.62% Accuracy (Acc) and 92.51% F1-score (F1) with TF-IDF, and 92.29% Acc and 92.41% F1 with BoW. Hard Voting with BoW yielded the highest Recall (95.23%). Classical models like Logistic Regression (LR) and Support Vector Machine (SVM) performed better with BoW, reaching 90.98% and 90.51% Acc, respectively. Overall, TF-IDF produced balanced outcomes, while BoW offered higher Recall and Precision in specific cases. These results highlight the significance of both model and text representation choices in achieving optimal classification performance.

Konular

  • Text and Document Classification Technologies
  • Sentiment Analysis and Opinion Mining
  • Advanced Text Analysis Techniques

Birincil konu Text and Document Classification Technologies

Yazarlar

  1. MUHAMMET SİNAN BAŞARSLAN
  2. FATİH BAL KIRKLARELİ ÜNİVERSİTESİ