Article detail · 2026
Evaluation of Cutting-Edge Object Detection Architectures on Multi-Object and Single-Object Datasets
Black Sea Journal of Engineering and Science
- Year
- 2026
- ISSN
2619-8991- Type
- article
Data source split
- YÖKSİS YÖKSİS article record
- OpenAlex OpenAlex enrichment (abstract, citations, topics)
Abstract
English (OpenAlex)
This study focuses on the performance evaluation of mainstream object detection model, namely, YOLO12, Mask-RCNN, RT-DETR, and RF-DETR on the Open Images (Multi-Object) and LaSOT (Single-Object) datasets. Object detection is one of the key technologies in computer vision tasks and very significant developments have been made during the last decades. Current cutting-edge trend applications involve CNN-based and transformer-based object detection models. CNN-based models can use one-pass (YOLO family) or two-pass (R-CNN family) implementations. One-pass object detection models can be faster but suffer from accuracy compared to the two-pass models. Transformer-based models can use Detection Transformers or Vision Transformers. Transformer-based models are gaining popularity, and their performance surpasses CNN-based models. Transformer-based models are also improving their speed. This study evaluates YOLO12 and Mask R-CNN from CNN-based family, and RT-DETR and RF-DETR transformer-based architectures in terms of accuracy and time on the Open Images and LaSOT datasets. All models are the largest models provided by the owners and pretrained on COCO dataset. Transformer-based models incorporate special types of self-attention and pose significant improvement both on accuracy and speed. The experimental results demonstrate that attention and transformer-based models perform better than the traditional CNN-based object detectors.
Topics
- Advanced Neural Network Applications
- Internet of Things and AI
- Adversarial Robustness in Machine Learning
Primary topic Advanced Neural Network Applications