Performance comparison of Vision transformers and CNN architectures for satellite object detection
2024 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE credits
Student thesis
Abstract [en]
Background This study explores how AI and machine learning are transforming satellite imagery analysis, focusing on object detection for environmental monitoring and disaster management. This thesis conducts a comparative analysis between Vision Transformers and traditional CNN architectures for satellite imagery object detection, aiming to determine their performance in terms of accuracy, execution time, and parameter fine-tuning implications.
Objective This thesis investigates the comparative performance of Vision Transformers and traditional CNN architectures for satellite imagery object detection, aiming to determine their suitability in replacing CNNs and potentially setting new benchmarks in accuracy and efficiency. The study addresses a critical gap in the literature by evaluating key metrics such as accuracy, IOU, and execution time, offering valuable insights for advancing satellite imagery analysis in fields like environmental monitoring and disaster management.
Method The proposed method involves doing a literature review to find the suitable CNN-based models and the Vision Transformers and then we perform the experi-mentation by training a YOLO v3, YOLO v5, Vision Transformer, and Swim Transformers and then comparing their performance using metrics such as IOU, execution time, and accuracy.
Results The results demonstrate improved performance metrics across YOLO v3,YOLO v5, Vision Transformer, and Swin Transformer models after fine-tuning, including higher IOU scores, enhanced accuracy, and reduced execution times. These findings underscore the effectiveness of fine-tuning in optimizing model performance for satellite imagery object detection tasks, with Vision Transformer exhibiting the highest IOU of 0.85 and Swin Transformer achieving a notable accuracy increase to 93%.
Conclusion The predictive modeling of Our study shows the significance of fine-tuning in enhancing object detection models using satellite images. Among the models we tested, the Vision Transformer stood out, by getting better enhancements in IOU and accuracy. This highlights the importance of choosing the right model and fine-tuning approach for satellite image tasks, suggesting there’s more to explore for better performance in different situations.
Place, publisher, year, edition, pages
2024. , p. 63
Keywords [en]
satellite imagery analysis, Object detection, Vision Transformers, Tra- ditional CNN architectures, YOLO, accuracy, execution time, parameter fine-tuning, literature review
National Category
Computer graphics and computer vision
Identifiers
URN: urn:nbn:se:bth-26843OAI: oai:DiVA.org:bth-26843DiVA, id: diva2:1892403
Subject / course
DV2572 Master´s Thesis in Computer Science
Educational program
DVADA Master Qualification Plan in Computer Science
Presentation
J3202, BTH, Karlskrona (English)
Supervisors
Examiners
2024-08-272024-08-262025-09-30Bibliographically approved