Change search
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf
Performance comparison of Vision transformers and CNN architectures for satellite object detection
Blekinge Institute of Technology.
Blekinge Institute of Technology, Faculty of Computing, Department of Computer Science.
2024 (English)Independent thesis Advanced level (degree of Master (Two Years)), 20 credits / 30 HE creditsStudent thesis
Abstract [en]

Background This study explores how AI and machine learning are transforming satellite imagery analysis, focusing on object detection for environmental monitoring and disaster management. This thesis conducts a comparative analysis between Vision Transformers and traditional CNN architectures for satellite imagery object detection, aiming to determine their performance in terms of accuracy, execution time, and parameter fine-tuning implications.

Objective This thesis investigates the comparative performance of Vision Transformers and traditional CNN architectures for satellite imagery object detection, aiming to determine their suitability in replacing CNNs and potentially setting new benchmarks in accuracy and efficiency. The study addresses a critical gap in the literature by evaluating key metrics such as accuracy, IOU, and execution time, offering valuable insights for advancing satellite imagery analysis in fields like environmental monitoring and disaster management.

Method The proposed method involves doing a literature review to find the suitable CNN-based models and the Vision Transformers and then we perform the experi-mentation by training a YOLO v3, YOLO v5, Vision Transformer, and Swim Transformers and then comparing their performance using metrics such as IOU, execution time, and accuracy.

Results The results demonstrate improved performance metrics across YOLO v3,YOLO v5, Vision Transformer, and Swin Transformer models after fine-tuning, including higher IOU scores, enhanced accuracy, and reduced execution times. These findings underscore the effectiveness of fine-tuning in optimizing model performance for satellite imagery object detection tasks, with Vision Transformer exhibiting the highest IOU of 0.85 and Swin Transformer achieving a notable accuracy increase to 93%.

Conclusion The predictive modeling of Our study shows the significance of fine-tuning in enhancing object detection models using satellite images. Among the models we tested, the Vision Transformer stood out, by getting better enhancements in IOU and accuracy. This highlights the importance of choosing the right model and fine-tuning approach for satellite image tasks, suggesting there’s more to explore for better performance in different situations.

Place, publisher, year, edition, pages
2024. , p. 63
Keywords [en]
satellite imagery analysis, Object detection, Vision Transformers, Tra- ditional CNN architectures, YOLO, accuracy, execution time, parameter fine-tuning, literature review
National Category
Computer graphics and computer vision
Identifiers
URN: urn:nbn:se:bth-26843OAI: oai:DiVA.org:bth-26843DiVA, id: diva2:1892403
Subject / course
DV2572 Master´s Thesis in Computer Science
Educational program
DVADA Master Qualification Plan in Computer Science
Presentation
J3202, BTH, Karlskrona (English)
Supervisors
Examiners
Available from: 2024-08-27 Created: 2024-08-26 Last updated: 2025-09-30Bibliographically approved

Open Access in DiVA

fulltext(1364 kB)1076 downloads
File information
File name FULLTEXT01.pdfFile size 1364 kBChecksum SHA-512
1fb16aec5f6289133db910de04804fe0d90c883217bada85e0d5f5a6dcd260385a8d734f5e8f57c898f0c1972a38571175a41f3b6152f6e66d75c090240ed4e1
Type fulltextMimetype application/pdf

By organisation
Blekinge Institute of TechnologyDepartment of Computer Science
Computer graphics and computer vision

Search outside of DiVA

GoogleGoogle Scholar
Total: 1081 downloads
The number of downloads is the sum of all downloads of full texts. It may include eg previous versions that are now no longer available

urn-nbn

Altmetric score

urn-nbn
Total: 617 hits
CiteExportLink to record
Permanent link

Direct link
Cite
Citation style
  • apa
  • ieee
  • modern-language-association-8th-edition
  • vancouver
  • Other style
More styles
Language
  • de-DE
  • en-GB
  • en-US
  • fi-FI
  • nn-NO
  • nn-NB
  • sv-SE
  • Other locale
More languages
Output format
  • html
  • text
  • asciidoc
  • rtf