An Optimized Composite YOLO Model for Transmission Tower Detection in Satellite Optical Remote Sensing Imagery.

Saved in:
Bibliographic Details
Title: An Optimized Composite YOLO Model for Transmission Tower Detection in Satellite Optical Remote Sensing Imagery.
Authors: Leng, Runming1,2,3 (AUTHOR), Zhang, Guo2,3 (AUTHOR), Hao, Weifeng1,2,3 (AUTHOR) haowf@whu.edu.cn, Guo, Bingxuan1,3 (AUTHOR), Zhu, Chunyang2,3 (AUTHOR)
Source: Remote Sensing. May2026, Vol. 18 Issue 10, p1499. 21p.
Subjects: Satellite-based remote sensing, Object recognition (Computer vision), High resolution imaging, Ensemble learning
Abstract: Highlights: What are the main findings? A multi-source, multi-resolution satellite-only transmission tower dataset (HRS-PTD) is constructed, and statistical analysis reveals that over 75% of targets occupy less than 3.26% of the image area, with nearly two-thirds exhibiting slender, randomly oriented bounding boxes. An optimized composite YOLO model integrating CARAFE upsampling and a direction-aware deformable convolution module (C_DCA) achieves 92.28% mAP on HRS-PTD, improving by 5.41 percentage points over RetinaNet while achieving 102.6 FPS, demonstrating superior accuracy–efficiency trade-off over representative classical detectors. What are the implications of the main findings? The proposed method demonstrates practical feasibility for large-scale transmission tower detection in real satellite imagery, achieving correct detection rates of 88% on Google Earth and 76% on Gaofen-7 imagery, with particularly pronounced gains under lower-resolution and weak-feature conditions. The complementary design of CARAFE and C_DCA offers a transferable framework for detecting small, slender, and randomly oriented objects in high-resolution satellite remote sensing imagery beyond transmission towers. Safe low-altitude flight requires precise perception of obstacles like widespread transmission towers. Traditional inspection is often costly and inefficient. While satellite remote sensing enables automated detection, transmission towers exhibit small scales, slender structures, and random orientations, causing feature loss and receptive field mismatch. This study constructs HRS-PTD, a multi-source, multi-resolution satellite optical dataset, and analyzes target morphology. We then propose an optimized composite YOLO model using a streamlined three-stage baseline with C3k2 and SPPF modules. To enhance small object feature reconstruction, CARAFE is integrated into the upsampling path for content-aware dynamic kernels. Furthermore, a direction-aware C_DCA module, incorporating deformable convolutions, utilizes multi-directional strip branches and adaptive attention to improve slender target representation. Ablation experiments show the model achieves 92.28% mAP, with precision and recall increasing by 1.53 and 12.12 percentage points over the baseline. Comparative experiments against representative classical detectors further demonstrate that the proposed model achieves superior overall performance in both detection accuracy and inference efficiency. Tests on Google Earth and Gaofen-7 imagery yield 88% and 76% accuracy, confirming real-world feasibility. [ABSTRACT FROM AUTHOR]
Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:Highlights: What are the main findings? A multi-source, multi-resolution satellite-only transmission tower dataset (HRS-PTD) is constructed, and statistical analysis reveals that over 75% of targets occupy less than 3.26% of the image area, with nearly two-thirds exhibiting slender, randomly oriented bounding boxes. An optimized composite YOLO model integrating CARAFE upsampling and a direction-aware deformable convolution module (C_DCA) achieves 92.28% mAP on HRS-PTD, improving by 5.41 percentage points over RetinaNet while achieving 102.6 FPS, demonstrating superior accuracy–efficiency trade-off over representative classical detectors. What are the implications of the main findings? The proposed method demonstrates practical feasibility for large-scale transmission tower detection in real satellite imagery, achieving correct detection rates of 88% on Google Earth and 76% on Gaofen-7 imagery, with particularly pronounced gains under lower-resolution and weak-feature conditions. The complementary design of CARAFE and C_DCA offers a transferable framework for detecting small, slender, and randomly oriented objects in high-resolution satellite remote sensing imagery beyond transmission towers. Safe low-altitude flight requires precise perception of obstacles like widespread transmission towers. Traditional inspection is often costly and inefficient. While satellite remote sensing enables automated detection, transmission towers exhibit small scales, slender structures, and random orientations, causing feature loss and receptive field mismatch. This study constructs HRS-PTD, a multi-source, multi-resolution satellite optical dataset, and analyzes target morphology. We then propose an optimized composite YOLO model using a streamlined three-stage baseline with C3k2 and SPPF modules. To enhance small object feature reconstruction, CARAFE is integrated into the upsampling path for content-aware dynamic kernels. Furthermore, a direction-aware C_DCA module, incorporating deformable convolutions, utilizes multi-directional strip branches and adaptive attention to improve slender target representation. Ablation experiments show the model achieves 92.28% mAP, with precision and recall increasing by 1.53 and 12.12 percentage points over the baseline. Comparative experiments against representative classical detectors further demonstrate that the proposed model achieves superior overall performance in both detection accuracy and inference efficiency. Tests on Google Earth and Gaofen-7 imagery yield 88% and 76% accuracy, confirming real-world feasibility. [ABSTRACT FROM AUTHOR]
ISSN:20724292
DOI:10.3390/rs18101499