Salient Object Detection for Optical Remote Sensing Images Based on Gated Differential Unit.

Saved in:
Bibliographic Details
Title: Salient Object Detection for Optical Remote Sensing Images Based on Gated Differential Unit.
Authors: Sun, Mingsi1 (AUTHOR), Lan, Ting2 (AUTHOR), Wang, Wei1,3 (AUTHOR) wangwei@ccu.edu.cn, Liu, Pingping3,4 (AUTHOR)
Source: Remote Sensing. Feb2026, Vol. 18 Issue 3, p389. 27p.
Subjects: Optical remote sensing, Transformer models, Image segmentation, Object recognition (Computer vision), Signal denoising
Abstract: Highlights: What are the main findings? Proposes the GDUFormer model, which combines the ViT backbone with Gated Differential Units to effectively address core issues such as noise suppression and scale differences in Optical Remote Sensing Image Salient Object Detection. The GDU incorporates Full-Dimensional Gated Attention and Hierarchical Differential Dynamic Convolution, enabling full-dimensional feature purification and synergistic modeling of long-range dependencies and local details. What are the implications of the main findings? The innovative design of the dual-branch gated attention and hierarchical differential mechanism provides an efficient paradigm for the integration of Transformers and convolutions, enhancing the foreground-background discrimination ability in complex scenarios. It achieves the optimal comprehensive performance on the EORSSD and ORSI-4199 datasets, with a balanced trade-off between parameters and computational complexity, offering reliable preprocessing support for downstream remote sensing image tasks. Salient object detection in optical remote sensing images has attracted extensive research interest in recent years. However, CNN-based methods are generally limited by local receptive fields, while ViT-based methods suffer from common defects in noise suppression, channel selection, foreground-background distinction, and detail enhancement. To address these issues and integrate long-distance contextual dependencies, we introduce GDUFormer, an ORSI-SOD detection method based on the ViT backbone and Gated Differential Units (GDU). Specifically, the GDU consists of two key components—Full-Dimensional Gated Attention (FGA) and Hierarchical Differential Dynamic Convolution (HDDC). FGA consists of two branches aimed at filtering effective features from the information flow. The first branch focuses on aggregating spatial local information under multiple receptive fields and filters the local feature maps via a grouping mechanism. The second branch imitates the Vision Mamba to acquire high-level reasoning and abstraction capabilities, enabling weak channel filtering. HDDC primarily utilizes distance decay and hierarchical intensity difference capture mechanisms to generate dynamic kernel spatial weights, thereby facilitating the convolution kernel to fully mix long-range contextual dependencies. Among these, the intensity difference capture mechanism can adaptively divide hierarchies and allocate parameters according to kernel size, thus realizing varying levels of difference capture in the kernel space. Extensive quantitative and qualitative experiments demonstrate the effectiveness and rationality of GDUFormer and its internal components. [ABSTRACT FROM AUTHOR]
Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:Highlights: What are the main findings? Proposes the GDUFormer model, which combines the ViT backbone with Gated Differential Units to effectively address core issues such as noise suppression and scale differences in Optical Remote Sensing Image Salient Object Detection. The GDU incorporates Full-Dimensional Gated Attention and Hierarchical Differential Dynamic Convolution, enabling full-dimensional feature purification and synergistic modeling of long-range dependencies and local details. What are the implications of the main findings? The innovative design of the dual-branch gated attention and hierarchical differential mechanism provides an efficient paradigm for the integration of Transformers and convolutions, enhancing the foreground-background discrimination ability in complex scenarios. It achieves the optimal comprehensive performance on the EORSSD and ORSI-4199 datasets, with a balanced trade-off between parameters and computational complexity, offering reliable preprocessing support for downstream remote sensing image tasks. Salient object detection in optical remote sensing images has attracted extensive research interest in recent years. However, CNN-based methods are generally limited by local receptive fields, while ViT-based methods suffer from common defects in noise suppression, channel selection, foreground-background distinction, and detail enhancement. To address these issues and integrate long-distance contextual dependencies, we introduce GDUFormer, an ORSI-SOD detection method based on the ViT backbone and Gated Differential Units (GDU). Specifically, the GDU consists of two key components—Full-Dimensional Gated Attention (FGA) and Hierarchical Differential Dynamic Convolution (HDDC). FGA consists of two branches aimed at filtering effective features from the information flow. The first branch focuses on aggregating spatial local information under multiple receptive fields and filters the local feature maps via a grouping mechanism. The second branch imitates the Vision Mamba to acquire high-level reasoning and abstraction capabilities, enabling weak channel filtering. HDDC primarily utilizes distance decay and hierarchical intensity difference capture mechanisms to generate dynamic kernel spatial weights, thereby facilitating the convolution kernel to fully mix long-range contextual dependencies. Among these, the intensity difference capture mechanism can adaptively divide hierarchies and allocate parameters according to kernel size, thus realizing varying levels of difference capture in the kernel space. Extensive quantitative and qualitative experiments demonstrate the effectiveness and rationality of GDUFormer and its internal components. [ABSTRACT FROM AUTHOR]
ISSN:20724292
DOI:10.3390/rs18030389