Cooperative Hybrid Domain Network for Salient Object Detection in Optical Remote Sensing Images.

Saved in:
Bibliographic Details
Title: Cooperative Hybrid Domain Network for Salient Object Detection in Optical Remote Sensing Images.
Authors: Gu, Yi1 (AUTHOR), Zhou, Jianhang2 (AUTHOR) jhzhou22@mails.jlu.edu.cn, Yan, Lelei3 (AUTHOR)
Source: Remote Sensing. Apr2026, Vol. 18 Issue 7, p1087. 31p.
Subjects: Optical remote sensing, Object recognition (Computer vision), Artificial neural networks
Abstract: Highlights: What are the main findings? A cooperative hybrid domain network is proposed to address semantic misalignment in cross-domain feature aggregation, leveraging the cross-domain multi-head self-attention and multi-branch cooperative decoder for synergistic collaboration between frequency and spatial domains. CHDNet achieves state-of-the-art performance on ORSSD and EORSSD datasets, with superior precision in salient object boundary delineation and strong robustness against complex backgrounds and extreme scale variations. What are the implications of the main findings? The cooperative hybrid paradigm breaks the limitation of passive feature concatenation, providing a new design direction for cross-domain feature fusion in the optical remote sensing image salient object detection task. The high-quality saliency detection results of CHDNet enhance the reliability of downstream applications such as urban planning and disaster assessment, expanding the application value of frequency-domain learning. Salient Object Detection (SOD) in Optical Remote Sensing Images (ORSIs) aims to localize and segment visually prominent objects amidst complex backgrounds and extreme scale variations. However, we observe that current frequency-aware methods typically rely on a naive feature aggregation paradigm, merging frequency and spatial features via simple concatenation, addition, or direct combination. This shallow interaction overlooks the inherent semantic misalignment between the two domains, resulting in feature redundancy and poor boundary delineation. To address this limitation, we propose the Cooperative Hybrid Domain Network (CHDNet), a framework designed to facilitate synergistic cooperation between heterogeneous domains. Specifically, we propose the Cross-Domain Multi-Head Self-Attention (CD-MHSA) mechanism as a semantic bridge following the encoder. It employs a dimension expansion strategy to construct a Unified Interaction Manifold and utilizes a Frequency Anchor Interaction mechanism to achieve precise modulation of spatial textures using global spectral cues. Furthermore, to address the dual challenges of lacking explicit interpretation mechanisms for semantic co-occurrence and the susceptibility of topological structures to fracture in complex scenes during the decoding phase, we design a Multi-Branch Cooperative Decoder (MBCD) comprising three parallel paths: edge semantics, global relations, and reverse correction. This module dynamically integrates these heterogeneous clues through a Cooperative Fusion Strategy, combining explicit global dependency modeling with dual-domain reverse mining. Extensive experiments on multiple benchmark datasets demonstrate that the proposed CHDNet achieves performance superior to state-of-the-art (SOTA) methods. [ABSTRACT FROM AUTHOR]
Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:Highlights: What are the main findings? A cooperative hybrid domain network is proposed to address semantic misalignment in cross-domain feature aggregation, leveraging the cross-domain multi-head self-attention and multi-branch cooperative decoder for synergistic collaboration between frequency and spatial domains. CHDNet achieves state-of-the-art performance on ORSSD and EORSSD datasets, with superior precision in salient object boundary delineation and strong robustness against complex backgrounds and extreme scale variations. What are the implications of the main findings? The cooperative hybrid paradigm breaks the limitation of passive feature concatenation, providing a new design direction for cross-domain feature fusion in the optical remote sensing image salient object detection task. The high-quality saliency detection results of CHDNet enhance the reliability of downstream applications such as urban planning and disaster assessment, expanding the application value of frequency-domain learning. Salient Object Detection (SOD) in Optical Remote Sensing Images (ORSIs) aims to localize and segment visually prominent objects amidst complex backgrounds and extreme scale variations. However, we observe that current frequency-aware methods typically rely on a naive feature aggregation paradigm, merging frequency and spatial features via simple concatenation, addition, or direct combination. This shallow interaction overlooks the inherent semantic misalignment between the two domains, resulting in feature redundancy and poor boundary delineation. To address this limitation, we propose the Cooperative Hybrid Domain Network (CHDNet), a framework designed to facilitate synergistic cooperation between heterogeneous domains. Specifically, we propose the Cross-Domain Multi-Head Self-Attention (CD-MHSA) mechanism as a semantic bridge following the encoder. It employs a dimension expansion strategy to construct a Unified Interaction Manifold and utilizes a Frequency Anchor Interaction mechanism to achieve precise modulation of spatial textures using global spectral cues. Furthermore, to address the dual challenges of lacking explicit interpretation mechanisms for semantic co-occurrence and the susceptibility of topological structures to fracture in complex scenes during the decoding phase, we design a Multi-Branch Cooperative Decoder (MBCD) comprising three parallel paths: edge semantics, global relations, and reverse correction. This module dynamically integrates these heterogeneous clues through a Cooperative Fusion Strategy, combining explicit global dependency modeling with dual-domain reverse mining. Extensive experiments on multiple benchmark datasets demonstrate that the proposed CHDNet achieves performance superior to state-of-the-art (SOTA) methods. [ABSTRACT FROM AUTHOR]
ISSN:20724292
DOI:10.3390/rs18071087