Spatial–Frequency Inductive Bias-Guided Cross-Domain Representation Learning for Infrared Small Object Detection.
Saved in:
| Title: | Spatial–Frequency Inductive Bias-Guided Cross-Domain Representation Learning for Infrared Small Object Detection. |
|---|---|
| Authors: | Cheng, Quanrun1,2,3 (AUTHOR), Zeng, Cao1,2,3 (AUTHOR) czeng@mail.xidian.edu.cn, He, Qi3,4 (AUTHOR), Zhang, Yuhong1,3,4 (AUTHOR), Ning, Hailong5 (AUTHOR) |
| Source: | Remote Sensing. May2026, Vol. 18 Issue 10, p1645. 24p. |
| Subjects: | Infrared imaging, Object recognition (Computer vision) |
| Abstract: | Highlights: What are the main findings? A spatial–frequency inductive bias-guided network (SFI-Net) is proposed to enable effective cross-domain representation learning from a vision foundation model for infrared small object detection. The designed spatial–frequency hybrid adapter and feature compensation strategy significantly improve detection accuracy while maintaining high computational efficiency on NUDT-SIRST and IRSTD-1K datasets. What are the implications of the main findings? The introduced infrared-specific spatial–frequency inductive biases provides an effective paradigm for adapting vision foundation models to specialized sensing domains. The proposed framework demonstrates that parameter-efficient adaptation can simultaneously enhance detection performance, robustness, and cross-scenario generalization in low-SNR infrared environments. Infrared small object detection (ISOD) plays a crucial role in military reconnaissance, security surveillance, and remote sensing monitoring, where weak thermal responses and complex backgrounds impose significant challenges. The recent self-supervised vision foundation model DINOv3 has demonstrated remarkable generalization ability across various visual tasks. However, directly transferring it to ISOD still remains challenging due to substantial cross-domain discrepancy between visible and infrared imagery, as well as the limited granularity of foundation features in capturing subtle thermal variations. To address these issues, this study proposes a spatial–frequency inductive bias-guided network (SFI-Net) based on DINOv3 for cross-domain representation learning in infrared small object detection. Instead of conventional domain adaptation strategies, SFI-Net explicitly models infrared-specific inductive biases in both spatial and frequency domains to enhance transferred representations. First, a spatio-frequency hybrid adapter (SFHA) is designed and embedded across multiple layers of the frozen backbone to learn infrared-specific inductive biases within distinct subspaces. Second, a feature compensation strategy with an auxiliary convolutional branch is devised to compensate for the limitation of DINOv3 in capturing multi-scale fine-grained features. Extensive experiments on the IRSTD-1K and NUDT-SIRST datasets demonstrate that the proposed SFI-Net outperforms state-of-the-art methods in both detection accuracy and computational efficiency while exhibiting strong cross-scenario generalization capability. [ABSTRACT FROM AUTHOR] |
| Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | Highlights: What are the main findings? A spatial–frequency inductive bias-guided network (SFI-Net) is proposed to enable effective cross-domain representation learning from a vision foundation model for infrared small object detection. The designed spatial–frequency hybrid adapter and feature compensation strategy significantly improve detection accuracy while maintaining high computational efficiency on NUDT-SIRST and IRSTD-1K datasets. What are the implications of the main findings? The introduced infrared-specific spatial–frequency inductive biases provides an effective paradigm for adapting vision foundation models to specialized sensing domains. The proposed framework demonstrates that parameter-efficient adaptation can simultaneously enhance detection performance, robustness, and cross-scenario generalization in low-SNR infrared environments. Infrared small object detection (ISOD) plays a crucial role in military reconnaissance, security surveillance, and remote sensing monitoring, where weak thermal responses and complex backgrounds impose significant challenges. The recent self-supervised vision foundation model DINOv3 has demonstrated remarkable generalization ability across various visual tasks. However, directly transferring it to ISOD still remains challenging due to substantial cross-domain discrepancy between visible and infrared imagery, as well as the limited granularity of foundation features in capturing subtle thermal variations. To address these issues, this study proposes a spatial–frequency inductive bias-guided network (SFI-Net) based on DINOv3 for cross-domain representation learning in infrared small object detection. Instead of conventional domain adaptation strategies, SFI-Net explicitly models infrared-specific inductive biases in both spatial and frequency domains to enhance transferred representations. First, a spatio-frequency hybrid adapter (SFHA) is designed and embedded across multiple layers of the frozen backbone to learn infrared-specific inductive biases within distinct subspaces. Second, a feature compensation strategy with an auxiliary convolutional branch is devised to compensate for the limitation of DINOv3 in capturing multi-scale fine-grained features. Extensive experiments on the IRSTD-1K and NUDT-SIRST datasets demonstrate that the proposed SFI-Net outperforms state-of-the-art methods in both detection accuracy and computational efficiency while exhibiting strong cross-scenario generalization capability. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 20724292 |
| DOI: | 10.3390/rs18101645 |