MMARNet: Two-Stage Remote Sensing Image Registration with Multimodal Attention Mechanism.
Saved in:
| Title: | MMARNet: Two-Stage Remote Sensing Image Registration with Multimodal Attention Mechanism. |
|---|---|
| Authors: | Liu, Xiangzeng1 (AUTHOR) xzliu@xidian.edu.cn, Shi, Guanglu1 (AUTHOR), Huang, Zhipeng1 (AUTHOR), Ji, Jian1 (AUTHOR), Miao, Qiguang1 (AUTHOR) |
| Source: | Remote Sensing. Jun2026, Vol. 18 Issue 12, p1983. 23p. |
| Subjects: | Image registration, Remote sensing, Feature extraction, Affine transformations, Drone aircraft |
| Abstract: | Highlights: What are the main findings? A geometric transformation prediction (GTP) module is developed by utilizing a dynamic adaptive sparse attention mechanism to capture prominent feature regions, thereby enabling accurate estimation and compensation for large-scale geometric transformations between the input images. A local feature refinement (LFR) module is constructed by leveraging a feature extraction network with a super token transformer attention mechanism. Therefore, high-precision keypoint-level features extracted by the module can be used to establish accurate correspondences across highly variable modalities. What are the implications of the main findings? The proposed model can be applied in the field of multimodal remote sensing image matching and registration for pixel-level spatial coordinate alignment in multimodal image fusion and change detection. The proposed method can also be applied in the field of unmanned aerial vehicle navigation in restricted environments, providing a technical foundation for its cross-modal geographic positioning. Multimodal image registration is a fundamental yet challenging task, particularly in remote sensing scenarios involving cross-platform, multi-temporal, and cross-modal data. The primary difficulty arises from the coexistence of large-scale geometric distortions and complex local appearance variations across modalities, which makes it difficult for a single-stage model to achieve both global alignment and fine-grained correspondence simultaneously. To address this issue, we propose MMARNet, a task-driven coarse-to-fine registration framework that explicitly decomposes multimodal registration into global geometric alignment and local correspondence refinement. Instead of treating registration as a unified problem, the proposed framework sequentially resolves distinct sources of error, leading to improved robustness and accuracy under challenging conditions. In the first stage, MMARNet learns geometry-aware global alignment by identifying structurally reliable regions across modalities and estimating large-scale transformations, effectively reducing the initial misalignment and normalizing the geometric space. In the second stage, the model focuses on residual local discrepancies by learning context-enhanced feature representations, enabling robust keypoint-level matching even under severe modality differences and nonlinear distortions. The two stages are designed to work in a complementary manner, where global alignment significantly simplifies the subsequent local matching process. Extensive experiments on three challenging multimodal datasets demonstrate that MMARNet achieves superior performance in both accuracy and robustness compared to existing methods. The results validate the effectiveness of the proposed problem decomposition and highlight the advantage of the coarse-to-fine optimization strategy for multimodal remote sensing image registration. [ABSTRACT FROM AUTHOR] |
| Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| Abstract: | Highlights: What are the main findings? A geometric transformation prediction (GTP) module is developed by utilizing a dynamic adaptive sparse attention mechanism to capture prominent feature regions, thereby enabling accurate estimation and compensation for large-scale geometric transformations between the input images. A local feature refinement (LFR) module is constructed by leveraging a feature extraction network with a super token transformer attention mechanism. Therefore, high-precision keypoint-level features extracted by the module can be used to establish accurate correspondences across highly variable modalities. What are the implications of the main findings? The proposed model can be applied in the field of multimodal remote sensing image matching and registration for pixel-level spatial coordinate alignment in multimodal image fusion and change detection. The proposed method can also be applied in the field of unmanned aerial vehicle navigation in restricted environments, providing a technical foundation for its cross-modal geographic positioning. Multimodal image registration is a fundamental yet challenging task, particularly in remote sensing scenarios involving cross-platform, multi-temporal, and cross-modal data. The primary difficulty arises from the coexistence of large-scale geometric distortions and complex local appearance variations across modalities, which makes it difficult for a single-stage model to achieve both global alignment and fine-grained correspondence simultaneously. To address this issue, we propose MMARNet, a task-driven coarse-to-fine registration framework that explicitly decomposes multimodal registration into global geometric alignment and local correspondence refinement. Instead of treating registration as a unified problem, the proposed framework sequentially resolves distinct sources of error, leading to improved robustness and accuracy under challenging conditions. In the first stage, MMARNet learns geometry-aware global alignment by identifying structurally reliable regions across modalities and estimating large-scale transformations, effectively reducing the initial misalignment and normalizing the geometric space. In the second stage, the model focuses on residual local discrepancies by learning context-enhanced feature representations, enabling robust keypoint-level matching even under severe modality differences and nonlinear distortions. The two stages are designed to work in a complementary manner, where global alignment significantly simplifies the subsequent local matching process. Extensive experiments on three challenging multimodal datasets demonstrate that MMARNet achieves superior performance in both accuracy and robustness compared to existing methods. The results validate the effectiveness of the proposed problem decomposition and highlight the advantage of the coarse-to-fine optimization strategy for multimodal remote sensing image registration. [ABSTRACT FROM AUTHOR] |
|---|---|
| ISSN: | 20724292 |
| DOI: | 10.3390/rs18121983 |