Bi-Level Meta-Learning for Reliable Remote Sensing Image Registration.

Saved in:
Bibliographic Details
Title: Bi-Level Meta-Learning for Reliable Remote Sensing Image Registration.
Authors: Shi, Lin1 (AUTHOR), Wang, Renzhen1,2 (AUTHOR), Zhu, Xiaofeng2,3 (AUTHOR), An, Cong1 (AUTHOR), Zhao, Kai1,2 (AUTHOR), Shu, Jun1,3 (AUTHOR), Yang, Dongfang3 (AUTHOR) yangdf@xjtu.edu.cn, Meng, Deyu1 (AUTHOR)
Source: Remote Sensing. Jun2026, Vol. 18 Issue 12, p2007. 30p.
Subjects: Image registration, Remote sensing, Mathematical optimization, Drone aircraft, Machine learning
Abstract: Highlights: What are the main findings? A saliency-aware bi-level optimization framework is proposed to address the challenges of extreme radiometric and temporal variations in heterogeneous remote sensing image matching. A novel Saliency Judgment module is developed to intelligently prioritize geometrically stable landmarks while suppressing spurious matches in weakly informative regions. What are the implications of the main findings? The bi-level meta-learning strategy establishes a new paradigm for UAV localization by enabling effective model training under weakly labeled data regimes, reducing the reliance on large-scale expert annotations. The framework significantly enhances registration accuracy and reduces ghosting artifacts, providing a robust "quality-over-quantity" correspondence selection strategy for autonomous aerial localization. Unmanned aerial vehicle (UAV) visual navigation relies critically on robust image matching between UAV-acquired aerial imagery and pre-existing satellite reference maps. However, extreme cross-domain heterogeneity—encompassing temporal, radiometric, viewpoint, and sensor variations—causes severe performance degradation in existing deep learning-based matchers trained on conventional benchmarks. Furthermore, manual annotation of ground-truth correspondences is prohibitively expensive. This paper proposes a semi-supervised saliency-aware image matching framework with bi-level meta-learning. Our approach comprises two synergistic stages: (1) automated dense correspondence generation via parameterized geometric synthesis, which constructs a large-scale coarse dataset D c (approximately 50,000 pairs) without dense manual point annotation, serving as the primary training corpus for the feature matching network; (2) expert-validated meta-data curation producing a high-quality meta-dataset D m (500 pairs) that supervises the training of a Saliency Judgment Network through bi-level meta-optimization, enabling the network to identify and prioritize geometrically reliable correspondences. Experimental results on the proposed RS-Hetero-50K benchmark and cross-domain FuJian-Mountain dataset demonstrate substantial improvements over representative sparse and detector-free matchers, including LoFTR, SuperGlue, and LightGlue. The complete CNN-attention and saliency-aware framework achieves 95.4% matching precision, which is consistent with the best result reported in the experimental section. The plug-and-play experiments further confirm that the proposed saliency module consistently improves representative sparse and detector-free matchers, indicating that the performance gain stems from both stronger feature representation and saliency-guided correspondence selection. The largest terrain-specific gain is observed in gobi scenes, where the AUC@5 px improves by 16.8% relative to the LoFTR baseline, demonstrating improved robustness in weakly textured remote sensing environments. [ABSTRACT FROM AUTHOR]
Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:Highlights: What are the main findings? A saliency-aware bi-level optimization framework is proposed to address the challenges of extreme radiometric and temporal variations in heterogeneous remote sensing image matching. A novel Saliency Judgment module is developed to intelligently prioritize geometrically stable landmarks while suppressing spurious matches in weakly informative regions. What are the implications of the main findings? The bi-level meta-learning strategy establishes a new paradigm for UAV localization by enabling effective model training under weakly labeled data regimes, reducing the reliance on large-scale expert annotations. The framework significantly enhances registration accuracy and reduces ghosting artifacts, providing a robust "quality-over-quantity" correspondence selection strategy for autonomous aerial localization. Unmanned aerial vehicle (UAV) visual navigation relies critically on robust image matching between UAV-acquired aerial imagery and pre-existing satellite reference maps. However, extreme cross-domain heterogeneity—encompassing temporal, radiometric, viewpoint, and sensor variations—causes severe performance degradation in existing deep learning-based matchers trained on conventional benchmarks. Furthermore, manual annotation of ground-truth correspondences is prohibitively expensive. This paper proposes a semi-supervised saliency-aware image matching framework with bi-level meta-learning. Our approach comprises two synergistic stages: (1) automated dense correspondence generation via parameterized geometric synthesis, which constructs a large-scale coarse dataset D c (approximately 50,000 pairs) without dense manual point annotation, serving as the primary training corpus for the feature matching network; (2) expert-validated meta-data curation producing a high-quality meta-dataset D m (500 pairs) that supervises the training of a Saliency Judgment Network through bi-level meta-optimization, enabling the network to identify and prioritize geometrically reliable correspondences. Experimental results on the proposed RS-Hetero-50K benchmark and cross-domain FuJian-Mountain dataset demonstrate substantial improvements over representative sparse and detector-free matchers, including LoFTR, SuperGlue, and LightGlue. The complete CNN-attention and saliency-aware framework achieves 95.4% matching precision, which is consistent with the best result reported in the experimental section. The plug-and-play experiments further confirm that the proposed saliency module consistently improves representative sparse and detector-free matchers, indicating that the performance gain stems from both stronger feature representation and saliency-guided correspondence selection. The largest terrain-specific gain is observed in gobi scenes, where the AUC@5 px improves by 16.8% relative to the LoFTR baseline, demonstrating improved robustness in weakly textured remote sensing environments. [ABSTRACT FROM AUTHOR]
ISSN:20724292
DOI:10.3390/rs18122007