Spectral–Spatial Masked Auto-Encoder with Central Pixel Reconstruction for Semi-Supervised Hyperspectral Image Classification.

Saved in:
Bibliographic Details
Title: Spectral–Spatial Masked Auto-Encoder with Central Pixel Reconstruction for Semi-Supervised Hyperspectral Image Classification.
Authors: Wang, Heng1 (AUTHOR), Hong, Ruize1,2 (AUTHOR), Wu, Yanxia1 (AUTHOR) wuyanxia@hrbeu.edu.cn, Lin, Dan1,2 (AUTHOR), Wang, Liguo2 (AUTHOR)
Source: Remote Sensing. May2026, Vol. 18 Issue 10, p1571. 31p.
Subjects: Image reconstruction, Machine learning, Artificial neural networks
Abstract: Highlights: What are the main findings? In the proposed spectral–spatial masked auto-encoder with central pixel reconstruction (SSMAE-CR), we introduced central pixel reconstruction in both the spectral and spatial branches, effectively improving the separability of the representations learned by the model, and thereby enhancing the optimization performance of the model. Due to the introduction of the central pixel reconstruction strategy, the representations learned by the proposed SSMAE-CR have a larger mean inter-class distance, thereby proving that the representations it has learned have good separability. What is the implication of the main finding? In the semi-supervised hyperspectral image classification method based on the MAE, the importance of reconstructing the central pixel should be emphasized. It can help the model enhance the separability of the representations it has learned, thereby improving the subsequent tuning performance of the model. It is necessary to conduct representation learning in both spectral and spatial branches. This strategy can prevent the model from learning the representation of samples only from one dimension and ignoring the complementary representation of the other dimension. Masked auto-encoders (MAEs) have been extensively employed in the field of semi-supervised hyperspectral image classification (HSIC). However, the developed models encounter significant challenges in learning separable representations, as they do not sufficiently prioritize the reconstruction of the central pixel, which hinders their ability to learn separable representations. To address this limitation, we propose spectral–spatial MAE with central pixel reconstruction (SSMAE-CR), a novel self-supervised framework tailored for HSIC. To capture more comprehensive representations, SSMAE-CR employs a dual-branch architecture comprising a spectral MAE with central pixel reconstruction (SpecMAE-CR) and a spatial MAE with central pixel reconstruction (SpatMAE-CR). SpecMAE-CR highlights the significance of central pixel reconstruction by measuring the deviation between the central pixels of reconstructed and original samples. To preserve the holism of the learned latent representations, SpatMAE-CR maps the central pixels of the reconstructed samples back to their original counterparts through the introduction of an additional linear layer. Rigorous comparative experiments conducted on four publicly available datasets fully demonstrate that SSMAE-CR outperforms state-of-the-art methods. Furthermore, we validate the effectiveness of SSMAE-CR by evaluating the mean intra-class and inter-class distances of the learned representations. Experimental results demonstrate that prioritizing central pixel reconstruction yields a statistically significant increase in the mean inter-class distance, suggesting enhanced class separability in the representation space. [ABSTRACT FROM AUTHOR]
Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:Highlights: What are the main findings? In the proposed spectral–spatial masked auto-encoder with central pixel reconstruction (SSMAE-CR), we introduced central pixel reconstruction in both the spectral and spatial branches, effectively improving the separability of the representations learned by the model, and thereby enhancing the optimization performance of the model. Due to the introduction of the central pixel reconstruction strategy, the representations learned by the proposed SSMAE-CR have a larger mean inter-class distance, thereby proving that the representations it has learned have good separability. What is the implication of the main finding? In the semi-supervised hyperspectral image classification method based on the MAE, the importance of reconstructing the central pixel should be emphasized. It can help the model enhance the separability of the representations it has learned, thereby improving the subsequent tuning performance of the model. It is necessary to conduct representation learning in both spectral and spatial branches. This strategy can prevent the model from learning the representation of samples only from one dimension and ignoring the complementary representation of the other dimension. Masked auto-encoders (MAEs) have been extensively employed in the field of semi-supervised hyperspectral image classification (HSIC). However, the developed models encounter significant challenges in learning separable representations, as they do not sufficiently prioritize the reconstruction of the central pixel, which hinders their ability to learn separable representations. To address this limitation, we propose spectral–spatial MAE with central pixel reconstruction (SSMAE-CR), a novel self-supervised framework tailored for HSIC. To capture more comprehensive representations, SSMAE-CR employs a dual-branch architecture comprising a spectral MAE with central pixel reconstruction (SpecMAE-CR) and a spatial MAE with central pixel reconstruction (SpatMAE-CR). SpecMAE-CR highlights the significance of central pixel reconstruction by measuring the deviation between the central pixels of reconstructed and original samples. To preserve the holism of the learned latent representations, SpatMAE-CR maps the central pixels of the reconstructed samples back to their original counterparts through the introduction of an additional linear layer. Rigorous comparative experiments conducted on four publicly available datasets fully demonstrate that SSMAE-CR outperforms state-of-the-art methods. Furthermore, we validate the effectiveness of SSMAE-CR by evaluating the mean intra-class and inter-class distances of the learned representations. Experimental results demonstrate that prioritizing central pixel reconstruction yields a statistically significant increase in the mean inter-class distance, suggesting enhanced class separability in the representation space. [ABSTRACT FROM AUTHOR]
ISSN:20724292
DOI:10.3390/rs18101571