RPAFormer: Building Extraction with Relative Position Aggregated Transformer.

Saved in:
Bibliographic Details
Title: RPAFormer: Building Extraction with Relative Position Aggregated Transformer.
Authors: Xing, Juehui1,2,3 (AUTHOR), Yao, Siyuan2,4 (AUTHOR), Zhu, Zhongyi3 (AUTHOR), Zhang, Lingxin1,2,4 (AUTHOR) zhanglingxin@iem.ac.cn
Source: Remote Sensing. Jun2026, Vol. 18 Issue 11, p1849. 21p.
Subjects: Transformer models, Remote sensing, Spatial arrangement, Urban planning, Image segmentation
Abstract: Highlights: What are the main findings? Develops a novel pure transformer-based building extraction framework named RPAFormer, which is capable of flexibly adapting to the diverse structure variations of buildings and producing accurate local details in complex scenarios. Conducts experiments on public building extraction datasets to verify the effectiveness of RPAFormer. The experimental results demonstrate that RPAFormer achieves a more competitive performance than other state-of-the-art methods. What are the implications of the main findings? Relative Position-aware Self-attention (RPSA) block learns the token dependencies within the local window and can flexibly adapt to the intricate background regions and varied structure patterns of buildings. Transformer decoder network consisting of multiple Cross Masked Attention (CMA) blocks fused with the multi-scale features is capable of modeling the relative position dependencies of the buildings. Automatic building extraction plays an important role in various remote sensing applications, such as seismic disaster investigation, seismic hazard risk assessment, urban planning, and photogrammetry. Despite the substantial progress, state-of-the-art building extraction methods are still limited by two issues: (i) existing approaches leverage convolutional layers or non-local self-attention to encode the position-aware dependencies, while they cannot flexibly adapt to the complex background contexts and varied structure patterns of buildings; and (ii) the local details cannot be well preserved by existing hierarchical decoders due to the imperfect feature aggregation, yielding unsatisfactory segmentation outputs in the local adjacent region. To address these issues, we propose Relative Position Aggregated Transformer (RPAFormer), which is capable of modeling the relative position dependencies of buildings and producing accurate local details using a dual attention transformer network. Specifically, we propose a Relative Position-aware Self-attention (RPSA) framework to learn the token dependencies within the local window. A transformer decoder network consisting of multiple Cross Masked Attention (CMA) blocks is also introduced to fuse the multi-scale features. Extensive experiments demonstrate the superior performance of the proposed method and its great promise for real-world engineering deployment. [ABSTRACT FROM AUTHOR]
Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:Highlights: What are the main findings? Develops a novel pure transformer-based building extraction framework named RPAFormer, which is capable of flexibly adapting to the diverse structure variations of buildings and producing accurate local details in complex scenarios. Conducts experiments on public building extraction datasets to verify the effectiveness of RPAFormer. The experimental results demonstrate that RPAFormer achieves a more competitive performance than other state-of-the-art methods. What are the implications of the main findings? Relative Position-aware Self-attention (RPSA) block learns the token dependencies within the local window and can flexibly adapt to the intricate background regions and varied structure patterns of buildings. Transformer decoder network consisting of multiple Cross Masked Attention (CMA) blocks fused with the multi-scale features is capable of modeling the relative position dependencies of the buildings. Automatic building extraction plays an important role in various remote sensing applications, such as seismic disaster investigation, seismic hazard risk assessment, urban planning, and photogrammetry. Despite the substantial progress, state-of-the-art building extraction methods are still limited by two issues: (i) existing approaches leverage convolutional layers or non-local self-attention to encode the position-aware dependencies, while they cannot flexibly adapt to the complex background contexts and varied structure patterns of buildings; and (ii) the local details cannot be well preserved by existing hierarchical decoders due to the imperfect feature aggregation, yielding unsatisfactory segmentation outputs in the local adjacent region. To address these issues, we propose Relative Position Aggregated Transformer (RPAFormer), which is capable of modeling the relative position dependencies of buildings and producing accurate local details using a dual attention transformer network. Specifically, we propose a Relative Position-aware Self-attention (RPSA) framework to learn the token dependencies within the local window. A transformer decoder network consisting of multiple Cross Masked Attention (CMA) blocks is also introduced to fuse the multi-scale features. Extensive experiments demonstrate the superior performance of the proposed method and its great promise for real-world engineering deployment. [ABSTRACT FROM AUTHOR]
ISSN:20724292
DOI:10.3390/rs18111849