RPAFormer: Building Extraction with Relative Position Aggregated Transformer.

Saved in:
Bibliographic Details
Title: RPAFormer: Building Extraction with Relative Position Aggregated Transformer.
Authors: Xing, Juehui1,2,3 (AUTHOR), Yao, Siyuan2,4 (AUTHOR), Zhu, Zhongyi3 (AUTHOR), Zhang, Lingxin1,2,4 (AUTHOR) zhanglingxin@iem.ac.cn
Source: Remote Sensing. Jun2026, Vol. 18 Issue 11, p1849. 21p.
Subjects: Transformer models, Remote sensing, Spatial arrangement, Urban planning, Image segmentation
Abstract: Highlights: What are the main findings? Develops a novel pure transformer-based building extraction framework named RPAFormer, which is capable of flexibly adapting to the diverse structure variations of buildings and producing accurate local details in complex scenarios. Conducts experiments on public building extraction datasets to verify the effectiveness of RPAFormer. The experimental results demonstrate that RPAFormer achieves a more competitive performance than other state-of-the-art methods. What are the implications of the main findings? Relative Position-aware Self-attention (RPSA) block learns the token dependencies within the local window and can flexibly adapt to the intricate background regions and varied structure patterns of buildings. Transformer decoder network consisting of multiple Cross Masked Attention (CMA) blocks fused with the multi-scale features is capable of modeling the relative position dependencies of the buildings. Automatic building extraction plays an important role in various remote sensing applications, such as seismic disaster investigation, seismic hazard risk assessment, urban planning, and photogrammetry. Despite the substantial progress, state-of-the-art building extraction methods are still limited by two issues: (i) existing approaches leverage convolutional layers or non-local self-attention to encode the position-aware dependencies, while they cannot flexibly adapt to the complex background contexts and varied structure patterns of buildings; and (ii) the local details cannot be well preserved by existing hierarchical decoders due to the imperfect feature aggregation, yielding unsatisfactory segmentation outputs in the local adjacent region. To address these issues, we propose Relative Position Aggregated Transformer (RPAFormer), which is capable of modeling the relative position dependencies of buildings and producing accurate local details using a dual attention transformer network. Specifically, we propose a Relative Position-aware Self-attention (RPSA) framework to learn the token dependencies within the local window. A transformer decoder network consisting of multiple Cross Masked Attention (CMA) blocks is also introduced to fuse the multi-scale features. Extensive experiments demonstrate the superior performance of the proposed method and its great promise for real-world engineering deployment. [ABSTRACT FROM AUTHOR]
Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 194587070
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: RPAFormer: Building Extraction with Relative Position Aggregated Transformer.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Xing%2C+Juehui%22">Xing, Juehui</searchLink><relatesTo>1,2,3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Yao%2C+Siyuan%22">Yao, Siyuan</searchLink><relatesTo>2,4</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zhu%2C+Zhongyi%22">Zhu, Zhongyi</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zhang%2C+Lingxin%22">Zhang, Lingxin</searchLink><relatesTo>1,2,4</relatesTo> (AUTHOR)<i> zhanglingxin@iem.ac.cn</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Remote+Sensing%22">Remote Sensing</searchLink>. Jun2026, Vol. 18 Issue 11, p1849. 21p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Transformer+models%22">Transformer models</searchLink><br /><searchLink fieldCode="DE" term="%22Remote+sensing%22">Remote sensing</searchLink><br /><searchLink fieldCode="DE" term="%22Spatial+arrangement%22">Spatial arrangement</searchLink><br /><searchLink fieldCode="DE" term="%22Urban+planning%22">Urban planning</searchLink><br /><searchLink fieldCode="DE" term="%22Image+segmentation%22">Image segmentation</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Highlights: What are the main findings? Develops a novel pure transformer-based building extraction framework named RPAFormer, which is capable of flexibly adapting to the diverse structure variations of buildings and producing accurate local details in complex scenarios. Conducts experiments on public building extraction datasets to verify the effectiveness of RPAFormer. The experimental results demonstrate that RPAFormer achieves a more competitive performance than other state-of-the-art methods. What are the implications of the main findings? Relative Position-aware Self-attention (RPSA) block learns the token dependencies within the local window and can flexibly adapt to the intricate background regions and varied structure patterns of buildings. Transformer decoder network consisting of multiple Cross Masked Attention (CMA) blocks fused with the multi-scale features is capable of modeling the relative position dependencies of the buildings. Automatic building extraction plays an important role in various remote sensing applications, such as seismic disaster investigation, seismic hazard risk assessment, urban planning, and photogrammetry. Despite the substantial progress, state-of-the-art building extraction methods are still limited by two issues: (i) existing approaches leverage convolutional layers or non-local self-attention to encode the position-aware dependencies, while they cannot flexibly adapt to the complex background contexts and varied structure patterns of buildings; and (ii) the local details cannot be well preserved by existing hierarchical decoders due to the imperfect feature aggregation, yielding unsatisfactory segmentation outputs in the local adjacent region. To address these issues, we propose Relative Position Aggregated Transformer (RPAFormer), which is capable of modeling the relative position dependencies of buildings and producing accurate local details using a dual attention transformer network. Specifically, we propose a Relative Position-aware Self-attention (RPSA) framework to learn the token dependencies within the local window. A transformer decoder network consisting of multiple Cross Masked Attention (CMA) blocks is also introduced to fuse the multi-scale features. Extensive experiments demonstrate the superior performance of the proposed method and its great promise for real-world engineering deployment. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=194587070
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.3390/rs18111849
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 21
        StartPage: 1849
    Subjects:
      – SubjectFull: Transformer models
        Type: general
      – SubjectFull: Remote sensing
        Type: general
      – SubjectFull: Spatial arrangement
        Type: general
      – SubjectFull: Urban planning
        Type: general
      – SubjectFull: Image segmentation
        Type: general
    Titles:
      – TitleFull: RPAFormer: Building Extraction with Relative Position Aggregated Transformer.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Xing, Juehui
      – PersonEntity:
          Name:
            NameFull: Yao, Siyuan
      – PersonEntity:
          Name:
            NameFull: Zhu, Zhongyi
      – PersonEntity:
          Name:
            NameFull: Zhang, Lingxin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 20724292
          Numbering:
            – Type: volume
              Value: 18
            – Type: issue
              Value: 11
          Titles:
            – TitleFull: Remote Sensing
              Type: main
ResultId 1