R2SCAT-LPR: Rotation-Robust Network with Self- and Cross-Attention Transformers for LiDAR-Based Place Recognition.

Saved in:
Bibliographic Details
Title: R2SCAT-LPR: Rotation-Robust Network with Self- and Cross-Attention Transformers for LiDAR-Based Place Recognition.
Authors: Jiang, Weizhong1 (AUTHOR), Xue, Hanzhang2 (AUTHOR), Si, Shubin1,3 (AUTHOR), Xiao, Liang1 (AUTHOR) xiaoliang@nudt.edu.cn, Zhao, Dawei1,2 (AUTHOR), Zhu, Qi1,3 (AUTHOR), Nie, Yiming1 (AUTHOR), Dai, Bin1 (AUTHOR)
Source: Remote Sensing. Mar2025, Vol. 17 Issue 6, p1057. 25p.
Subjects: Recognition (Psychology), Graphical projection, Transformer models, Mobile robots, Compact spaces (Topology)
Abstract: LiDAR-based place recognition (LPR) is crucial for the navigation and localization of autonomous vehicles and mobile robots in large-scale outdoor environments and plays a critical role in loop closure detection for simultaneous localization and mapping (SLAM). Existing LPR methods, which utilize 2D bird's-eye view (BEV) projections of 3D point clouds, achieve competitive performance in efficiency and recognition accuracy. However, these methods often struggle with capturing global contextual information and maintaining robustness to viewpoint variations. To address these challenges, we propose R2SCAT-LPR, a novel, transformer-based model that leverages self-attention and cross-attention mechanisms to extract rotation-robust place feature descriptors from BEV images. R2SCAT-LPR consists of three core modules: (1) R2MPFE, which employs weight-shared cascaded multi-head self-attention (MHSA) to extract multi-level spatial contextual patch features from both the original BEV image and its randomly rotated counterpart; (2) DSCA, which integrates dual-branch self-attention and multi-head cross-attention (MHCA) to capture intrinsic correspondences between multi-level patch features before and after rotation, enhancing the extraction of rotation-robust local features; and (3) a combined NetVLAD module, which aggregates patch features from both the original feature space and the rotated interaction space into a compact and viewpoint-robust global descriptor. Extensive experiments conducted on the KITTI and NCLT datasets validate the effectiveness of the proposed model, demonstrating its robustness to rotation variations and its generalization ability across diverse scenes and LiDAR sensors types. Furthermore, we evaluate the generalization performance and computational efficiency of R2SCAT-LPR on our self-constructed OffRoad-LPR dataset for off-road autonomous driving, verifying its deployability on resource-constrained platforms. [ABSTRACT FROM AUTHOR]
Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 184100634
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: R2SCAT-LPR: Rotation-Robust Network with Self- and Cross-Attention Transformers for LiDAR-Based Place Recognition.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Jiang%2C+Weizhong%22">Jiang, Weizhong</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Xue%2C+Hanzhang%22">Xue, Hanzhang</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Si%2C+Shubin%22">Si, Shubin</searchLink><relatesTo>1,3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Xiao%2C+Liang%22">Xiao, Liang</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> xiaoliang@nudt.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Zhao%2C+Dawei%22">Zhao, Dawei</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Zhu%2C+Qi%22">Zhu, Qi</searchLink><relatesTo>1,3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Nie%2C+Yiming%22">Nie, Yiming</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Dai%2C+Bin%22">Dai, Bin</searchLink><relatesTo>1</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Remote+Sensing%22">Remote Sensing</searchLink>. Mar2025, Vol. 17 Issue 6, p1057. 25p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Recognition+%28Psychology%29%22">Recognition (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Graphical+projection%22">Graphical projection</searchLink><br /><searchLink fieldCode="DE" term="%22Transformer+models%22">Transformer models</searchLink><br /><searchLink fieldCode="DE" term="%22Mobile+robots%22">Mobile robots</searchLink><br /><searchLink fieldCode="DE" term="%22Compact+spaces+%28Topology%29%22">Compact spaces (Topology)</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: LiDAR-based place recognition (LPR) is crucial for the navigation and localization of autonomous vehicles and mobile robots in large-scale outdoor environments and plays a critical role in loop closure detection for simultaneous localization and mapping (SLAM). Existing LPR methods, which utilize 2D bird's-eye view (BEV) projections of 3D point clouds, achieve competitive performance in efficiency and recognition accuracy. However, these methods often struggle with capturing global contextual information and maintaining robustness to viewpoint variations. To address these challenges, we propose R2SCAT-LPR, a novel, transformer-based model that leverages self-attention and cross-attention mechanisms to extract rotation-robust place feature descriptors from BEV images. R2SCAT-LPR consists of three core modules: (1) R2MPFE, which employs weight-shared cascaded multi-head self-attention (MHSA) to extract multi-level spatial contextual patch features from both the original BEV image and its randomly rotated counterpart; (2) DSCA, which integrates dual-branch self-attention and multi-head cross-attention (MHCA) to capture intrinsic correspondences between multi-level patch features before and after rotation, enhancing the extraction of rotation-robust local features; and (3) a combined NetVLAD module, which aggregates patch features from both the original feature space and the rotated interaction space into a compact and viewpoint-robust global descriptor. Extensive experiments conducted on the KITTI and NCLT datasets validate the effectiveness of the proposed model, demonstrating its robustness to rotation variations and its generalization ability across diverse scenes and LiDAR sensors types. Furthermore, we evaluate the generalization performance and computational efficiency of R2SCAT-LPR on our self-constructed OffRoad-LPR dataset for off-road autonomous driving, verifying its deployability on resource-constrained platforms. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Remote Sensing is the property of MDPI and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=184100634
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.3390/rs17061057
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 25
        StartPage: 1057
    Subjects:
      – SubjectFull: Recognition (Psychology)
        Type: general
      – SubjectFull: Graphical projection
        Type: general
      – SubjectFull: Transformer models
        Type: general
      – SubjectFull: Mobile robots
        Type: general
      – SubjectFull: Compact spaces (Topology)
        Type: general
    Titles:
      – TitleFull: R2SCAT-LPR: Rotation-Robust Network with Self- and Cross-Attention Transformers for LiDAR-Based Place Recognition.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Jiang, Weizhong
      – PersonEntity:
          Name:
            NameFull: Xue, Hanzhang
      – PersonEntity:
          Name:
            NameFull: Si, Shubin
      – PersonEntity:
          Name:
            NameFull: Xiao, Liang
      – PersonEntity:
          Name:
            NameFull: Zhao, Dawei
      – PersonEntity:
          Name:
            NameFull: Zhu, Qi
      – PersonEntity:
          Name:
            NameFull: Nie, Yiming
      – PersonEntity:
          Name:
            NameFull: Dai, Bin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 15
              M: 03
              Text: Mar2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 20724292
          Numbering:
            – Type: volume
              Value: 17
            – Type: issue
              Value: 6
          Titles:
            – TitleFull: Remote Sensing
              Type: main
ResultId 1