Dynamic window transformer for three-dimensional indoor scene segmentation.

Saved in:
Bibliographic Details
Title: Dynamic window transformer for three-dimensional indoor scene segmentation.
Authors: Kim, Hyebin1 (AUTHOR), Yoon, Jungho1 (AUTHOR), Yoon, Sang Min1,2 (AUTHOR) smyoon@kookmin.ac.kr
Source: Neurocomputing. Mar2026, Vol. 671, pN.PAG-N.PAG. 1p.
Subjects: Point cloud, Image segmentation, Transformer models, Geometric modeling, Computer graphics
Abstract: Segmenting three-dimensional indoor scenes with complex layouts and object arrangements remains a core challenge in computer graphics and computational photography. We propose a Transformer-based architecture designed for semantic segmentation on point clouds in complex indoor scenes. It explicitly addresses the inherent challenges of data through a dynamic, multi-scale attention mechanism. At the core of the proposed approach is the dynamic window multi-head self-attention (DW-MSA3D) module, which adaptively fuses features captured at varying window scales. Unlike prior approaches that rely on fixed-window attention, our method dynamically adjusts the receptive field to local scene complexity, enabling expressive encoding of sparse volumes across scales. We achieve competitive performance on public datasets, validating the effectiveness of scale-adaptive attention for representing geometric detail in geometry-aware vision tasks. The source code is released at https://github.com/hyebinny/Dawin3D. [ABSTRACT FROM AUTHOR]
Copyright of Neurocomputing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 191350746
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Dynamic window transformer for three-dimensional indoor scene segmentation.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Kim%2C+Hyebin%22">Kim, Hyebin</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Yoon%2C+Jungho%22">Yoon, Jungho</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Yoon%2C+Sang+Min%22">Yoon, Sang Min</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<i> smyoon@kookmin.ac.kr</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Neurocomputing%22">Neurocomputing</searchLink>. Mar2026, Vol. 671, pN.PAG-N.PAG. 1p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Point+cloud%22">Point cloud</searchLink><br /><searchLink fieldCode="DE" term="%22Image+segmentation%22">Image segmentation</searchLink><br /><searchLink fieldCode="DE" term="%22Transformer+models%22">Transformer models</searchLink><br /><searchLink fieldCode="DE" term="%22Geometric+modeling%22">Geometric modeling</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+graphics%22">Computer graphics</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Segmenting three-dimensional indoor scenes with complex layouts and object arrangements remains a core challenge in computer graphics and computational photography. We propose a Transformer-based architecture designed for semantic segmentation on point clouds in complex indoor scenes. It explicitly addresses the inherent challenges of data through a dynamic, multi-scale attention mechanism. At the core of the proposed approach is the dynamic window multi-head self-attention (DW-MSA3D) module, which adaptively fuses features captured at varying window scales. Unlike prior approaches that rely on fixed-window attention, our method dynamically adjusts the receptive field to local scene complexity, enabling expressive encoding of sparse volumes across scales. We achieve competitive performance on public datasets, validating the effectiveness of scale-adaptive attention for representing geometric detail in geometry-aware vision tasks. The source code is released at https://github.com/hyebinny/Dawin3D. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Neurocomputing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=191350746
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.neucom.2026.132746
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 1
        StartPage: N.PAG
    Subjects:
      – SubjectFull: Point cloud
        Type: general
      – SubjectFull: Image segmentation
        Type: general
      – SubjectFull: Transformer models
        Type: general
      – SubjectFull: Geometric modeling
        Type: general
      – SubjectFull: Computer graphics
        Type: general
    Titles:
      – TitleFull: Dynamic window transformer for three-dimensional indoor scene segmentation.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Kim, Hyebin
      – PersonEntity:
          Name:
            NameFull: Yoon, Jungho
      – PersonEntity:
          Name:
            NameFull: Yoon, Sang Min
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 28
              M: 03
              Text: Mar2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 09252312
          Numbering:
            – Type: volume
              Value: 671
          Titles:
            – TitleFull: Neurocomputing
              Type: main
ResultId 1