TrDPNet: A transformer-based diffusion model for single-image 3D point cloud reconstruction.

Saved in:
Bibliographic Details
Title: TrDPNet: A transformer-based diffusion model for single-image 3D point cloud reconstruction.
Authors: Li, Fei1 (AUTHOR), Li, Tiansong1 (AUTHOR), Xiao, Ke1 (AUTHOR) xiaoke@cqnu.edu.cn, Wang, Lin1 (AUTHOR), Yu, Li2 (AUTHOR)
Source: Journal of Visual Communication & Image Representation. Sep2025, Vol. 111, pN.PAG-N.PAG. 1p.
Subjects: Multilayer perceptrons, Point cloud, Transformer models, Deep learning, Diffusion control
Abstract: The conditional diffusion model has shown great promise in 3D point cloud reconstruction from single-view image. Nevertheless, it is extremely challenging to effectively utilize the only image information to conditionally control the diffusion model to generate 3D point clouds. Previous methods heavily relied on projecting image information onto 3D point clouds and using PointNet to extract features from them. However, due to the locality of the projection method, PointNet may insufficiently fuse point clouds and image features. In this paper, we present TrDPNet, a novel Transformer-based diffusion model for single-image 3D point cloud reconstruction. TrDPNet integrates image features and point clouds for conditional control using the Transformer to achieve high-quality 3D reconstruction. Firstly, farthest point sampling is applied to identify key points, a sub-point cloud is established within the specified radius, and then the features are mapped to tokens in the high-dimensional space. Secondly, a series of cascaded Transformer blocks is utilized to fuse the image and point cloud information via attention mechanisms, conditionally guiding the diffusion model. This design not only integrates image information across the entire point cloud but also strengthens connections between point clouds. Finally, multi-layer perceptrons and linear interpolation restore the tokens to the original point cloud size, producing the final noisy prediction. The experimental results show that TrDPNet achieves over a 20% improvement on synthetic benchmarks compared to previous state-of-the-art methods. Our code and weights are available at https://github.com/TLab512/TrDPNet. • Proposed TrDPNet for single-view 3D point cloud reconstruction. • Cross-attention enables stable conditional control in diffusion denoising. • Achieves >20% improvement on synthetic benchmarks, generating high-quality outputs. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Visual Communication & Image Representation is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 187462214
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: TrDPNet: A transformer-based diffusion model for single-image 3D point cloud reconstruction.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Li%2C+Fei%22">Li, Fei</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Li%2C+Tiansong%22">Li, Tiansong</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Xiao%2C+Ke%22">Xiao, Ke</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> xiaoke@cqnu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Wang%2C+Lin%22">Wang, Lin</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Yu%2C+Li%22">Yu, Li</searchLink><relatesTo>2</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Visual+Communication+%26+Image+Representation%22">Journal of Visual Communication & Image Representation</searchLink>. Sep2025, Vol. 111, pN.PAG-N.PAG. 1p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Multilayer+perceptrons%22">Multilayer perceptrons</searchLink><br /><searchLink fieldCode="DE" term="%22Point+cloud%22">Point cloud</searchLink><br /><searchLink fieldCode="DE" term="%22Transformer+models%22">Transformer models</searchLink><br /><searchLink fieldCode="DE" term="%22Deep+learning%22">Deep learning</searchLink><br /><searchLink fieldCode="DE" term="%22Diffusion+control%22">Diffusion control</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The conditional diffusion model has shown great promise in 3D point cloud reconstruction from single-view image. Nevertheless, it is extremely challenging to effectively utilize the only image information to conditionally control the diffusion model to generate 3D point clouds. Previous methods heavily relied on projecting image information onto 3D point clouds and using PointNet to extract features from them. However, due to the locality of the projection method, PointNet may insufficiently fuse point clouds and image features. In this paper, we present TrDPNet, a novel Transformer-based diffusion model for single-image 3D point cloud reconstruction. TrDPNet integrates image features and point clouds for conditional control using the Transformer to achieve high-quality 3D reconstruction. Firstly, farthest point sampling is applied to identify key points, a sub-point cloud is established within the specified radius, and then the features are mapped to tokens in the high-dimensional space. Secondly, a series of cascaded Transformer blocks is utilized to fuse the image and point cloud information via attention mechanisms, conditionally guiding the diffusion model. This design not only integrates image information across the entire point cloud but also strengthens connections between point clouds. Finally, multi-layer perceptrons and linear interpolation restore the tokens to the original point cloud size, producing the final noisy prediction. The experimental results show that TrDPNet achieves over a 20% improvement on synthetic benchmarks compared to previous state-of-the-art methods. Our code and weights are available at https://github.com/TLab512/TrDPNet. • Proposed TrDPNet for single-view 3D point cloud reconstruction. • Cross-attention enables stable conditional control in diffusion denoising. • Achieves >20% improvement on synthetic benchmarks, generating high-quality outputs. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of Visual Communication & Image Representation is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=187462214
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.jvcir.2025.104503
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 1
        StartPage: N.PAG
    Subjects:
      – SubjectFull: Multilayer perceptrons
        Type: general
      – SubjectFull: Point cloud
        Type: general
      – SubjectFull: Transformer models
        Type: general
      – SubjectFull: Deep learning
        Type: general
      – SubjectFull: Diffusion control
        Type: general
    Titles:
      – TitleFull: TrDPNet: A transformer-based diffusion model for single-image 3D point cloud reconstruction.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Li, Fei
      – PersonEntity:
          Name:
            NameFull: Li, Tiansong
      – PersonEntity:
          Name:
            NameFull: Xiao, Ke
      – PersonEntity:
          Name:
            NameFull: Wang, Lin
      – PersonEntity:
          Name:
            NameFull: Yu, Li
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: Sep2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 10473203
          Numbering:
            – Type: volume
              Value: 111
          Titles:
            – TitleFull: Journal of Visual Communication & Image Representation
              Type: main
ResultId 1