TrDPNet: A transformer-based diffusion model for single-image 3D point cloud reconstruction.

Saved in:
Bibliographic Details
Title: TrDPNet: A transformer-based diffusion model for single-image 3D point cloud reconstruction.
Authors: Li, Fei1 (AUTHOR), Li, Tiansong1 (AUTHOR), Xiao, Ke1 (AUTHOR) xiaoke@cqnu.edu.cn, Wang, Lin1 (AUTHOR), Yu, Li2 (AUTHOR)
Source: Journal of Visual Communication & Image Representation. Sep2025, Vol. 111, pN.PAG-N.PAG. 1p.
Subjects: Multilayer perceptrons, Point cloud, Transformer models, Deep learning, Diffusion control
Abstract: The conditional diffusion model has shown great promise in 3D point cloud reconstruction from single-view image. Nevertheless, it is extremely challenging to effectively utilize the only image information to conditionally control the diffusion model to generate 3D point clouds. Previous methods heavily relied on projecting image information onto 3D point clouds and using PointNet to extract features from them. However, due to the locality of the projection method, PointNet may insufficiently fuse point clouds and image features. In this paper, we present TrDPNet, a novel Transformer-based diffusion model for single-image 3D point cloud reconstruction. TrDPNet integrates image features and point clouds for conditional control using the Transformer to achieve high-quality 3D reconstruction. Firstly, farthest point sampling is applied to identify key points, a sub-point cloud is established within the specified radius, and then the features are mapped to tokens in the high-dimensional space. Secondly, a series of cascaded Transformer blocks is utilized to fuse the image and point cloud information via attention mechanisms, conditionally guiding the diffusion model. This design not only integrates image information across the entire point cloud but also strengthens connections between point clouds. Finally, multi-layer perceptrons and linear interpolation restore the tokens to the original point cloud size, producing the final noisy prediction. The experimental results show that TrDPNet achieves over a 20% improvement on synthetic benchmarks compared to previous state-of-the-art methods. Our code and weights are available at https://github.com/TLab512/TrDPNet. • Proposed TrDPNet for single-view 3D point cloud reconstruction. • Cross-attention enables stable conditional control in diffusion denoising. • Achieves >20% improvement on synthetic benchmarks, generating high-quality outputs. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Visual Communication & Image Representation is the property of Academic Press Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:The conditional diffusion model has shown great promise in 3D point cloud reconstruction from single-view image. Nevertheless, it is extremely challenging to effectively utilize the only image information to conditionally control the diffusion model to generate 3D point clouds. Previous methods heavily relied on projecting image information onto 3D point clouds and using PointNet to extract features from them. However, due to the locality of the projection method, PointNet may insufficiently fuse point clouds and image features. In this paper, we present TrDPNet, a novel Transformer-based diffusion model for single-image 3D point cloud reconstruction. TrDPNet integrates image features and point clouds for conditional control using the Transformer to achieve high-quality 3D reconstruction. Firstly, farthest point sampling is applied to identify key points, a sub-point cloud is established within the specified radius, and then the features are mapped to tokens in the high-dimensional space. Secondly, a series of cascaded Transformer blocks is utilized to fuse the image and point cloud information via attention mechanisms, conditionally guiding the diffusion model. This design not only integrates image information across the entire point cloud but also strengthens connections between point clouds. Finally, multi-layer perceptrons and linear interpolation restore the tokens to the original point cloud size, producing the final noisy prediction. The experimental results show that TrDPNet achieves over a 20% improvement on synthetic benchmarks compared to previous state-of-the-art methods. Our code and weights are available at https://github.com/TLab512/TrDPNet. • Proposed TrDPNet for single-view 3D point cloud reconstruction. • Cross-attention enables stable conditional control in diffusion denoising. • Achieves >20% improvement on synthetic benchmarks, generating high-quality outputs. [ABSTRACT FROM AUTHOR]
ISSN:10473203
DOI:10.1016/j.jvcir.2025.104503