A multi-view 3D perception network based on spatio-temporal fusion.

Saved in:
Bibliographic Details
Title: A multi-view 3D perception network based on spatio-temporal fusion.
Authors: LI, He1,2, CHEN, Pintong1, YU, Rong1, TAN, Beihai3
Source: Computer Engineering & Science / Jisuanji Gongcheng yu Kexue. Nov2025, Vol. 47 Issue 11, p2019-2028. 10p.
Subjects: Autonomous vehicles, Temporal integration, Object recognition (Computer vision), Metadata, Depth perception, Safety, Computer performance
Abstract: As a critical component of autonomous driving, the perception system directly influences a vehicle's comprehension of its surrounding environment and serves as the foundation for achieving safe and reliable autonomous driving. Traditional 2D image detection perception techniques can only provide limited information. While 3D perception offers richer perceptual data, it faces key challenges, including insufficient fusion of spatial information and inadequate utilization of temporal information. This paper proposes and designs a multi-view 3D perception network that integrates spatio-temporal information. This network comprises a multi-view surround 3D perception network and a spatio-temporal fusion network MVSPNet. The multi-view surround perception network efficiently fuses multi-camera image data through precise spatial perspective transformation, constructing a unified bird's eye view spatial representation. This achieves spatial alignment and fusion of data from multiple cameras. Compared to the current advanced monocular baseline model FCOS3D, it achieves a mean average precision (mAP) of 0.343, representing a performance improvement of 14.7%. The spatio-temporal fusion network MVSPNet enables the temporal fusion of multi-view images by integrating multi-frame data. This further significantly enhances the network's performance, and fusing 2 frames of temporal data results in an additional mAP improvement of 10.2%. The experiment results fully demonstrate the advancement of the designed network in effectively fusing multi-view spatial information and temporal information. This study provides an effective solution for enhancing 3D perception of autonomous driving systems in dynamic and complex scenarios, holding significant implications for advancing the development of safe and reliable autonomous driving technology. [ABSTRACT FROM AUTHOR]
Copyright of Computer Engineering & Science / Jisuanji Gongcheng yu Kexue is the property of Computer Engineering & Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 190505683
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A multi-view 3D perception network based on spatio-temporal fusion.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22LI%2C+He%22">LI, He</searchLink><relatesTo>1,2</relatesTo><br /><searchLink fieldCode="AR" term="%22CHEN%2C+Pintong%22">CHEN, Pintong</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22YU%2C+Rong%22">YU, Rong</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22TAN%2C+Beihai%22">TAN, Beihai</searchLink><relatesTo>3</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Computer+Engineering+%26+Science+%2F+Jisuanji+Gongcheng+yu+Kexue%22">Computer Engineering & Science / Jisuanji Gongcheng yu Kexue</searchLink>. Nov2025, Vol. 47 Issue 11, p2019-2028. 10p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Autonomous+vehicles%22">Autonomous vehicles</searchLink><br /><searchLink fieldCode="DE" term="%22Temporal+integration%22">Temporal integration</searchLink><br /><searchLink fieldCode="DE" term="%22Object+recognition+%28Computer+vision%29%22">Object recognition (Computer vision)</searchLink><br /><searchLink fieldCode="DE" term="%22Metadata%22">Metadata</searchLink><br /><searchLink fieldCode="DE" term="%22Depth+perception%22">Depth perception</searchLink><br /><searchLink fieldCode="DE" term="%22Safety%22">Safety</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+performance%22">Computer performance</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: As a critical component of autonomous driving, the perception system directly influences a vehicle's comprehension of its surrounding environment and serves as the foundation for achieving safe and reliable autonomous driving. Traditional 2D image detection perception techniques can only provide limited information. While 3D perception offers richer perceptual data, it faces key challenges, including insufficient fusion of spatial information and inadequate utilization of temporal information. This paper proposes and designs a multi-view 3D perception network that integrates spatio-temporal information. This network comprises a multi-view surround 3D perception network and a spatio-temporal fusion network MVSPNet. The multi-view surround perception network efficiently fuses multi-camera image data through precise spatial perspective transformation, constructing a unified bird's eye view spatial representation. This achieves spatial alignment and fusion of data from multiple cameras. Compared to the current advanced monocular baseline model FCOS3D, it achieves a mean average precision (mAP) of 0.343, representing a performance improvement of 14.7%. The spatio-temporal fusion network MVSPNet enables the temporal fusion of multi-view images by integrating multi-frame data. This further significantly enhances the network's performance, and fusing 2 frames of temporal data results in an additional mAP improvement of 10.2%. The experiment results fully demonstrate the advancement of the designed network in effectively fusing multi-view spatial information and temporal information. This study provides an effective solution for enhancing 3D perception of autonomous driving systems in dynamic and complex scenarios, holding significant implications for advancing the development of safe and reliable autonomous driving technology. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Computer Engineering & Science / Jisuanji Gongcheng yu Kexue is the property of Computer Engineering & Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=190505683
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.3969/j.issn.1007-130X.2025.11.012
    Languages:
      – Code: chi
        Text: Chinese
    PhysicalDescription:
      Pagination:
        PageCount: 10
        StartPage: 2019
    Subjects:
      – SubjectFull: Autonomous vehicles
        Type: general
      – SubjectFull: Temporal integration
        Type: general
      – SubjectFull: Object recognition (Computer vision)
        Type: general
      – SubjectFull: Metadata
        Type: general
      – SubjectFull: Depth perception
        Type: general
      – SubjectFull: Safety
        Type: general
      – SubjectFull: Computer performance
        Type: general
    Titles:
      – TitleFull: A multi-view 3D perception network based on spatio-temporal fusion.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: LI, He
      – PersonEntity:
          Name:
            NameFull: CHEN, Pintong
      – PersonEntity:
          Name:
            NameFull: YU, Rong
      – PersonEntity:
          Name:
            NameFull: TAN, Beihai
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 11
              Text: Nov2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1007130X
          Numbering:
            – Type: volume
              Value: 47
            – Type: issue
              Value: 11
          Titles:
            – TitleFull: Computer Engineering & Science / Jisuanji Gongcheng yu Kexue
              Type: main
ResultId 1