Multi-document localization method based on bottom-up architecture.

Saved in:
Bibliographic Details
Title: Multi-document localization method based on bottom-up architecture.
Authors: Xu, Kun1 (AUTHOR) xkun@chd.edu.cn, Tan, Qiuman1 (AUTHOR) 2023224023@chd.edu.cn, Cheng, Xin1 (AUTHOR) xincheng@chd.edu.cn, Jing, Wancheng1 (AUTHOR) 2024124067@chd.edu.cn, Hu, WenSheng1 (AUTHOR) 2021224049@chd.edu.cn
Source: Multimedia Systems. Aug2026, Vol. 32 Issue 4, p1-14. 14p.
Subjects: Document imaging systems
Abstract: Document localization is a primary step in intelligent document analysis. The document images captured by smartphones in natural scenario inevitably contain multiple documents, but previous methods rarely focus on multi-document localization. In this paper, we propose a multi-document localization method based on bottom-up architecture for unconstrained environments and a comprehensive multi-document dataset for the first time. Specifically, we design parallel high-to-low resolution branches to extract multi-scale features, enhancing the spatial accuracy of corner localization and perform repeated adaptive spatial fusion during feature aggregation to mitigate the inconsistency across different resolutions. We supervise the network at multiple resolutions to handle scale variation, and a unique embedding tag for each corner is generated through an grouping approach. For evaluation, we collect a multi-document dataset for unconstrained environments, including 24,738 document images with various annotations. Extensive experiments on SmartDoc2015 dataset, Desired dataset and our dataset demonstrate that our method outperforms other state-of-the-art methods for both single document and multi-document localization. Source code is available at https://github.com/TanQiuman/Bottom-Up-Multi-document-Localization. [ABSTRACT FROM AUTHOR]
Copyright of Multimedia Systems is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 193529346
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Multi-document localization method based on bottom-up architecture.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Xu%2C+Kun%22">Xu, Kun</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> xkun@chd.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Tan%2C+Qiuman%22">Tan, Qiuman</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> 2023224023@chd.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Cheng%2C+Xin%22">Cheng, Xin</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> xincheng@chd.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Jing%2C+Wancheng%22">Jing, Wancheng</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> 2024124067@chd.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Hu%2C+WenSheng%22">Hu, WenSheng</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> 2021224049@chd.edu.cn</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Multimedia+Systems%22">Multimedia Systems</searchLink>. Aug2026, Vol. 32 Issue 4, p1-14. 14p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Document+imaging+systems%22">Document imaging systems</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Document localization is a primary step in intelligent document analysis. The document images captured by smartphones in natural scenario inevitably contain multiple documents, but previous methods rarely focus on multi-document localization. In this paper, we propose a multi-document localization method based on bottom-up architecture for unconstrained environments and a comprehensive multi-document dataset for the first time. Specifically, we design parallel high-to-low resolution branches to extract multi-scale features, enhancing the spatial accuracy of corner localization and perform repeated adaptive spatial fusion during feature aggregation to mitigate the inconsistency across different resolutions. We supervise the network at multiple resolutions to handle scale variation, and a unique embedding tag for each corner is generated through an grouping approach. For evaluation, we collect a multi-document dataset for unconstrained environments, including 24,738 document images with various annotations. Extensive experiments on SmartDoc2015 dataset, Desired dataset and our dataset demonstrate that our method outperforms other state-of-the-art methods for both single document and multi-document localization. Source code is available at https://github.com/TanQiuman/Bottom-Up-Multi-document-Localization. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Multimedia Systems is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=193529346
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s00530-026-02314-w
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 14
        StartPage: 1
    Subjects:
      – SubjectFull: Document imaging systems
        Type: general
    Titles:
      – TitleFull: Multi-document localization method based on bottom-up architecture.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Xu, Kun
      – PersonEntity:
          Name:
            NameFull: Tan, Qiuman
      – PersonEntity:
          Name:
            NameFull: Cheng, Xin
      – PersonEntity:
          Name:
            NameFull: Jing, Wancheng
      – PersonEntity:
          Name:
            NameFull: Hu, WenSheng
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 08
              Text: Aug2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 09424962
          Numbering:
            – Type: volume
              Value: 32
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Multimedia Systems
              Type: main
ResultId 1