UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification.

Saved in:
Bibliographic Details
Title: UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification.
Authors: Hong, Xiaobin1,2 (AUTHOR), Adam, Tarmizi1 (AUTHOR) tarmizi.adam@utm.my, Ghazali, Masitah1 (AUTHOR)
Source: Image & Vision Computing. Sep2025, Vol. 161, pN.PAG-N.PAG. 1p.
Subjects: Comparative studies, Classification, Cameras
Abstract: Person re-identification focuses on recognizing and matching the same pedestrian across different camera views, which has important applications in intelligent surveillance, intelligent transportation, and other fields. Due to the significant differences within and between modalities, Visible-Infrared Person Re-Identification (VI-ReID) is a highly challenging task. However, existing methods primarily map the two modalities into a shared feature space, relying solely on modality-shared features, which limits the discriminative power of feature representations and overlooks useful information in modality-specific features, such as color and contrast. Therefore, we propose a novel Unified Harmonisation and Dependency Network (UHDNet). Specifically, the Multimodal Fusion Harmonizer (MFH) dynamically models modality-specific and modality-shared features in an instance-guided manner, effectively mitigating style variations within specific modalities while preserving discriminative information. Given the critical role of relationships between different body parts in distinguishing occluded and overlapping pedestrians, we design a Hierarchical Relational Attention Module (HRAM) to capture hierarchical relationships between different parts using a differential global–local similarity matrix. Finally, to distinguish modality-specific and modality-shared feature representations, we introduce instance-level classification and alignment losses, along with modality-level alignment loss, enabling precise identity alignment and modality consistency. To verify the effectiveness of the proposed architecture, we conducted comparative and ablation studies on the public datasets SYSU-MM01, RegDB, and LLCM. Our method achieved superior performance, with mAP scores of 80.5% for all-search and 89.9% for indoor-search on SYSU-MM01, 90.1% for VIS-to-IR and 88.9% for IR-to-VIS on RegDB, and 66.7% for IR-to-VIS and 67.2% for VIS-to-IR on LLCM. • Dynamic fusion harmonizes modality-specific and shared features adaptively. • Hierarchical attention models body part relations for occluded pedestrians. • Novel multimodal losses enhance robustness in cross-modality matching. • Achieves top performance on cross-modality re-ID benchmark datasets. [ABSTRACT FROM AUTHOR]
Copyright of Image & Vision Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 187563159
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Hong%2C+Xiaobin%22">Hong, Xiaobin</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Adam%2C+Tarmizi%22">Adam, Tarmizi</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> tarmizi.adam@utm.my</i><br /><searchLink fieldCode="AR" term="%22Ghazali%2C+Masitah%22">Ghazali, Masitah</searchLink><relatesTo>1</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Image+%26+Vision+Computing%22">Image & Vision Computing</searchLink>. Sep2025, Vol. 161, pN.PAG-N.PAG. 1p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Comparative+studies%22">Comparative studies</searchLink><br /><searchLink fieldCode="DE" term="%22Classification%22">Classification</searchLink><br /><searchLink fieldCode="DE" term="%22Cameras%22">Cameras</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Person re-identification focuses on recognizing and matching the same pedestrian across different camera views, which has important applications in intelligent surveillance, intelligent transportation, and other fields. Due to the significant differences within and between modalities, Visible-Infrared Person Re-Identification (VI-ReID) is a highly challenging task. However, existing methods primarily map the two modalities into a shared feature space, relying solely on modality-shared features, which limits the discriminative power of feature representations and overlooks useful information in modality-specific features, such as color and contrast. Therefore, we propose a novel Unified Harmonisation and Dependency Network (UHDNet). Specifically, the Multimodal Fusion Harmonizer (MFH) dynamically models modality-specific and modality-shared features in an instance-guided manner, effectively mitigating style variations within specific modalities while preserving discriminative information. Given the critical role of relationships between different body parts in distinguishing occluded and overlapping pedestrians, we design a Hierarchical Relational Attention Module (HRAM) to capture hierarchical relationships between different parts using a differential global–local similarity matrix. Finally, to distinguish modality-specific and modality-shared feature representations, we introduce instance-level classification and alignment losses, along with modality-level alignment loss, enabling precise identity alignment and modality consistency. To verify the effectiveness of the proposed architecture, we conducted comparative and ablation studies on the public datasets SYSU-MM01, RegDB, and LLCM. Our method achieved superior performance, with mAP scores of 80.5% for all-search and 89.9% for indoor-search on SYSU-MM01, 90.1% for VIS-to-IR and 88.9% for IR-to-VIS on RegDB, and 66.7% for IR-to-VIS and 67.2% for VIS-to-IR on LLCM. • Dynamic fusion harmonizes modality-specific and shared features adaptively. • Hierarchical attention models body part relations for occluded pedestrians. • Novel multimodal losses enhance robustness in cross-modality matching. • Achieves top performance on cross-modality re-ID benchmark datasets. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Image & Vision Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=187563159
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.imavis.2025.105628
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 1
        StartPage: N.PAG
    Subjects:
      – SubjectFull: Comparative studies
        Type: general
      – SubjectFull: Classification
        Type: general
      – SubjectFull: Cameras
        Type: general
    Titles:
      – TitleFull: UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Hong, Xiaobin
      – PersonEntity:
          Name:
            NameFull: Adam, Tarmizi
      – PersonEntity:
          Name:
            NameFull: Ghazali, Masitah
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: Sep2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 02628856
          Numbering:
            – Type: volume
              Value: 161
          Titles:
            – TitleFull: Image & Vision Computing
              Type: main
ResultId 1