UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification.
Saved in:
| Title: | UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification. |
|---|---|
| Authors: | Hong, Xiaobin1,2 (AUTHOR), Adam, Tarmizi1 (AUTHOR) tarmizi.adam@utm.my, Ghazali, Masitah1 (AUTHOR) |
| Source: | Image & Vision Computing. Sep2025, Vol. 161, pN.PAG-N.PAG. 1p. |
| Subjects: | Comparative studies, Classification, Cameras |
| Abstract: | Person re-identification focuses on recognizing and matching the same pedestrian across different camera views, which has important applications in intelligent surveillance, intelligent transportation, and other fields. Due to the significant differences within and between modalities, Visible-Infrared Person Re-Identification (VI-ReID) is a highly challenging task. However, existing methods primarily map the two modalities into a shared feature space, relying solely on modality-shared features, which limits the discriminative power of feature representations and overlooks useful information in modality-specific features, such as color and contrast. Therefore, we propose a novel Unified Harmonisation and Dependency Network (UHDNet). Specifically, the Multimodal Fusion Harmonizer (MFH) dynamically models modality-specific and modality-shared features in an instance-guided manner, effectively mitigating style variations within specific modalities while preserving discriminative information. Given the critical role of relationships between different body parts in distinguishing occluded and overlapping pedestrians, we design a Hierarchical Relational Attention Module (HRAM) to capture hierarchical relationships between different parts using a differential global–local similarity matrix. Finally, to distinguish modality-specific and modality-shared feature representations, we introduce instance-level classification and alignment losses, along with modality-level alignment loss, enabling precise identity alignment and modality consistency. To verify the effectiveness of the proposed architecture, we conducted comparative and ablation studies on the public datasets SYSU-MM01, RegDB, and LLCM. Our method achieved superior performance, with mAP scores of 80.5% for all-search and 89.9% for indoor-search on SYSU-MM01, 90.1% for VIS-to-IR and 88.9% for IR-to-VIS on RegDB, and 66.7% for IR-to-VIS and 67.2% for VIS-to-IR on LLCM. • Dynamic fusion harmonizes modality-specific and shared features adaptively. • Hierarchical attention models body part relations for occluded pedestrians. • Novel multimodal losses enhance robustness in cross-modality matching. • Achieves top performance on cross-modality re-ID benchmark datasets. [ABSTRACT FROM AUTHOR] |
| Copyright of Image & Vision Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 187563159 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Hong%2C+Xiaobin%22">Hong, Xiaobin</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Adam%2C+Tarmizi%22">Adam, Tarmizi</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> tarmizi.adam@utm.my</i><br /><searchLink fieldCode="AR" term="%22Ghazali%2C+Masitah%22">Ghazali, Masitah</searchLink><relatesTo>1</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Image+%26+Vision+Computing%22">Image & Vision Computing</searchLink>. Sep2025, Vol. 161, pN.PAG-N.PAG. 1p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Comparative+studies%22">Comparative studies</searchLink><br /><searchLink fieldCode="DE" term="%22Classification%22">Classification</searchLink><br /><searchLink fieldCode="DE" term="%22Cameras%22">Cameras</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Person re-identification focuses on recognizing and matching the same pedestrian across different camera views, which has important applications in intelligent surveillance, intelligent transportation, and other fields. Due to the significant differences within and between modalities, Visible-Infrared Person Re-Identification (VI-ReID) is a highly challenging task. However, existing methods primarily map the two modalities into a shared feature space, relying solely on modality-shared features, which limits the discriminative power of feature representations and overlooks useful information in modality-specific features, such as color and contrast. Therefore, we propose a novel Unified Harmonisation and Dependency Network (UHDNet). Specifically, the Multimodal Fusion Harmonizer (MFH) dynamically models modality-specific and modality-shared features in an instance-guided manner, effectively mitigating style variations within specific modalities while preserving discriminative information. Given the critical role of relationships between different body parts in distinguishing occluded and overlapping pedestrians, we design a Hierarchical Relational Attention Module (HRAM) to capture hierarchical relationships between different parts using a differential global–local similarity matrix. Finally, to distinguish modality-specific and modality-shared feature representations, we introduce instance-level classification and alignment losses, along with modality-level alignment loss, enabling precise identity alignment and modality consistency. To verify the effectiveness of the proposed architecture, we conducted comparative and ablation studies on the public datasets SYSU-MM01, RegDB, and LLCM. Our method achieved superior performance, with mAP scores of 80.5% for all-search and 89.9% for indoor-search on SYSU-MM01, 90.1% for VIS-to-IR and 88.9% for IR-to-VIS on RegDB, and 66.7% for IR-to-VIS and 67.2% for VIS-to-IR on LLCM. • Dynamic fusion harmonizes modality-specific and shared features adaptively. • Hierarchical attention models body part relations for occluded pedestrians. • Novel multimodal losses enhance robustness in cross-modality matching. • Achieves top performance on cross-modality re-ID benchmark datasets. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Image & Vision Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=187563159 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1016/j.imavis.2025.105628 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 1 StartPage: N.PAG Subjects: – SubjectFull: Comparative studies Type: general – SubjectFull: Classification Type: general – SubjectFull: Cameras Type: general Titles: – TitleFull: UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Hong, Xiaobin – PersonEntity: Name: NameFull: Adam, Tarmizi – PersonEntity: Name: NameFull: Ghazali, Masitah IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 09 Text: Sep2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 02628856 Numbering: – Type: volume Value: 161 Titles: – TitleFull: Image & Vision Computing Type: main |
| ResultId | 1 |