Bibliographic Details
| Title: |
UHDNet: Unified multimodal fusion harmonization and hierarchical dependency learning for visible-infrared person re-identification. |
| Authors: |
Hong, Xiaobin1,2 (AUTHOR), Adam, Tarmizi1 (AUTHOR) tarmizi.adam@utm.my, Ghazali, Masitah1 (AUTHOR) |
| Source: |
Image & Vision Computing. Sep2025, Vol. 161, pN.PAG-N.PAG. 1p. |
| Subjects: |
Comparative studies, Classification, Cameras |
| Abstract: |
Person re-identification focuses on recognizing and matching the same pedestrian across different camera views, which has important applications in intelligent surveillance, intelligent transportation, and other fields. Due to the significant differences within and between modalities, Visible-Infrared Person Re-Identification (VI-ReID) is a highly challenging task. However, existing methods primarily map the two modalities into a shared feature space, relying solely on modality-shared features, which limits the discriminative power of feature representations and overlooks useful information in modality-specific features, such as color and contrast. Therefore, we propose a novel Unified Harmonisation and Dependency Network (UHDNet). Specifically, the Multimodal Fusion Harmonizer (MFH) dynamically models modality-specific and modality-shared features in an instance-guided manner, effectively mitigating style variations within specific modalities while preserving discriminative information. Given the critical role of relationships between different body parts in distinguishing occluded and overlapping pedestrians, we design a Hierarchical Relational Attention Module (HRAM) to capture hierarchical relationships between different parts using a differential global–local similarity matrix. Finally, to distinguish modality-specific and modality-shared feature representations, we introduce instance-level classification and alignment losses, along with modality-level alignment loss, enabling precise identity alignment and modality consistency. To verify the effectiveness of the proposed architecture, we conducted comparative and ablation studies on the public datasets SYSU-MM01, RegDB, and LLCM. Our method achieved superior performance, with mAP scores of 80.5% for all-search and 89.9% for indoor-search on SYSU-MM01, 90.1% for VIS-to-IR and 88.9% for IR-to-VIS on RegDB, and 66.7% for IR-to-VIS and 67.2% for VIS-to-IR on LLCM. • Dynamic fusion harmonizes modality-specific and shared features adaptively. • Hierarchical attention models body part relations for occluded pedestrians. • Novel multimodal losses enhance robustness in cross-modality matching. • Achieves top performance on cross-modality re-ID benchmark datasets. [ABSTRACT FROM AUTHOR] |
|
Copyright of Image & Vision Computing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) |
| Database: |
Engineering Source |