Cross-Modal Hashing via Diverse Instances Matching.

Saved in:
Bibliographic Details
Title: Cross-Modal Hashing via Diverse Instances Matching.
Authors: Tu, Junfeng1 ganlantee@gmail.com, Liu, Xueliang1 liuxueliang1982@gmail.com, Huang, Zhen2 huangzhen@nudt.edu.cn, Hao, Yanbin1 haoyanbin@hotmail.com, Hong, Richang1 hongrc.hfut@gmail.com, Wang, Meng1 eric.mengwang@gmail.com
Source: IEEE Transactions on Image Processing. 2025, Vol. 34, p2737-2749. 13p.
Subjects: Hashing, Estimation theory, Information retrieval, Electronic file management, Semantics, Labels
Abstract: Cross-modal hashing is a highly effective technique for searching relevant data across different modalities, owing to its low storage costs and fast similarity retrieval capability. While significant progress has been achieved in this area, prior investigations predominantly concentrate on a one-to-one feature alignment approach, where a singular feature is derived for similarity retrieval. However, the singular feature in these methods fails to adequately capture the varied multi-instance information inherent in the original data across disparate modalities. Consequently, the conventional one-to-one methodology is plagued by a semantic mismatch issue, as the rigid one-to-one alignment inhibits effective multi-instance matching. To address this issue, we propose a novel Diverse Instances Matching for Cross-modal Hashing (DIMCH), which explores the relevance between multiple instances in different modalities using a multi-instance learning algorithm. Specifically, we design a novel diverse instances learning module to extract a multi-feature set, which enables our model to capture detailed multi-instance semantics. To evaluate the similarity between two multi-feature sets, we adopt the smooth chamfer distance function, which enables our model to incorporate the conventional similarity retrieval structure. Moreover, to sufficiently exploit the supervised information from the semantic label, we adopt the weight cosine triplet loss as the objective function, which incorporates the multilevel similarity among the multi-labels into the training procedure and enables the model to mine the multi-label correlation effectively. Extensive experiments demonstrate that our diverse hashing embedding method achieves state-of-the-art performance in supervised cross-modal hashing retrieval tasks. [ABSTRACT FROM AUTHOR]
Copyright of IEEE Transactions on Image Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 191897016
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Cross-Modal Hashing via Diverse Instances Matching.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Tu%2C+Junfeng%22">Tu, Junfeng</searchLink><relatesTo>1</relatesTo><i> ganlantee@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Liu%2C+Xueliang%22">Liu, Xueliang</searchLink><relatesTo>1</relatesTo><i> liuxueliang1982@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Huang%2C+Zhen%22">Huang, Zhen</searchLink><relatesTo>2</relatesTo><i> huangzhen@nudt.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Hao%2C+Yanbin%22">Hao, Yanbin</searchLink><relatesTo>1</relatesTo><i> haoyanbin@hotmail.com</i><br /><searchLink fieldCode="AR" term="%22Hong%2C+Richang%22">Hong, Richang</searchLink><relatesTo>1</relatesTo><i> hongrc.hfut@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Wang%2C+Meng%22">Wang, Meng</searchLink><relatesTo>1</relatesTo><i> eric.mengwang@gmail.com</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22IEEE+Transactions+on+Image+Processing%22">IEEE Transactions on Image Processing</searchLink>. 2025, Vol. 34, p2737-2749. 13p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Hashing%22">Hashing</searchLink><br /><searchLink fieldCode="DE" term="%22Estimation+theory%22">Estimation theory</searchLink><br /><searchLink fieldCode="DE" term="%22Information+retrieval%22">Information retrieval</searchLink><br /><searchLink fieldCode="DE" term="%22Electronic+file+management%22">Electronic file management</searchLink><br /><searchLink fieldCode="DE" term="%22Semantics%22">Semantics</searchLink><br /><searchLink fieldCode="DE" term="%22Labels%22">Labels</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Cross-modal hashing is a highly effective technique for searching relevant data across different modalities, owing to its low storage costs and fast similarity retrieval capability. While significant progress has been achieved in this area, prior investigations predominantly concentrate on a one-to-one feature alignment approach, where a singular feature is derived for similarity retrieval. However, the singular feature in these methods fails to adequately capture the varied multi-instance information inherent in the original data across disparate modalities. Consequently, the conventional one-to-one methodology is plagued by a semantic mismatch issue, as the rigid one-to-one alignment inhibits effective multi-instance matching. To address this issue, we propose a novel Diverse Instances Matching for Cross-modal Hashing (DIMCH), which explores the relevance between multiple instances in different modalities using a multi-instance learning algorithm. Specifically, we design a novel diverse instances learning module to extract a multi-feature set, which enables our model to capture detailed multi-instance semantics. To evaluate the similarity between two multi-feature sets, we adopt the smooth chamfer distance function, which enables our model to incorporate the conventional similarity retrieval structure. Moreover, to sufficiently exploit the supervised information from the semantic label, we adopt the weight cosine triplet loss as the objective function, which incorporates the multilevel similarity among the multi-labels into the training procedure and enables the model to mine the multi-label correlation effectively. Extensive experiments demonstrate that our diverse hashing embedding method achieves state-of-the-art performance in supervised cross-modal hashing retrieval tasks. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of IEEE Transactions on Image Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=191897016
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1109/TIP.2025.3561659
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
        StartPage: 2737
    Subjects:
      – SubjectFull: Hashing
        Type: general
      – SubjectFull: Estimation theory
        Type: general
      – SubjectFull: Information retrieval
        Type: general
      – SubjectFull: Electronic file management
        Type: general
      – SubjectFull: Semantics
        Type: general
      – SubjectFull: Labels
        Type: general
    Titles:
      – TitleFull: Cross-Modal Hashing via Diverse Instances Matching.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Tu, Junfeng
      – PersonEntity:
          Name:
            NameFull: Liu, Xueliang
      – PersonEntity:
          Name:
            NameFull: Huang, Zhen
      – PersonEntity:
          Name:
            NameFull: Hao, Yanbin
      – PersonEntity:
          Name:
            NameFull: Hong, Richang
      – PersonEntity:
          Name:
            NameFull: Wang, Meng
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Text: 2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 10577149
          Numbering:
            – Type: volume
              Value: 34
          Titles:
            – TitleFull: IEEE Transactions on Image Processing
              Type: main
ResultId 1