Describe, Spot and Explain: Interpretable Representation Learning for Discriminative Visual Reasoning.

Saved in:
Bibliographic Details
Title: Describe, Spot and Explain: Interpretable Representation Learning for Discriminative Visual Reasoning.
Authors: Lin, Ci-Siang1 (AUTHOR) d08942011@ntu.edu.tw, Wang, Yu-Chiang Frank2 (AUTHOR) ycwang@ntu.edu.tw
Source: IEEE Transactions on Image Processing. 2023, Vol. 32, p2481-2492. 12p.
Subjects: Artificial neural networks, Visual learning, Computer vision, Data visualization, Transformer models, Deep learning
Abstract: Despite the recent success achieved by deep neural networks (DNNs), it remains challenging to disclose/explain the decision-making process from the numerous parameters and complex non-linear functions. To address the problem, explainable AI (XAI) aims to provide explanations corresponding to the learning and prediction processes for deep learning models. In this paper, we propose a novel representation learning framework of Describe, Spot and eXplain (DSX). Based on the architecture of Transformer, our proposed DSX framework is composed of two learning stages, descriptive prototype learning and discriminative prototype discovery. Given an input image, the former stage is designed to derive a set of descriptive representations, while the latter stage further identifies a discriminative subset, offering semantic interpretability for the corresponding classification tasks. While our DSX does not require any ground truth attribute supervision during training, the derived visual representations can be practically associated with physical attributes provided by domain experts. Extensive experiments on fine-grained classification and person re-identification tasks qualitatively and quantitatively verify the use our DSX model for offering semantically practical interpretability with satisfactory recognition performances. [ABSTRACT FROM AUTHOR]
Copyright of IEEE Transactions on Image Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 182093152
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Describe, Spot and Explain: Interpretable Representation Learning for Discriminative Visual Reasoning.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Lin%2C+Ci-Siang%22">Lin, Ci-Siang</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> d08942011@ntu.edu.tw</i><br /><searchLink fieldCode="AR" term="%22Wang%2C+Yu-Chiang+Frank%22">Wang, Yu-Chiang Frank</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> ycwang@ntu.edu.tw</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22IEEE+Transactions+on+Image+Processing%22">IEEE Transactions on Image Processing</searchLink>. 2023, Vol. 32, p2481-2492. 12p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+neural+networks%22">Artificial neural networks</searchLink><br /><searchLink fieldCode="DE" term="%22Visual+learning%22">Visual learning</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+vision%22">Computer vision</searchLink><br /><searchLink fieldCode="DE" term="%22Data+visualization%22">Data visualization</searchLink><br /><searchLink fieldCode="DE" term="%22Transformer+models%22">Transformer models</searchLink><br /><searchLink fieldCode="DE" term="%22Deep+learning%22">Deep learning</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Despite the recent success achieved by deep neural networks (DNNs), it remains challenging to disclose/explain the decision-making process from the numerous parameters and complex non-linear functions. To address the problem, explainable AI (XAI) aims to provide explanations corresponding to the learning and prediction processes for deep learning models. In this paper, we propose a novel representation learning framework of Describe, Spot and eXplain (DSX). Based on the architecture of Transformer, our proposed DSX framework is composed of two learning stages, descriptive prototype learning and discriminative prototype discovery. Given an input image, the former stage is designed to derive a set of descriptive representations, while the latter stage further identifies a discriminative subset, offering semantic interpretability for the corresponding classification tasks. While our DSX does not require any ground truth attribute supervision during training, the derived visual representations can be practically associated with physical attributes provided by domain experts. Extensive experiments on fine-grained classification and person re-identification tasks qualitatively and quantitatively verify the use our DSX model for offering semantically practical interpretability with satisfactory recognition performances. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of IEEE Transactions on Image Processing is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=182093152
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1109/TIP.2023.3268001
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 12
        StartPage: 2481
    Subjects:
      – SubjectFull: Artificial neural networks
        Type: general
      – SubjectFull: Visual learning
        Type: general
      – SubjectFull: Computer vision
        Type: general
      – SubjectFull: Data visualization
        Type: general
      – SubjectFull: Transformer models
        Type: general
      – SubjectFull: Deep learning
        Type: general
    Titles:
      – TitleFull: Describe, Spot and Explain: Interpretable Representation Learning for Discriminative Visual Reasoning.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Lin, Ci-Siang
      – PersonEntity:
          Name:
            NameFull: Wang, Yu-Chiang Frank
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Text: 2023
              Type: published
              Y: 2023
          Identifiers:
            – Type: issn-print
              Value: 10577149
          Numbering:
            – Type: volume
              Value: 32
          Titles:
            – TitleFull: IEEE Transactions on Image Processing
              Type: main
ResultId 1