Graph Representation and Prototype Learning for webly supervised fine-grained image recognition.

Saved in:
Bibliographic Details
Title: Graph Representation and Prototype Learning for webly supervised fine-grained image recognition.
Authors: Lin, Jiantao1,2 (AUTHOR) jlin695@connect.hkust-gz.edu.cn, Chen, Tianshui1 (AUTHOR) chentianshui@gdut.edu.cn, Chen, Yingcong2 (AUTHOR) yingcongchen@ust.hk, Yang, Zhijing1 (AUTHOR) yzhj@gdut.edu.cn, Gao, YueFang3 (AUTHOR) gaoyuefang@scau.edu.cn
Source: Pattern Recognition Letters. Jul2024, Vol. 183, p78-85. 8p.
Subjects: Representations of graphs, Machine learning, Image recognition (Computer vision), Supervised learning, Prototypes, Holistic education
Abstract: Webly supervised fine-grained image recognition (FGIR) learns to distinguish sub-ordinate categories based on webly-retrieved data, which can dramatically alleviate the dependency on manually annotated labels. This is quite a challenging task due to the heavy noise labels and the inherent dilemma of small inter-class variance and large intra-class variance. Current webly supervised algorithms learn holistic category prototypes to help correct noisy labels but ignore local features that can distinguish different sub-ordinate categories. In this work, we propose a Graph Representation and Prototype Learning (GRPL) framework to automatically mine discriminative local regions and their interactions with holistic image to learn instance graph representation both category graph prototype to help correct noisy labels and retrieve out-of-distribution (OOD) samples. Specifically, an attention-focused module is designed to extract the discriminative regions and then build a structured graph to correlate them with the holistic image for each instance and an identical graph to model holistic-local correlations for each category. Next, we apply two stacked graph convolution networks to explore holistic-local interaction within each graph and across two graphs to learn graph representation for each instance and graph prototype for each category. Finally, the similarities between the instance-level and prototype-level graph representation are learned to help correct noisy labels and exclude OOD samples. Extensive experiments conducted on several datasets show the proposed approach achieves superior performance compared with current leading algorithms. • Framework learning graph-level similarity for webly supervised fine-grained image recognition. • Unified strategy to identify and correct noisy labels. • Superior performance across multiple datasets. [ABSTRACT FROM AUTHOR]
Copyright of Pattern Recognition Letters is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:Webly supervised fine-grained image recognition (FGIR) learns to distinguish sub-ordinate categories based on webly-retrieved data, which can dramatically alleviate the dependency on manually annotated labels. This is quite a challenging task due to the heavy noise labels and the inherent dilemma of small inter-class variance and large intra-class variance. Current webly supervised algorithms learn holistic category prototypes to help correct noisy labels but ignore local features that can distinguish different sub-ordinate categories. In this work, we propose a Graph Representation and Prototype Learning (GRPL) framework to automatically mine discriminative local regions and their interactions with holistic image to learn instance graph representation both category graph prototype to help correct noisy labels and retrieve out-of-distribution (OOD) samples. Specifically, an attention-focused module is designed to extract the discriminative regions and then build a structured graph to correlate them with the holistic image for each instance and an identical graph to model holistic-local correlations for each category. Next, we apply two stacked graph convolution networks to explore holistic-local interaction within each graph and across two graphs to learn graph representation for each instance and graph prototype for each category. Finally, the similarities between the instance-level and prototype-level graph representation are learned to help correct noisy labels and exclude OOD samples. Extensive experiments conducted on several datasets show the proposed approach achieves superior performance compared with current leading algorithms. • Framework learning graph-level similarity for webly supervised fine-grained image recognition. • Unified strategy to identify and correct noisy labels. • Superior performance across multiple datasets. [ABSTRACT FROM AUTHOR]
ISSN:01678655
DOI:10.1016/j.patrec.2024.05.002