Loading [a11y]/accessibility-menu.js
Attend and Imagine: Multi-Label Image Classification With Visual Attention and Recurrent Neural Networks | IEEE Journals & Magazine | IEEE Xplore

Attend and Imagine: Multi-Label Image Classification With Visual Attention and Recurrent Neural Networks


Abstract:

Real images often have multiple labels, i.e., each image is associated with multiple objects or attributes. Compared to single-label image classification, the multilabel ...Show More

Abstract:

Real images often have multiple labels, i.e., each image is associated with multiple objects or attributes. Compared to single-label image classification, the multilabel classification problem is much more challenging due to several issues. At first, multiple objects can be anywhere in the image. Second, the importance of different regions in an image is different, and the regions of interest in a multilabel image can be very different from another one. Finally, multiple labels of an image can have label dependencies due to complex image structures. To address these challenges, in this paper, we propose to predict the labels sequentially by applying the recurrent neural networks (RNNs), which are used to encode the label dependencies. When predicting a specific label, we introduce a dynamic attention mechanism to enable the model to focus on only regions of interest in the image. Two benchmark datasets (i.e., Pascal VOC and MS-COCO) are adopted to demonstrate the effectiveness of our work. Moreover, we construct a new dataset, which includes many semantic dependent labels in each image, to verify the effectiveness of our model. Experimental results show that our method outperforms several state-of-the-arts, especially when predicting some semantic relative labels.
Published in: IEEE Transactions on Multimedia ( Volume: 21, Issue: 8, August 2019)
Page(s): 1971 - 1981
Date of Publication: 23 January 2019

ISSN Information:

Funding Agency:


References

References is not available for this document.