An image retrieval method based on semantic matching with multiple positional representations

Li, Chunye; Zhou, Zhiping; Zhang, Wei

doi:10.1007/s11042-019-08165-0

An image retrieval method based on semantic matching with multiple positional representations

Published: 10 October 2019

Volume 78, pages 35607–35631, (2019)
Cite this article

Multimedia Tools and Applications Aims and scope Submit manuscript

Chunye Li¹,
Zhiping Zhou^1,2 &
Wei Zhang¹

218 Accesses
2 Citations
Explore all metrics

Abstract

Text-based image retrieval requires manual annotation or automatic labeling of the machine. Manual annotation is time-consuming, and simple text description is difficult to fully express the content of the image. Existing deep models rely on the representation of a single sentence, and such methods cannot well capture the contextualized local information in the matching process. In response to these problems, this paper presents a new retrieval idea based on image caption. First, the image description sentences of images are generated by using the image caption model. Then, for the sentence matching model, we propose a multiple positional representations semantic matching model. We use two interrelated Bi-LSTMs and the attention mechanism to match sentences. the matching score is finally produced by aggregating interactions between these different positional sentence representations. The sentence matching model is used to match the retrieval sentence with the image description sentences in the image library. In our experiments, the accuracy of the proposed image caption model and the sentence matching model are all improved compared with the competitive models, and our method can complete the image retrieval task.

This is a preview of subscription content, log in via an institution to check access.

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

Deep Learning-Based Image Retrieval System with Clustering on Attention-Based Representations

Article 30 March 2021

Deep Convolutional Neural Network for Correlating Images and Sentences

Key Words Extraction and Semantic-Based Image Retrieval on RNNs

References

Bahdanau D, Cho K, Bengio Y (2014) Neural machine translation by jointly learning to align and translate. Comput Sci
Berger A, Caruana R, Cohn D, Freitag D, Mittal V (2000) Bridging the lexical chasm: statistical approaches to answer-finding. In: International ACM SIGIR conference on research and development in information retrieval, pp 192–199
Blacoe W, Lapata M (2012) A comparison of vector-based representations for semantic composition. In: Proceedings of the 2012 joint conference on empirical methods in natural language processing and computational natural language learning, EMNLP-CoNLL 2012, July 12-14, 2012, Jeju Island, Korea, pp 546–556
Collobert R, Weston J (2008) A unified architecture for natural language processing: deep neural networks with multitask learning. In: Machine learning, Proceedings of the twenty-fifth international conference (ICML 2008), Helsinki, Finland, June 5-9, 2008, pp 160–167
Ding G, Chen M, Zhao S, Chen H, Han J, Liu Q (2018) Neural image caption generation with weighted training and reference. Cogn Comput
Dolan B, Quirk C, Brockett C (2004) Unsupervised construction of large paraphrase corpora: exploiting massively parallel news sources. In: International conference on computational linguistics, p 350
Eakins JP (1996) Automatic image content retrieval - are we getting anywhere? De Montfort University Milton Keynes (1): 123–135
Fang H, Gupta S, Iandola FN, Srivastava RK, Deng L, Dollár P, Gao J, He X, Mitchell M, Platt JC, Zitnick CL, Zweig G (2015) From captions to visual concepts and back. In: IEEE conference on computer vision and pattern recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pp 1473–1482
Ferreira R, Cavalcanti GDC, Freitas F, Lins RD, Simske SJ, Riss M (2018) Combining sentence similarities measures to identify paraphrases. Comput Speech Lang 47:59–73
Article Google Scholar
Harmandas V, Sanderson M, Dunlop MD (1997) Image retrieval by hypertext links. Acm Sigir Forum 31(SI):296–303
Article Google Scholar
He K, Zhang X, Ren S, Sun J (2016) Deep residual learning for image recognition. In: 2016 IEEE conference on computer vision and pattern recognition, CVPR 2016, las vegas, NV, USA, June 27-30, 2016, pp 770–778
Hermann KM, Kočiský T, Grefenstette E, Espeholt L, Kay W, Suleyman M, Blunsom P (2015) Teaching machines to read and comprehend : 1693–1701
Hu B, Lu Z, Li H, Chen Q (2015) Convolutional neural network architectures for matching natural language sentences. Adv Neural Inf Proces Syst 3:2042–2050
Google Scholar
Huang PS, He X, Gao J, Deng L, Acero A, Heck L (2013) Learning deep structured semantic models for web search using clickthrough data. In: ACM international conference on conference on information & knowledge management, pp 2333–2338
Jia X, Gavves E, Fernando B, Tuytelaars T (2015) Guiding the long-short term memory model for image caption generation. In: 2015 IEEE international conference on computer vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pp 2407–2415
Karpathy A, Fei-Fei L (2017) Deep visual-semantic alignments for generating image descriptions. IEEE Trans Pattern Anal Mach Intell 39(4):664–676
Article Google Scholar
Kim Y (2014) Convolutional neural networks for sentence classification. Eprint Arxiv
Kiros R, Zhu Y, Salakhutdinov R, Zemel RS, Urtasun R, Torralba A, Fidler S (2015) Skip-thought vectors. In: Advances in neural information processing systems 28: annual conference on neural information processing systems 2015, December 7-12, 2015, Montreal, Quebec, Canada, pp 3294–3302
Li H, Xu J (2014) Semantic matching in search. Found Trends Inf Retr 7 (5):343–469
Article Google Scholar
Li YN, Wang P, Su YT (2015) Robust image hashing based on selective quaternion invariance. IEEE Signal Process Lett 22(12):2396–2400
Article Google Scholar
Liang X, Shen X, Feng J, Lin L, Yan S (2016) Semantic object parsing with graph LSTM. In: Computer vision - ECCV 2016 - 14th European conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part I, pp 125–143
Chapter Google Scholar
Lin T, Maire M, Belongie SJ, Hays J, Perona P, Ramanan D, Dollár P, Zitnick CL (2014) Microsoft COCO: common objects in context. In: Computer vision - ECCV 2014 - 13th European conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V, pp 740–755
Chapter Google Scholar
Liu L, Finch AM, Utiyama M, Sumita E (2016) Agreement on target-bidirectional lstms for sequence-to-sequence learning. In: Proceedings of the thirtieth AAAI conference on artificial intelligence, February 12-17, 2016, Phoenix, Arizona, USA, pp 2630–2637
Liu B, Zhang T, Han FX, Niu D, Lai K, Xu Y (2018) Matching natural language sentences with hierarchical sentence factorization. In: Proceedings of the 2018 world wide web conference on world wide web, WWW 2018, Lyon, France, April 23-27, 2018, pp 1237–1246
Mao J, Xu W, Yang Y, Wang J, Yuille AL (2015) Deep captioning with multimodal recurrent neural networks (m-rnn). In: 3Rd international conference on learning representations, ICLR 2015, San Diego, CA, USA, May 7-9, 2015, Conference Track Proceedings
Mikolov T, Sutskever I, Chen K, Corrado GS, Dean J (2013) Distributed representations of words and phrases and their compositionality. In: Advances in neural information processing systems 26: 27th annual conference on neural information processing systems 2013. Proceedings of a meeting held December 5-8, 2013, Lake Tahoe, Nevada, USA, pp 3111–3119
Palangi H, Deng L, Shen Y, Gao J, He X, Chen J, Song X, Ward R (2016) Deep sentence embedding using long short-term memory networks: analysis and application to information retrieval. IEEE/ACM Trans Audio Speech Lang Process 24(4):694–707
Article Google Scholar
Pennington J, Socher R, Manning C (2014) Glove: global vectors for word representation. In: Conference on empirical methods in natural language processing, pp 1532–1543
Piplani T, Bamman D (2018) Deepseek: content based image search & retrieval. CoRR arXiv:1801.03406
Plummer BA, Wang L, Cervantes CM, Caicedo JC, Hockenmaier J, Lazebnik S (2015) Flickr30k entities: collecting region-to-phrase correspondences for richer image-to-sentence models. In: 2015 IEEE International conference on computer vision, ICCV 2015, Santiago, Chile, December 7-13, 2015, pp 2641–2649
Qin C, Chen X, Luo X, Zhang X, Sun X (2018) Perceptual image hashing via dual-cross pattern encoding and salient structure detection. Inf Sci 423:284–302
Article MathSciNet Google Scholar
Qiu X, Huang X (2015) Convolutional neural tensor network architecture for community-based question answering. In: International conference on artificial intelligence, pp 1305–1311
Qu S, Xi Y, Ding S (2017) Visual attention based on long-short term memory model for image caption generation. In: 2017 29th Chinese control and decision conference (CCDC), pp 4789–4794
Rocktäschel T, Grefenstette E, Hermann KM, Kočiský T, Blunsom P (2015) Reasoning about entailment with neural attention. CoRR
Shen Y, He X, Gao J, Deng L, Mesnil G (2014) Learning semantic representations using convolutional neural network for web search. Proc Www: 373–374
Shetty R, Rohrbach M, Hendricks LA, Fritz M, Schiele B (2017) Speaking the same language: matching machine to human captions by adversarial training. In: IEEE International conference on computer vision, ICCV 2017, Venice, Italy, October 22-29, 2017, pp 4155–4164
Smeulders AWM, Worring M, Santini S, Gupta A, Jain R (2000) Content-based image retrieval: the end of the early years. IEEE Trans Pattern Anal Mach Intell 22(12):1348
Article Google Scholar
Socher R, Chen D, Manning CD, Ng AY (2013) Reasoning with neural tensor networks for knowledge base completion. In: International conference on neural information processing systems, pp 926–934
Vinyals O, Toshev A, Bengio S, Erhan D (2015) Show and tell: a neural image caption generator. In: IEEE Conference on computer vision and pattern recognition, CVPR 2015, Boston, MA, USA, June 7-12, 2015, pp 3156–3164
Wan J, Wang D, Hoi SCH, Wu P, Zhu J, Zhang Y, Li J (2014) Deep learning for content-based image retrieval: a comprehensive study (FullPaper) 157–166
Wan S, Lan Y, Guo J, Xu J, Pang L, Cheng X (2015) A deep architecture for semantic matching with multiple positional sentence representations. CoRR, 2835–2841
Wan S, Lan Y, Guo J, Xu J, Pang L, Cheng X (2016) A deep architecture for semantic matching with multiple positional sentence representations. In: Proceedings of the thirtieth AAAI conference on artificial intelligence, February 12-17, 2016, Phoenix, Arizona, USA, pp 2835–2841
Wu Q, Shen C, Liu L, Dick AR, van den Hengel A (2016) What value do explicit high level concepts have in vision to language problems?. In: 2016 IEEE Conference on computer vision and pattern recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp 203–212
Xiaojun BI, Pan T (2017) Image retrieval method with relevance feedback based on improved teaching-learning-based optimization algorithm. Syst Eng Electron 39(10):2359–2367
Google Scholar
Xu K, Ba J, Kiros R, Cho K, Courville AC, Salakhutdinov R, Zemel RS, Bengio Y (2015) Show, attend and tell: neural image caption generation with visual attention. In: Proceedings of the 32nd international conference on machine learning, ICML 2015, Lille, France, 6-11 July 2015, pp 2048–2057
Yang J, Yu K, Gong Y, Huang T (2009) Linear spatial pyramid matching using sparse coding for image classification. Cvpr, 1794–1801
Yang Z, Yuan Y, Wu Y, Salakhutdinov R, Cohen WW (2016) Encode, review, and decode: reviewer module for caption generation. CoRR arXiv:1605.07912
Yao H, Liu H, Zhang P (2018) A novel sentence similarity model with word embedding based on convolutional neural network. Concurrency and Computation: Practice and Experience. 30(23)
Yin W, Schütze H (2015) Multigrancnn: an architecture for general matching of text chunks on multiple levels of granularity. In: Meeting of the association for computational linguistics and the international joint conference on natural language processing, pp 63–73
Yin W, Schütze H (2015) Convolutional neural network for paraphrase identification. In: NAACL HLT 2015, The 2015 conference of the north american chapter of the association for computational linguistics: human language technologies, Denver, Colorado, USA, May 31 - June 5, 2015, pp 901–911
Yin W, Schütze H, Xiang B, Zhou B (2015) Abcnn: attention-based convolutional neural network for modeling sentence pairs. Comput Sci
You Q, Jin H, Wang Z, Fang C, Luo J (2016) Image captioning with semantic attention. In: 2016 IEEE Conference on computer vision and pattern recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016, pp 4651–4659

Download references

Acknowledgements

This work is supported by Postgraduate Research & Practice Innovation Program of Jiangsu Province of the People’s Republic of China under SJCX19_0797.

Author information

Authors and Affiliations

School of Internet of Things Engineering, Jiangnan University, Wuxi, 214122, People’s Republic of China
Chunye Li, Zhiping Zhou & Wei Zhang
Engineering Research Center of Internet of Things Technology Applications Ministry of Education, Jiangnan University, Wuxi, 214122, People’s Republic of China
Zhiping Zhou

Authors

Chunye Li
View author publications
You can also search for this author in PubMed Google Scholar
Zhiping Zhou
View author publications
You can also search for this author in PubMed Google Scholar
Wei Zhang
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Zhiping Zhou.

Additional information

Publisher’s note

Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.

Rights and permissions

Reprints and permissions

About this article

Cite this article

Li, C., Zhou, Z. & Zhang, W. An image retrieval method based on semantic matching with multiple positional representations. Multimed Tools Appl 78, 35607–35631 (2019). https://doi.org/10.1007/s11042-019-08165-0

Download citation

Received: 08 October 2018
Revised: 14 July 2019
Accepted: 02 September 2019
Published: 10 October 2019
Issue Date: December 2019
DOI: https://doi.org/10.1007/s11042-019-08165-0

Keywords

Access this article

Log in via an institution

Price excludes VAT (USA)
Tax calculation will be finalised during checkout.

Instant access to the full article PDF.

Institutional subscriptions

An image retrieval method based on semantic matching with multiple positional representations

Abstract

Access this article

Similar content being viewed by others

Deep Learning-Based Image Retrieval System with Clustering on Attention-Based Representations

Deep Convolutional Neural Network for Correlating Images and Sentences

Key Words Extraction and Semantic-Based Image Retrieval on RNNs

References

Acknowledgements

Author information

Authors and Affiliations

Corresponding author

Additional information

Publisher’s note

Rights and permissions

About this article

Cite this article

Keywords

Navigation

An image retrieval method based on semantic matching with multiple positional representations

Abstract

Access this article

Similar content being viewed by others

Deep Learning-Based Image Retrieval System with Clustering on Attention-Based Representations

Deep Convolutional Neural Network for Correlating Images and Sentences

Key Words Extraction and Semantic-Based Image Retrieval on RNNs

References

Acknowledgements

Author information

Authors and Affiliations

Corresponding author

Additional information

Publisher’s note

Rights and permissions

About this article

Cite this article

Share this article

Keywords

Search

Navigation