Weakly-Supervised Grounding for VQA with Dual Visual-Linguistic Interaction

Liu, Yi; Pan, Junwen; Wang, Qilong; Chen, Guanlin; Nie, Weiguo; Zhang, Yudong; Gao, Qian; Hu, Qinghua; Zhu, Pengfei

doi:10.1007/978-981-99-8850-1_13

Yi Liu^11,12,
Junwen Pan¹¹,
Qilong Wang¹¹,
Guanlin Chen¹¹,
Weiguo Nie¹²,
Yudong Zhang¹²,
Qian Gao¹²,
Qinghua Hu¹¹ &
…
Pengfei Zhu¹¹

Part of the book series: Lecture Notes in Computer Science ((LNAI,volume 14473))

Included in the following conference series:

CAAI International Conference on Artificial Intelligence

166 Accesses

Abstract

Visual question answer (VQA) grounding, aimed at locating the visual evidence associated with the answers while answering questions, has attracted increasing research interest. To locate the evidence, most existing methods extract attention maps in an unsupervised manner from pretrained VQA models. As only the text-related objective is considered during training, the attention map coarsely depicts the grounding region, resulting in poor interpretability. A straightforward solution for improving grounding accuracy is leveraging pixel-wise masks as strong supervision. However, precise per-pixel annotation is time-consuming and labor-intensive. To address above issues, this paper presents the weakly-supervised grounding for VQA, which learns an end-to-end Dual Visual-Linguistic Interaction (DaVi) network in a unified architecture with various low-cost annotations, such as click-, scribble- and box-level grounding labels. Specifically, to enable the visual mask prediction, DaVi proposes a language-based visual decoder that extends the previous VQA network. Since the visual decoder is guided with weak labels, we also present a Pseudo Grounding Refinement Module (PGRM) to refine the relatively coarse predictions as an additional constraint. Extensive experiments demonstrate that our weakly supervised DaVi significantly improves grounding performance even under the click-level supervision with one pixel annotation. Scribble-level supervision achieves 92% performance at a dramatically reduced annotation cost compared to its fully supervised counterpart. More essentially, weak visual grounding usually boosts the accuracy of text answers despite using inaccurate supervision.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Log in via an institution

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 69.99; Price excludes VAT (USA)

Softcover Book: USD 89.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

References

Agrawal, A., Batra, D., Parikh, D., Kembhavi, A., Don’t just assume; look and answer: overcoming priors for visual question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4971–4980 (2018)
Google Scholar
Antol, S., et al.: VQA: visual question answering. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 2425–2433 (2015)
Google Scholar
Chen, C., Anjum, S., Gurari, D.: Grounding answers for visual questions asked by visually impaired people. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19098–19107 (2022)
Google Scholar
Deng, J., Dong, W., Socher, R., Li, L.-J., Li, K., Fei-Fei, L.: Imagenet: a large-scale hierarchical image database. In: 2009 IEEE Conference on Computer Vision and Pattern Recognition, pp. 248–255. IEEE (2009)
Google Scholar
Devlin, J., Chang, M.-W., Lee, K., Toutanova, K.: BERT: pre-training of deep bidirectional transformers for language understanding. In: North American Chapter of the Association for Computational Linguistics (2018)
Google Scholar
Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., Parikh, D.: Making the V in VQA matter: elevating the role of image understanding in visual question answering. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 6904–6913 (2017)
Google Scholar
Gurari, D., et al.: Vizwiz-priv: a dataset for recognizing the presence and purpose of private visual information in images taken by blind people. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 939–948 (2019)
Google Scholar
Johnson, J., Hariharan, B., Van Der Maaten, L., Fei-Fei, L., Zitnick, C.L., Girshick, R.: ClevR: a diagnostic dataset for compositional language and elementary visual reasoning. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2901–2910 (2017)
Google Scholar
Khan, A.U., Kuehne, H., Gan, C., Da Vitoria Lobo, N., Shah, M.: Weakly supervised grounding for VQA in vision-language transformers. In: Avidan, S., Brostow, G., Cissé, M., Farinella, G.M., Hassner, T. (eds.) ECCV 2022. LNCS, vol. 13695, pp. 652–670. Springer, Cham (2022). https://doi.org/10.1007/978-3-031-19833-5_38
Chapter Google Scholar
Krishna, R., et al.: Visual genome: connecting language and vision using crowdsourced dense image annotations. Int. J. Comput. Vision 123(1), 32–73 (2017)
Google Scholar
Li, J., Li, D., Xiong, C., Hoi, S.: Blip: bootstrapping language-image pre-training for unified vision-language understanding and generation. In: ICML (2022)
Google Scholar
Li, J., Selvaraju, R., Gotmare, A., Joty, S., Xiong, C., Hoi, S.C.H.: Align before fuse: Vision and language representation learning with momentum distillation. In: Advances in Neural Information Processing Systems, vol. 34, pp. 9694–9705 (2021)
Google Scholar
Li, X., et al.: Oscar: object-semantics aligned pre-training for vision-language tasks. In: Vedaldi, A., Bischof, H., Brox, T., Frahm, J.-M. (eds.) ECCV 2020. LNCS, vol. 12375, pp. 121–137. Springer, Cham (2020). https://doi.org/10.1007/978-3-030-58577-8_8
Chapter Google Scholar
Loshchilov, I., Hutter, F.: Fixing weight decay regularization in Adam (2017)
Google Scholar
Pan, J., et al.: Tell me the evidence? Dual visual-linguistic interaction for answer grounding. arXiv preprint arXiv:2207.05703 (2022)
Ramakrishnan, S.K., Pal, A., Sharma, G., Mittal, A.: An empirical evaluation of visual question answering for novel objects. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4392–4401 (2017)
Google Scholar
Ren, S., He, K., Girshick, R., Sun, J.: Faster R-CNN: towards real-time object detection with region proposal networks. In: Advances in Neural Information Processing Systems, vol. 28 (2015)
Google Scholar
Ronneberger, O., Fischer, P., Brox, T.: U-net: convolutional networks for biomedical image segmentation. In: Navab, N., Hornegger, J., Wells, W.M., Frangi, A.F. (eds.) MICCAI 2015. LNCS, vol. 9351, pp. 234–241. Springer, Cham (2015). https://doi.org/10.1007/978-3-319-24574-4_28
Chapter Google Scholar
Schuhmann, C., et al.: Laion-5b: an open large-scale dataset for training next generation image-text models. arXiv preprint arXiv:2210.08402 (2022)
Schuhmann, C., et al.: Laion-400m: open dataset of clip-filtered 400 million image-text pairs. In: NeurIPS Workshop Datacentric AI, number FZJ-2022-00923. Jülich Supercomputing Center (2021)
Google Scholar
Selvaraju, R.R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., Batra, D.: Grad-CAM: visual explanations from deep networks via gradient-based localization. In: Proceedings of the IEEE International Conference on Computer Vision, pp. 618–626 (2017)
Google Scholar
Su, H., Jampani, V., Sun, D., Gallo, O., Learned-Miller, E., Kautz, J.: Pixel-adaptive convolutional neural networks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11166–11175 (2019)
Google Scholar
Tan, H., Bansal, M.: LxMERT: learning cross-modality encoder representations from transformers. In: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 5100–5111 (2019)
Google Scholar
Urooj, A., Kuehne, H., Duarte, K., Gan, C., Lobo, N., Shah, M.: Found a reason for me? Weakly-supervised grounded visual question answering using capsules. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8465–8474 (2021)
Google Scholar
Vaswani, A., et al.: Attention is all you need. In: Advances in Neural Information Processing Systems, vol. 30 (2017)
Google Scholar
Xu, D., Zhu, Y., Choy, C.B., Fei-Fei, L., Scene graph generation by iterative message passing. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5410–5419 (2017)
Google Scholar
Xu, K., et al.: Show, attend and tell: neural image caption generation with visual attention. In: International Conference on Machine Learning, pp. 2048–2057. PMLR (2015)
Google Scholar
Yang, Z., Wang, J., Tang, Y., Chen, K., Zhao, H., Torr, P.H.S.: LAVT: language-aware vision transformer for referring image segmentation. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 18155–18165 (2022)
Google Scholar

Download references

Author information

Authors and Affiliations

College of Intelligence and Computing, Tianjin University, Tianjin, China
Yi Liu, Junwen Pan, Qilong Wang, Guanlin Chen, Qinghua Hu & Pengfei Zhu
Baidu Inc, Beijing, China
Yi Liu, Weiguo Nie, Yudong Zhang & Qian Gao

Authors

Yi Liu
View author publications
You can also search for this author in PubMed Google Scholar
Junwen Pan
View author publications
You can also search for this author in PubMed Google Scholar
Qilong Wang
View author publications
You can also search for this author in PubMed Google Scholar
Guanlin Chen
View author publications
You can also search for this author in PubMed Google Scholar
Weiguo Nie
View author publications
You can also search for this author in PubMed Google Scholar
Yudong Zhang
View author publications
You can also search for this author in PubMed Google Scholar
Qian Gao
View author publications
You can also search for this author in PubMed Google Scholar
Qinghua Hu
View author publications
You can also search for this author in PubMed Google Scholar
Pengfei Zhu
View author publications
You can also search for this author in PubMed Google Scholar

Corresponding author

Correspondence to Pengfei Zhu .

Editor information

Editors and Affiliations

Tsinghua University, Beijing, China
Lu Fang
Duke University, Durham, NC, USA
Jian Pei
Shanghai Jiao Tong Univeristy, Shanghai, China
Guangtao Zhai
Chinese Academy of Sciences, Beijing, China
Ruiping Wang

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Liu, Y. et al. (2024). Weakly-Supervised Grounding for VQA with Dual Visual-Linguistic Interaction. In: Fang, L., Pei, J., Zhai, G., Wang, R. (eds) Artificial Intelligence. CICAI 2023. Lecture Notes in Computer Science(), vol 14473. Springer, Singapore. https://doi.org/10.1007/978-981-99-8850-1_13

Download citation

DOI: https://doi.org/10.1007/978-981-99-8850-1_13
Published: 04 February 2024
Publisher Name: Springer, Singapore
Print ISBN: 978-981-99-8849-5
Online ISBN: 978-981-99-8850-1
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics

Weakly-Supervised Grounding for VQA with Dual Visual-Linguistic Interaction