poster

Convolutional Method for Modeling Video Temporal Context Effectively in Transformer

Authors:

Hae Sung Park,

Yong Suk ChoiAuthors Info & Claims

SAC '23: Proceedings of the 38th ACM/SIGAPP Symposium on Applied Computing

Pages 1205 - 1208

https://doi.org/10.1145/3555776.3578481

Published: 07 June 2023 Publication History

Get Access

Abstract

Video understanding remains a challenging task because video understanding models have many parameters to be trained and should capture detailed spatiotemporal contexts in video effectively. Recent methods have typically employed 3D convolution modules or else self-attention modules. However, we identify that when the self-attention mechanism captures temporal semantics, it often struggles to find out proper temporal context for video understanding. In this paper, we propose a new method for enhancing temporal modeling by incorporating 3D convolution modules into attention-based model, transformer. In particular, we replace the temporal attention of the TimeSformer with a 3D convolution module to improve temporal context learning. In contrast to the TimeSformer, our proposed method can focus on modeling temporal details at the low-level encoders, while gradually getting to focus on temporal contexts more globally at the high-level encoders. Our method surpasses the TimeSformer by 2.2% margin on Something-Something v2, which is required complex temporal modeling for getting high performance.

References

[1]

Anurag Arnab, Mostafa Dehghani, Georg Heigold, Chen Sun, Mario Lučić, and Cordelia Schmid. 2021. Vivit: A video vision transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision. 6836--6846.

Crossref

Google Scholar

[2]

Gedas Bertasius, Heng Wang, and Lorenzo Torresani. 2021. Is space-time attention all you need for video understanding?. In ICML, Vol. 2. 4.

Google Scholar

[3]

Adrian Bulat, Juan Manuel Perez Rua, Swathikiran Sudhakaran, Brais Martinez, and Georgios Tzimiropoulos. 2021. Space-time mixing attention for video transformer. Advances in Neural Information Processing Systems 34 (2021), 19594--19607.

Google Scholar

[4]

Joao Carreira, Eric Noland, Andras Banki-Horvath, Chloe Hillier, and Andrew Zisserman. 2018. A short note about kinetics-600. arXiv preprint arXiv:1808.01340 (2018).

Google Scholar

[5]

Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. 2009. Imagenet: A large-scale hierarchical image database. In 2009 IEEE conference on computer vision and pattern recognition. Ieee, 248--255.

Crossref

Google Scholar

[6]

Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, et al. 2020. An image is worth 16×16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020).

Google Scholar

[7]

Christoph Feichtenhofer, Haoqi Fan, Jitendra Malik, and Kaiming He. 2019. Slow-fast networks for video recognition. In Proceedings of the IEEE/CVF international conference on computer vision. 6202--6211.

Crossref

Google Scholar

[8]

Raghav Goyal, Samira Ebrahimi Kahou, Vincent Michalski, Joanna Materzynska, Susanne Westphal, Heuna Kim, Valentin Haenel, Ingo Fruend, Peter Yianilos, Moritz Mueller-Freitag, et al. 2017. The" something something" video database for learning and evaluating visual common sense. In Proceedings of the IEEE international conference on computer vision. 5842--5850.

Crossref

Google Scholar

[9]

Andrew G Howard, Menglong Zhu, Bo Chen, Dmitry Kalenichenko, Weijun Wang, Tobias Weyand, Marco Andreetto, and Hartwig Adam. 2017. Mobilenets: Efficient convolutional neural networks for mobile vision applications. arXiv preprint arXiv:1704.04861 (2017).

Google Scholar

[10]

Shuiwang Ji, Wei Xu, Ming Yang, and Kai Yu. 2012. 3D convolutional neural networks for human action recognition. IEEE transactions on pattern analysis and machine intelligence 35, 1 (2012), 221--231.

Digital Library

Google Scholar

[11]

Yanghao Li, Chao-Yuan Wu, Haoqi Fan, Karttikeya Mangalam, Bo Xiong, Jitendra Malik, and Christoph Feichtenhofer. 2022. MViTv2: Improved Multiscale Vision Transformers for Classification and Detection. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 4804--4814.

Crossref

Google Scholar

[12]

Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedantam, Devi Parikh, and Dhruv Batra. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE international conference on computer vision. 618--626.

Crossref

Google Scholar

[13]

Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017).

Google Scholar

Index Terms

Convolutional Method for Modeling Video Temporal Context Effectively in Transformer
1. Computing methodologies
  1. Artificial intelligence

Recommendations

Dual-Spatial Normalized Transformer for image captioning
Abstract
Self-attention modules have shown dominance in image captioning. However, current self-attention modules insufficiently consider spatial correlations between objects in an image and easily suffer from distribution shifts. In this work, we aim to ...
Highlights
- A Spatial-Enhanced Attention (SEA) module is proposed to enhance spatial relevance.
- A Gated-Normalized Attention (GNA) module is proposed to fix distributions.
- SEA and GNA module are applied to a Transformer architecture for image ...
Token Shift Transformer for Video Classification
MM '21: Proceedings of the 29th ACM International Conference on Multimedia

Transformer achieves remarkable successes in understanding 1 and 2-dimensional signals (e.g., NLP and Image Content Understanding). As a potential alternative to convolutional neural networks, it shares merits of strong interpretability, high ...
Self-attention-based long temporal sequence modeling method for temporal action detection
Abstract
Temporal Action Detection (TAD) is a basic and complex task in video understanding. It aims at detecting both the localization and category of actions in a video. The anchor-free TAD methods directly predict the action classes at each location ...
Highlights
- Model long temporal sequence for TAD.
- Build a novel framework to model temporal semantics.
- Save computational resource overhead for untrimmed videos in TAD.
- Expensive experiments prove the effectiveness of our method.

Comments

Information & Contributors

Information

Published In

SAC '23: Proceedings of the 38th ACM/SIGAPP Symposium on Applied Computing

March 2023

1932 pages

ISBN:9781450395175

DOI:10.1145/3555776

Conference Chairs:
Jiman Hong
Soongsil University, South Korea
,
Maart Lanperne
Tallinn University, Estonia
,
Program Chairs:
Juw Won Park
University of Louisville, USA
,
Tomas Cerny
Baylor University, USA
,
Publication Chair:
Hossain Shahriar
Kennesaw State University, USA

Permission to make digital or hard copies of part or all of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for third-party components of this work must be honored. For all other uses, contact the owner/author(s).

Publisher

Association for Computing Machinery

New York, NY, United States

Publication History

Published: 07 June 2023

Check for updates

Author Tags

Qualifiers

Poster

Conference

SAC '23

Sponsor:

SIGAPP

SAC '23: 38th ACM/SIGAPP Symposium on Applied Computing

March 27 - 31, 2023

Tallinn, Estonia

Acceptance Rates

Overall Acceptance Rate 1,650 of 6,669 submissions, 25%

Upcoming Conference

SAC '25

Sponsor:
sigapp

The 40th ACM/SIGAPP Symposium on Applied Computing

March 31 - April 4, 2025

Catania , Italy

Contributors

Other Metrics

View Article Metrics

Bibliometrics & Citations

Bibliometrics

Article Metrics

0
Total Citations
69
Total Downloads

Downloads (Last 12 months)37
Downloads (Last 6 weeks)0

Reflects downloads up to 16 Feb 2025

Other Metrics

View Author Metrics

Citations

View Options

Login options

Check if you have access through your login credentials or your institution to get full access on this article.

Full Access

Get this Publication

View options

PDF

View or Download as a PDF file.

PDF

eReader

View online with eReader.

eReader

Abstract

References

Index Terms

Recommendations

Dual-Spatial Normalized Transformer for image captioning

Token Shift Transformer for Video Classification

Self-attention-based long temporal sequence modeling method for temporal action detection

Comments

Information

Published In

Sponsors

Publisher

Publication History

Check for updates

Author Tags

Qualifiers

Conference

Acceptance Rates

Upcoming Conference

Contributors

Other Metrics

Bibliometrics

Article Metrics

Other Metrics

Citations

Login options

Full Access

View options

PDF

eReader

Share

Share this Publication link

Share on social media

Affiliations