Efficient Non-fused Winograd on GPUs

Wei, Hui; Liu, Enjie; Zhao, Youbing; Yu, Hongqing

doi:10.1007/978-3-030-61864-3_35

Efficient Non-fused Winograd on GPUs

Hui Wei¹⁶,
Enjie Liu¹⁶,
Youbing Zhao¹⁷ &
…
Hongqing Yu¹⁶

Conference paper
First Online: 18 October 2020

1922 Accesses
3 Citations

Part of the book series: Lecture Notes in Computer Science ((LNIP,volume 12221))

Abstract

This paper presents an optimized implementation for Winograd non-fused convolution. Our optimizations comprise application-independent grouped producer-consumer chains and a set of Winograd-specific software techniques, including specialized interface-kernels data format which enhances memory access efficiency; warp specialization and double buffer prefetching which effectively exploit computational resources and memory bandwidth; utilizing “shuffle” instruction which conserves hardware resources. The paper also provides supplementary explanation of Winograds’ tile extraction, which saves memory and computing resources.

The proposed techniques has been evaluated head to head by kernel level in GTX 980 GPU, CUDA 9.2 with a wide range of parameters which meet CNN layers benchmark. Compared with the state-of-the-art Winograd Non-fused convolution in CuDnn 7.6.4 (released in Sept, 2019), our implementation achieves a total speedup of 1.64x.

This is a preview of subscription content, log in via an institution.

Chapter: USD 29.95; Price excludes VAT (USA)

eBook: USD 89.00; Price excludes VAT (USA)

Softcover Book: USD 119.99; Price excludes VAT (USA)

Tax calculation will be finalised at checkout

Purchases are for personal use only

Learn about institutional subscriptions

References

Lavin, A., Gray, S.: Fast algorithms for convolutional neural networks. In: Proceedings of the CVPR 2016, pp. 4013–4021 (2016)
Google Scholar
Xygkis, A., Soudris, D., Papadopoulos, L., Yous, S., Moloney, D.: Efficient winograd-based convolution kernel implementation on edge devices. In: 55th DAC, pp. 1–6 (2018)
Google Scholar
Jordà, M., Valero-Lara, P., Peña, A.J.: Performance evaluation of cuDNN convolution algorithms on NVIDIA Volta GPUs. IEEE Access 7, 70461–70473 (2019)
Article Google Scholar
Xiao, Q., Liang, Y., Lu, L., Yan, S., Tai, Y.W.: Exploring heterogeneous algorithms for accelerating deep convolutional neural networks on FPGAs. In: Proceedings of the 54th Annual Design Automation Conference, p. 62. ACM (2017)
Google Scholar
Abdelfattah, A., Haidar, A., Tomov, S., Dongarra, J.: Performance, design, and autotuning of batched GEMM for GPUs. In: Kunkel, J.M., Balaji, P., Dongarra, J. (eds.) ISC High Performance 2016. LNCS, vol. 9697, pp. 21–38. Springer, Cham (2016). https://doi.org/10.1007/978-3-319-41321-1_2
Chapter Google Scholar
Tan, G., Li, L., Triechle, S., Phillips, E., Bao, Y., Sun, N.: Fast implementation of DGEMM on Fermi GPU. In: HiPC, Networking, Storage and Analysis, pp. 1–11 (2011)
Google Scholar
Jia, L., Liang, Y., Li, X., Lu, L., Yan, S.: Enabling efficient fast convolution algorithms on GPUs via MegaKernels. IEEE Trans. Comput. 69(7), 986–997 (2020)
MathSciNet MATH Google Scholar
Yan, D., Wang, W., Chu, X.: Optimizing batched winograd convolution on GPUs. In: PPoPP, Main Conference (2020)
Google Scholar
Michael, B., Sean, T., Alex, A.: Singe: leveraging warp specialization for high performance on GPUs. In: ACM SIGPLAN Notices, pp. 119–130 (2014)
Google Scholar
Bauer, M., et al.: CudaDMA: optimizing GPU memory bandwidth via warp specialization. In: HiPC, Networking Storage and Analysis (2011)
Google Scholar

Download references

Author information

Authors and Affiliations

University of Bedfordshire, Luton, UK
Hui Wei, Enjie Liu & Hongqing Yu
Communication University of Zhejiang, Hangzhou, China
Youbing Zhao

Authors

Hui Wei
View author publications
You can also search for this author in PubMed Google Scholar
Enjie Liu
View author publications
You can also search for this author in PubMed Google Scholar
Youbing Zhao
View author publications
You can also search for this author in PubMed Google Scholar
Hongqing Yu
View author publications
You can also search for this author in PubMed Google Scholar

Editor information

Editors and Affiliations

University of Geneva, Geneva, Switzerland
Nadia Magnenat-Thalmann
University of Crete, Heraklion, Greece
Constantine Stephanidis
University of Macau, Macau, China
Enhua Wu
Swiss Federal Institute of Technology, Lausanne, Switzerland
Daniel Thalmann
Shanghai Jiao Tong University, Shanghai, China
Bin Sheng
University of Sydney, Sydney, Australia
Jinman Kim
University of Crete, Heraklion, Greece
George Papagiannakis
University of Calgary, Calgary, AB, Canada
Marina Gavrilova

Rights and permissions

Reprints and permissions

Copyright information

About this paper

Cite this paper

Wei, H., Liu, E., Zhao, Y., Yu, H. (2020). Efficient Non-fused Winograd on GPUs. In: Magnenat-Thalmann, N., et al. Advances in Computer Graphics. CGI 2020. Lecture Notes in Computer Science(), vol 12221. Springer, Cham. https://doi.org/10.1007/978-3-030-61864-3_35

Download citation

DOI: https://doi.org/10.1007/978-3-030-61864-3_35
Published: 18 October 2020
Publisher Name: Springer, Cham
Print ISBN: 978-3-030-61863-6
Online ISBN: 978-3-030-61864-3
eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics