research-article

Logic Shrinkage: Learned Connectivity Sparsification for LUT-Based Neural Networks

Authors:

Erwei Wang,

Marie Auffret,

Georgios-Ilias Stavrou,

Peter Y. K. Cheung,

George A. Constantinides,

Mohamed S. Abdelfattah,

James J. DavisAuthors Info & Claims

ACM Transactions on Reconfigurable Technology and Systems, Volume 16, Issue 4

Article No.: 57, Pages 1 - 25

https://doi.org/10.1145/3583075

Published: 01 September 2023 Publication History

Get Access

Abstract

Field-programmable gate array (FPGA)–specific deep neural network (DNN) architectures using native lookup tables (LUTs) as independently trainable inference operators have been shown to achieve favorable area-accuracy and energy-accuracy trade-offs. The first work in this area, LUTNet, exhibited state-of-the-art performance for standard DNN benchmarks. In this article, we propose the learned optimization of such LUT-based topologies, resulting in higher-efficiency designs than via the direct use of off-the-shelf, hand-designed networks. Existing implementations of this class of architecture require the manual specification of the number of inputs per LUT, K. Choosing appropriate K a priori is challenging. Doing so at even high granularity, for example, per layer, is a time-consuming and error-prone process that leaves FPGAs’ spatial flexibility underexploited. Furthermore, prior works see LUT inputs connected randomly, which does not guarantee a good choice of network topology. To address these issues, we propose logic shrinkage, a fine-grained netlist pruning methodology enabling K to be automatically learned for every LUT in a neural network targeted for FPGA inference. By removing LUT inputs determined to be of low importance, our method increases the efficiency of the resultant accelerators. Our GPU-friendly solution to LUT input removal is capable of processing large topologies during their training with negligible slowdown. With logic shrinkage, we improve the area and energy efficiency of the best-performing LUTNet implementation of the CNV network classifying CIFAR-10 by 1.54× and 1.31×, respectively, while matching its accuracy. This implementation also reaches 2.71× the area efficiency of an equally accurate, heavily pruned binary neural network (BNN). On ImageNet, with the Bi-Real Net architecture, employment of logic shrinkage results in a post-synthesis area reduction of 2.67× vs. LUTNet, allowing for implementation that was previously impossible on today’s largest FPGAs. We validate the benefits of logic shrinkage in the context of real application deployment by implementing a face mask detection DNN using a BNN, LUTNet, and logic-shrunk layers. Our results show that logic shrinkage results in area gains versus LUTNet (up to 1.20×) and equally pruned BNNs (up to 1.08×), along with accuracy improvements.

References

[1]

Tushar Agrawal, K. Imran, Matteo Figus, and C. Kirkpatrick. 2020. Automatically Detecting Personal Protective Equipment on Persons in Images Using Amazon Rekognition. https://aws.amazon.com/cn/blogs/machine-learning/automaticallydetecting-personal-protective-equipment-on-persons-in-images-using-amazon-rekognition/.

Abstract

References

Cited By

Index Terms

Recommendations

Logic Shrinkage: Learned FPGA Netlist Sparsity for Efficient Neural Network Inference

FracBNN: Accurate and FPGA-Efficient Binary Neural Networks with Fractional Activations

Resource-Aware Saliency-Guided Differentiable Pruning for Deep Neural Networks

Comments

Information

Published In

Publisher

Publication History

Permissions

Check for updates

Author Tags

Qualifiers

Funding Sources

Contributors

Other Metrics

Bibliometrics

Article Metrics

Other Metrics

Citations

Cited By

Login options

Full Access

View options

PDF

eReader

Full Text

Figures

Other

Share

Share this Publication link

Share on social media

Affiliations