Loading [MathJax]/extensions/TeX/ietmacros.js
Heterogeneity-Aware Gradient Coding for Tolerating and Leveraging Stragglers | IEEE Journals & Magazine | IEEE Xplore

Heterogeneity-Aware Gradient Coding for Tolerating and Leveraging Stragglers


Abstract:

Distributed gradient descent has been widely adopted in the machine learning field because considerable computing resources are available when facing the huge volume of d...Show More

Abstract:

Distributed gradient descent has been widely adopted in the machine learning field because considerable computing resources are available when facing the huge volume of data. Specifically, the gradient over the whole data is cooperatively computed by multiple workers. However, its performance can be severely affected by slow workers, namely stragglers. Recently, coding-based approaches have been introduced to mitigate the straggler problem, but they could hardly deal with the heterogeneity among workers. Besides, they always discard the results of stragglers causing huge resource waste. In this article, we first investigate how to tolerate stragglers by discarding their results and then seek to leverage the stragglers. For tolerating stragglers, we propose a heterogeneity-aware coding scheme that encodes gradients adaptive to the computing capability of workers. Theoretically, this scheme is optimal for stragglers tolerance. Relying on the scheme, we further propose an algorithm called DHeter-aware to exploit the gradients of stragglers which we called delayed gradients. Moreover, theoretical results characterized for DHeter-aware exhibits the same convergence rate as the gradient descent without delayed gradients. Experiments on various tasks and clusters demonstrate that our coding scheme outperforms all the state-of-the-art methods and the DHeter-aware further accelerates the coding scheme by achieving 25 percent time savings.
Published in: IEEE Transactions on Computers ( Volume: 71, Issue: 4, 01 April 2022)
Page(s): 779 - 794
Date of Publication: 03 March 2021

ISSN Information:

Funding Agency:


Contact IEEE to Subscribe

References

References is not available for this document.