An empirical comparison of sampling techniques for matrix column subset selection | IEEE Conference Publication | IEEE Xplore

An empirical comparison of sampling techniques for matrix column subset selection


Abstract:

Column subset selection (CSS) is the problem of selecting a small portion of columns from a large data matrix as one form of interpretable data summarization. Leverage sc...Show More

Abstract:

Column subset selection (CSS) is the problem of selecting a small portion of columns from a large data matrix as one form of interpretable data summarization. Leverage score sampling, which enjoys both sound theoretical guarantee and superior empirical performance, is widely recognized as the state-of-the-art algorithm for column subset selection. In this paper, we revisit iterative norm sampling, another sampling based CSS algorithm proposed even before leverage score sampling, and demonstrate its competitive performance under a wide range of experimental settings. We also compare iterative norm sampling with several of its other competitors and show its superior performance in terms of both approximation accuracy and computational efficiency. We conclude that further theoretical investigation and practical consideration should be devoted to iterative norm sampling in column subset selection.
Date of Conference: 29 September 2015 - 02 October 2015
Date Added to IEEE Xplore: 07 April 2016
ISBN Information:
Conference Location: Monticello, IL, USA

Contact IEEE to Subscribe

References

References is not available for this document.