Conferences >Proceedings 19th Internationa...

Joining massive high-dimensional datasets

Download PDF
Download References
Request Permissions
Save to
Alerts

Abstract:

We consider the problem of joining massive datasets. We propose two techniques for minimizing disk I/O cost of join operations for both spatial and sequence data. Our tec...Show More

Metadata

Abstract:

We consider the problem of joining massive datasets. We propose two techniques for minimizing disk I/O cost of join operations for both spatial and sequence data. Our techniques optimize the available buffer space using a global view of the datasets. We build a boolean matrix on the pages of the given datasets using a lower bounding distance predictor. The marked entries of this matrix represent candidate page pairs to be joined. Our first technique joins the marked pages iteratively. Our second technique clusters the marked entries using rectangular dense regions that have minimal perimeter and fit into buffer. These clusters are then ordered so that the total number of common pages between consecutive clusters is maximal. The clusters are then read from disk and joined. Our experimental results on various real datasets show that our techniques are 2 to 86 times faster than the competing techniques for spatial datasets, and 13 to 133 times faster than the competing techniques for sequence datasets.

Published in: Proceedings 19th International Conference on Data Engineering (Cat. No.03CH37405)

Date of Conference: 05-08 March 2003

Date Added to IEEE Xplore: 21 January 2004

Print ISBN:0-7803-7665-X

DOI: 10.1109/ICDE.2003.1260798

Conference Location: Bangalore, India