Skip to main content

Multi-way Theta-Join Based on CMD Storage Method

  • Conference paper
Book cover Database Systems for Advanced Applications (DASFAA 2014)

Part of the book series: Lecture Notes in Computer Science ((LNISA,volume 8421))

Included in the following conference series:

  • 1707 Accesses

Abstract

In the era of the Big Data, how to analyze such a vast quantity of data is a challenging problem, and conducting a multi-way theta-join query is one of the most time consuming operations. MapReduce has been mentioned most in the massive data processing area and some join algorithms based on it have been raised in recent years. However, MapReduce paradigm itself may not be suitable to some scenarios and multi-way theta-join seems to be one of them. Many multi- way theta-join algorithms on traditional parallel database have been raised for many years, but no algorithm has been mentioned on the CMD (coordinate modulo distribution) storage method, although some algorithms on equal-join have been proposed. In this paper, we proposed a multi-way theta-join method based on CMD, which takes the advantage of the CMD storage method. Experiments suggest that it’s a valid and efficient method which achieves significant improvement compared to those applied on the MapReduce.

This is a preview of subscription content, log in via an institution to check access.

Access this chapter

Chapter
USD 29.95
Price excludes VAT (USA)
  • Available as PDF
  • Read on any device
  • Instant download
  • Own it forever
eBook
USD 39.99
Price excludes VAT (USA)
  • Available as PDF
  • Read on any device
  • Instant download
  • Own it forever
Softcover Book
USD 54.99
Price excludes VAT (USA)
  • Compact, lightweight edition
  • Dispatched in 3 to 5 business days
  • Free shipping worldwide - see info

Tax calculation will be finalised at checkout

Purchases are for personal use only

Institutional subscriptions

Preview

Unable to display preview. Download preview PDF.

Unable to display preview. Download preview PDF.

References

  1. Li, J., Srivastava, J., Rotem, D.: CMD: a multidimensional de-clustering method for parallel database systems. In: Proceedings of the 18th International Conference on Very Large Data Bases Conference, Canada (1992)

    Google Scholar 

  2. Okcan, A., Riedewald, M.: Processing Theta-Joins using MapReduce. In: SIGMOD 2011, Athens, Greece, June 12-16 (2011)

    Google Scholar 

  3. Zhang, X., Chen, L., Wang, M.: Efficient Multiway Theta Join Processing Using MapReduce. Proceedings of the VLDB Endowment 5(11) (August 27-31, 2012)

    Google Scholar 

  4. Li, J.-Z., Wei, D.: Parallel CMD- Join Algorithms on Parallel Databases. Journal of Software 9(4) (1998)

    Google Scholar 

  5. Li, J.-Z.: A Dynamic and Multidimensional Declustering Method for Parallel Databases. Journal of Software 10(9) (1999)

    Google Scholar 

  6. DeWitt, D.: MapReduce: A major step backwards (January 8, 2008), http://databasecolumn.verca.com/2008/01/mapreduce-a-major-step-back.html

  7. Dean, J., Ghemawat, S.: Proc. 2004. MapReduce: simplified data processing on large clusters. In: The 6th Symposium on Operating System Design and Implementation, OSDI 2004 (2004)

    Google Scholar 

  8. Chang, F., Dean, J., Ghemawat, S.: Proc. 2006. Bigtable: a distributed storage system for structured data. In: The 7th Symp. Operating System Design and Implementation, pp. 205–218. Usenix Assoc. (2006)

    Google Scholar 

  9. Apache hadoop, http://hadoop.apache.org

  10. Kitsuregawa, M., Tanaka, H., Moto-Oka, T.: Application of Hash to Data Base Machine and Its Architecture. New Generation Comput. 1(1), 63–74 (1983)

    Article  Google Scholar 

  11. Boral, H., et al.: Join on a cube: analysis, simulation and implementation. In: Kitsuregawa, M., Tanaka, H. (eds.) Database Machines and Knowledge Base Machines, pp. 61–74. Kluwer, Boston (1988)

    Google Scholar 

  12. Schneider, D.A., DeWitt, D.J.: A performance evaluation of parallel in algorithms in a shared-nothing multiprocessor environment. In: Maier, D. (ed.) Proc. of ACM SIGMOD 1989, USA, pp. 110–121. ACM Press, M Baltimore (1989)

    Chapter  Google Scholar 

  13. Kitsuregawa, M., Tanaka, H., Moto-oka, T.: Application of hash to data base machine and its architecture. New Generation Computing 1(1), 25–39 (1983)

    Article  Google Scholar 

  14. DeWitt, D.J., Gerber, R.: Multiprocessor hash-based join algorithms. In: Proceedings of VLDB 1985, pp. 151–164. Morgan kaufmann Publishers, Inc., Stockholm (1985)

    Google Scholar 

  15. Li, J.: A Dynamic and Multidimensional Declustering Method for Parallel Databases. Journal of Software 10(9) (1999)

    Google Scholar 

Download references

Author information

Authors and Affiliations

Authors

Editor information

Editors and Affiliations

Rights and permissions

Reprints and permissions

Copyright information

© 2014 Springer International Publishing Switzerland

About this paper

Cite this paper

Li, L., Gao, H., Zhu, M., Zou, Z. (2014). Multi-way Theta-Join Based on CMD Storage Method. In: Bhowmick, S.S., Dyreson, C.E., Jensen, C.S., Lee, M.L., Muliantara, A., Thalheim, B. (eds) Database Systems for Advanced Applications. DASFAA 2014. Lecture Notes in Computer Science, vol 8421. Springer, Cham. https://doi.org/10.1007/978-3-319-05810-8_5

Download citation

  • DOI: https://doi.org/10.1007/978-3-319-05810-8_5

  • Publisher Name: Springer, Cham

  • Print ISBN: 978-3-319-05809-2

  • Online ISBN: 978-3-319-05810-8

  • eBook Packages: Computer ScienceComputer Science (R0)

Publish with us

Policies and ethics