Easy and reliable cluster management: the self-management experience of Fire Phoenix | IEEE Conference Publication | IEEE Xplore

Easy and reliable cluster management: the self-management experience of Fire Phoenix


Abstract:

High-Performance clusters are rapidly becoming an important computing platform for both scientific and business applications. To fulfil the new demands and challenges, cl...Show More

Abstract:

High-Performance clusters are rapidly becoming an important computing platform for both scientific and business applications. To fulfil the new demands and challenges, cluster system software is inevitably complex. Even for experienced administrators, the management of a cluster system is an exhausting job. This paper introduces Fire Phoenix, a scalable and self-managing cluster system software that supports both scientific and commercial applications. With the self-configuring and self-healing features, much of the machine configuration and error recovery can be done automatically. Our design has been proven effective in the operations of the Dawning 4000A supercomputer, which is the biggest cluster system in China
Date of Conference: 25-29 April 2006
Date Added to IEEE Xplore: 26 June 2006
Print ISBN:1-4244-0054-6
Print ISSN: 1530-2075
Conference Location: Rhodes Island

Contact IEEE to Subscribe

References

References is not available for this document.