Abstract:
High-Performance clusters are rapidly becoming an important computing platform for both scientific and business applications. To fulfil the new demands and challenges, cl...Show MoreMetadata
Abstract:
High-Performance clusters are rapidly becoming an important computing platform for both scientific and business applications. To fulfil the new demands and challenges, cluster system software is inevitably complex. Even for experienced administrators, the management of a cluster system is an exhausting job. This paper introduces Fire Phoenix, a scalable and self-managing cluster system software that supports both scientific and commercial applications. With the self-configuring and self-healing features, much of the machine configuration and error recovery can be done automatically. Our design has been proven effective in the operations of the Dawning 4000A supercomputer, which is the biggest cluster system in China
Date of Conference: 25-29 April 2006
Date Added to IEEE Xplore: 26 June 2006
Print ISBN:1-4244-0054-6
Print ISSN: 1530-2075