skip to main content
OSTI.GOV title logo U.S. Department of Energy
Office of Scientific and Technical Information

Title: Analysis and Modeling of the End-to-End I/O Performance on OLCF's Titan Supercomputer

Conference ·

With the increase of scale and complexity seen in a variety of leadership-class scientific computation and simulation applications, it has become more important to understand their I/O performance characteristics. The user-observed performance is a combination of properties of how the application is using the HPC facility, as well as how others' use of the facility causes variability in the static machine capabilities. Our work leverages statistical analysis of I/O performance data gathered with fine time resolution over a full week from Titan supercomputer. Based on observed properties of the distribution of I/O latencies, we build a three-state hidden Markov model (HMM) to characterize the end-to-end I/O performance on Titan. We parameterize our model using part of the field-gathered I/O performance data and validate it against the rest. The validation results demonstrate that our model can capture the dynamics of end-to-end I/O performance on Titan accurately.

Research Organization:
Oak Ridge National Lab. (ORNL), Oak Ridge, TN (United States)
Sponsoring Organization:
USDOE Office of Science (SC), Advanced Scientific Computing Research (ASCR)
DOE Contract Number:
AC05-00OR22725
OSTI ID:
1569392
Resource Relation:
Conference: IEEE International Conference on High Performance Computing and Communications; IEEE International Conference on Smart City; IEEE International Conference on Data Science and Systems (HPCC/SmartCity/DSS) - Bangkok, , Thailand - 12/18/2017 10:00:00 AM-12/20/2017 5:00:00 AM
Country of Publication:
United States
Language:
English

Similar Records

Learning from Five-year Resource-Utilization Data of Titan System
Conference · Sun Sep 01 00:00:00 EDT 2019 · OSTI ID:1569392

Learning from Five-year Resource-Utilization Data of Titan System
Conference · Sun Sep 01 00:00:00 EDT 2019 · OSTI ID:1569392

Integration of PanDA workload management system with Titan supercomputer at OLCF
Journal Article · Wed Dec 23 00:00:00 EST 2015 · Journal of Physics. Conference Series · OSTI ID:1569392

Related Subjects