skip to main content
OSTI.GOV title logo U.S. Department of Energy
Office of Scientific and Technical Information

Title: Scalable load-balance measurement for SPMD codes

Conference ·

Good load balance is crucial on very large parallel systems, but the most sophisticated algorithms introduce dynamic imbalances through adaptation in domain decomposition or use of adaptive solvers. To observe and diagnose imbalance, developers need system-wide, temporally-ordered measurements from full-scale runs. This potentially requires data collection from multiple code regions on all processors over the entire execution. Doing this instrumentation naively can, in combination with the application itself, exceed available I/O bandwidth and storage capacity, and can induce severe behavioral perturbations. We present and evaluate a novel technique for scalable, low-error load balance measurement. This uses a parallel wavelet transform and other parallel encoding methods. We show that our technique collects and reconstructs system-wide measurements with low error. Compression time scales sublinearly with system size and data volume is several orders of magnitude smaller than the raw data. The overhead is low enough for online use in a production environment.

Research Organization:
Lawrence Livermore National Lab. (LLNL), Livermore, CA (United States)
Sponsoring Organization:
USDOE
DOE Contract Number:
W-7405-ENG-48
OSTI ID:
945658
Report Number(s):
LLNL-CONF-406045; TRN: US200903%%706
Resource Relation:
Conference: Presented at: SC08, Austin, TX, United States, Nov 15 - Nov 21, 2008
Country of Publication:
United States
Language:
English

Similar Records

Parallel tetrahedral mesh adaptation with dynamic load balancing
Journal Article · Wed Jun 28 00:00:00 EDT 2000 · Parallel Computing Journal · OSTI ID:945658

Scalable Performance Measurement and Analysis
Thesis/Dissertation · Thu Jan 01 00:00:00 EST 2009 · OSTI ID:945658

Quantum Monte Carlo Endstation for Petascale Computing
Technical Report · Wed Mar 02 00:00:00 EST 2011 · OSTI ID:945658