PROCEEDINGS OF THE CRAY USER GROUP, CUG 2025(2025)
Argonne Natl Lab
被引用1|浏览1
摘要
We analyze hardware errors over the seven-year lifetime of the Theta supercomputer, a large-scale Cray XC40 system at the Argonne Leadership Computing Facility. To ensure accurate interpretation of the logs, we leverage expert knowledge to clean the dataset and remove redundant information. Temporal and spatial analysis techniques are then used to expose how failures and errors trend over time and across components in the system. Additionally, we correlate hardware error logs to system downtime logs to capture the relationship between critical errors and outages over the lifetime of the system. The results in this work represent a state-ofthe-practice report highlighting how severe error types vary over time and across different component types, such as on-node and offnode (network) components. We also demonstrate the effectiveness of our technique in simplifying log analysis by using a unified error classification across components from different vendors, providing valuable insights into normal and anomalous system behaviors.