Materials research is an area that is expected to strongly benefit from the growing performance capabilities of future supercomputers towards exascale. Density functional theory (DFT) has become one of the most important methods for numerical materials science. In this paper we present results of a performance model based analysis of a particular, scalable DFT-based application on GPU-accelerated compute nodes with POWER8 processors. These technologies are part of a future roadmap for pre-exascale architectures. With power consumption becoming a major design constraint, we also determine the energy required for executing the most performance critical kernel.
QPACE is a special purpose supercomputer for Lattice Quantum Chromodynamics (LQCD) calculations designed with focus on low power consumption. To open QPACE for a broader range of High Performance Computing (HPC) applications, the High Performance Linpack benchmark (HPL), whose calculations are common for some class of HPC applications, is chosen for porting to the QPACE system. In this respect, this paper discusses features and limitations of the QPACE architecture. It starts with analysing the requirements of HPL and, based on that, presents different approaches to meet these requirements. The paper describes the different stages of porting HPL to QPACE ranging from low-level communication interface to MPI collective operations. Finally, HPL benchmark results show to what extent HPC applications other than LQCD could benefit from QPACE.
Application-driven computers for Lattice Gauge Theory simulations have often been based on system-on-chip designs, but the development costs can be prohibitive for academic project budgets. An alternative approach uses compute nodes based on a commercial processor tightly coupled to a custom-designed network processor. Preliminary analysis shows that this solution offers good performance, but it also entails several challenges, including those arising from the processor's multicore structure and from implementing the network processor on a field-programmable gate array.