Staggered Dslash Performance on Intel Xeon Phi Architecture

Proceedings of The 32nd International Symposium on Lattice Field Theory — PoS(LATTICE2014)（2014）

引用 4|浏览1

暂无评分

摘要

The conjugate gradient (CG) algorithm is among the most essential and time consuming parts of lattice calculations with staggered quarks. We test the performance of CG and dslash, the key step in the CG algorithm, on the Intel Xeon Phi, also known as the Many Integrated Core (MIC) architecture. We try different parallelization strategies using MPI, OpenMP, and the vector processing units (VPUs).

查看译文

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要