On sequential computers, the prime factor algorithm (PFA) allows the computation of the discrete Fourier transform (DFT) with a higher efficiency than the traditional Cooley-Tukey FFT algorithm (CTA). However, the PFA requires substantial data movement, which poses a challenging problem for distributed-memory multi-processor systems. In this paper, two approaches for a concurrent implementation of the PFA on these structures are presented. In the first approach, the concurrent PFA runs on all nodes or the multi-processor system, which is inefficient on large configurations due to the large communication overhead. A second approach developed to reduce this bottleneck is also presented. These solutions have been benchmarked on Caltech hypercubes, and the performances achieved are reported. In both approaches, the crystal-router algorithm was exploited as a concurrent technique for communicating data among nodes.
Geoffrey Fox合作论文数Department of Physics, College of Arts and Sciences, Indiana University;Department of Intelligent Systems Engineering, Indiana University;Community Grid Laboratory, Indiana University;Digital Science Center of Pervasive Technology Institute;School of Engineering and Applied Science, University of Virginia1
Giovanni Aloisio合作论文数Engineering Faculty of the University of Salento1