Sparse Matrix-Vector Multiplication (SpMV) is a critical operation for the iterative solver of Finite Element Methods on computer simulation. Since the SpMV operation is a memory-bound algorithm, the efficiency of data movements heavily influenced the performance of the SpMV on GPU. In recent years, many research is conducted in accelerating the performance of SpMV on the graphic processing units (GPU). The performance optimization methods used in existing studies focus on the following areas: improve the load balancing between GPU processors, and reduce the execution divergence between GPU threads. Although some studies have made preliminary optimization on the input vector fetching, the effect of explicitly caching the input vector on GPU base SpMV has not been studied in depth yet. In this study, we are trying to minimize the data movements cost for GPU-based SpMV using a new framework named"explicit caching Hybrid (EHYB)". The EHYB framework achieved significant performance improvement by using the following methods: 1. Improve the speed of data movements by partitioning and explicitly caching the input vector to the shared memory of the CUDA kernel. 2. Reduce the volume of data movements by storing the major part of the column index with a compact format. We tested our implementation with sparse matrices derived from FEM applications in different areas. The experiment results show that our implementation can overperform the state-of-the-arts implementation with significant speedup, and leads to higher FLOPs than the theoryperformance up-boundary of the existing GPU-based SpMV implementations.
SummaryFor hyperspectral imaging, the vector bilateral filter usually leads to better performance when compared with the traditional 2D bilateral filter. However, the large computation complexity of vector bilateral filtering makes it an extremely time cost algorithm. To overcome this challenge, a GPU‐based acceleration for vector bilateral filtering called vBF_GPU was proposed in this paper. To improve the efficiency of the cache memory usage, multiple CUDA threads were utilized to processing one pixel of the hyperspectral image in vBF_GPU. The memory access operation of vBF_GPU was fully optimized to reduce the memory access cost of the GPU program. The experiment results indicated that vBF_GPU can provide more than 30× speedup when compared with an octa‐core CPU implementation and more than 20× speedup when compared with a naïve GPU implementation of vector bilateral filtering.
We propose a stochastic bilateral filter (SBF) - fast image filtering aimed at processing high dimensional images (such as color and hy-perspectral images). SBF is comprised of an efficient randomized process, where it agrees with conventional bilateral filter (BF) on average. By Monte-Carlo, we repeat this process a few times with different random instantiations so that they can be averaged to attain the correct BF output. The computational bottleneck of the SBF is constant with respect to the color dimension, meaning the complexity for filtering hyperspectral images is nearly the same as the grayscale images. It is considerably faster than the conventional and existing “fast” bilateral filter implementations.
We present Task-D, a task-based distributed programming framework. Traditionally, programming for distributed programs requires using either low-level MPI or high-level pattern based models such as Hadoop/Spark. Task based models are frequently and well used for multicore and heterogeneous environment rather than distributed. Our Task-D tries to bridge this gap by creating a higher-level abstraction than MPI, while providing more flexibility than Hadoop/Spark for task-based distributed programming. The Task-D framework alleviates programmers from considering the complexities involved in distributed programming. We provide a set of APIs that can be directly embedded into user code to enable the program to run in a distributed fashion across heterogeneous computing nodes. We also explore the design space and necessary features the runtime should support, including data communication among tasks, data sharing among programs, resource management, memory transfers, job scheduling, automatic workload balancing and fault tolerance, etc. A prototype system is realized as one implementation of Task-D. A distributed ALS algorithm is implemented using the Task-D APIs, and achieved significant performance gains compared to Spark based implementation. We conclude that task-based models can be well suitable to distributed programming. Our Task-D is not only able to improve the programmability for distributed environment, but also able to leverage the performance with effective runtime support.
Finite Element Methods (FEM) are widely used in academia and industry, especially in the fields of mechanical engineering, civil engineering, aerospace, and electrical engineering. These methods usually convert partial difference equations into large sparse linear systems. For complex problems, solving these large sparse linear systems is a time consuming process. This paper presents a parallelized iterative solver for large sparse linear systems implemented on a GPGPU cluster. Traditionally, these problems do not scale well on GPGPU clusters. This paper presents an approach to reduce the communications between cluster compute nodes for these solvers. Additionally, computation and communication are overlapped to reduce the impact of data exchange. The parallelized system achieved a speedup of up to 15.3 times on 16 NVIDIA Tesla GPUs, compared to a single GPU. An analytical evaluation of the algorithm is conducted in this paper, and the analytical equations for predicting the performance are presented and validated.
Radial Basis Function (RBF) neural networks have strong engineering applications. The training of these networks however can be time consuming. In this paper, we examine the calibration of a MOTOMAN industrial robot using an RBF based network. Additionally we examine the acceleration of RBF network using general purpose graphical processing units (GPGPUs). On a data set of 1989 calibration points, we are able to achieve a speedup of over 300 times compared to a MATLAB simulation of the algorithm. Given that MATLAB's simulation required over a week of runtime, the GPGPU acceleration enables reasonable training time for datasets with more calibration points, thus providing better precision.
In this paper we examine the acceleration of two spiking neural network models on three clusters of multicore processors representing three categories of processors: x86, STI Cell, and NVIDIA GPGPUs. The x86 cluster utilized consists of 352 dualcore AMD Opterons, the Cell cluster consists of 320 Sony Playstation 3s, while the GPGPU cluster contains 32 NVIDIA Tesla S1070 systems. The results indicate that the GPGPU platform can dominate in performance compared to the Cell and x86 platforms examined. From a cost perspective, the GPGPU is more expensive in terms of neuron/s throughput. If the cost of GPGPUs go down in the future, this platform will become very cost effective for these models.
There is a significant interest in the research community to develop large scale, high performance implementations of neuromorphic models. These have the potential to provide significantly stronger information processing capabilities than current computing algorithms. In this paper we present the implementation of five neuromorphic models on a 50 TeraFLOPS 336 node Playstation 3 cluster at the Air Force Research Laboratory. The five models examined span two classes of neuromorphic algorithms: hierarchical Bayesian and spiking neural networks. Our results indicate that the models scale well on this cluster and can emulate between 108 to 1010 neurons. In particular, our study indicates that a cluster of Playstation 3s can provide an economical, yet powerful, platform for simulating large scale neuromorphic models.
High frequency, as well as automation and day light ranging, is a signify feature of new generation Satellite Laser Ranging (SLR) systems. In spite of increase the quantity of observation data, the high frequency SLR can also significantly improve the SP and NP precision. These trends of SLR technology lead to new requirement of control circuit. In this paper, an implementation of control circuit in single FPGA chip was present. SOPC (system on programmable chip) system was proposed to solve these problems. To realize the system, a control circuit custom component was designed and simulated by us. Then, the component was integrated into a SOPC system. Cooperated with software, the circuit has the ability to control the SLR system running at high frequency. Finally, the system was simulated in the Quartus software and NIOS IDE provided by Altera and implemented in an Altera EP1S10 development kit.
Highly automation is a direction of Satellite Laser Ranging (SLR). The optical path of the laser in the domestic SLR system often shows instability, which causes errors in the laser orientation and greatly lowers the laser echo rate. In order to improve the laser echo rate, it's essential to calibrate the laser orientation real-timely. Firstly it's requisite to obtain precise pointing of laser beam-peak, whose effective way is to locate the position of the beam-peak which is imaged on the CCD because of the effect of backward scattering. Then the target-missing quantity of the position of the beam-peak is calculated and transmitted to the servo system of the adjustable mounts for calibrating the laser orientation. Using Harris corner detecting algorithm, this paper detects the beam-peak of laser, gives the precise point of beam-peak and lays a foundation for improving the automation of SLR system.
The daylight tracking observation is necessary and the tendency of Satellite Laser Ranging (SLR) in the future. More than half of the SLR stations in the world can take the daylight tracking observations. From the experience of the most successful stations around the world, Changchun station has been working on the daylight tracking technique in recent years. This paper introduces the performance and progress for SLR system daylight tracking in Changchun station. It first introduces the problems and difficulties facing this system for daylight tracking-mount model, the separation of emitting and receiving parts of the telescope, control range gate, and installing narrower filter. Secondly, it presents some work which has been done in the system for daylight tracking: system stability improvement, laser stability improvement, mount model adoption, control system, etc. From these analysis and work which has been done, the system performance has been greatly improved. A routine operation system for daylight tracking observation has been set up.
The daylight tracking observation is necessary and the tendency of Satellite Laser Ranging (SLR) in the future. More than half of the SLR stations in the world can take the daylight tracking observations. From the experience of the most successful stations around the world, Changchun station has been working on the daylight tracking technique in recent years. This paper introduces the performance and progress for SLR system daylight tracking in Changchun station. It first introduces the problems and difficulties facing this system for daylight tracking - mount model, the separation of emitting and receiving parts of the telescope, control range gate, and installing narrower filter. Secondly, it presents some work which has been done in the system for daylight tracking: system stability improvement, laser stability improvement, mount model adoption, control system, etc. From these analysis and work which has been done, the system performance has been greatly improved. A routine operation system for daylight tracking observation has been set up.