While using High-Performance Computing (HPC) for precise and accurate air quality forecasts is a common issue, similar services devoted to marine pollution in coastal areas remain challenging. This paper presents Water quality Community Model Plus Plus (WaComM++) leveraging a parallelization schema enabling the users to run it on heterogeneous parallel architectures. We evaluated the proposed model under several execution approaches using a real-world application for pollutants forecast in the Gulf of Napoli (Campania, Italy). As a result, WaComM++ has produced results 657K times faster than the sequential run (taking into account the Particles' Outer Cycle and not considering the particle domain distribution) when using distributed and shared memory with multi-GPUs dealing with about 25 million particles.
The purpose of this paper is to provide a parallel acceleration of peer methods for the numerical solution of systems of Ordinary Differential Equations (ODEs) arising from the space discretization of Partial Differential Equations (PDEs) modeling the growth of vegetation in semi-arid climatic zones. The parallel algorithm is implemented by using the CUDA environment for Graphics Processing Units (GPUs) architectures. Numerical experiments, showing the performance gain of the proposed strategy, are provided.
The use of hardware accelerators, based on code and data offloading devoted to overcoming the CPU limitations in cores, is one of the main distinctive trends in high-end computing and related applications in the last decade. However, while code offloading is convenient for performance improvement, becoming a commonly used paradigm, memory access and management are a source of bottlenecks due to the need to interact with different address spaces. In this regard, NVidia introduced the CUDA Unified Memory model to avoid explicit memory copies between the machine hosting the accelerator device and the device itself and vice-versa. This paper shows a novel design and implementation of the support to the CUDA Unified Memory in open-source GPGPU virtualization services. The performance evaluation demonstrates that the overhead due to the virtualization and remoting is acceptable considering the possibility of sharing CUDA-enabled GPUs between various and heterogeneous machines hosted at the edge, in cloud infrastructures, or as accelerator nodes in an HPC scenario. A prototype implementation of the proposed solution is available as open-source.
In this paper we are interested in fitting data arising from environmental problems. To this aim, several procedures and methods are available in literature, and all of them involve high computational complexity when real dataset are considered. In this work, we propose a novel GPU parallel algorithm, specifically designed for fitting environmental and bathymetric data, which is based on the Kriging method. The implementation exploits the capabilities of advanced parallel computing architectures for efficiently solving large size problems. We obtain remarkable gain in terms of execution times and memory usage, as confirmed by experimental tests, by combining suitable parallel numerical libraries and ad hoc parallel kernels in CUDA environment.
Machine Learning algorithms try to provide an adequate forecast for predicting and understanding a multitude of phenomena. However, due to the chaotic nature of real systems, it is very difficult to predict data: a small perturbation from initial state can generate serious errors. Data Assimilation is used to estimate the best initial state of a system in order to predict carefully the future states. Therefore, an accurate and fast Data Assimilation can be considered a fundamental step for the entire Machine Learning process. Here, we deal with the Gaussian convolution operation which is a central step of the Data Assimilation approach and, in general, in several data analysis procedures. In particular, we propose a parallel algorithm, based on the use of Recursive Filters to approximate the Gaussian convolution in a very fast way. Tests and experiments confirm the efficiency of the proposed implementation.
In this work we deal with the solution of a two-dimensional inverse time fractional diffusion equation, involving a Caputo fractional derivative in his expression. Since we deal with a huge practical problem with a large domain, by starting from an accurate meshless localized collocation method using RBFs, here we propose a fast algorithm, implemented in a multicore architecture, which exploits suitable parallel computational kernels. More in detail, we firstly developed, a C code based on the numerical library LAPACK to perform the basic linear algebra operations and to solve linear systems, then, due to the high computational complexity and the large size of the problem, we propose a parallel algorithm specifically designed for multicore architectures and based on the Pthreads library. Performance analysis will show accuracy and reliability of our parallel implementation.
Data from sensors incorporated into mobile devices, such as networked navigational sensors, can be used to capture detailed environmental information. We describe here a workflow and framework for using sensors on boats to construct unique new datasets of underwater topography (bathymetry). Starting with a large number of measurements of position, depth, etc., obtained from such an Internet of Floating Things, we illustrate how, with a specialized protocol, data can be communicated to cloud resources, even when using delayed, intermittent, or disconnected networks. We then propose a method for automatic sensor calibration based on a novel reputation approach. Sampled depth data are interpolated efficiently on a cloud computing platform in order to provide a continuously updated bathymetric database. Our prototype implementation uses the FACE-IT Galaxy workflow engine to manage network communication and exploits the computational power of GPGPUs in a virtualized cloud environment, working with a CUDA-parallel algorithm, for efficient data processing. We report on an initial evaluation involving data from a sailing vessel in Italian coastal waters.
The data crowdsourcing paradigm applied in coastal and marine monitoring and management has been developed only recently due to the challenges of the marine environment. The pervasive internet of things technology is contributing to increase the number of connected instrumented devices available for data crowd-sourcing. A main issue in the fog/edge/cloud paradigm is that collected data need to be moved from tiny low power devices to cloud resources in order to be processed. This paper is about the DYNAMO data transfer framework enabling the data transfer feature in a internet of floating things scenario. The proposed framework is our solution to mitigate the effects of extreme and delay tolerant environments.
Summary Fast technology development has influenced the widespread use of low‐power devices in different scientific, environmental, and everyday life areas, giving birth to the Internet of Things. In this paper, we focus on the context of marine studies, addressing the problem of marine bathymetry data processing and analysis via pervasive and Internet‐connected sensors and low‐power distributed devices. Pervasive and Internet‐connected low‐power devices (as the components involved in the sensing and processing actions) made diverse and different “things” as a worldwide‐distributed system. Given the high complexity of the algorithms involved in these studies, which usually involve general‐purpose graphic processing unit (GPGPU) computation, it is impossible for the limited devices to perform the required calculations. To overcome these limitations, in this paper, we propose and implement a vertical application of GVirtuS, the open‐source GPGPU virtualization and remoting service, for achieving high performance geographical data interpolation in a high performance cloud computing scenario. We present an innovative implementation by comparing, in terms of performance and accuracy, the inverse distance weighting and kriging interpolation methods in their parallel implementations leveraging on CUDA‐enabled GPGPUs. We present a real‐world use case related to high‐resolution bathymetry interpolation in a crowdsource data context in the Bay of Pozzuoli, Italy.
In this work, we describe and implement a data assimilation approach for PM10 pollution data in Northern Italy. This was done by combining the best available information from observations and chemical transport models. Specifically, by (1) incorporating PM10 surface daily concentrations and model results from the CAMS (Copernicus Atmosphere Monitoring Service) ensemble; and (2) spreading the forecast corrections from the observation locations to the entire gridded domain covered by model forecasts by means of a data regularization approach. Results were verified against independent PM10 observations measured at 169 stations by local Environmental Protection Agencies. Twelve months of observations were matched in time and space, from January to December 2017, with air pollution model results. The studied domain encompassed the Po Valley, one of the most polluted areas in Europe, and that still does not meet the air quality criteria for the annual average concentration and the maximum number of exceedances allowed for the particulate matter. Raw model data were found to be affected by a bias with a strong seasonal dependency: a large negative bias in winter and a small bias in the summer months. The data assimilation approach, embedded into a Bayesian hierarchical approach, was able to drastically reduce the bias. Furthermore, an advanced computational approach, based on the variational Bayes method coupled with the minimization of the Kullback Leibler divergence to approximate the optimal solution, made it possible to cost-effectively assimilate data throughout the period under consideration. By using stratified cross-validation to test the accuracy of our predictions, we found high out-of-sample R-2 ( = 0.83) and an average decrease of about two-thirds of the root mean square error. Assimilated data were used to produce daily resolved cumulative population exposures. The Po Valley, in relation to the interim targets (ITs) defined by the World Health Organization, accomplishes the IT-2 level, that is to say that the average annual concentration is lower than 50 mu g/m(3), but it is still very far from the IT-3 level, corresponding to an average annual concentration of less than 30 mu g/m(3). Moreover, most of the Po Valley still has a high number of days in which the average daily concentration is higher than 50 mu g/m(3), well above the maximum limit of 35 days established by European and Italian legislation. Our results demonstrate that PM10 can be reproduced reliably using this assimilation approach, combining different sources of information, so as to perform a thorough diagnosis of air quality over a spatially and temporally uniform area.
The astonishing development of diverse and different hardware platforms is twofold: on one side, the challenge for the exascale performance for big data processing and management; on the other side, the mobile and embedded devices for data collection and human machine interaction. This drove to a highly hierarchical evolution of programming models. GVirtuS is the general virtualization system developed in 2009 and firstly introduced in 2010 enabling a completely transparent layer among GPUs and VMs. This paper shows the latest achievements and developments of GVirtuS, now supporting CUDA 6.5, memory management and scheduling. Thanks to the new and improved remoting capabilities, GVirtus now enables GPU sharing among physical and virtual machines based on x86 and ARM CPUs on local workstations, computing clusters and distributed cloud appliances.
In many applications, the Gaussian convolution is approximately computed by means of recursive filters, with a significant improvement of computational efficiency. We are interested in theoretical and numerical issues related to such an use of recursive filters in a three-dimensional variational data assimilation (3Dvar) scheme as it appears in the software OceanVar. In that context, the main numerical problem consists in solving large linear systems with high efficiency, so that an iterative solver, namely the conjugate gradient method, is equipped with a recursive filter in order to compute matrix-vector multiplications that in fact are Gaussian convolutions. Here we present an error analysis that gives effective bounds for the perturbation on the solution of such linear systems, when is computed by means of recursive filters. We first prove that such a solution can be seen as the exact solution of a perturbed linear system. Then we study the related perturbation on the solution and we demonstrate that it can be bounded in terms of the difference between the two linear operators associated to the Gaussian convolution and the recursive filter, respectively. Moreover, we show through numerical experiments that the error on the solution, which exhibits a kind of edge effect, i. e. most of the error is localized in the first and last few entries of the computed solution, is due to the structure of the difference of the two linear operators.
SummaryLow‐power devices are usually highly constrained in terms of CPU computing power, memory, and GPGPU resources for real‐time applications to run. In this paper, we describe RAPID, a complete framework suite for computation offloading to help low‐powered devices overcome these limitations. RAPID supports CPU and GPGPU computation offloading on Linux and Android devices. Moreover, the framework implements lightweight secure data transmission of the offloading operations. We present the architecture of the framework, showing the integration of the CPU and GPGPU offloading modules. We show by extensive experiments that the overhead introduced by the security layer is negligible. We present the first benchmark results showing that Java/Android GPGPU code offloading is possible. Finally, we show the adoption of the GPGPU offloading into BioSurveillance, a commercial real‐time face recognition application. The results show that, thanks to RAPID, BioSurveillance is being successfully adapted to run on low‐power devices. The proposed framework is highly modular and exposes a rich application programming interface to developers, making it highly versatile while hiding the complexity of the underlying networking layer.
Recursive Filters (RFs) are a well-known way to approximate the Gaussian convolution and, due to their computational efficiency, are intensively used in several technical and scientific fields. The accuracy of the RFs can be improved by means of the repeated application of the filter, which gives rise to the so-called K-iterated Gaussian recursive filter. In this work we propose a parallel algorithm for the implementation of the K — iterated first-order Gaussian RF for multicore architectures. This algorithm is based on a domain decomposition with overlapping strategy. The presented implementation is tailored for multicore architectures and makes use of the Pthreads library. We will show through extensive numerical tests that our parallel implementation is very efficient for large one-dimensional signals and guarantees the same accuracy level of the sequential K-iterated first-order Gaussian RF.
The acceleration of inexpensive ARM-based computing nodes with high-end CUDA enabled GPGPUs hosted on x86 64 machines using the GVirtuS general-purpose virtualization service is a novel approach to hierarchical parallelism. In this paper we draw the vision of a possible hierarchical remote workload distribution among different devices. Preliminary, but promising, performance evaluation data suggests that the developed technology is suitable for real world applications.
Representation of curves and surfaces is a basic topic in computer graphic and computer aided design (CAD). In this paper we focus on theoretical and practical issues in using radial basis functions (RBF) for reconstructing implicit curves and surfaces from point clouds. We study the conditioning of the problem and give some insight on how the problem parameters and the results have to be taken in order to achieve meaningful solutions and avoid artifacts. Moreover, a strategy for decreasing the conditioning of the problem is suggested and a general framework for preconditioning and solving the problem, even for large datasets, is also provided.
Recursive filters (RFs) have achieved a central role in several research fields over the last few years. For example, they are used in image processing, in data assimilation and in electrocardiogram denoising. More in particular, among RFs, the Gaussian RFs are an efficient computational tool for approximating Gaussian-based convolutions and are suitable for digital image processing and applications of the scale-space theory. As is a common knowledge, the Gaussian RFs, applied to signals with support in a finite domain, generate distortions and artifacts, mostly localized at the boundaries. Heuristic and theoretical improvements have been proposed in literature to deal with this issue (namely boundary conditions). They include the case in which a Gaussian RF is applied more than once, i.e. the so called K-iterated Gaussian RFs. In this paper, starting from a summary of the comprehensive mathematical background, we consider the case of the K-iterated first-order Gaussian RF and provide the study of its num...
Nowadays, recursive filters (RFs) are frequently used in several research fields. More in particular, Gaussian RFs offer a more efficient way for computing approximate Gaussian filters and Gaussian-based convolutions. The use of such recursive filters introduces many sources of errors. Among them, here we consider the filter truncation error, that is the error due to the transition from the starting filter operator to the RF approximating it. Since input and output signals have infinite dimensions, the analysis of the related filter operator involves infinite matrices. In this paper, starting from a summary of the comprehensive mathematical background, we consider the case of the first-order Gaussian recursive filter. Then, taking into account the matrix form of the related operator, we perform the error analysis and provide theoretical results that estimate the filter truncation error.
Gaussian recursive filters (RFs) are frequently used in several research fields with th aim to approximate in an efficient way Gaussian filters and Gaussian-based convolutions. Among them, the first-order Gaussian RF, also in its K-iterated form, has been recently used in data assimilation. However, a recent study has proved that in the base case (K = 1) this method is not able to well approximate the Gaussian convolution for all values of the standard deviation. Here we propose a new way to construct a second order RF whose smoothing coefficients are chosen in order to enhance the accuracy of the approximation.
The aim of this study is to introduce a new scheme, based on a compressive sampling technique, for the reconstruction of lost data in multimedia streaming. The audio streaming data are encapsulated in different packets, at the sender, by using an interleaving technique. The compressive sampling technique is used to recover audio information in case of lost packets, at the receiver. Experimental results are presented for speech and musical audio signals which illustrate the performances and the capabilities of the proposed methodology.