The forecast of particulate matter PM10 concentration is crucial due to its impacts on public health and the environment. Chemical Transport Models (CTM) are used to predict air quality. However, these models are subject to bias because of the precision of inputs. This paper explores a hybrid approach combining CTM (WRF-CHIMERE) predictions with machine learning (ML) to forecast PM10 concentrations. Five ML algorithms were developed: Multiple Linear Regression (MLR), Random Forests (RF), Extreme Gradient Boosting (XGB), Support Vector Regression (SVR), and Artificial Neural Networks (ANN). This hybrid system was trained using hourly data from September to December 2020 on seven Moroccan sites, incorporating nine parameters including meteorological variables, chemical concentrations, and other spatiotemporal variables. The hybrid model was evaluated against PM10 measurements. The results reveal that CHIMERE combined with RF and with XGB presented the best accuracy of predictions of PM10, when compared to the CHIMERE model. These two hybrid models achieved high correlation coefficients of 0.756 and 0.747, and determination coefficients of 57
We present a supervised machine learning (ML) approach to improve the accuracy of the regional horizontal distribution of the aerosol optical depth (AOD) simulated by the CHIMERE chemistry transport model over North Africa and the Arabian Peninsula using Moderate Resolution Imaging Spectroradiometer (MODIS) AOD satellite observations. Our method produces daily AOD maps with enhanced precision and full spatial domain coverage, which is particularly relevant for regions with a high aerosol abundance, such as the Sahara Desert, where there is a dramatic lack of ground-based measurements for validating chemistry transport simulations. We use satellite observations and some geophysical variables to train four popular regression models, namely multiple linear regression (MLR), random forests (RF), gradient boosting (XGB), and artificial neural networks (NN). We evaluate their performances against satellite and independent ground-based AOD observations. The results indicate that all models perform similarly, with RF exhibiting fewer spatial artifacts. While the regression slightly overcorrects extreme AODs, it remarkably reduces biases and absolute errors and significantly improves linear correlations with respect to the independent observations. We analyze a case study to illustrate the importance of the geophysical input variables and demonstrate the regional significance of some of them.
The Madden–Julian Oscillation (MJO) is one of the main sources of sub-seasonal atmospheric predictability in the tropical region. The MJO affects precipitation over highly populated areas, especially around southern India. Therefore, predicting its phase and intensity is important as it has a high societal impact. Indices of the MJO can be derived from the first principal components of zonal wind and outgoing longwave radiation (OLR) in the tropics (RMM1 and RMM2 indices). The amplitude and phase of the MJO are derived from those indices. Our goal is to forecast these two indices on a sub-seasonal timescale. This study aims to provide an ensemble forecast of MJO indices from analogs of the atmospheric circulation, computed from the geopotential at 500 hPa (Z500) by using a stochastic weather generator (SWG). We generate an ensemble of 100 members for the MJO amplitude for sub-seasonal lead times (from 2 to 4 weeks). Then we evaluate the skill of the ensemble forecast and the ensemble mean using probabilistic scores and deterministic skill scores. According to score-based criteria, we find that a reasonable forecast of the MJO index could be achieved within 40 d lead times for the different seasons. We compare our SWG forecast with other forecasts of the MJO. The comparison shows that the SWG forecast has skill compared to ECMWF forecasts for lead times above 20 d and better skill compared to machine learning forecasts for small lead times.
The role of modelling the atmospheric dispersion of pollutants at microscale, the scale that allows to resolve explicitly the presence of obstacles, is becoming increasingly important for performing air quality assessments in cities, as well as for regulatory purposes and for the design of pollution control strategies. However, the use of microscale models can be computationally demanding, both in terms of time and CPUs required, especially if the computational domain considers wide spatial extension and the simulation considers long time periods. This article proposes the application of a kernel method as the concentration calculation methodology inside microscale Lagrangian particle dispersion models (LPDMs) in order to reduce the required computational time. In these models, the concentration is normally estimated with the box-counting method, while the use of this alternative method, based on the use of the statistical technique of kernel density estimation, allows for a reduction of numerical particles emitted during the simulation, while guaranteeing a similar accuracy to that of the box-counting method. It therefore enables an optimization of computational efficiency. In an earlier manuscript, the kernel method was applied inside the LPDM of the PMSS (Parallel-Micro-SWIFT-SPRAY) system to perform high-resolution simulations of line sources, enabling an 80% simulation time reduction. In this article, additional features of this method are developed within the Micro-SPRAY model and tested through two test cases. The kernel method has been applied to estimate the pollutant concentrations of point sources as well as to compute the corresponding deposition at building-resolving scale. The results with tiled and nested configurations of domains are also verified.