Accurate prediction of fine particulate matter (PM2.5) concentration is crucial for improving environmental conditions and effectively controlling air pollution. However, some existing studies could ignore the nonlinearity and spatial correlation of time series data observed from stations, and it is difficult to avoid the redundancy between features during feature selection. To further improve the accuracy, this study proposes a hybrid model based on empirical mode decomposition (EMD), minimal-redundancy-maximal-relevance (mRMR), and geographically weighted neural network (GWNN) for hourly PM2.5 concentration prediction, named EMD-mRMR-GWNN. Firstly, the original PM2.5 concentration sequence with distinct nonlinearity and non-stationarity is decomposed into multiple intrinsic mode functions (IMFs) and a residual component using EMD. IMFs are further classified and reconstructed into high-frequency and low-frequency components using the one-sample t-test. Secondly, the optimal feature subset is selected from high-frequency and low-frequency components with mRMR for the prediction model, thus holding the correlation between features and the target variable and reducing the redundancy among features. Thirdly, the residual component is predicted with the simple moving average (SMA) due to its strong trend and autocorrelation, and GWNN is used to predict the high-frequency and low-frequency components. The final prediction of the PM2.5 concentration value is calculated by an artificial neural network (ANN) composed of the predictive values of each component. PM2.5 concentration prediction experiments in three representational cities, such as Beijing, Wuhan, and Kunming were carried out. The proposed model achieved high accuracy with a coefficient of determination greater than 0.92 in forecasting PM2.5 concentration for the next 1 h. We compared this model with four baseline models in forecasting PM2.5 concentration for the next few hours and found it performed the best in PM2.5 concentration prediction. The experimental results indicated the proposed model can improve prediction accuracy.
Precisely delineating the spatiotemporal heterogeneity of water conservation services function (WCF) holds paramount importance for watershed management. However, the existing assessment techniques exhibit common limitations, such as utilizing only multi-year average values for spatial changes and relying solely on the spatial average values for temporal changes. Moreover, traditional research does not encompass all WCF values at each time step and spatial grid, hindering quantitative analysis of spatial heterogeneity in WCF. This study addresses these limitations by utilizing an improved water balance model based on ecosystem type and soil type (ESM-WBM) and employing the EFAST and Sobol’ method for parameter sensitivity analysis. Furthermore, a space–time cube of WCF, constructed using remote-sensing data, is further explored by Emerging Hot Spot Analysis for the expression of WCF spatial heterogeneity. Additionally, this study investigates the impact of two core parameters: neighborhood distance and spatial relationship conceptualization type. The results reveal that (1) the ESM-WBM model demonstrates high sensitivity toward ecosystem types and soil data, facilitating the accurate assessment of the impacts of ecosystem and soil pattern alterations on WCF; (2) the EHSA categorizes WCF into 17 patterns, which in turn allows for adjustments to ecological compensation policies in related areas based on each pattern; and (3) neighborhood distance and the type of spatial relationships conceptualization significantly impacts the results of EHSA. In conclusion, this study offers references for analyzing the spatial heterogeneity of WCF, providing a theoretical foundation for regional water resource management and ecological restoration policies with tailored strategies.
Surface air temperature (Ta), as an important climate variable, has been used in a wide range of fields such as ecology, hydrology, climatology, epidemiology, and environmental science. However, ground measurements are limited by poor spatial representation and inconsistency, and reanalysis and meteorological forcing datasets suffer from coarse spatial resolution and inaccuracy. Previous studies using satellite data have mainly estimated Ta under clear-sky conditions or with limited temporal and spatial coverage. In this study, an all-sky daily mean land Ta product at a 1 km spatial resolution over mainland China for 2003–2019 has been generated mainly from the Moderate Resolution Imaging Spectroradiometer (MODIS) products and the Global Land Data Assimilation System (GLDAS) dataset. Three Ta estimation models based on random forest were trained using ground measurements from 2384 stations for three different clear-sky and cloudy-sky conditions. The random sample validation results showed that the R2 and root-mean-square error (RMSE) values of the three models ranged from 0.984 to 0.986 and from 1.342 to 1.440 K, respectively. We examined the spatiotemporal patterns and land cover type dependences of model accuracy. Two cross-validation (CV) strategies of leave-time-out (LTO) CV and leave-location-out (LLO) CV were also used to evaluate the models. Finally, we developed the all-sky Ta dataset from 2003 to 2009 and compared it with the China Land Data Assimilation System (CLDAS) dataset at a 0.0625∘ spatial resolution, the China Meteorological Forcing Data (CMFD) dataset at a 0.1∘ spatial resolution, and the GLDAS dataset at a 0.25∘ spatial resolution. Validation accuracy of our product in 2010 was significantly better than other datasets, with R2 and RMSE values of 0.992 and 1.010 K, respectively. In summary, the developed all-sky daily mean land Ta dataset has achieved satisfactory accuracy and high spatial resolution simultaneously, which fills the current dataset gap in this field and plays an important role in the studies of climate change and the hydrological cycle. This dataset is currently freely available at https://doi.org/10.5281/zenodo.4399453 (Chen et al., 2021b) and the University of Maryland (http://glass.umd.edu/Ta_China/, last access: 24 August 2021). A sub-dataset that covers Beijing generated from this dataset is also publicly available at https://doi.org/10.5281/zenodo.4405123 (Chen et al., 2021a).
Land surface temperature (LST) is a crucial parameter for hydrology, climate monitoring, and ecological and environmental research. LST products from thermal infrared (TIR) satellite data have been widely used for that. However, TIR information cannot provide LST data under cloudy-sky conditions. All-sky LST can be estimated from microwave measurements, but their coarse spatial resolution, narrow swaths, and short temporal range make it impossible to generate a long-term, high-resolution, accurate global all-sky LST global. This study proposes a methodology for generating the all-sky LST product by combining multiple data from Moderate Resolution Imaging Spectroradiometer (MODIS), reanalysis, and ground in situ measurements using a random forest. Field measurements from the AmeriFlux and Surface Radiation Budget (SURFRAD) networks were used for model training and validation. Cloudy-sky and clear-sky LST models were developed separately. To further improve the accuracy of the cloudy-sky LST model, the conventional RF model was extended to incorporate temporal information. The models were validated using in situ LST measurements from 2010, 2011, and 2017 that were not used for the model training. For the cloudy-sky and clear-sky models, root-mean-square-error (RMSE) = 2.767 and 2.756 K, R2 = 0.943 and 0.963, and bias = -0.143 and - 0.138 K, respectively. The same validation samples were used to validate both the MODIS LST product under clear-sky conditions and allsky Global Land Data Assimilation System (GLDAS) LST product at 0.25 degrees spatial resolution, with RMSE = 3.033 and 4.157 K, bias = -0.362 and - 0.224 K, and R2 = 0.904 and 0.955, respectively. Additionally, the 10-fold cross-validation results using all the training datasets further indicate the model stability. The models were applied to generate the all-sky LST product from 2000 to 2015 over the conterminous United States (CONUS). Our product shows similar spatial patterns to the MODIS and GLDAS LST products, but it is more accurate. Both validation and product comparisons demonstrated the robustness of our proposed models in generating the allsky LST product.