Rapid urbanization leads the building sector to embrace a strategic role in energy conservation and carbon reduction. While the urban 3D compact form influences building carbon emissions by modulating the microclimate, existing studies lack a systematic quantification of the "form-microclimate-carbon emissions" pathway and effective spatial optimization strategies. This study developed a 1-km gridded database of urban 3D compact form indicators (NVCI), Urban Heat Accumulation (UHA), and building energy consumption carbon emissions (BECCE) from 1995 to 2020 in Xiamen. A Bayesian-parameter-convolution-optimized Random Forest model (RFBPC) was proposed to analyze the pathway. The results showed that (1) the RFBPC model consistently demonstrated superior performance over RF and XGBoost across all test cases, with testing R2 values reaching as high as 0.993 (vs. 0.879 for RF) and RMSE reductions of up to 75% (from 0.583 degrees C to 0.144 degrees C); (2) UHA exhibited a significant mediating effect between NVCI and BECCE, accounting for 13.88% of the variation in BECCE; (3) the optimal compactness transition path was identified using a geographical detector, showing that shifting NVCI from Level 5 (high compactness) to Level 2 (low compactness) could yield BECCE reduction of 401.74 kg per grid unit; and (4) scenario projections indicated that such compactness optimization could advance the building sector's carbon peak in Xiamen by 2.56 years. This study provides a quantifiable and actionable modeling framework and planning reference for coastal cities to achieve carbon-peak goals.
Accurately quantifying the effectiveness of air pollution control policies is critical for the regional mitigation of fine particulate matter (PM2.5). However, most existing assessments are confined to the urban scale, thereby overlooking the fine-grained spatial impacts of policy implementation. This study integrated the 1 km-resolution PM2.5 dataset from the Multiple Air Pollutants (MuAP) with PM2.5 control policies across the Greater Bay Area (GBA) to quantitatively analyze the nonlinear relationship between the two at the 1 km pixel scale. We processed and counted the number of PM2.5 control policies, and quantitatively scored the quality of each policy based on Large Language Models (LLMs), such as Deepseek. We use the quartile segmentation method applied to population data to reasonably distinguish the entire GBA and main urban region for multidimensional difference analysis. At the urban scale, through the K-means clustering and Pearson correlation analysis, it was found that the PM2.5 concentration difference (time lag scenario) was significantly positively correlated with the quantity and quality of policies (best scenario R = 0.691, p < 0.01). At the pixel scale, the Light Gradient Boosting Machine (LightGBM) model was utilized to construct a robust multi-factor nonlinear relationship (0.86 <= R-2 <= 0.88) effectively fitted between PM2.5 and the quantity and quality of PM2.5 control policies. Further, the SHapley Additive exPlanations (SHAP) framework was used to analyze the nonlinear spatial heterogeneity between multi-dimensional factors and the quantity and quality of PM2.5 control policies. Crucially, our findings also clarified that PM2.5 concentration changes between cities can be affected and propagated through the diffusion and dissemination of policies across different cities. This research provides new analytical paths and key insights for promoting environmental policy assessment, air quality evolution modeling, and health risk assessment.
Air pollution poses significant risks to public health and the environment, highlighting the need for long-term, spatially comprehensive air quality datasets. Here, we present a large-scale ground-based air pollutant dataset comprising 4,453,372 daily records from 2,040 monitoring stations across mainland China during 2015-2020. The dataset includes concentrations of six major air pollutants: ozone (O3), carbon monoxide (CO), nitrogen dioxide (NO2), sulfur dioxide (SO2), and particulate matter (PM2.5 and PM10). To address missing observations in long-term monitoring records, we applied a LightGBM-based imputation framework that integrates model selection, feature enrichment, and iterative imputation to leverage spatiotemporal dependencies and inter-pollutant relationships. Each record is enriched with 128 features, including meteorological variables, remote sensing products, geographic attributes, anthropogenic indicators, and temporal descriptors.Comprehensive technical validation demonstrates that the reconstructed dataset preserves realistic temporal dynamics, spatial patterns, statistical distributions, and inter-pollutant correlation structures without introducing systematic biases. Benchmark experiments further illustrate the dataset’s usability for air pollution prediction tasks. This dataset is designed with machine-learning readiness, providing a comprehensive, high-quality, and standardized resource with spatiotemporally explicit data to advance air quality research, develop and benchmark machine learning models, conduct exposure assessment, and enable data-driven environmental management across China.
The synergistic effects of large-scale surface compound ozone and heat (SCOH) present a more extensive and persistent risk to population exposure and environmental safety compared to isolated extreme heat or ozone events. Quantifying the spatiotemporal mechanisms and diffusivity of SCOH in urban areas is therefore critical for risk mitigation. This study integrates the air pollutants spatiotemporal dataset named Multiple Air Pollutants dataset (MuAP) and surface heat datasets to map the 1 km-scale time delay correlation between surface ozone and heat. Combining BayesConvLightGBM and SHapley Additive exPlanations (SHAP), the quantitative influence of urban factors such as building/canopy height and road length on SCOH in predominant urban is examined through scene analysis and diffusion potential analysis. The results show that SCOH has significant temporal and spatial distribution characteristics. Based on more effective spatiotemporal response BayesConvLightGBM modeling of SCOH (The BayesConvLightGBM's R2 is 0.03-0.07 higher than LightGBM), we found that buildings, roads, and trees have the ability to significantly affect SCOH in urban, locally or globally. Meanwhile, more compact planning of urban areas will help reduce the complex risk of SCOH. Even so, it is still important to be aware of the risk of exposure of SCOH to the population at a range of 4 km or more during a 30-day time delay period. This study deepens the quantification of nonlinear interactions between urban infrastructure and SCOH propagation, the understanding of surface compound ozone and heat, and strengthens key elements and quantification approaches using optimized machine learning. This is of great significance to explain the spatiotemporal response of SCOH, and provides an important reference for the study of compound exposure.
Formaldehyde (HCHO) is a significant indoor pollutant found in various sources and poses potential health risks to humans. Noble metal catalysts show efficient and stable catalytic activity for ambient-temperature HCHO oxidation, yet suffer from low metal utilization. Efforts focus on designing catalysts with enhanced intrinsic activity and reduced noble metal loading. In this study, we developed a simple pretreatment method using ammonia solution on SiO2 carrier to enhance the activity of the Pd/SiO2 catalyst for HCHO oxidation. After the carrier was pretreated with an ammonia solution, a significant promoting effect was observed on the Pd/SiO2(NH3 H2O)-R catalyst. It achieved almost complete oxidation of 150 ppmV of HCHO at 25 degrees C, much better than the Pd/SiO2-R (5% HCHO conversion rate). Multiple characterization results indicated that the ammonia solution pretreatment of the SiO2 carrier increased the surface defects, facilitating the anchoring of Pd nanoparticles and increasing their dispersion. The increase dispersion of Pd resulted in the generation of additional oxygen vacancies on the catalyst surfaces. The increased in oxygen vacancies on the catalyst was beneficial for enhancing the catalyst's ability to activate H2O to form surface hydroxyl groups, thereby accelerating the catalytic oxidation process of HCHO. The reaction mechanism of HCHO on the Pd/SiO2(NH3 H2O)-R catalyst mainly follows an efficient pathway: firstly, the HCHO being oxidized by surface active hydroxyl groups to formate; subsequently, the formate being oxidized by hydroxyl groups to H2O and CO2. This study provides a promising strategy for designing high-performance noble metal catalysts for HCHO catalytic oxidation. (c) 2025 The Research Center for Eco-Environmental Sciences, Chinese Academy of Sciences. Published by Elsevier B.V.
Understanding the influence of urban 3D compact form on the urban thermal environment is crucial for mitigating urban heat risks. However, conventional machine learning models fail to capture the complex nonlinear relationships involved. To address this issue, we develop a Random Forest model with Bayesian Parametric Convolutional optimization (RFBPC) to quantify the relationship between the Normalized 3D Compactness Index (NVCI) and the Urban Heat Accumulation (UHA) index. The model integrates feature importance analysis to quantify the contribution of each morphological variable, partial dependence analysis to identify nonlinear effects, and cross-validation to enhance robustness. Using building form and monthly temperature data from 1995 to 2020 in Xiamen, the RFBPC model achieved a coefficient of determination (R2) of 0.998, significantly outperforming conventional machine learning models in both predictive accuracy and stability. The results demonstrate that NVCI contributes statistically and quantitatively (8.09 %) to UHA. Based on this quantification, further simulations of different NVCI levels show that reducing urban compactness from level 5 (mean: 1.2 x 10-3) to level 2 (mean: 2.7 x 10-7) can decrease UHA by approximately 0.89 degrees C. These findings confirm both the measurable impact and regulatory potential of NVCI on UHA, providing a robust scientific foundation for optimizing urban spatial form and supporting climate-adaptive urban planning strategies.
Recently, the issue of near-surface ozone pollution has become a growing concern. To effectively manage and control ozone pollution, emerging deep learning (DL) techniques have been applied for future ozone concentration trend prediction, generating promising outcomes. However, existing studies employ various DL models and rely on diverse datasets to predict ozone concentrations. This leads to a lack of comprehensive evaluations of how the architecture and depth of different DL models influence the predictive accuracy of ozone concentration trends when assessed using a unified dataset. This lack of uniformity in evaluations creates a gap in our understanding of the influence of different neural network architectures and depths on ozone concentration predictions. In this work, we aim to address this research gap by conducting a systematic performance evaluation that benchmarks six prominent DL architectures, each with varying depths, to evaluate their effectiveness for predicting ozone concentrations across diverse geographical regions. Our findings indicate that the best-performing DL model in the nationwide prediction task is the one-layer bidirectional long short-term memory (Bi-LSTM) model, which achieves an R2 of 0.66, an RMSE of 15.32μg⋅m−3, and an MAE of 11.51μg⋅m−3. In contrast, the poorest-performing model in the same prediction task is the one-block transformer-based model, with an R2 of 0.57, an RMSE of 17.34μg⋅m−3, and an MAE of 13.3μg⋅m−3. Furthermore, fully connected networks (FCNs) demonstrate robust and efficient predictive performance across both nationwide and regional prediction tasks. Notably, our study reveals that no single DL model consistently performs well across all prediction tasks, emphasizing the need for tailored approaches that cater to the specific attributes of each region. Additionally, we observe that DL models with more than two hidden layers frequently suffer from overfitting. Particularly for the Bi-LSTM architecture, as the number of hidden layers increases from 1 to 7, we observe a 12% reduction in R2 performance. Our analysis also identifies the most influential meteorological factors among the top-performing DL models, offering insights for feature selection and optimization in model development. This research contributes to a deeper understanding of the design and selection of appropriate DL architectures for predicting the concentrations of ozone and other air pollutants.
With the acceleration of urbanization, building energy consumption carbon emissions (BECCE) were increasing, weightily influencing global warming and the social-economic developments in cities. The detailed exploration of the most effective distance of urban three-dimensional (3D) compact form on BECCE guides urban space optimization. In this study, 135 People's Banks of China (PBC) were taken as samples. Based on the banks building features, socioeconomic conditions, macroclimate, and urban 3D compact form at different buffers, the most effective distance on the BECCE was determined. According to Partial least-squares regression (PLSR), the results show that the most effective distance of the urban compact form on the BECCE was 150 m. The study divided the normalized 3D compactness index (NVCI) values from 135 buildings for the distance of 150 m into five categories and used the geographic detector to recognize statistically obvious differences between the subregions. The natural breakpoint method was used to categorize the study area into five classes, specifically: low, medium-low, medium, medium-high, and high compact form. Based on the geographic detector method result, the compact form of medium, medium high, and high compact form was optimized to low, medium low. The optimization can effectively reduce the BECCE due to changing the compact form of the building by 65.2% or 65.7%. The study combined mathematical statistics with spatial analysis approaches to identify the most effective distance for reducing energy consumption. Our results will contribute toward considerable reductions in the BECCE. Innovatively used mathematical statistics to determine the most effective distance of urban compact forms.Demonstrated the influence of on BECCE to urban compact forms.Determined the most effective distance was 150 m.Verified the most effective distance was 150 m in different building climate zones.Optimized the urban 3D compact form within 150 m to reduce the BECCE.
The spatial distribution characteristics of multi-air pollutants and their impacts are difficult to quantify effectively. As PM2.5 and NO2 are the main air pollutants, it is of great significance to explore the spatial causes of their pollution and their interaction mechanism. This study used machine learning (LightGBM) and hot spot analysis to map the spatial distribution of PM2.5 and NO2 in Southwest Fujian (SWFJ) in 2018 and their key pollution areas. Then, the factors and interactive detection of geographical detectors were used to conduct a detailed analysis of the quantitative impact of potential factors such as human activities, terrain, air pollutants, and meteorology on PM2.5 and NO2 pollution. From this we can learn that 1. LightGBM has good stability for drawing the spatial distribution of PM2.5 and NO2. 2. The spatial mechanism of PM2.5 and NO2 can be effectively interpreted from a massive data and macro perspective. 3. A large amount of evidence shows that potential factors such as human activities, topography, air pollutants and meteorology have direct or indirect effects on PM2.5 and NO2 pollution in the SWFJ area. This includes the direct impact of local road traffic emissions on the distribution of PM2.5 and NO2 pollution, the digestion of both by vegetation, the mutual transformation of atmospheric pollutants themselves, and the impact of meteorological conditions. This study not only confirms the effectiveness of machine learning combined with geographical detectors to promote the study of regional air pollution mechanisms, but also confirms the feasibility of exploring the spatial distribution mechanisms of various air pollutants. Therefore, this study is of great significance for explaining the spatial distribution of PM2.5 and NO2, and can also provide reference for policy formulation to reduce regional PM2.5 and NO2 concentrations.
Cloud coverage poses a prevalent challenge in optical remote sensing image processing, significantly hindering the visibility and interpretability of ground-level information. The application of deep learning technology in remote sensing image cloud detection has witnessed a remarkable advancement, rendering it an increasingly pervasive and potent solution to the challenge of cloud coverage. To thoroughly investigate the effectiveness of diverse deep learning architectures in cloud detection tasks, this research selected five seminal models: AlexNet, VGG16, GoogLeNet, ResNet34, and the cutting-edge Swin Transformer, and performed a comprehensive comparative analysis of their application in identifying clouds within medium-resolution Landsat 8 satellite imagery. During the training phase, each model distinctly demonstrated unique strengths in enhancing accuracy, refining loss functions, optimizing resource utilization, and accelerating training processes. In the test phase, the models' performance was rigorously evaluated using overall accuracy (OA) and other quantitative metrics, providing solid empirical evidence for the variations in their performance. Furthermore, an insightful analysis of the visualization outcomes and a meticulous comparison of the nuanced differences between model predictions and actual cloud images underscored the individual models' exceptional abilities in capturing intricate cloud details and executing precise edge processing. To alleviate the cumbersome process of pixel-level labeling, this study adopts the effective block-level labeling approach. This approach drastically streamlines annotation processes, ultimately translating into substantial cost savings and a marked reduction in time required for data preparation. The experimental results demonstrate that the Swin Transformer network performs particularly well in the cloud detection task of Landsat 8 images, achieving an OA of 90.73% under block-level annotation, significantly higher than other comparison models, showcasing excellent cloud detection capabilities. Furthermore, the study also found that the size of input image blocks has a significant impact on detection results. Under the same network architecture, using 32 x 32 image blocks for cloud detection improves the OA by at least 10 percentage points compared to using 64 x 64 image blocks. This finding provides important insights for optimizing cloud detection algorithms. In summary, this study not only verifies the superiority of Swin Transformer in cloud detection in remote sensing images but also reveals the impact of image block size on detection performance, providing new ideas and directions for future advancements in remote sensing image processing and cloud detection technology.
Nitrogen dioxide (NO2) is a critical air pollutant affecting health and the environment. However, existing monitoring station networks often fail to adequately capture the regional distribution of NO2, indicating a need for enhanced sampling strategies. This study focuses on Southwest Fujian in China and introduces a high-precision NO2 background map to delineate spatial stratified heterogeneity. Subsequently, the Mean of Surface with Non-homogeneity (MSN) method was employed to optimize the design of the NO2 monitoring network, proposing the addition of 125 new stations. The Bayesian Kriging analysis, utilized to evaluate the optimized network, resulted in a coefficient of determination (R2) of 0.87, a mean absolute percentage error (MAPE) of 9.61%, a mean absolute error (MAE) of 2.75 μg/m3 and a root mean square error (RMSE) of 2.47 μg/m3. The improved accuracy and efficiency of the NO2 monitoring network were validated against the background map. This research underscores the effectiveness of integrating precise mapping techniques with strategic network optimization for superior environmental monitoring outcomes.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
The short-term risks associated with atmospheric trace gases, particularly carbon monoxide (CO), are critical for ecological security and human health. Traditional statistical methods, which still dominate the assessment of these risks, limit the potential for enhanced accuracy and reliability. This study evaluates the performance of traditional models (ARIMA), machine learning models (LightGBM, ConvLSTM2D), and optimized machine learning solutions (Bayes residual optimization ConvLSTM2D LightGBM, Bayes_CL) in predicting Sentinel 5P columnar CO levels. This study findings demonstrate that machine learning models and their optimized versions significantly outperform traditional ARIMA models in cross-validation (CV), visualization, and overall prediction performance. Notably, machine learning model based on Bayes and residual optimization (Bayes_CL) achieved the highest CV score (Bayes_CL R2 = 0.8, LightGBM R2 = 0.79, ConvLSTM2D R2 = 0.75, ARIMA R2 = 0.61), along with superior visualization and other metrics. Using Bayes_CL, we effectively quantified a 2.4% increase in columnar CO levels in mainland China in the second half of 2023, following the complete lifting of COVID-19 lockdowns. This study confirms that machine learning models can effectively replace traditional methods for short-term risk assessment of Sentinel 5P columnar CO. This transition holds significant implications for policy formulation, greenhouse effect assessment, and population health risk evaluation, especially in uncertain situations where human activities are severely disrupted, thereby affecting environmental safety.
The pursuit of higher-resolution and more reliable spatial distribution simulation results for air pollutants is important to human health and environmental safety. However, the lack of high-resolution remote sensing retrieval parameters for gaseous pollutants (sulfur dioxide and ozone) limits the simulation effect to a 1 km resolution. To address this issue, we sequentially generated and optimized the spatial distributions of near-surface PM2.5, SO2, and ozone at a 1 km resolution in China through two approaches. First, we employed spatial sampling, random ID, and parameter convolution methods to jointly optimize a tree-based machine-learning gradient-boosting framework, LightGBM, and improve the performance of spatial air pollutant simulations. Second, we simulated PM2.5, used the simulated PM2.5 result to simulate SO2, and then used the simulated SO2 to simulate ozone. We improved the stability of 1 km-resolution SO2 and ozone products through the proposed sequence of multiple-pollutant simulations. The cross-validation (CV) of the random sample yielded an R2 of 0.90 and an RMSE of 9.62 µg∙m−3 for PM2.5, an R2 of 0.92 and an RMSE of 3.9 µg∙m−3 for SO2, and an R2 of 0.94 and an RMSE of 5.9 µg∙m−3 for ozone, which are values better than those in previous related studies. In addition, we tested the reliability of PM2.5, SO2, and ozone products in China through spatial distribution reliability analysis and parameter importance reliability analysis. The PM2.5, SO2, and ozone simulation models and multiple-air-pollutant (MuAP) products generated by the two optimization methods proposed in this study are of great value for long-term, large-scale, and regional-scale air pollution monitoring and predictions, as well as population health assessments.