Forest ecosystems are increasingly stressed through heatwaves, drought periods, and other factors such as ozone pollution or insect infestations. These stressors have a profound impact on the emissions of biogenic volatile organic compounds (BVOC) from trees, which in turn influence aerosol formation and atmospheric oxidation cycles and thus feedback on the atmospheric cleansing capacity and climate change itself. While previous studies have investigated the impacts of specific stressors on BVOC emissions, analyses of combined stress effects are rare, even though the stressors seldomly occur in isolation. This study investigates the impact of heat and (nighttime) ozone stress, both individually and in combination, on BVOC emissions from two ecologically significant temperate tree species: European beech (Fagus sylvatica L.) and English oak (Quercus robur L.). In a climate-controlled chamber, both tree species were subjected to heat stress (38 +/- 3.3 degrees C) and ozone stress (similar to 120 ppb), separately and in combination. BVOC emission rates were measured using proton transfer reaction time-of-flight mass spectrometry, and the results were compared across pre-stress, heat, ozone, and combined heat-ozone conditions.Heat stress elicited the strongest emission increases of isoprene, monoterpene, and green leaf volatiles in both species, while ozone suppressed the emissions of most BVOCs. Combined stress led to non-additive responses different from those in single-stress scenarios. Both machine learning and positive matrix factorization analyses were performed to identify key VOC fingerprint markers that may be applied to identify stress-impacted emissions from field data, and both methods showed good agreement. The OH reactivity of the emissions, which serves as a measure for their atmospheric chemistry and ozone formation impacts, was consistently highest under heat stress for both species. However, nighttime ozone stress led to reduced OH reactivity of emissions (by 10 %-18 %).Our results underscore that the study of realistic combinations of stressors is crucial to understand future BVOC emissions and indicate that BVOC emissions could alter atmospheric chemistry and feedback with air quality and climate as heatwaves and pollutant-induced stress become more frequent due to climate change.
Accurate characterization of station locations is crucial for reliable air quality assessments such as the Tropospheric Ozone Assessment Report (TOAR). While urban and rural areas are relatively well-defined, the boundaries and identity of suburban areas remain ambiguous, overlapping with both urban and rural zones and varying due to cultural and social factors. This study investigates a machine learning approach to classify 24 348 stations in the unique global TOAR database as urban, suburban, or rural. We tested two different approaches: unsupervised K-means clustering with three clusters, and an ensemble of supervised learning classifiers including random forest, CatBoost, and LightGBM. We integrate these classifiers into a robust voting model, leveraging their collective predictive power. To address the inherent ambiguity of suburban areas, we implement a grid-search adjusted threshold probability technique. Our models, trained on the TOAR station metadata, are evaluated on 1979 unseen data points. K-means clustering achieves 71.88 % and 87.67 % accuracy for urban and rural areas respectively, but only 15.84 % for suburban zones. The supervised classifiers surpass this performance, reaching over 84 % accuracy for urban and rural categories, and 66 %-72 % for suburban areas. The adjusted threshold technique significantly enhances overall model accuracy, particularly for suburban classification. The good separation of our model is confirmed through evaluation with NOx and PM2.5 concentration measurements, which were not included in the training data. Furthermore, manual inspection of 30 randomly selected sites with Google maps reveals that our method provides a better label for the station type than the labels that were reported by data providers and used in the model evaluation. The objective station classification proposed in this paper therefore provides a robust foundation for type-of-area-specific air quality assessments in TOAR and elsewhere.
Data-driven weather prediction models based on artificial intelligence (AI) have rapidly advanced in recent years and are frequently reported to outperform traditional physics-based numerical weather prediction (NWP) models for selected verification scores. However, optimization with respect to a specific loss function can adversely affect other metrics, potentially leading to unrealistic forecast characteristics, such as overly smooth spatial structures when mean-squared or mean-absolute error–based loss functions are used.In Bonavita & Geer, 2026 an orthogonal decomposition of the Mean Squared Error (MSE) into Information Error and Noise Error is introduced to unravel different strategies for minimizing this commonly used accuracy metric. These insights allow for a more meaningful interpretation of the MSE as accuracy measure.Additionally, a scale dependent analysis of model performance can help to reveal systematic differences between AI and NWP models specifically on smaller scales such as the effective resolution.The combination of both approaches – the decomposition of the scale dependent MSE into Information and Noise Error – has been derived and applied to different AI and NWP models.First results show the potential of this method to disentangle different effects allowing for a fairer, more comprehensive comparison between AI and NWP weather prediction models.The analysis is partly based on forecasts from the Weather Prediction Model Intercomparison Project (WP MIP), which provides a collection of NWP and AI-model forecasts from multiple national weather services and research institutions.The work is conducted within the RAINA project, which aims to develop a foundation model for the atmosphere with a particular focus on reliable, high-resolution forecasts of extreme wind and precipitation events.Bonavita, M. & Geer, A.J. (2026) Forecast verification using information and noise. Quarterly Journal of the Royal Meteorological Society, e70109. Available from: https://doi.org/10.1002/qj.70109
Events of extreme air pollution pose threats to humans and the environment. To investigate air quality under extreme atmospheric situations, the DestinE air quality use case developed a comprehensive user interface that enables high-resolution air quality forecasts with diverse analysis options. The user interface encapsulates the two state-of-the-art approaches that are physics-based numerical simulations with the chemistry transport model EURAD-IM (European Air pollution Dispersion – Inverse Model) and data-driven machine learning forecasts with MLAir (Machine Learning on Air data). The EURAD-IM simulations are coupled to the meteorological output of the DestinE digital twin for weather extremes, which provides high-resolution information (~4.4 km). An additionally implemented machine learning based postprocessing even allows for the downscaling of the EURAD-IM forecast output to a resolution of 1 km. MLAir produces 4-day point forecasts at station sites using data from the Tropospheric Ozone Assessment Report (TOAR) data base. The developed system is complemented by an efficient module that enables emission scenario simulations to investigate and develop air pollution mitigation strategies for future extreme events under realistic conditions.The established user interface is demonstrated by two selected air quality extreme events in early 2017 and summer 2018. It aims to provide a new quality of air pollution information that supports the core users, i.e., environment agencies, in decision making. For the near future, it is planned to fully embed the system to the Destination Earth Service Platform (DESP) such that it will be available to a wider community. Besides assisting policy making, the air quality products help to answer scientific questions on air quality and atmospheric chemical processes under extreme weather conditions that are expected to increase in future.
Ozone (O3), a short-lived climate pollutant, continues to increase despite policies aimed at suppressing its precursors in South Korea. The government operates approximately 500 observatories to monitor O3 and trace gases. Researchers use these data to address the ongoing issue of increasing O3 levels. However, challenges in data retrieval from observatories may introduce biases in O3 studies. In this study, we developed a graph-based machine learning model to simulate missing O3 concentrations for mitigate bias. The model incorporates spatiotemporal distribution characteristics using a merged observation dataset from South Korea in 2021. Regardless of region or length of missing data, the model effectively simulates O3 variations with R2 of up to 0.9 and RMSE of 3.6. To determine the influence of input parameters on O3 interpolation, we used eXplainable AI methods. The results indicated that NO2 is the most important factor in cities, while photochemical indicators are more influential in provinces.
In recent years, deep neural networks (DNN) to enhance the resolution of meteorological data, known as statistical downscaling, have surpassed classical statistical methods that have been developed previously with respect to several validation metrics. The prevailing approach for DNN downscaling is to train deep learning models in an end-to-end manner. However, foundation models trained on very large datasets in a self-supervised way have proven to provide new SOTA results for various applications in natural language processing and computer vision. To investigate the benefit of foundation models in Earth Science applications, we deploy the large-scale representation model for atmospheric dynamics AtmoRep (Lessig et al., 2023) for statistical downscaling of the 2m temperature over Central Europe. AtmoRep has been trained on almost 40 years of ERA5 data from 1979 to 2017 and has shown promising skill in several intrinsic and downstream applications. By extending AtmoRep’s encoder-decoder with a tail network for downscaling, we super-resolve the coarse-grained 2 m temperature field from ERA5-data (Δx = 25 km) to attain the high spatial resolution (Δx = 6 km) of the COSMO REA6 dataset. Different coupling approaches between the core and tail network (e.g. with and without fine-tuning the core model) are tested and analyzed in terms of accuracy and computational efficiency. Preliminary results show that downscaling with a task-specific extension of the foundation model AtmoRep can improve the downscaled product in terms of standard evaluation metrics such as the RMSE compared to a task-specific deep learning model. However, deficiencies in the spatial variability of the downscaled product are also revealed, highlighting the need for future work to focus especially on target data that inhibit a high degree of spatial variability and intrinsic uncertainty such as precipitation.
As the High Performance Computing (HPC) marches into the exascale era, earth system models have transformed into a numerical regime wherein simulations with a 1 km spatial resolution on a global scale are a reality and are currently being performed at various HPC centers across the globe. In this contribution, we provide an overview of the strategy and plans to adapt the data handling services and workflows available at the German Climate Computing Center (DKRZ) and the Jülich Supercomputing Center (JSC) to enable efficient data access, processing and sharing of output from such simulations using current and next generation Earth System Models. These activities are carried out in the framework of projects funded on an EU as well as national level, such as NextGEMS, WarmWorld and EERIE. With the increase in spatial resolution comes the inevitable jump in the volume of the output data. In particular, the throughput due to the enhanced computing power always surpasses the capacity of single-tier storage systems made up of homogeneous hardware and necessitates multi-tier storage systems consisting of heterogeneous hardware . As a consequence, new issues arise for an efficient, user-friendly data management within each site. Sharing of model outputs that may be produced at different data centers and stored across different multi-tier storage systems poses additional challenges, both in terms of technical aspects (efficient data handling, data formats, reduction of unnecessary transfers) and semantic aspects (data discovery and selection across sites). Furthermore, there is an increasing need for scientifically operational solutions, which requires the development of long-term strategies that can be sustained within the different data centers. To achieve all of this, existing workflows need to be analyzed and largely rewritten. On the upside, this will allow the introduction of new concepts and technologies, for example using the recent zarr file format instead of the more traditional netCDF format.More specifically, in WarmWorld, the strategy is to create an overarching user interface, to enable the discovery of the federated data, and implement the backend infrastructure for handling the movement of the data, across the storage tiers (SSD, HDD, tape, cloud), within as well as across the HPC centers, as necessitated by the analytical tasks. This approach will also leverage the benefits of community efforts in redesigning the way km-scale models provide their output, i.e. on hierarchical grids and in relatively small chunks.We present specific ongoing work to implement this data handling strategy across HPC centers and outline the vision for the handling of high-volume climate model simulation output in the exascale era to enable the efficient analysis of the information content from these simulations.
Ground-level ozone is a significant air pollutant that detrimentally affects human health and agriculture. Global ground-level ozone concentrations have been estimated using chemical reanalyses, geostatistical methods, and machine learning, but these datasets have not been compared systematically. We compare six global ground-level ozone datasets (three chemical reanalyses, two machine learning, one geostatistics) relative to observations and against one another, for the ozone season daily maximum 8 h average mixing ratio, for 2006 to 2016. Comparing with global ground-level observations, most datasets overestimate ozone, particularly at lower observed concentrations. In 2016, across all stations, grid-to-grid R2 ranges from 0.50 to 0.75 and RMSE 4.25 to 12.22 ppb. Agreement with observed distributions is reduced at ozone concentrations above 50 ppb. Results show significant differences among datasets in global average ozone, as large as 5–10 ppb, multi-year trends, and regional distributions. For example, in Europe, the two chemical reanalyses show an increasing trend while other datasets show no increase. Among the six datasets, the share of population exposed to over 50 ppb varies from 61 % [28 %, 94 %] to 99 % [62 %, 100 %] in East Asia, 17 % [4 %, 72 %] to 88 % [53 %, 99 %] in North America, and 9 % [0 %, 58 %] to 76 % [22 %, 96 %] in Europe (2006–2016 average). Although sharing some of the same input data, we found important differences, likely from variations in approaches, resolution, and other input data, highlighting the importance of continued research on global ozone distributions. These discrepancies are large enough to impact assessments of health impacts and other applications.
Explainable machine learning has gained substantial attention for its role in enhancing transparency and trust in computer vision applications. Attribution methods like Grad-CAM and occlusion sensitivity analysis are frequently used to identify how features contribute to predictions of neural networks. However, a key challenge is that different attribution methods often produce different outcomes undermining trust in their results. Furthermore, the unique characteristics of remote sensing imagery pose additional challenges for attribution interpretation: it primarily comprises continuous “stuff” classes rather than objects, exhibits fine-grained spatial variability, contains mixed pixels, is often multispectral, and exhibits spatially heterogeneity. To tackle this challenge, we present a novel methodology that harmonizes attributions, resulting in: 1. greater consistency across different attribution methods; 2. more meaningful explanations when validated against known segmentation ground truth; and 3. enhanced transparency and traceability. This is achieved by coherently linking feature representations to attributions derived from analyzing the training data, enabling direct attribution assignment to features in (unseen) images. We evaluate our methodology using two satellite-based land cover classification datasets, three convolutional neural network architectures, and nine attribution methods. Harmonizing attributions increases the Pearson correlation coefficient between different attribution methods by an average of 0.18 across all datasets, models, and methods; and improves the micro F1-score — a measure of accuracy — by 12%. We demonstrate that Grad-CAM attributions are inherently well-aligned with the features, whereas other gradient-based attribution methods exhibit significant noise, mitigated through harmonization. It further enhances the resolution of occlusion-based attribution maps and adjusts misleading explanations.
The emergence of exascale computing and artificial intelligence offer tremendous potential to significantly advance Earth system prediction capabilities. However, enormous challenges must be overcome to adapt models and prediction systems to use these new technologies effectively. A 2022 WMO report on exascale computing recommends "urgency in dedicating efforts and attention to disruptions associated with evolving computing technologies that will be increasingly difficult to overcome, threatening continued advancements in weather and climate prediction capabilities." Further, the explosive growth in data from observations, model and ensemble output, and postprocessing threatens to overwhelm the ability to deliver timely, accurate, and precise information needed for decision-making. Artificial intelligence (AI) offers untapped opportunities to alter how models are developed, observations are processed, and predictions are analyzed and extracted for decision-making. Given the extraordinarily high cost of computing, growing complexity of prediction systems, and increasingly unmanageable amount of data being produced and consumed, these challenges are rapidly becoming too large for any single institution or country to handle. This paper describes key technical and budgetary challenges, identifies gaps and ways to address them, and makes a number of recommendations.
In remote sensing, plenty of multi-spectral images are publicly available from various landcover satellite missions. Contrastive self-supervised learning is commonly applied to unlabeled data but relies on domain-specific transformations used for learning. When focusing on vegetation, standard transformations from image processing cannot be applied to the NIR channel, which carries valuable information about the vegetation state. Therefore, we use contrastive learning, relying on different views of unlabelled, multi-spectral images to obtain a pre-trained model to improve the accuracy scores on small-sized remote sensing datasets. This study presents the generation of additional views tailored to remote sensing images using atmospheric correction as an alternative transformation to color jittering. The purpose of the atmospheric transformation is to provide a physically consistent transformation. The proposed transformation can be easily integrated with multiple channels to exploit spectral signatures of objects. Our approach can be applied to other remote sensing tasks. Using this transformation leads to improved classification accuracy of up to 6%.
Earth Observation (EO) data processing faces challenges due to large volumes, multiple sources, and diverse formats. To address this issue, this paper presents a scalable and parallelizable workflow using Apache Airflow, capable of integrating Machine Learning (ML) and Deep Learning (DL) models with Modular Supercomputing Architecture (MSA) systems. To test the workflow, we considered the production of large-scale Land-Cover (LC) maps as a case study. The workflow manager, Airflow, offers scalability, extensibility, and programmable task definition in Python. It allows us to execute different steps of the workflow in different High-Performance Computing (HPC) systems. The workflow is demonstrated on the Dynamical Exascale Entry Platform (DEEP) and Jülich Research on Exascale Cluster Architectures (JURECA) hosted at the Jülich Supercomputing Centre (JSC), a platform that incorporates heterogeneous JSC systems.
Gaps in the measurement series of atmospheric pollutants can impede the reliable assessment of their impacts and trends. We propose a new method for missing data imputation of the air pollutant tropospheric ozone by using the graph machine learning algorithm "correct and smooth". This algorithm uses auxiliary data that characterize the measurement location and, in addition, ozone observations at neighboring sites to improve the imputations of simple statistical and machine learning models. We apply our method to data from 278 stations of the year 2011 of the German Environment Agency (Umweltbundesamt - UBA) monitoring network. The preliminary version of these data exhibits three gap patterns: shorter gaps in the range of hours, longer gaps of up to several months in length, and gaps occurring at multiple stations at once. For short gaps of up to 5 h, linear interpolation is most accurate. Longer gaps at single stations are most effectively imputed by a random forest in connection with the correct and smooth. For longer gaps at multiple stations, the correct and smooth algorithm improved the random forest despite a lack of data in the neighborhood of the missing values. We therefore suggest a hybrid of linear interpolation and graph machine learning for the imputation of tropospheric ozone time series.
Extreme air pollution events of high concentrated surface ozone (O3) or particulate matter (PM) pose a lethal threat to humans worldwide. To investigate air quality under extreme atmospheric situations, the DestinE-AQ use case develops a comprehensive user interface that enables high resolution air quality forecasts and analysis by combining numerical simulations, machine learning approaches and observations. The core of the system encloses access to the open database of global air quality observations (i.e. the Tropospheric Ozone Assessment Report data base, TOAR), innovative machine learning workflows (e.g. MLAir, IntelliO3-ts) including downscaling modules, and high resolution numerical simulations using the state of the art chemistry transport model EURAD-IM (EURopean Air pollution Dispersion- Inverse Model). The aspired air quality forecasts and analyses will dynamically be driven by the DestinE Digital Twin for weather extremes. Thus, the pursued horizontal resolution of the air quality simulations is identical to the resolution of the DestinE digital twin (< 1 km2).To provide reliable air quality analyses, the system makes use of observational data in both the machine learning tools and EURAD-IM by enabling 3D-var data assimilation. While focusing on Europe, the system will demonstrate how observations and physics-based and data-driven models can be woven together to achieve enhanced realism and finer resolution of air pollution information and thus provide better support to decision makers. The system is complemented by an efficient ensemble module that enables emission scenario simulations to test and develop air pollution mitigation strategies for future extreme events under realistic conditions. The development of the user interface is done in close cooperation with the German and North Rhine-Westphalian Environment Agencies to meet the end users’ needs. The system will provide detailed information about air quality, its underlying chemical processes, the influence of meteorological extreme events, and the impacts of anthropogenic emissions on air quality. Hence, it will also serve the scientific community to answer questions on air quality and atmospheric chemical processes under extreme weather conditions that are expected to increase in future. To allow for the investigation of the human impact of extreme events, the DestinE-AQ focuses on the key air pollutants PM2.5, nitrogen oxides (NOx), and O3 in the planetary boundary layer. The potential combination of the system with socio-economic and medical models will be evaluated.
<p>The representation of the atmospheric state at high spatial resolution is of particular relevance in various domains of Earth science. While global reanalysis datasets such as ERA5 provide comprehensive repositories of meteorological data, their spatial resolution (&#8710;x&#8805;25 km) is too coarse to capture relevant local features, mainly over complex terrain (e.g. cold pools in valleys, low-level jets, local heavy precipitation events).<br />Recently, various studies have started to apply deep neural networks adapted from computer vision to increase the spatial resolution of meteorological fields. Although these studies reveal great potential in the domain of statistical downscaling, intercomparison of the approaches is impeded due to a large variety of methods and deployed datasets. Comparisons to classical downscaling methods developed for decades in the meteorological community are also often underrepresented.<br /><br />Inspired by the available benchmark datasets for various computer vision tasks and for weather forecasting (e.g. WeatherBench and WeatherBench Probability), our study aims to provide a benchmark dataset for statistical downscaling of meteorological fields. We choose the coarse-grained ERA5 reanalysis (&#8710;x<sub>ERA5</sub>&#8771;30 km) and the fine-scaled COSMO-REA6 (&#8710;x<sub>CREA6</sub>&#8771;6km) as input and target datasets. Both datasets enable the formulation of a real downscaling task: super-resolve the data and correct for model biases.<br />The benchmark dataset provides a collection of predictors and predictands for a couple of standard downscaling tasks. These comprise downscaling of the 2m temperature, the surface irradiance, the near-surface wind field and precipitation. Along with the dataset, benchmark deep neural networks, namely variants of U-Nets and GANs, will be provided. Well-chosen sets of evaluation metrics including baseline scores of the benchmarked deep neural networks are presented to enable comparison between different methods.<br />The envisioned benchmark dataset will provide a comprehensive basis for comparing neural network approaches on statistical downscaling of meteorological fields. This, in turn, is considered to enhance confidence and transparency in the application of deep learning methods on Earth system problems.</p>
Over the last years, several repositories with curated environmental datasets have been created so that scientific communities have gained access to large collections of data from various domains. The level of data harmonisation and FAIRness, technical readiness and scalability of these repositories differs substantially. This restricts data exploration opportunities and limits scientific exploration with modern data science methods, such as machine learning. In the domain of air quality research, we have pioneered a data infrastructure for global observations of surface ozone and other air pollutant measurements that comes with rich possibilities for online data analysis. The data in the Tropospheric Ozone Assessment Report (TOAR) database is collected from about 40 different resource providers, from national and international environmental agencies to individual research groups around the world.One of these data providers is OpenAQ, the world's first open, real-time air quality platform. Due to the higher standards of curation, the need for data harmonization, and the enriched metadata in the TOAR database, we had to develop an automated workflow to transport archived and real-time data from this provider to the TOAR database. The primary step is to clean and format all the OpenAQ records, according to the TOAR database schema, and concurrently, refine the metadata. The workflow includes tests for data sanity and checks if time series and station metadata can be amended, or whether new time series or station records must be created. The automation manager triggers the workflow hourly, so the database provides clean and updated air quality data at any time. The presentation describes the automated workflow and its design principles and discusses how such a workflow might be re-used in other environmental domains. All TOAR-related codes are open source.
The atmosphere affects humans in a multitude of ways, from loss of life due to adverse weather effects to long-term social and economic impacts on societies. Computer simulations of atmospheric dynamics are, therefore, of great importance for the well-being of our and future generations. Here, we propose AtmoRep, a novel, task-independent stochastic computer model of atmospheric dynamics that can provide skillful results for a wide range of applications. AtmoRep uses large-scale representation learning from artificial intelligence to determine a general description of the highly complex, stochastic dynamics of the atmosphere from the best available estimate of the system's historical trajectory as constrained by observations. This is enabled by a novel self-supervised learning objective and a unique ensemble that samples from the stochastic model with a variability informed by the one in the historical record. The task-independent nature of AtmoRep enables skillful results for a diverse set of applications without specifically training for them and we demonstrate this for nowcasting, temporal interpolation, model correction, and counterfactuals. We also show that AtmoRep can be improved with additional data, for example radar observations, and that it can be extended to tasks such as downscaling. Our work establishes that large-scale neural networks can provide skillful, task-independent models of atmospheric dynamics. With this, they provide a novel means to make the large record of atmospheric observations accessible for applications and for scientific inquiry, complementing existing simulations based on first principles.
Estimates of ground-level ozone concentrations have been improved through data fusion of observations and atmospheric chemistry models. Our previous global ozone estimates for the Global Burden of Disease study corrected for bias uniformly across continents and then corrected near monitoring stations using the Bayesian Maximum Entropy (BME) framework for data fusion. Here, we use the Regionalized Air Quality Model Performance (RAMP) framework to correct model bias over a much larger spatial range than BME can, accounting for the spatial inhomogeneity of bias and nonlinearity as a function of modeled ozone. RAMP bias correction is applied to a composite of 9 global chemistry-climate models, based on the nearest set of monitors. These estimates are then fused with observations using BME, which matches observations at measurement stations, with the influence of observations declining with distance in space and time. We create global ozone maps for each year from 1990 to 2017 at fine spatial resolution. RAMP is shown to create unrealistic discontinuities due to the spatial clustering of ozone monitors, which we overcome by applying a weighting for RAMP based on the number of monitors nearby. Incorporating RAMP before BME has little effect on model performance near stations, but strongly increases R2 by 0.15 at locations farther from stations, shown through a checkerboard cross-validation. Corrections to estimates differ based on location in space and time, confirming heterogeneity. We quantify the likelihood of exceeding selected ozone levels, finding that parts of the Middle East, India, and China are most likely to exceed 55 parts per billion (ppb) in 2017. About 96% of the global population was exposed to ozone levels above the World Health Organization guideline of 60 µg m−3 (30 ppb) in 2017. Our annual fine-resolution ozone estimates may be useful for several applications including epidemiology and assessments of impacts on health, agriculture, and ecosystems.