The year 2023 represents a significant milestone in climate history: it was indeed confirmed by the Copernicus Climate Change Service (C3S) as the warmest calendar year in global temperature data records since 1850. With a deviation of 1.48ºC from the 1850-1900 pre-industrial level, 2023 largely surpasses 2016, 2019, 2020, previously identified as the warmest years on record. As expected, this sustained warmth leads to an increase in frequency and intensity of Extreme Events (EE) with dramatic environmental and societal consequences.To assess the evolution of these EE and establish adaptation and mitigation strategies, it is crucial to evaluate the trends of extreme indices (EI). However, the observational climate data that are commonly used for the calculation of these indices frequently contains missing values, resulting in partial and inaccurate EI. As we delve deeper into the past, this issue becomes more pronounced due to the scarcity of historical measurements.To circumvent the lack of information, we are using a deep learning technique based on a U-Net made of partial convolutional layers [1]. Models are trained with Earth system model data from CMIP6 and has the capability to reconstruct large and irregular regions of missing data using minimal computational resources. This approach has shown its ability to outperform traditional statistical methods such as Kriging by learning intricate patterns in climate data [2].In this study, we have applied our technique to the reconstruction of gridded land surface EI from an intermediate product of the HadEX3 dataset [3]. This intermediate product is obtained by combining station measurements without interpolation, resulting in numerous missing values that varies in both space and time. These missing values affect significantly the calculation of the long-term linear trend (1901-2018), especially if we consider solely the grid boxes containing values for the whole time period. The trend calculated for the TX90p index that measures the monthly (or annual) frequency of warm days (defined as a percentage of days where daily maximum temperature is above the 90th percentile) is presented for the European continent on the left panel of the figure. It illustrates the resulting amount of missing values indicated by the gray pixels. With our AI method, we have been able to reconstruct the TX90p values for all the time steps and calculate the long-term trend shown on the right panel of the figure. The reconstructed dataset is being prepared for the community in the framework of the H2020 CLINT project [4] for further detection and attribution studies.[1] Liu G. et al., Lecture Notes in Computer Science, 11215, 19-35 (2018)[2] Kadow C. et al., Nat. Geosci., 13, 408-413 (2020)[3] Dunn R. J. H. et al., J. Geophys. Res. Atmos., 125, 1 (2020)[4] https://climateintelligence.eu/
Obtaining accurate estimates of uncertainty in climate scenarios often requires generating large ensembles of high-resolution climate simulations, a computationally expensive and memory intensive process. To address this challenge, we train a novel generative deep learning approach on extensive sets of climate simulations. The model consists of two components: a variational autoencoder for dimensionality reduction and a denoising diffusion probabilistic model that generates multiple ensemble members. We validate our model on the Max Planck Institute Grand Ensemble and show that it achieves good agreement with the original ensemble in terms of variability. By leveraging the latent space representation, our model can rapidly generate large ensembles on-the-fly with minimal memory requirements, which can significantly improve the efficiency of uncertainty quantification in climate simulations.
Weather radars are a significant component of modern precipitation recordings,as they provide information with high spatial and temporal resolution. However, radars as a tool for weather applications emerged only after the 1950s. AI/ML methods have proven to be successful when it comes to determining patterns and connections between related fields in space and time. Moreover, AI/ML methods have exhibited remarkable skill in infilling missing climate information (see Kadow et al. 2020). Desired outcomes of the project include using these AI/ML techniques to build a spatial precipitation field by combining station and radar data. We will use data from two well-known datasets: RADOLAN and COSMO-REA2. The validity of this digital twin will be investigated by comparing its output with other reanalysis data (e.g. ERA5). Further evaluation can be carried out by testing the radar field’s accuracy in detecting extreme precipitation events in the past (e.g. heavy rain events in the summer of 2021 in Western Germany). We aim for the creation of a radar field that will be successfully projected in the past. Moreover, it will uncover new information on regional climatology, especially in areas where station data is sparse.
Survivors of childhood cancer are at risk for therapy-related subsequent malignant neoplasms (SMN). Exposure to radiation therapy and certain chemotherapy molecules are known risk factors for SMN development. However, there remains inter-individual variability in these treatment related SMN that is attributed to genetic variations. The aim of our study is to identify rare genetic variants associated with risk of SMN and their interactions with treatment of the primary cancer in childhood. We conducted a nested case-control study of SMN within the French Childhood Cancer Survivors Study (FCCSS) cohort with 163 cases and 287 controls. Whole exome sequencing was realized on these 450 cancer survivors. The mean radiation dose at the site of interest and doses of each chemotherapeutic agent were calculated for all the survivors. In order to have sufficient power to detect an association between rare variants and the risk of SMN, variants were grouped by genetic unit to identify genes enriched in variants in cases and controls. Three statistical methods accounting for the specific characteristics of rare variant studies were tested, all of them are gene-based association tests: a) Burden tests, more powerful when the variants are causal and in the same direction; b) Variant-component tests (e.g., SKAT), more powerful in the presence of a mixture of variants with deleterious and protective effects; c) Combined tests of the two previous approaches (e.g., SKATO). All tests were adjusted for sex, age at diagnosis and year of diagnosis of primary cancer, type of first cancer, length of follow-up, and the first 4 principal components of the PCA to account for population stratification. The results will be presented as odds ratios (ORs) with their 95% confidence intervals (CIs). Genes significantly associated with SMN status will be further analyzed for interaction with radiotherapy and chemotherapy. The mean age at diagnosis of the first cancer was 7 years for cases and controls. The length of follow-up is similar for cases and controls with a median of 24 years. A total of 151 cases (92.7 %) and 194 controls (67.5 %) had received radiation therapy. Association tests identified 19 genes enriched in rare variants in cases with promising associations with the risk of SMN. Further analysis of these genes is underway, as well as analysis focusing on the most common SMN type (thyroid: 40.7 % and breast: 29.1 %). Interaction studies with spleen, active bone marrow radiation doses and radiation doses received at the site of the second cancer will be performed as well as with doses of selected chemotherapy drugs. This study will provide new evidence regarding the involvement of genetics in the development of SMN in interaction with the first cancer treatment received. This could allow the identification of survivors at higher risk for SMN development in order to adapt treatment and/or follow-up.
Computational climate science depends on ever stronger high performance computers combined with high volume data storage. However, the development of new technology slows down. The fast increase in computational performance that we were used to is no longer possible with current semiconductor technology. The talk will address the requirements of computational climate science and the current bottlenecks that limit science. We will discuss the latest trends in the field of high performance computing and have a look on the future of HPC systems. Recently, the method of machine learning also gets deployed in climate science, too. We will see how this might help us to more efficiently use existing and upcoming hardware. As Moore's law is dead for semiconductors, perhaps machine learning will bring us the necessary speedup.
Missing climate data is a widespread problem in climate science and leads to uncertainty of prediction models that rely on these data resources. So far, existing approaches for infilling missing precipitation data are mostly numerical or statistical techniques that require considerable computational resources and are not suitable for large regions with missing data. Most recently, there have been several approaches to infill missing climate data with machine learning methods such as convolutional neural networks or generative adversarial networks. They have proven to perform well on infilling missing temperature or satellite data. However, these techniques consider only spatial variability in the data whereas precipitation data is much more variable in both space and time. Rainfall extremes with high amplitudes play an important role. We propose a convolutional inpainting network that additionally considers a memory module. One approach investigates the temporal variability in the missing data regions using a long-short term memory. An attention-based module has also been added to the technology to consider further atmospheric variables provided by reanalysis data. The model was trained and evaluated on the RADOLAN data set which is based on radar precipitation recordings and weather station measurements. With the method we are able to complete gaps in this high quality, highly resolved spatial precipitation data set over Germany. In conclusion, we compare our approach to statistical techniques for infilling precipitation data as well as other state-of-the-art machine learning techniques. This well-combined technology of computer and atmospheric research components will be presented as a dedicated climate service component and data set.
Cyber-infrastructures have changed the process of research. Researchers can now access distributed data from all parts of the world with the help of cyber-infrastructures. User support services play an important role to facilitate researchers to accomplish their research goals with the help of cyber-infrastructures. However, the current user-support practices in cyber-infrastructures are being followed on intuitive basis (at least in climate e-infrastructures) thus over-burdening cyber-infrastructure employees. The main contribution of this paper is to present the snap-shot of the current user support practices in a cyber-infrastructure of a climate science known as Earth System Grid Federation (ESGF). ESGF is a leading distributed peer-to-peer (P2P) data-grid system in Earth System Modelling (ESM) having around 2700 users distributed worldwide. The questionnaire conducted with the climate cyber-infrastructure projects’ employees presents the picture of the current user support situation by highlighting their profile, utilization of various communication media, user-request service time, attributes of incoming user problems and information requests. The respondents of the questionnaire were 26 support staffs of cyber-infrastructure projects, from different parts of the world. The paper then presents the critique of the current user support process in ESGF and finally emphasizes on the need to streamline user-support in cyber-infrastructures.
Computing power and complexity of HPC systems are steadily increasing. This leads to an increasing demand for a good education of their users so that they can use such systems adequately. A special challenge is to provide users with skills according to their scientific backgrounds and specific demands in terms of the usage of the system. Users in the role of testers, who want to simply run a parallel program for benchmark purposes, must e.g. have a solid knowledge of operating system basics and should be able to use a workload manager like SLURM [SLUR 17] or TORQUE [TORQ 17], but in general they do not need a deeper understanding of the technical refinements of the parallelization of the program. A user who wants to develop a parallel program will usually already be able to use the operating system and the workload manager but will need further skills to apply parallelization techniques like OpenMP [OpMP 17] or GPU-computing based on CUDA [NVID 17] at the intra-node level, MPI [MPI 17] at the inter-node level or even combinations of such techniques in the sense of a hybrid or multi-level approach.
Scientific discovery increasingly depends on complex workflows consisting of multiple phases and sometimes millions of parallelizable tasks or pipelines. These workflows access storage resources for a variety of purposes, including preprocessing, simulation output, and postprocessing steps. Unfortunately, most workflow models focus on the scheduling and allocation of computational resources for tasks while the impact on storage systems remains a secondary objective and an open research question. I/O performance is not usually accounted for in workflow telemetry reported to users. In this paper, we present an approach to augment the I/O efficiency of the individual tasks of workflows by combining workflow description frameworks with system I/O telemetry data. A conceptual architecture and a prototype implementation for HPC data center deployments are introduced. We also identify and discuss challenges that will need to be addressed by workflow management and monitoring systems for HPC in the future. We demonstrate how real-world applications and workflows could benefit from the approach, and we show how the approach helps communicate performance-tuning guidance to users.
—Both energy and storage are becoming key issues in high-performance (HPC) systems, especially when thinking about upcoming Exascale systems. The amount of energy consumption and storage capacity needed to solve future problems is growing in a marked curve that the HPC community must face in cost-/energy-efficient ways. In this paper we provide a power-performance evaluation of HPC storage servers that take over tasks other than simply storing the data to disk. We use the Lustre parallel distributed file system with its ZFS back-end, which natively supports compression, to show that data compression can help to alleviate capacity and energy problems. In the first step of our analysis we study different compression algorithms with regards to their CPU and power overhead with a real dataset. Then, we use a modified version of the IOR benchmark to verify our claims for the HPC environment. The results demonstrate that the energy consumption can be reduced by up to 30% in the write phase of the application and 7% for write-intensive applications. At the same time, the required storage capacity can be reduced by approximately 50%. These savings can help in designing more power-efficient and leaner storage systems.
The performance of parallel distributed file systems suffers from many clients executing a large number of operations in parallel, because the I/O subsystem can be easily overwhelmed by the sheer amount of incoming I/O operations. Many optimizations exist that try to alleviate this problem. Client-side optimizations perform preprocessing to minimize the amount of work the file servers have to do. Server-side optimizations use server-internal knowledge to improve performance. The HD Trace framework contains components to simulate, trace and visualize applications. It is used as a test bed to evaluate optimizations that could later be implemented in real-life projects. This paper compares existing client-side optimizations and newly implemented server-side optimizations and evaluates their usefulness for I/O patterns commonly found in HPC. Server-directed I/O chooses the order of non-contiguous I/O operations and tries to aggregate as many operations as possible to decrease the load on the I/O subsystem and improve overall performance. The results show that server-side optimizations beat client-side optimizations in terms of performance for many use cases. Integrating such optimizations into parallel distributed file systems could alleviate the need for sophisticated client-side optimizations. Due to their additional knowledge of internal workflows server-side optimizations may be better suited to provide high performance in general.
Th. Bemmerl合作论文数Lehrstuhl für Betriebssysteme
RWTH Aachen
Gebäude AVZ2