Large-scale simulations or scientific experiments produce petabytes of data per run. This poses massive challenges for I/O and storage when scientific analysis workflows are run manually offline. Unsupervised deep learning-based techniques to extract patterns and non-linear relations from these large amounts of data provide a way to build scientific understanding from raw data, reducing the need for manual pre-selection of analysis steps, but require exascale compute and memory to process the full dataset available. In this paper, we demonstrate a heterogeneous streaming workflow in which plasma simulation data is streamed directly to a Machine Learning (ML) application training a model on the simulation data in-transit, completely circumventing the capacity-constrained filesystem bottleneck. This workflow employs openPMD to provide a high level interface to describe scientific data and also uses ADIOS2, to transfer volumes of data that exceed the capabilities of the filesystem. We employ experience replay to avoid catastrophic forgetting in learning from this non-steady state process in a continual manner and adapt it to improve model convergence while learning in-transit. As a proof-of-concept, we approach the ill-posed inverse problem of predicting particle dynamics from radiation in a particle-in-cell (PIConGPU) simulation of the Kelvin-Helmholtz instability (KHI). We detail hardware-software co-design challenges as we scale PIConGPU to full Frontier, the Top-1 system as of June 2024 Top500 list.
We present two multilingual LLMs, Teuken 7B-base and Teuken 7B-instruct, designed to embrace Europe’s linguistic diversity by supporting all 24 official languages of the European Union. Trained on a dataset comprising around 60% non-English data and utilizing a custom multilingual tokenizer, our models address the limitations of existing Large Language Models (LLMs) that predominantly focus on English or a few high-resource languages. We detail the models’ development principles, i.e., data composition, tokenizer optimization, and training methodologies. The models demonstrate strong performance across multilingual benchmarks, as evidenced by their performance on European versions of ARC, HellaSwag, and TruthfulQA.
The recent success of Large Language Models (LLMs) has been predominantly driven by curating the training dataset composition, scaling of model architectures and dataset sizes and advancements in pretraining objectives, leaving tokenizer influence as a blind spot. Shedding light on this underexplored area, we conduct a comprehensive study on the influence of tokenizer choice on LLM downstream performance by training 24 mono- and multilingual LLMs at a 2.6B parameter scale, ablating different tokenizer algorithms and parameterizations. Our studies highlight that the tokenizer choice can significantly impact the model's downstream performance and training costs. In particular, we find that the common tokenizer evaluation metrics fertility and parity are not always predictive of model downstream performance, rendering these metrics a questionable proxy for the model's downstream performance. Furthermore, we show that multilingual tokenizers trained on the five most frequent European languages require vocabulary size increases of factor three in comparison to English. While English-centric tokenizers have been applied to the training of multi-lingual LLMs in the past, we find that this approach results in a severe downstream performance degradation and additional training costs of up to 68%, due to an inefficient tokenization vocabulary.
Ion Beam Analysis (IBA) utilizing MeV ion beams provides valuable insights into surface elemental composition across the entire periodic table. While ion beam measurements have advanced towards high throughput for mapping applications, data analysis has lagged behind due to the challenges posed by large volumes of data and multiple detectors providing diverse analytical information. Traditional physics-based fitting algorithms for these spectra can be time-consuming and prone to local minima traps, often taking days or weeks to complete. This study presents an approach employing a Mixture Density Network (MDN) to model the posterior distribution of Elemental Depth Profiles (EDP) from input spectra. Our MDN architecture includes an encoder module (EM), leveraging a Convolutional Neural Network-Gated Recurrent Unit (CNN-GRU), and a Mixture Density Head (MDH) employing a Multi-Layer Perceptron (MLP). Validation across three datasets with varying complexities demonstrates that for simple and intermediate cases, the MDN performs comparably to the conventional automatic fitting method (Autofit). However, for more complex datasets, Autofit still outperforms the MDN. Additionally, our integrated approach, combining MDN with the automatic fit method, significantly enhances accuracy while still reducing computational time, offering a promising avenue for improved analysis in IBA.
The adaption of multilingual pre-trained LLMs into eloquent and helpful assistants is essential to facilitate their use across different language regions. In that spirit, we are the first to conduct an extensive study of the performance of multilingual models instruction-tuned on different language compositions on parallel instruction-tuning benchmarks across a selection of the most spoken Indo-European languages. We systematically examine the effects of language and instruction dataset size on a mid-sized and a large, multilingual LLMs by instruction-tuning them on parallel instruction-tuning datasets. Our results demonstrate that instruction-tuning on parallel instead of monolingual corpora benefits cross-lingual instruction following capabilities by up to 9.9%. Furthermore, we show that the Superficial Alignment Hypothesis does not hold in general, as the investigated multilingual 7B parameter model presents a counter-example requiring large-scale instruction-tuning datasets. Finally, we conduct a human annotation study to understand the alignment between human-based and GPT-4-based evaluation within multilingual chat scenarios.
The BigCode community, an open-scientific collaboration working on the responsible development of Large Language Models for Code (Code LLMs), introduces StarCoder and StarCoderBase: 15.5B parameter models with 8K context length, infilling capabilities and fast large-batch inference enabled by multi-query attention. StarCoderBase is trained on 1 trillion tokens sourced from The Stack, a large collection of permissively licensed GitHub repositories with inspection tools and an opt-out process. We fine-tuned StarCoderBase on 35B Python tokens, resulting in the creation of StarCoder. We perform the most comprehensive evaluation of Code LLMs to date and show that StarCoderBase outperforms every open Code LLM that supports multiple programming languages and matches or outperforms the OpenAI code-cushman-001 model. Furthermore, StarCoder outperforms every model that is fine-tuned on Python and still retains its performance on other programming languages. We take several important steps towards a safe open-access model release, including an improved PII redaction pipeline and a novel attribution tracing tool, and make the StarCoder models publicly available under a more commercially viable version of the Open Responsible AI Model license.
One of the main challenges in optimal scaling of large language models (LLMs) is the prohibitive cost of hyperparameter tuning, particularly learning rate η and batch size B. While techniques like μP (Yang et al., 2022) provide scaling rules for optimal η transfer in the infinite model size limit, the optimal scaling behavior in the infinite data size limit (T →∞) remains unknown. We fill in this gap by observing for the first time an interplay of three optimal η scaling regimes: η∝√(T), η∝ 1, and η∝ 1/√(T) with transitions controlled by B and its relation to the time-evolving critical batch size B_crit∝ T. Furthermore, we show that the optimal batch size is positively correlated with B_crit: keeping it fixed becomes suboptimal over time even if learning rate is scaled optimally. Surprisingly, our results demonstrate that the observed optimal η and B dynamics are preserved with μP model scaling, challenging the conventional view of B_crit dependence solely on loss value. Complementing optimality, we examine the sensitivity of loss to changes in learning rate, where we find the sensitivity to decrease with T →∞ and to remain constant with μP model scaling. We hope our results make the first step towards a unified picture of the joint optimal data and model scaling.
Concentrating solar power plants are a clean energy source capable of competitive electricity generation even during night time, as well as the production of carbon-neutral fuels, offering a complementary role alongside photovoltaic plants. In these power plants, thousands of mirrors (heliostats) redirect sunlight onto a receiver, potentially generating temperatures exceeding 1000°C. Practically, such efficient temperatures are never attained. Several unknown, yet operationally crucial parameters, e.g., misalignment in sun-tracking and surface deformations can cause dangerous temperature spikes, necessitating high safety margins. For competitive levelized cost of energy and large-scale deployment, in-situ error measurements are an essential, yet unattained factor. To tackle this, we introduce a differentiable ray tracing machine learning approach that can derive the irradiance distribution of heliostats in a data-driven manner from a small number of calibration images already collected in most solar towers. By applying gradient-based optimization and a learning non-uniform rational B-spline heliostat model, our approach is able to determine sub-millimeter imperfections in a real-world setting and predict heliostat-specific irradiance profiles, exceeding the precision of the state-of-the-art and establishing full automatization. The new optimization pipeline enables concurrent training of physical and data-driven models, representing a pioneering effort in unifying both paradigms for concentrating solar power plants and can be a blueprint for other domains.
The continuous stream of high spatial resolution satellite data offers the opportunity to regularly produce land cover (LC) maps. To this end, Transformer deep learning (DL) models have recently proven their effectiveness in accurately classifying long time series (TS) of satellite images. The continual generation of regularly updated LC maps can be used to analyze dynamic phenomena and extract multi-temporal information. However, several challenges need to be addressed. Our paper aims to study how the performance of a Transformer model changes when classifying TS of satellite images acquired in years later than those in the training set. In particular, the behavior of the attention in the Transformer model is analyzed to determine when the information provided by the initial training set needs to be updated to keep generating accurate LC products. Preliminary results show that: (i) the selection of the positional encoding strategy used in the Transformer has a significant impact on the classification accuracy obtained with multi-year TS, and (ii) the most affected classes are the seasonal ones.
Land Cover (LC) maps generated by the classification of Remote Sensing (RS) data allow for the monitoring of Earth processes and the dynamics of objects and phenomena. Environmental monitoring applications can implement accurate quantification of LC variability when maps are spatiotemporally consistent and are continuously updated, as they provide information on consistent and permanent LC changes. However, the production of frequent and spatiotemporally consistent LC maps is challenging because it involves balancing the need for temporal consistency with the risk of missing real changes. In this work, we propose a scalable and semi-automatic method for generating annual maps with labels that are consistently applied from one year to the next. It uses a Transformer Deep Learning (DL) model as a classifier, which is trained on satellite Time Series (TS) data using High Performance Computing (HPC). The trained model is able to generate stable land cover maps by shifting the prediction window along the temporal direction. We test the method on a Sentinel-2 dataset acquired over a three-year period and demonstrate that: (1) the annual maps can be directly compared to detect changes and, (2) the accuracy of previously generated maps can be improved via a backpropagation strategy.
Solar tower power plants play a key role to facilitate the ongoing energy transition as they deliver climate neutral electricity and direct heat for chemical processes. These plants generate temperatures over 1000 °C by reflecting sunlight with thousands of mirrors (heliostats) to a receiver. The temper- ature achievable in practice is limited due to the system’s susceptibility to small surface defects and misalignments of individual mirrors, hindering the plant’s full efficiency. We present an inverse render- ing technique that predicts the incident power distribution of each heliostat, including the inaccuracies, based solely on focal spot images that are already acquired in most solar power plants. The method allows reconstructing flawed mirror shapes within sub-mm precision. Applied at the solar tower plant in Juelich, our approach outperforms all alternatives in accuracy and reliability. Our data-driven method is a key ingredient to building digital twins of solar power plants. It can be integrated into the existing infrastructure and plant control at low cost, leading to increased efficiency of existing and decreased expenses for future power plants, the key factors of success in the competitive market. For other fields, our approach can be a blueprint, as we present the option for the very first large-scale indus- trial deployment of differentiable ray tracing. Merging data-intensive Machine Learning with physical modeling creates flexible, data-efficient and trustworthy solutions applicable in science and industry.
Each solar tower power plant is designed for a pre-calculated optimal flux density distribution.Any deviation from this has a direct impact on the output power as well as the durability of the components.An accurate knowledge of the current and predicted flux density is therefore essential.Also, because this is one of the most important input variables for all subsequent power plant processes.But due to individual errors of each heliostat, this theoretical flux density is very difficult to obtain.This includes, that common methods for measuring the flux density are either inaccurate, complicated, or expensive.Although raytracers exists, which can predict a flux density by analytical calculations, such an approach does not reflect reality sufficiently.We present a novel AI based method to predict the flux density map, which is capable to include heliostat specific errors, without having to measure the heliostats surface.Furthermore, we compare the advantages and disadvantages of different network structures for this approach and show first results archiving at best a Peak Signal to Noise Ratio (PSNR) value of up to 27.8 using neural radiance fields (NeRFs) .Prediction Target
The BigCode community, an open-scientific collaboration working on the responsible development of Large Language Models for Code (Code LLMs), introduces StarCoder and StarCoderBase: 15.5B parameter models with 8K context length, infilling capabilities and fast large-batch inference enabled by multi-query attention. StarCoderBase is trained on 1 trillion tokens sourced from The Stack, a large collection of permissively licensed GitHub repositories with inspection tools and an opt-out process. We fine-tuned StarCoderBase on 35B Python tokens, resulting in the creation of StarCoder. We perform the most comprehensive evaluation of Code LLMs to date and show that StarCoderBase outperforms every open Code LLM that supports multiple programming languages and matches or outperforms the OpenAI code-cushman-001 model. Furthermore, StarCoder outperforms every model that is fine-tuned on Python, can be prompted to achieve 40\% pass@1 on HumanEval, and still retains its performance on other programming languages. We take several important steps towards a safe open-access model release, including an improved PII redaction pipeline and a novel attribution tracing tool, and make the StarCoder models publicly available under a more commercially viable version of the Open Responsible AI Model license.
The Vlasov-Poisson system is employed in its reduced form version (1D1V) as a test bed for the applicability of Physics Informed Neural Network (PINN) to the wave-particle resonance. Two examples are explored: the Landau damping and the bump-on-tail instability. PINN is first tested as a compression method for the solution of the Vlasov-Poisson system and compared to the standard neural networks. Second, the application of PINN to solving the Vlasov-Poisson system is also presented with the special emphasis on the integral part, which motivates the implementation of a PINN variant, called Integrable PINN (I-PINN), based on the automatic-differentiation to solve the partial differential equation and on the automatic-integration to solve the integral equation.
The camera target method is the most commonly used calibration method for heliostats at solar tower power plants to minimize their sun tracking errors. In this method, individual heliostats are moved to a white surface and their deviation from the targeted position is measured. A regression is used to calculate errors in a geometry model from the tabular data obtained in this way. For modern aim point strategies, or simply heliostats in the rearmost end of the field, extremely high accuracies are needed, which can only be achieved by many degrees of freedom in the geometry model. The problem here is that the camera target method produces only a very small data set per heliostat, which limits the number of free variables and thus the accuracy. In this work, we extend existing ray tracing methods for solar towers with a differentiable description, allowing for the first time a data-driven optimization of object parameters within the ray tracing environment. Therefore, the heliostat calibration can take place directly within the ray tracing environment. Thus, the image data acquired during the measurement can be processed directly and more information about the orientation of the heliostat can be obtained. Within a simple example we show the advantages of the method, which converges faster and corrects errors that could not be considered before. Without any disadvantages or additional costs, the state-of-the-art calibration method can be improved.
Amidst the COVID-19 pandemic, the authors of this paper organized a Reinforcement Learning (RL) course for a graduate school in the field of data science. We describe the strategy and materials for creating an exciting learning experience despite the ubiquitous Zoom fatigue and evaluate the course qualitatively. The key organizational features are a focus on a competitive hands-on setting in teams, supported by a minimum of lectures providing the essential background on RL. The practical part of the course revolved around Hearts Gym, an RL environment for the card game Hearts that we developed as an entry-level tutorial to RL. Participants were tasked with training agents to explore reward shaping and other RL hyperparameters. For a final evaluation, the agents of the participants competed against each other.
In this article, we present JUWELS Booster, a recently commissioned high-performance computing system at the Jülich Supercomputing Center. With its system architecture, most importantly its large number of powerful Graphics Processing Units (GPUs) and its fast interconnect via InfiniBand, it is an ideal machine for large-scale Artificial Intelligence (AI) research and applications. We detail its system architecture, parallel, distributed model training, and benchmarks indicating its outstanding performance. We exemplify its potential for research application by presenting large-scale AI research highlights from various scientific fields that require such a facility.