Predicting the extrinsic properties from hysteresis loops of a magnetic grain, namely the coercive field, remanent magnetisation, and maximum energy product, from its intrinsic micromagnetic parameters is a central problem in permanent-magnet modelling. Established analytical models provide useful estimates but often neglect nonuniform magnetisation processes, whereas direct micromagnetic simulations are computationally expensive. In this work, we train machine-learning models on 12012 micromagnetic simulations of an idealised cubic grain, spanning broad ranges of the saturation magnetisation, exchange constant, and uniaxial anisotropy constant. Benchmarked against the analytical models on identical held-out data, the machine-learning models predict all three extrinsic properties with substantially lower errors. Symbolic regression recovers the Kronmüller form of the coercive field, with an effective demagnetising factor that depends on the material, and finds new closed-form expressions for the remanence and maximum energy product. Each law contains at most two fitted constants yet approaches the accuracy of the machine-learning models. We also investigate the inverse problem of recovering the intrinsic parameters from the three extrinsic properties. The saturation magnetisation and anisotropy constant are recovered accurately, whereas the exchange constant is not, because it influences the extrinsic properties only weakly. The trained models are released through the mammos-ai Python package, enabling thousands of candidate parameter sets to be screened in seconds rather than the hours or days required by direct micromagnetic simulation.
Electronic-structure theory is the foundation of the description of materials including multiscale modeling of their properties and functions. Obviously, without sufficient accuracy at the base, reliable predictions are unlikely at any level that follows. The software package FHI-aims has proven to be a game changer for accurate free-energy calculations because of its scalability, numerical precision, and its efficient handling of density functional theory (DFT) with hybrid functionals and van der Waals interactions. It treats molecules, clusters, and extended systems (solids and liquids) on an equal footing. Besides DFT, FHI-aims also includes quantum-chemistry methods, descriptions for excited states and vibrations, and calculations of various types of transport. Recent advancements address the integration of FHI-aims into an increasing number of workflows and various artificial intelligence (AI) methods. This Roadmap describes the state-of-the-art of FHI-aims and advancements that are currently ongoing or planned.
In this contribution, we give an overview of the ELPA library and ELSI interface, which are crucial elements for large-scale electronic structure calculations in FHI-aims. ELPA is a key solver library that provides efficient solutions for both standard and generalized eigenproblems, which are central to the Kohn-Sham formalism in density functional theory (DFT). It supports CPU and GPU architectures, with full support for NVIDIA and AMD GPUs, and ongoing development for Intel GPUs. Here we also report the results of recent optimizations, leading to significant improvements in GPU performance for the generalized eigenproblem. ELSI is an open-source software interface layer that creates a well-defined connection between "user" electronic structure codes and "solver" libraries for the Kohn-Sham problem, abstracting the step between Hamilton and overlap matrices (as input to ELSI and the respective solvers) and eigenvalues and eigenvectors or density matrix solutions (as output to be passed back to the "user" electronic structure code). In addition to ELPA, ELSI supports solvers including LAPACK and MAGMA, the PEXSI and NTPoly libraries (which bypass an explicit eigenvalue solution), and several others.
Science is and always has been based on data, but the terms "data-centric" and the "4th paradigm of" materials research indicate a radical change in how information is retrieved, handled and research is performed. It signifies a transformative shift towards managing vast data collections, digital repositories, and innovative data analytics methods. The integration of Artificial Intelligence (AI) and its subset Machine Learning (ML), has become pivotal in addressing all these challenges. This Roadmap on Data-Centric Materials Science explores fundamental concepts and methodologies, illustrating diverse applications in electronic-structure theory, soft matter theory, microstructure research, and experimental techniques like photoemission, atom probe tomography, and electron microscopy. While the roadmap delves into specific areas within the broad interdisciplinary field of materials science, the provided examples elucidate key concepts applicable to a wider range of topics. The discussed instances offer insights into addressing the multifaceted challenges encountered in contemporary materials research.
Quantitative analysis of microstructural features on the nanoscale, including precipitates, local chemical orderings (LCOs) or structural defects (e.g. stacking faults) plays a pivotal role in understanding the mechanical and physical responses of engineering materials. Atom probe tomography (APT), known for its exceptional combination of chemical sensitivity and sub-nanometer resolution, primarily identifies microstructures through compositional segregations. However, this fails when there is no significant segregation, as can be the case for LCOs and stacking faults. Here, we introduce a 3D deep learning approach, AtomNet, designed to process APT point cloud data at the single-atom level for nanoscale microstructure extraction, simultaneously considering compositional and structural information. AtomNet is showcased in segmenting L12-type nanoprecipitates from the matrix in an AlLiMg alloy, irrespective of crystallographic orientations, which outperforms previous methods. AtomNet also allows for 3D imaging of L10-type LCOs in an AuCu alloy, a challenging task for conventional analysis due to their small size and subtle compositional differences. Finally, we demonstrate the use of AtomNet for revealing 2D stacking faults in a Co-based superalloy, without any stacking-faults-relevant samples in the training dataset, expanding the capabilities for automated exploration of hidden microstructures in APT data. AtomNet can thus recognize challenging microstructures, including nanoprecipitates with diameters above 2 nm, LCOs with diameters of about 1–2 nm without obvious compositional segregation, and even unforeseen planar defects by analyzing atom-atom environments. AtomNet pushes the boundaries of APT analysis, and holds promise in establishing precise quantitative microstructure-property relationships across a diverse range of metallic materials.
In solids, chemical short-range order (CSRO) refers to the self-organization of atoms of certain species occupying specific crystal sites. CSRO is increasingly being envisaged as a lever to tailor the mechanical and functional properties of materials. Yet quantitative relationships between properties and the morphology, number density, and atomic configurations of CSRO domains remain elusive. Herein, it is showcased how machine learning-enhanced atom probe tomography (APT) can mine the near-atomically resolved APT data and jointly exploit the technique's high elemental sensitivity to provide a 3D quantitative analysis of CSRO in a CoCrNi medium-entropy alloy. Multiple CSRO configurations are revealed, with their formation supported by state-of-the-art Monte-Carlo simulations. Quantitative analysis of these CSROs allows establishing relationships between processing parameters and physical properties. The unambiguous characterization of CSRO will help refine strategies for designing advanced materials by manipulating atomic-scale architectures.
Chemical short-range order (CSRO) refers to atoms of specific elements self-organising within a disordered crystalline matrix to form particular atomic neighbourhoods. CSRO is typically characterized indirectly, using volume-averaged or through projection microscopy techniques that fail to capture the three-dimensional atomistic architectures. Here, we present a machine-learning enhanced approach to break the inherent resolution limits of atom probe tomography enabling three-dimensional imaging of multiple CSROs. We showcase our approach by addressing a long-standing question encountered in body-centred-cubic Fe-Al alloys that see anomalous property changes upon heat treatment. We use it to evidence non-statistical B2-CSRO instead of the generally-expected D03-CSRO. We introduce quantitative correlations among annealing temperature, CSRO, and nano-hardness and electrical resistivity. Our approach is further validated on modified D03-CSRO detected in Fe-Ga. The proposed strategy can be generally employed to investigate short/medium/long-range ordering phenomena in different materials and help design future high-performance materials.
Thanks to the availability of huge amounts of data and improved computational resources, AI methods are gaining importance in scientific workflows, from image recognition and natural language processing to materials science. In many domains the usage of AI is under active investigation and first results show a tremendous potential, suggesting that AI will have significant impact way beyond the currently dominating examples of image and language processing.
Computational methods and machine learning algorithms for automatic information extraction are crucial to enable data-driven materials science. These approaches are changing materials characterization and analytics, which often require a user-specified threshold to e.g. detect structure or symmetries in structures with defects. Here, we present a machine learning-based approach that directly works on the original periodic arrangements of atoms based on a three-dimensional convolutional neural network without any transformation of descriptors. Our approach shows a high classification accuracy and tolerance to the presence of random displacements and missing atoms. Experimentally, we successfully reconstruct the ordered L12 precipitates extracted from atom probe tomography data, consistent with segmentation based on isocomposition surfaces. The convolutional layers are essential for the simultaneous identification of compositional and structural information, which also give rise to its high tolerance. Our work advances machine learning-based crystal structure identification for incomplete crystal structural data.
Nanoscale L1 2 -type ordered structures are widely used in face-centered cubic (FCC) alloys to exploit their hardening capacity and thereby improve mechanical properties. These fine-scale particles are typically fully coherent with matrix with the same atomic configuration disregarding chemical species, which makes them challenging to be characterized. Spatial distribution maps (SDMs) are used to probe local order by interrogating the three-dimensional (3D) distribution of atoms within reconstructed atom probe tomography (APT) data. However, it is almost impossible to manually analyze the complete point cloud (>10 million) in search for the partial crystallographic information retained within the data. Here, we proposed an intelligent L1 2 -ordered structure recognition method based on convolutional neural networks (CNNs). The SDMs of a simulated L1 2 -ordered structure and the FCC matrix were firstly generated. These simulated images combined with a small amount of experimental data were used to train a CNN-based L1 2 -ordered structure recognition model. Finally, the approach was successfully applied to reveal the 3D distribution of L1 2 –type δ′–Al 3 (LiMg) nanoparticles with an average radius of 2.54 nm in a FCC Al-Li-Mg system. The minimum radius of detectable nanodomain is even down to 5 Å. The proposed CNN-APT method is promising to be extended to recognize other nanoscale ordered structures and even more-challenging short-range ordered phenomena in the near future.
The solution of eigenproblems is often a key computational bottleneck that limits the tractable system size of numerical algorithms, among them electronic structure theory in chemistry and in condensed matter physics. Large eigenproblems can easily exceed the capacity of a single compute node, thus must be solved on distributed-memory parallel computers. We here present GPU-oriented optimizations of the ELPA two-stage tridiagonalization eigensolver (ELPA2). On top of cuBLAS-based GPU offloading, we add a CUDA kernel to speed up the back-transformation of eigenvectors, which can be the computationally most expensive part of the two-stage tridiagonalization algorithm. We benchmark the performance of this GPU-accelerated eigensolver on two hybrid CPU–GPU architectures, namely a compute cluster based on Intel Xeon Gold CPUs and NVIDIA Volta GPUs, and the Summit supercomputer based on IBM POWER9 CPUs and NVIDIA Volta GPUs. Consistent with previous benchmarks on CPU-only architectures, the GPU-accelerated two-stage solver exhibits a parallel performance superior to the one-stage counterpart. Finally, we demonstrate the performance of the GPU-accelerated eigensolver developed in this work for routine semi-local KS-DFT calculations comprising thousands of atoms.
For symmetric (hermitian) (dense or banded) matrices the computation of eigenvalues and eigenvectors Ax = lambda Bx is an important task, e.g. in electronic structure calculations. If a larger number of eigenvectors are needed, often direct solvers are applied. On parallel architectures the ELPA implementation has proven to be very efficient, also compared to other parallel solvers like EigenExa or MAGMA. The main improvement that allows better parallel efficiency in ELPA is the two-step transformation of dense to band to tridiagonal form. This was the achievement of the ELPA project. The continuation of this project has been targeting at additional improvements like allowing monitoring and autotuning of the ELPA code, optimizing the code for different architectures, developing curtailed algorithms for banded A and B, and applying the improved code to solve typical examples in electronic structure calculations. In this paper we will present the outcome of this project.
We present a high-performance solver for dense skew-symmetric matrix eigenvalue problems. Our work is motivated by applications in computational quantum physics, where one solution approach to solve the so-called Bethe-Salpeter equation involves the solution of a large, dense, skew-symmetric eigenvalue problem. The computed eigenpairs can be used to compute the optical absorption spectrum of molecules and crystalline systems. One state-of-the art high-performance solver package for symmetric matrices is the ELPA (Eigenvalue SoLvers for Petascale Applications) library. We extend the methods available in ELPA to skew-symmetric matrices. This way, the presented solution method can benefit from the optimizations available in ELPA that make it a well-established, efficient and scalable library, such as GPU support. We compare performance and scalability of our method to the only available high-performance approach for skew-symmetric matrices, an indirect route involving complex arithmetic. In total, we achieve a performance that is up to 3.67 higher than the reference method using Intel's ScaLAPACK implementation. The runtime to solve the Bethe-Salpeter-Eigenvalue problem can be improved by a factor of 10. Our method is freely available in the current release of the ELPA library.
We first briefly report on the status and recent achievements of the ELPA-AEO (Eigen value Solvers for Petaflop ApplicationsAlgorithmic Extensions and Optimizations) and ESSEXII (Equipping Sparse Solvers for Exascale) projects. In both collaboratory efforts, scientists from the application areas, mathematicians, and computer scientists work together to develop and make available efficient highly parallel methods for the solution of eigenvalue problems. Then we focus on a topic addressed in both projects, the use of mixed precision computations to enhance efficiency. We give a more detailed description of our approaches for benefiting from either lower or higher precision in three selected contexts and of the results thus obtained.
The solution of (generalized) eigenvalue problems for symmetric or Hermitian matrices is a common subtask of many numerical calculations in electronic structure theory or materials science. Solving the eigenvalue problem can easily amount to a sizeable fraction of the whole numerical calculation. For researchers in the field of computational materials science, an efficient and scalable solution of the eigenvalue problem is thus of major importance. The ELPA-library is a well-established dense direct eigenvalue solver library, which has proven to be very efficient and scalable up to very large core counts. In this paper, we describe the latest optimizations of the ELPA-library for new HPC architectures of the Intel Skylake processor family with an AVX-512 SIMD instruction set, or for HPC systems accelerated with recent GPUs. We also describe a complete redesign of the API in a modern modular way, which, apart from a much simpler and more flexible usability, leads to a new path to access system-specific performance optimizations. In order to ensure optimal performance for a particular scientific setting or a specific HPC system, the new API allows the user to influence in straightforward way the internal details of the algorithms and of performance-critical parameters used in the ELPA-library. On top of that, we introduced an autotuning functionality, which allows for finding the best settings in a self-contained automated way. In situations where many eigenvalue problems with similar settings have to be solved consecutively, the autotuning process of the ELPA-library can be done "on-the-fly". Practical applications from materials science which rely on so-called self-consistency iterations can profit from the autotuning. On some examples of scientific interest, simulated with the FHI-aims application, the advantages of the latest optimizations of the ELPA-library are demonstrated.
We present a GPU implementation of a large-scale eigenvalue solver as a part of the ELPA library. We describe the methodology of utilizing the GPU accelerators within an already well optimized MPI-based code. We present numerical results using two different HPC systems equipped with modern GPU accelerators and show the performance benefits of the GPU version.
We present the first self-consistent, three-dimensional (3D) core-collapse supernova simulations performed with the Prometheus-Vertex code for a rotating progenitor star. Besides using the angular momentum of the 15 solar-mass model as obtained in the stellar evolution calculation with an angular frequency of about 0.001 rad/s (spin period of more than 6000 s) at the Si/Si-O interface, we also computed 2D and 3D cases with no rotation and with a ~300 times shorter rotation period and different angular resolutions. In 2D, only the nonrotating and slowly rotating models explode, while rapid rotation prevents an explosion within 500 ms after bounce because of lower radiated neutrino luminosities and mean energies and thus reduced neutrino heating. In contrast, only the fast rotating model develops an explosion in 3D when the Si/Si-O interface collapses through the shock. The explosion becomes possible by the support of a powerful SASI spiral mode, which compensates for the reduced neutrino heating and pushes strong shock expansion in the equatorial plane. Fast rotation in 3D leads to a "two-dimensionalization" of the turbulent energy spectrum (yielding roughly a -3 instead of a -5/3 power-law slope at intermediate wavelengths) with enhanced kinetic energy on the largest spatial scales. We also introduce a generalization of the "universal critical luminosity condition" of Summa et al. (2016) to account for the effects of rotation, and demonstrate its viability for a set of more than 40 core-collapse simulations including 9 and 20 solar-mass progenitors as well as black-hole forming cases of 40 and 75 solar-mass stars to be discussed in forthcoming papers.
Bruno Lang合作论文数Wuppertal
Fachbereich C - Mathematik und Naturwissenschaften
Angewandte Informatik - Algorithmik
Bergische Universit?t3