Recent studies have reported the experimental discovery that nanoscale specimens of even a natural material, such as diamond, can be deformed elastically to as much as 10% tensile elastic strain at room temperature without the onset of permanent damage or fracture. Computational work combining ab initio calculations and machine learning (ML) algorithms has further demonstrated that the bandgap of diamond can be altered significantly purely by reversible elastic straining. These findings open up unprecedented possibilities for designing materials and devices with extreme physical properties and performance characteristics for a variety of technological applications. However, a general scientific framework to guide the design of engineering materials through such elastic strain engineering (ESE) has not yet been developed. By combining first-principles calculations with ML, we present here a general approach to map out the entire phonon stability boundary in six-dimensional strain space, which can guide the ESE of a material without phase transitions. We focus on ESE of vibrational properties, including harmonic phonon dispersions, nonlinear phonon scattering, and thermal conductivity. While the framework presented here can be applied to any material, we show as an example demonstration that the room-temperature lattice thermal conductivity of diamond can be increased by more than 100% or reduced by more than 95% purely by ESE, without triggering phonon instabilities. Such a framework opens the door for tailoring of thermal-barrier, thermoelectric, and electro-optical properties of materials and devices through the purposeful design of homogeneous or inhomogeneous strains.
Large language models (LLMs) are notorious for hallucinating, i.e., producing erroneous claims in their output. Such hallucinations can be dangerous, as occasional factual inaccuracies in the generated text might be obscured by the rest of the output being generally factually correct, making it extremely hard for the users to spot them. Current services that leverage LLMs usually do not provide any means for detecting unreliable generations. Here, we aim to bridge this gap. In particular, we propose a novel fact-checking and hallucination detection pipeline based on token-level uncertainty quantification. Uncertainty scores leverage information encapsulated in the output of a neural network or its layers to detect unreliable predictions, and we show that they can be used to fact-check the atomic claims in the LLM output. Moreover, we present a novel tokenlevel uncertainty quantification method that removes the impact of uncertainty about what claim to generate on the current step and what surface form to use. Our method Claim Conditioned Probability (CCP) measures only the uncertainty of a particular claim value expressed by the model. Experiments on the task of biography generation demonstrate strong improvements for CCP compared to the baselines for seven LLMs and four languages. Human evaluation reveals that the fact-checking pipeline based on uncertainty quantification is competitive with a fact-checking tool that leverages external knowledge.
This paper introduces ClimateGPT, a model family of domain-specific large language models that synthesize interdisciplinary research on climate change. We trained two 7B models from scratch on a science-oriented dataset of 300B tokens. For the first model, the 4.2B domain-specific tokens were included during pre-training and the second was adapted to the climate domain after pre-training. Additionally, ClimateGPT-7B, 13B and 70B are continuously pre-trained from Llama~2 on a domain-specific dataset of 4.2B tokens. Each model is instruction fine-tuned on a high-quality and human-generated domain-specific dataset that has been created in close cooperation with climate scientists. To reduce the number of hallucinations, we optimize the model for retrieval augmentation and propose a hierarchical retrieval strategy. To increase the accessibility of our model to non-English speakers, we propose to make use of cascaded machine translation and show that this approach can perform comparably to natively multilingual models while being easier to scale to a large number of languages. Further, to address the intrinsic interdisciplinary aspect of climate change we consider different research perspectives. Therefore, the model can produce in-depth answers focusing on different perspectives in addition to an overall answer. We propose a suite of automatic climate-specific benchmarks to evaluate LLMs. On these benchmarks, ClimateGPT-7B performs on par with the ten times larger Llama-2-70B Chat model while not degrading results on general domain benchmarks. Our human evaluation confirms the trends we saw in our benchmarks. All models were trained and evaluated using renewable energy and are released publicly.
Uncertainty estimation (UE) of model predictions is a crucial step for a variety of tasks such as active learning, misclassification detection, adversarial attack detection, out-of-distribution detection, etc. Most of the works on modeling the uncertainty of deep neural networks evaluate these methods on image classification tasks. Little attention has been paid to UE in natural language processing. To fill this gap, we perform a vast empirical investigation of state-of-the-art UE methods for Transformer models on misclassification detection in named entity recognition and text classification tasks and propose two computationally efficient modifications, one of which approaches or even outperforms computationally intensive methods.
Uncertainty estimation for machine learning models is of high importance in many scenarios such as constructing the confidence intervals for model predictions and detection of out-of-distribution or adversarially generated points. In this work, we show that modifying the sampling distributions for dropout layers in neural networks improves the quality of uncertainty estimation. Our main idea consists of two main steps: computing data-driven correlations between neurons and generating samples, which include maximally diverse neurons. In a series of experiments on simulated and real-world data, we demonstrate that the diversification via determinantal point processes-based sampling achieves state-of-the-art results in uncertainty estimation for regression and classification tasks. An important feature of our approach is that it does not require any modification to the models or training procedures, allowing straightforward application to any deep learning model with dropout layers.
The controlled introduction of elastic strains is an appealing strategy for modulating the physical properties of semiconductor materials. With the recent discovery of large elastic deformation in nanoscale specimens as diverse as silicon and diamond, employing this strategy to improve device performance necessitates first-principles computations of the fundamental electronic band structure and target figures-of-merit, through the design of an optimal straining pathway. Such simulations, however, call for approaches that combine deep learning algorithms and physics of deformation with band structure calculations to custom-design electronic and optical properties. Motivated by this challenge, we present here details of a machine learning framework involving convolutional neural networks to represent the topology and curvature of band structures in k-space. These calculations enable us to identify ways in which the physical properties can be altered through “deep” elastic strain engineering up to a large fraction of the ideal strain. Algorithms capable of active learning and informed by the underlying physics were presented here for predicting the bandgap and the band structure. By training a surrogate model with ab initio computational data, our method can identify the most efficient strain energy pathway to realize physical property changes. The power of this method is further demonstrated with results from the prediction of strain states that influence the effective electron mass. We illustrate the applications of the method with specific results for diamonds, although the general deep learning technique presented here is potentially useful for optimizing the physical properties of a wide variety of semiconductor materials.
In this work, we consider the problem of uncertainty estimation for Transformer-based models. We investigate the applicability of uncertainty estimates based on dropout usage at the inference stage (Monte Carlo dropout). The series of experiments on natural language understanding tasks shows that the resulting uncertainty estimates improve the quality of detection of error-prone instances. Special attention is paid to the construction of computationally inexpensive estimates via Monte Carlo dropout and Determinantal Point Processes.
During his long and distinguished career, Subra Suresh has made crucial contributions to the field of engineering. While finishing up high school in India in the 1970s, however, Suresh was not even sure of going to college, let alone becoming an engineer. Nonetheless, Suresh decided to take a shot at the entrance examination for the prestigious Indian Institutes of Technology. “A month before my exam, I bought a book to prepare and worked through some practice questions and just thought, go try it,” Suresh says. “To my surprise, I got in.” His degree in mechanical engineering from Indian Institutes of Technology Madras would turn out to be the starting point of a wide-ranging research career. Suresh’s research interests would eventually span engineering, basic science, and medicine. His multidisciplinary work led to elected memberships in all three US National Academies: The National Academy of Engineering in 2002, the National Academy of Sciences in 2012, and the National Academy of Medicine in 2013. Suresh has held several prestigious positions, from being dean of Massachusetts Institute of Technology’s School of Engineering and president of Carnegie Mellon University to leading the National Science Foundation (NSF) of the United States. Now President of Nanyang Technological University, Singapore, Suresh continues to push forward in research with his recent work on deforming nanoscale diamond. In his Inaugural Article (1), Suresh and his colleagues show computationally that it is possible to make nanoscale diamond behave like a metal with respect to select properties, which would open up a wide array of applications in microelectronics, optoelectronics, and solar energy.
ENGINEERING Correction for “Deep elastic strain engineering of bandgap through machine learning,” by Zhe Shi, Evgenii Tsymbalov, Ming Dao, Subra Suresh, Alexander Shapeev, and Ju Li, which was first published February 15, 2019; 10.1073/pnas.1818555116 (Proc. Natl. Acad. Sci. U.S.A. 116, 4117–4122). The authors note that “In Fig. 2A, the 6D strain tensor for zerobandgap was not reported in the [100],[010],[001] coordinate frame, contrary to the rest of the article. The correct values of the strain tensor for the [001],[010],[001] frame along with the revised figure legend are included below. This correction does not alter any conclusions in the paper. We apologize for the error.” The corrected figure and its corrected legend appear below.
Active learning refers to collections of algorithms of systematically constructing the training dataset. It is closely related to uncertainty estimation—we, generally, do not need to train our model on samples on which our prediction already has low uncertainty. This chapter reviews active learning algorithms in the context of molecular modeling and illustrates their applications on practical problems.
Q1: The study addresses an ideal crystal without any defects. This is correct for the idealized case of zero temperature, T=0. Each real crystal however has inevitably equilibrium point defects for non-zero temperature and generally, non-equilibrium defects like dislocations. Hence, a question arises, how the presented results would change, if these important properties of real crystals were taken into account? I do not expect quantitative estimates, but just a qualitative answer – whether the observed decrease of the bandgap persists? Will it shift for larger or smaller strain? R1: To the best of my knowledge, the defects may introduce the “defect bands” (or intermediate bands) to the electronic bandstructure, which corresponds to the discrete eigenvalues of the Hamiltonian (in contrast to the continuous spectrum, which is related to the bands). Since the bandgap is defined as the difference between the conduction band minimum and the valence band maximum, technically, if these bands are not changed, then the bandgap value remains the same. However, the energy required to excite an electron to become a conduction electron decreases.
Experimental discovery of ultralarge elastic deformation in nanoscale diamond and machine learning of its electronic and phonon structures have created opportunities to address new scientific questions. Can diamond, with an ultrawide bandgap of 5.6 eV, be completely metallized, solely under mechanical strain without phonon instability, so that its electronic bandgap fully vanishes? Through first-principles calculations, finite-element simulations validated by experiments, and neural network learning, we show here that metallization/demetallization as well as indirect-to-direct bandgap transitions can be achieved reversibly in diamond below threshold strain levels for phonon instability. We identify the pathway to metallization within six-dimensional strain space for different sample geometries. We also explore phonon-instability conditions that promote phase transition to graphite. These findings offer opportunities for tailoring properties of diamond via strain engineering for electronic, photonic, and quantum applications.
Running complex test suites against a financial transaction system produces huge amounts of responses, both expected and unexpected. In this article, we outline our experience of using ML for reliable automatic extraction of "that" unexpected response from a big number of same type messages produced a by system under test. We describe classification approaches and data manipulations we have tried, and explain the final choices. Also we outline business constraints and final design decisions for the resultant tool.We also address the task of classifying difference patterns between expected and actual responses in attempt to provide automated pre-judgement on a reason for test failure. We outline clustering considerations and results achieved.
Testing of distributed systems is a complex task, which is hampered by the impossibility of guaranteed reproduction of errors associated with race conditions. Even minor instrumentation of the system significantly changes its characteristics, which becomes critical, especially for load testing. All of that increases the importance of quality control methods based on the system log analysis. In this paper, we present our experience of semi-automated analysis of the behavior of clearing and settlement system by utilizing its logs for the purpose of identifying and classifying errors.
Active learning methods for neural networks are usually based on greedy criteria, which ultimately give a single new design point for the evaluation. Such an approach requires either some heuristics to sample a batch of design points at one active learning iteration, or retraining the neural network after adding each data point, which is computationally inefficient. Moreover, uncertainty estimates for neural networks sometimes are overconfident for the points lying far from the training sample. In this work, we propose to approximate Bayesian neural networks (BNN) by Gaussian processes (GP), which allows us to update the uncertainty estimates of predictions efficiently without retraining the neural network while avoiding overconfident uncertainty prediction for out-of-sample points. In a series of experiments on real-world data, including large-scale problems of chemical and physical modeling, we show the superiority of the proposed approach over the state-of-the-art methods.
Nanoscale specimens of semiconductor materials as diverse as silicon and diamond are now known to be deformable to large elastic strains without inelastic relaxation. These discoveries harbinger a new age of deep elastic strain engineering of the band structure and device performance of electronic materials. Many possibilities remain to be investigated as to what pure silicon can do as the most versatile electronic material and what an ultrawide bandgap material such as diamond, with many appealing functional figures of merit, can offer after overcoming its present commercial immaturity. Deep elastic strain engineering explores full six-dimensional space of admissible nonlinear elastic strain and its effects on physical properties. Here we present a general method that combines machine learning and ab initio calculations to guide strain engineering whereby material properties and performance could be designed. This method invokes recent advances in the field of artificial intelligence by utilizing a limited amount of ab initio data for the training of a surrogate model, predicting electronic bandgap within an accuracy of 8 meV. Our model is capable of discovering the indirect-to-direct bandgap transition and semiconductor-to-metal transition in silicon by scanning the entire strain space. It is also able to identify the most energy-efficient strain pathways that would transform diamond from an ultrawide-bandgap material to a smaller-bandgap semiconductor. A broad framework is presented to tailor any target figure of merit by recourse to deep elastic strain engineering and machine learning for a variety of applications in microelectronics, optoelectronics, photonics, and energy technologies.
Active learning is relevant and challenging for high-dimensional regression models when the annotation of the samples is expensive. Yet most of the existing sampling methods cannot be applied to large-scale problems, consuming too much time for data processing. In this paper, we propose a fast active learning algorithm for regression, tailored for neural network models. It is based on uncertainty estimation from stochastic dropout output of the network. Experiments on both synthetic and real-world datasets show comparable or better performance (depending on the accuracy metric) as compared to the baselines. This approach can be generalized to other deep learning architectures. It can be used to systematically improve a machine-learning model as it offers a computationally efficient way of sampling additional data.
We develop a new compact scheme for the second-order PDE (parabolic and Schrodinger type) with a variable time-independent coefficient. It has a higher order and smaller error than classic implicit scheme. The Dirichlet and Neumann boundary problems are considered. The relative finite-difference operator is almost self-adjoint. (C) 2018 Elsevier Inc. All rights reserved.