As materials researchers increasingly embrace machine-learning (ML) methods, it is natural to wonder what lessons can be learned from other fields undergoing similar developments. In this Review, we comparatively assess the evolution of applied ML in materials research, gameplaying and robotics. We observe ML being integrated into each field in three phases: first into discrete hardware and software tools (toolset integration); second across different steps in a workflow (workflow integration); and third through the incorporation, generation and representation of generalizable knowledge beyond any one study (knowledge integration). We identify transferrable lessons from gameplaying and robotics to materials research, including adaptive and accessible automation, the gamification of grand challenges to focus community efforts on specific workflow integrations and motivate benchmarks and canonical datasets, and the adoption of hybrid (data-based and model-based) algorithms that combine domain expertise and current learning to economically address high-complexity tasks. We identify opportunities for researchers from different fields to collaborate, including novel ways to represent and integrate a rich but heterogeneous corpus of knowledge (such as heuristics, physical laws, literature or data) with ML algorithms to create new knowledge, and safe and equitable deployment of technologies with societally beneficial outcomes.
We present evidence that learned density functional theory (``DFT'') force fields are ready for ground state catalyst discovery. Our key finding is that relaxation using forces from a learned potential yields structures with similar or lower energy to those relaxed using the RPBE functional in over 50\% of evaluated systems, despite the fact that the predicted forces differ significantly from the ground truth. This has the surprising implication that learned potentials may be ready for replacing DFT in challenging catalytic systems such as those found in the Open Catalyst 2020 dataset. Furthermore, we show that a force field trained on a locally harmonic energy surface with the same minima as a target DFT energy is also able to find lower or similar energy structures in over 50\% of cases. This ``Easy Potential'' converges in fewer steps than a standard model trained on true energies and forces, which further accelerates calculations. Its success illustrates a key point: learned potentials can locate energy minima even when the model has high force errors. The main requirement for structure optimisation is simply that the learned potential has the correct minima. Since learned potentials are fast and scale linearly with system size, our results open the possibility of quickly finding ground states for large systems.
In this paper we show that simple noise regularisation can be an effective way to address GNN oversmoothing. First we argue that regularisers addressing oversmoothing should both penalise node latent similarity and encourage meaningful node representations. From this observation we derive "Noisy Nodes", a simple technique in which we corrupt the input graph with noise, and add a noise correcting node-level loss. The diverse node level loss encourages latent node diversity, and the denoising objective encourages graph manifold learning. Our regulariser applies well-studied methods in simple, straightforward ways which allow even generic architectures to overcome oversmoothing and achieve state of the art results on quantum chemistry tasks, and improve results significantly on Open Graph Benchmark (OGB) datasets. Our results suggest Noisy Nodes can serve as a complementary building block in the GNN toolkit.
Density functional theory describes matter at the quantum level, but all popular approximations suffer from systematic errors that arise from the violation of mathematical properties of the exact functional. We overcame this fundamental limitation by training a neural network on molecular data and on fictitious systems with fractional charge and spin. The resulting functional, DM21 (DeepMind 21), correctly describes typical examples of artificial charge delocalization and strong correlation and performs better than traditional functionals on thorough benchmarks for main-group atoms and molecules. DM21 accurately models complex systems such as hydrogen chains, charged DNA base pairs, and diradical transition states. More crucially for the field, because our methodology relies on data and constraints, which are continually improving, it represents a viable pathway toward the exact universal functional.
Machine Learning (ML) has the potential to accelerate discovery of new materials and shed light on useful properties of existing materials. A key difficulty when applying ML in Materials Science is that experimental datasets of material properties tend to be small. In this work we show how material descriptors can be learned from the structures present in large scale datasets of material simulations; and how these descriptors can be used to improve the prediction of an experimental property, the energy of formation of a solid. The material descriptors are learned by training a Graph Neural Network to regress simulated formation energies from a material's atomistic structure. Using these learned features for experimental property predictions outperforms existing methods that are based solely on chemical composition. Moreover, we find that the advantage of our approach increases as the generalization requirements of the task are made more stringent, for example when limiting the amount of training data or when generalizing to unseen chemical spaces.
Protein structure prediction can be used to determine the three-dimensional shape of a protein from its amino acid sequence1. This problem is of fundamental importance as the structure of a protein largely determines its function2; however, protein structures can be difficult to determine experimentally. Considerable progress has recently been made by leveraging genetic information. It is possible to infer which amino acid residues are in contact by analysing covariation in homologous sequences, which aids in the prediction of protein structures3. Here we show that we can train a neural network to make accurate predictions of the distances between pairs of residues, which convey more information about the structure than contact predictions. Using this information, we construct a potential of mean force4 that can accurately describe the shape of a protein. We find that the resulting potential can be optimized by a simple gradient descent algorithm to generate structures without complex sampling procedures. The resulting system, named AlphaFold, achieves high accuracy, even for sequences with fewer homologous sequences. In the recent Critical Assessment of Protein Structure Prediction5 (CASP13)-a blind assessment of the state of the field-AlphaFold created high-accuracy structures (with template modelling (TM) scores6 of 0.7 or higher) for 24 out of 43 free modelling domains, whereas the next best method, which used sampling and contact information, achieved such accuracy for only 14 out of 43 domains. AlphaFold represents a considerable advance in protein-structure prediction. We expect this increased accuracy to enable insights into the function and malfunction of proteins, especially in cases for which no structures for homologous proteins have been experimentally determined7.
Device variability is a bottleneck for the scalability of semiconductor quantum devices. Increasing device control comes at the cost of a large parameter space that has to be explored in order to find the optimal operating conditions. We demonstrate a statistical tuning algorithm that navigates this entire parameter space, using just a few modelling assumptions, in the search for specific electron transport features. We focused on gate-defined quantum dot devices, demonstrating fully automated tuning of two different devices to double quantum dot regimes in an up to eight-dimensional gate voltage space. We considered a parameter space defined by the maximum range of each gate voltage in these devices, demonstrating expected tuning in under 70 minutes. This performance exceeded a human benchmark, although we recognise that there is room for improvement in the performance of both humans and machines. Our approach is approximately 180 times faster than a pure random search of the parameter space, and it is readily applicable to different material systems and device architectures. With an efficient navigation of the gate voltage space we are able to give a quantitative measurement of device variability, from one device to another and after a thermal cycle of a device. This is a key demonstration of the use of machine learning techniques to explore and optimise the parameter space of quantum devices and overcome the challenge of device variability.
Protein structure prediction aims to determine the three-dimensional shape of a protein from 11 its amino acid sequence1. This problem is of fundamental importance to biology as the struc12 ture of a protein largely determines its function2 but can be hard to determine experimen13 tally. In recent years, considerable progress has been made by leveraging genetic informa14 tion: analysing the co-variation of homologous sequences can allow one to infer which amino 15 acid residues are in contact, which in turn can aid structure prediction3. In this work, we 16 show that we can train a neural network to accurately predict the distances between pairs 17 of residues in a protein which convey more about structure than contact predictions. With 18 this information we construct a potential of mean force4 that can accurately describe the 19 shape of a protein. We find that the resulting potential can be optimised by a simple gradient 20 descent algorithm, to realise structures without the need for complex sampling procedures. 21 The resulting system, named AlphaFold, has been shown to achieve high accuracy, even for 22 sequences with relatively few homologous sequences. In the most recent Critical Assessment 23 of Protein Structure Prediction5 (CASP13), a blind assessment of the state of the field of pro24 tein structure prediction, AlphaFold created high-accuracy structures (with TM-scores† of 25 0.7 or higher) for 24 out of 43 free modelling domains whereas the next best method, using 26 sampling and contact information, achieved such accuracy for only 14 out of 43 domains. 27 AlphaFold represents a significant advance in protein structure prediction. We expect the in28 creased accuracy of structure predictions for proteins to enable insights in understanding the 29 function and malfunction of these proteins, especially in cases where no homologous proteins 30 have been experimentally determined7. 31
While recent continual learning methods largely alleviate the catastrophic problem on toy-size datasets, there are issues that remain to be tackled in order to apply them to real-world problem domains. First, a continual learning model should effectively handle catastrophic forgetting and be efficient to train even with large number of tasks. Secondly, it needs to tackle the problem of order-sensitivity, where the performance of the tasks largely vary based on the order of the task arrival sequence, as it may cause serious problems where fairness plays a critical role (e.g. medical diagnosis). To tackle these practical challenges, we propose a novel continual learning method that is scalable as well as order-robust, which instead of learning a completely shared set of weights, represents the parameter for each task as a sum of task-shared and sparse task-adaptive parameters. With our hierarchically decomposed networks (HDN), the task-adaptive parameters for earlier tasks remain mostly unaffected, where we update them only to reflect the changes made to the task-shared parameters. This decomposition of parameters effectively prevents catastrophic forgetting and order-sensitivity, while being computation- and memory-efficient. Further, with hierarchical knowledge consolidation which clusters the task-adaptive parameters to obtain hierarchically shared parameters, HDN becomes highly scalable. We validate HDN on multiple benchmark datasets against state-of-the-art continual learning methods, which it largely outperforms in accuracy, efficiency, scalability, and order-robustness.
We describe AlphaFold, the protein structure prediction system that was entered by the group A7D in CASP13. Submissions were made by three free-modeling (FM) methods which combine the predictions of three neural networks. All three systems were guided by predictions of distances between pairs of residues produced by a neural network. Two systems assembled fragments produced by a generative neural network, one using scores from a network trained to regress GDT_TS. The third system shows that simple gradient descent on a properly constructed potential is able to perform on par with more expensive traditional search techniques and without requiring domain segmentation. In the CASP13 FM assessors' ranking by summed z-scores, this system scored highest with 68.3 vs 48.2 for the next closest group (an average GDT_TS of 61.4). The system produced high-accuracy structures (with GDT_TS scores of 70 or higher) for 11 out of 43 FM domains. Despite not explicitly using template information, the results in the template category were comparable to the best performing template-based methods.
MethodsThe systems tested all use multiple sequence alignments (MSA) and profiles generated from HHBlits [2] and PSIBLAST [3]. No templates were used, nor were server predictions. No manual intervention was made except for domain segmentation of T0999 and final decoy ranking in a handful of cases. In protein complexes, each chain was processed independently.
In our recent work on elastic weight consolidation (EWC) (1) we show that forgetting in neural networks can be alleviated by using a quadratic penalty whose derivation was inspired by Bayesian evidence accumulation. In his letter (2), Dr. Huszar provides an alternative form for this penalty by following the standard work on expectation propagation using the Laplace approximation (3). He correctly argues that in cases when more than two tasks are undertaken the two forms of the penalty are different. Dr. Huszar also shows that for a toy linear regression problem his expression appears to be better. We would like to thank Dr. Huszar for pointing out … [↵][1]1To whom correspondence should be addressed. Email: kirkpatrick@google.com. [1]: #xref-corresp-1-1
Electronic polarisation contributes to the electronic landscape as seen by separating charges in organic materials. The nature of electronic polarisation depends on the polarisability, density, and arrangement of polarisable molecules. In this paper, we introduce a microscopic, coarse-grained model in which we treat each molecule as a polarisable site, and use an array of such polarisable dipoles to calculate the electric field and associated energy of any arrangement of charges in the medium. The model incorporates chemical structure via the molecular polarisability and molecular packing patterns via the structure of the array. We use this model to calculate energies of charge pairs undergoing separation in finite fullerene lattices of different chemical and crystal structures. The effective dielectric constants that we estimate from this approach are in good quantitative agreement with those measured experimentally in C60 and phenyl-C61-butyric acid methyl ester (PCBM) films, but we find significant differences in dielectric constant depending on packing and on direction of separation, which we rationalise in terms of density of polarisable fullerene cages in regions of high field. In general, we find lattices containing molecules of more isotropic polarisability tensors exhibit higher dielectric constants. By exploring several model systems we conclude that differences in molecular polarisability (and therefore, chemical structure) appear to be less important than differences in molecular packing and separation direction in determining the energetic landscape for charge separation. We note that the results are relevant for finite lattices, but not necessarily for infinite systems. We propose that the model could be used to design molecular systems for effective electronic screening.
Grand Canonical Molecular Dynamics (GCMD) simulations were performed to investigate the intercalation of CO2 and, H2O molecules in the interlayers of the smectite clay, Na-hectorite, at temperatures and pressures relevant to petroleum reservoir and geological carbon sequestration conditions and in equilibrium with H2O-saturated CO2. The computed adsorption isotherms indicate that CO2 molecules enter the interlayer space of Na-hectorite only when it is hydrated with approximately three H2O molecules per unit cell. The computed immersion energies show that the bilayer hydrate structure (2WL) contains less CO2 than the monolayer structure (1WL) but that the 2WL hydrate is the most thermodynamically stable state, consistent with experimental results fora similar Namontmorillonite smectite. Under all T and. P conditions examined (323-368 K and 90-150 bar), the CO2 molecules are adsorbed at the midplane of clay interlayers for the 1WL structure and closer to one of the basal surfaces for the 2WL structure. Interlayer CO2 molecules are dynamically less restricted in the 2WL structures. The CO2 molecules are preferentially located near basal surface oxygen atoms and H2O molecules rather than in coordination with Na+ ions. Accounting for the orientation and flexibility of the structural -OH groups of the clay layer has a significant effect on the details of the computed structure and dynamics of H2O and CO2 molecules but does not affect the overall trends with changing basal spacing or the principal structural and dynamical conclusions. Temperature and pressure in the ranges examined have little effect on the principal structural and energetic conclusions, but the rates of dynamical processes increase with increasing temperature, as expected.
The ability to learn tasks in a sequential fashion is crucial to the development of artificial intelligence. Neural networks are not, in general, capable of this and it has been widely thought that catastrophic forgetting is an inevitable feature of connectionist models. We show that it is possible to overcome this limitation and train networks that can maintain expertise on tasks which they have not experienced for a long time. Our approach remembers old tasks by selectively slowing down learning on the weights important for those tasks. We demonstrate our approach is scalable and effective by solving a set of classification tasks based on the MNIST hand written digit dataset and by learning several Atari 2600 games sequentially.
Most deep reinforcement learning algorithms are data inefficient in complex and rich environments, limiting their applicability to many scenarios. One direction for improving data efficiency is multitask learning with shared neural network parameters, where efficiency may be improved through transfer across related tasks. In practice, however, this is not usually observed, because gradients from different tasks can interfere negatively, making learning unstable and sometimes even less data efficient. Another issue is the different reward schemes between tasks, which can easily lead to one task dominating the learning of a shared model. We propose a new approach for joint training of multiple tasks, which we refer to as Distral (Distill & transfer learning). Instead of sharing parameters between the different workers, we propose to share a "distilled" policy that captures common behaviour across tasks. Each worker is trained to solve its own task while constrained to stay close to the shared policy, while the shared policy is trained by distillation to be the centroid of all task policies. Both aspects of the learning process are derived by optimizing a joint objective function. We show that our approach supports efficient transfer on complex 3D environments, outperforming several related methods. Moreover, the proposed learning process is more robust and more stable---attributes that are critical in deep reinforcement learning.
Learning to solve complex sequences of tasks--while both leveraging transfer and avoiding catastrophic forgetting--remains a key obstacle to achieving human-level intelligence. The progressive networks approach represents a step forward in this direction: they are immune to forgetting and can leverage prior knowledge via lateral connections to previously learned features. We evaluate this architecture extensively on a wide variety of reinforcement learning tasks (Atari and 3D maze games), and show that it outperforms common baselines based on pretraining and finetuning. Using a novel sensitivity measure, we demonstrate that transfer occurs at both low-level sensory and high-level control layers of the learned policy.
Policies for complex visual tasks have been successfully learned with deep reinforcement learning, using an approach called deep Q-networks (DQN), but relatively large (task-specific) networks and extensive training are needed to achieve good performance. In this work, we present a novel method called policy distillation that can be used to extract the policy of a reinforcement learning agent and train a new network that performs at the expert level while being dramatically smaller and more efficient. Furthermore, the same method can be used to consolidate multiple task-specific policies into a single policy. We demonstrate these claims using the Atari domain and show that the multi-task distilled agent outperforms the single-task teachers as well as a jointly-trained DQN agent.