Long-acting injectables are considered one of the most promising therapeutic strategies for the treatment of chronic diseases as they can afford improved therapeutic efficacy, safety, and patient compliance. The use of polymer materials in such a drug formulation strategy can offer unparalleled diversity owing to the ability to synthesize materials with a wide range of properties. However, the interplay between multiple parameters, including the physicochemical properties of the drug and polymer, make it very difficult to intuitively predict the performance of these systems. This necessitates the development and characterization of a wide array of formulation candidates through extensive and time-consuming in vitro experimentation. Machine learning is enabling leap-step advances in a number of fields including drug discovery and materials science. The current study takes a critical step towards data-driven drug formulation development with an emphasis on long-acting injectables. Here we show that machine learning algorithms can be used to predict experimental drug release from these advanced drug delivery systems. We also demonstrate that these trained models can be used to guide the design of new long acting injectables. The implementation of the described data-driven approach has the potential to reduce the time and cost associated with drug formulation development.
Optimization strategies driven by machine learning, such as Bayesian optimization, are being explored across experimental sciences as an efficient alternative to traditional design of experiment. When combined with automated laboratory hardware and high-performance computing, these strategies enable next-generation platforms for autonomous experimentation. However, the practical application of these approaches is hampered by a lack of flexible software and algorithms tailored to the unique requirements of chemical research. One such aspect is the pervasive presence of constraints in the experimental conditions when optimizing chemical processes or protocols, and in the chemical space that is accessible when designing functional molecules or materials. Although many of these constraints are known a priori, they can be interdependent, non-linear, and result in non-compact optimization domains. In this work, we extend our experiment planning algorithms Phoenics and Gryffin such that they can handle arbitrary known constraints via an intuitive and flexible interface. We benchmark these extended algorithms on continuous and discrete test functions with a diverse set of constraints, demonstrating their flexibility and robustness. In addition, we illustrate their practical utility in two simulated chemical research scenarios: the optimization of the synthesis of o-xylenyl Buckminsterfullerene adducts under constrained flow conditions, and the design of redox active molecules for flow batteries under synthetic accessibility constraints. The tools developed constitute a simple, yet versatile strategy to enable model-based optimization with known experimental constraints, contributing to its applicability as a core component of autonomous platforms for scientific discovery.
An oracle that correctly predicts the outcome of every particle physics experiment, the products of every possible chemical reaction or the function of every protein would revolutionize science and technology. However, scientists would not be entirely satisfied because they would want to comprehend how the oracle made these predictions. This is scientific understanding, one of the main aims of science. With the increase in the available computational power and advances in artificial intelligence, a natural question arises: how can advanced computational systems, and specifically artificial intelligence, contribute to new scientific understanding or gain it autonomously? Trying to answer this question, we adopted a definition of 'scientific understanding' from the philosophy of science that enabled us to overview the scattered literature on the topic and, combined with dozens of anecdotes from scientists, map out three dimensions of computer-assisted scientific understanding. For each dimension, we review the existing state of the art and discuss future developments. We hope that this Perspective will inspire and focus research directions in this multidisciplinary emerging field.
Bayesian optimization has emerged as a powerful strategy to accelerate scientific discovery by means of autonomous experimentation. However, expensive measurements are required to accurately estimate materials properties, and can quickly become a hindrance to exhaustive materials discovery campaigns. Here, we introduce Gemini: a data-driven model capable of using inexpensive measurements as proxies for expensive measurements by correcting systematic biases between property evaluation methods. We recommend using Gemini for regression tasks with sparse data and in an autonomous workflow setting where its predictions of expensive to evaluate objectives can be used to construct a more informative acquisition function, thus reducing the number of expensive evaluations an optimizer needs to achieve desired target values. In a regression setting, we showcase the ability of our method to make accurate predictions of DFT calculated bandgaps of hybrid organic-inorganic perovskite materials. We further demonstrate the benefits that Gemini provides to autonomous workflows by augmenting the Bayesian optimizer Phoenics to yeild a scalable optimization framework leveraging multiple sources of measurement. Finally, we simulate an autonomous materials discovery platform for optimizing the activity of electrocatalysts for the oxygen evolution reaction. Realizing autonomous workflows with Gemini, we show that the number of measurements of a composition space comprising expensive and rare metals needed to achieve a target overpotential is significantly reduced when measurements from a proxy composition system with less expensive metals are available.
Superconducting circuits have emerged as a promising platform to build quantum processors. The challenge of designing a circuit is to compromise between realizing a set of performance metrics and reducing circuit complexity and noise sensitivity. At the same time, one needs to explore a large design space, and computational approaches often yield long simulation times. Here, we automate the circuit design task using SCILLA. The software SCILLA performs a parallelized, closed-loop optimization to design superconducting circuit diagrams that match predefined properties, such as spectral features and noise sensitivities. We employ it to design 4-local couplers for superconducting flux qubits and identify a circuit that outperforms an existing proposal with a similar circuit structure in terms of coupling strength and noise resilience for experimentally accessible parameters. This work demonstrates how automated design can facilitate the development of complex circuit architectures for quantum information processing.
Machine learning is enabling leap-step advances in a number of fields including drug discovery and materials science. The current study explores the application of machine learning to address a critical challenge in pharmaceutical formulation development: the prediction of drug release profiles from polymer-based long-acting injectables. Long acting injectables are considered one of the most promising therapeutic strategies for the treatment of chronic diseases as they can afford improved therapeutic efficacy, safety, and patient compliance. The use of polymer materials in such a drug formulation strategy can offer unparalleled diversity owing to the ability to synthesize materials with a wide range of properties. However, the interplay between multiple parameters, including the physicochemical properties of the drug and polymer, make it near to impossible to predict the performance of these systems a priori. This results in a need to develop and characterize a wide array of formulation candidates through extensive and time-consuming in vitro experimentation. In this study, various neural network architectures are constructed and trained, resulting in accurate predictions of drug release profiles that agree with experimental data. The simplicity with which these broadly applicable machine learning models are identified, using a limited amount of training data, is evidence of the promising potential of data-driven approaches in advanced pharmaceutical formulation development.
Machine learning (ML) has enabled ground-breaking advances in the healthcare and pharmaceutical sectors, from improvements in cancer diagnosis, to the identification of novel drugs and drug targets as well as protein structure prediction. Drug formulation is an essential stage in the discovery and development of new medicines. Through the design of drug formulations, pharmaceutical scientists can engineer important properties of new medicines, such as improved bioavailability and targeted delivery. The traditional approach to drug formulation development relies on iterative trial-and-error, requiring a large number of resource-intensive and time-consuming in vitro and in vivo experiments. This review introduces the basic concepts of ML-directed workflows and discusses how these tools can be used to aid in the development of various types of drug formulations. ML-directed drug formulation development offers unparalleled opportunities to fast-track development efforts, uncover new materials, innovative formulations, and generate new knowledge in drug formulation science. The review also highlights the latest artificial intelligence (AI) technologies, such as generative models, Bayesian deep learning, reinforcement learning, and self-driving laboratories, which have been gaining momentum in drug discovery and chemistry and have potential in drug formulation development.
Autonomous process optimization involves the human intervention-free exploration of a range of pre-defined process parameters in order to improve responses such as reaction yield and product selectivity. Utilizing off-the-shelf components, we developed a closed-loop system capable of carrying out parallel autonomous process optimization experiments in batch with significantly reduced cycle times. Upon implementation of our system in the autonomous optimization of a palladium-catalyzed stereoselective Suzuki-Miyaura coupling, we found that the definition of a set of meaningful, broad, and unbiased process parameters was the most critical aspect of a successful optimization. In addition, we found that categorical parameters such as phosphine ligand were vital to determining the reaction outcome. To date, categorical parameter selection has relied on chemical intuition, potentially introducing an element of bias into the experimental design. In seeking a systematic method for the selection of a diverse set of phosphine ligands fully representative of the chemical space, we developed a strategy that leveraged computed molecular descriptor clustering analysis. This strategy allowed for the successful autonomous optimization of a stereoselective Suzuki-Miyaura coupling between a vinyl sulfonate and an arylboronic acid to selectively generate the E -product isomer in high yield.
The choice of simulation methods in computational materials science is driven by a fundamental trade-off: bridging large time- and length-scales with highly accurate simulations at an affordable computational cost. Venturing the investigation of complex phenomena on large scales requires fast yet accurate computational methods. We review the emerging field of machine-learned potentials, which promises to reach the accuracy of quantum mechanical computations at a substantially reduced computational cost. This Review will summarize the basic principles of the underlying machine learning methods, the data acquisition process and active learning procedures. We highlight multiple recent applications of machine-learned potentials in various fields, ranging from organic chemistry and biomolecules to inorganic crystal structure predictions and surface science. We furthermore discuss the developments required to promote a broader use of ML potentials, and the possibility of using them to help solve open questions in materials science and facilitate fully computational materials design.
Research challenges encountered across science, engineering, and economics can frequently be formulated as optimization tasks. In chemistry and materials science, recent growth in laboratory digitization and automation has sparked interest in optimization-guided autonomous discovery and closed-loop experimentation. Experiment planning strategies based on off-the-shelf optimization algorithms can be employed in fully autonomous research platforms to achieve desired experimentation goals with the minimum number of trials. However, the experiment planning strategy that is most suitable to a scientific discovery task is a priori unknown while rigorous comparisons of different strategies are highly time and resource demanding. As optimization algorithms are typically benchmarked on low-dimensional synthetic functions, it is unclear how their performance would translate to noisy, higher-dimensional experimental tasks encountered in chemistry and materials science. We introduce Olympus , a software package that provides a consistent and easy-to-use framework for benchmarking optimization algorithms against realistic experiments emulated via probabilistic deep-learning models. Olympus includes a collection of experimentally derived benchmark sets from chemistry and materials science and a suite of experiment planning strategies that can be easily accessed via a user-friendly Python interface. Furthermore, Olympus facilitates the integration, testing, and sharing of custom algorithms and user-defined datasets. In brief, Olympus mitigates the barriers associated with benchmarking optimization algorithms on realistic experimental scenarios, promoting data sharing and the creation of a standard framework for evaluating the performance of experiment planning strategies.
Numerous challenges in science and engineering can be framed as optimization tasks, including the maximization of reaction yields, the optimization of molecular and materials properties, and the fine-tuning of automated hardware protocols. Design of experiment and optimization algorithms are often adopted to solve these tasks efficiently. Increasingly, these experiment planning strategies are coupled with automated hardware to enable autonomous experimental platforms. The vast majority of the strategies used, however, do not consider robustness against the variability of experiment and process conditions. In fact, it is generally assumed that these parameters are exact and reproducible. Yet some experiments may have considerable noise associated with some of their conditions, and process parameters optimized under precise control may be applied in the future under variable operating conditions. In either scenario, the optimal solutions found might not be robust against input variability, affecting the reproducibility of results and returning suboptimal performance in practice. Here, we introduce Golem, an algorithm that is agnostic to the choice of experiment planning strategy and that enables robust experiment and process optimization. Golem identifies optimal solutions that are robust to input uncertainty, thus ensuring the reproducible performance of optimized experimental protocols and processes. It can be used to analyze the robustness of past experiments, or to guide experiment planning algorithms toward robust solutions on the fly. We assess the performance and domain of applicability of Golem through extensive benchmark studies and demonstrate its practical relevance by optimizing an analytical chemistry protocol under the presence of significant noise in its experimental conditions.
The current Edisonian approach to discovery requires up to two decades of fundamental and applied research for materials technologies to reach the market. Such a slow and capital-intensive turnaround calls for disruptive strategies to expedite innovation. Self-driving laboratories have the potential to provide the means to revolutionize experimentation by empowering automation with artificial intelligence to enable autonomous discovery. However, the lack of adequate software solutions significantly impedes the development of self-driving laboratories. In this paper, we make progress towards addressing this challenge, and we propose and develop an implementation of ChemOS; a portable, modular and versatile software package which supplies the structured layers necessary for the deployment and operation of self-driving laboratories. ChemOS facilitates the integration of automated equipment, and it enables remote control of automated laboratories. ChemOS can operate at various degrees of autonomy; from fully unsupervised experimentation to actively including inputs and feedbacks from researchers into the experimentation loop. The flexibility of ChemOS provides a broad range of functionality as demonstrated on five applications, which were executed on different automated equipment, highlighting various aspects of the software package.
The increasing integration of software and automation in modern chemical laboratories prompts special emphasis on two important skills in the chemistry classroom. First, students need to learn the technical skills involved in modern scientific computing and automation. Second, applying these techniques in practice requires effective collaboration in teams. This work aims at developing a teaching module to help students gain both skills. In particular, we describe a modular and collaborative approach for introducing undergraduate students to scientific computing in the context of automated and autonomous chemical laboratories. Using online collaboration tools, students work in parallel teams to develop central components of an automated computer vision system that monitors color changes in ongoing chemical reactions. These components include three different aspects: image capture, communication, and data visualization. The image capture team collects and stores the images of the chemical reaction, the communication team processes the images, and the visualization team develops the tools for analyzing the processed image data. Using this educational framework, students built an open-source Python tool called AutoVis that enables the automated tracking of color and intensity changes in a liquid. The software is tested by simulating chemical reactions with dilute solutions of food coloring in water. It is shown that the system reliably tracks color and intensity, providing feedback to the experimentalist and enabling further computational analysis. Over the course of the project, students gain proficiency in scientific computing using Python and collaborate on software development using GitHub. In this way, they learn the role of software in chemical laboratories of the future.
Fundamental advances to increase the efficiency as well as stability of organic photovoltaics (OPVs) are achieved by designing ternary blends, which represents a clear trend toward multicomponent active layer blends. The development of high-throughput and autonomous experimentation methods is reported for the effective optimization of multicomponent polymer blends for OPVs. A method for automated film formation enabling the fabrication of up to 6048 films per day is introduced. Equipping this automated experimentation platform with a Bayesian optimization, a self-driving laboratory is constructed that autonomously evaluates measurements to design and execute the next experiments. To demonstrate the potential of these methods, a 4D parameter space of quaternary OPV blends is mapped and optimized for photostability. While with conventional approaches, roughly 100 mg of material would be necessary, the robot-based platform can screen 2000 combinations with less than 10 mg, and machine-learning-enabled autonomous experimentation identifies stable compositions with less than 1 mg.
Post-calculation analyses are often required to extract physical insights from ab initio molecular dynamics simulations. In the present work, we use different machine learning classifiers to take a new perspective on the decomposition reaction of dioxetane. Upon thermally activated decomposition, dioxetane can form products in an electronically excited state and can thus chemiluminesce. Simulated dynamics trajectories exhibit both successful and frustrated dissociations. As an exhaustive and systematic study of the decomposition mechanism "by hand" is beyond feasibility, machine learning models have been employed to study the relevant nuclear distortions governing molecular dissociation. According to all classifiers used in the study, the two sets of geometries differ by the in-phase planarisation of the two formaldehyde moieties. New insights are obtained from this analysis: if both moieties are not planar enough when the dissociation is attempted, it is frustrated and the molecule remains trapped. The postponing of the decomposition reaction by the so-called entropic trap enhances the chemiexcitation efficiency.
The digitalization of the economy is one of the drivers of the fourth industrial revolution. This trend is already heavily permeating biology laboratories and rapidly moving into chemistry as well. Notably, automated laboratories enhance process quality and intensification while freeing researchers from repetitive tasks. With these societal changes in place, students need to be prepared for the advanced digitization of chemistry and science by teaching fundamental chemistry concepts in combination with emerging Industry 4.0 technologies, including programming and automation. We describe an undergraduate classroom exercise at the interface of chemistry, computer science and engineering based on the development of an autonomous titration platform. Following an inquiry learning ansatz, the exercise focuses on standard titration experiments which are first executed manually, then automatically and finally in full autonomy by a student-designed robotic platform. We demonstrate that the exercise introduced in this work enables students to learn fundamental concepts in analytical chemistry, naturally integrates basic aspects of programming and automation, and as a consequence promotes and reinforces the detailed understanding of experimental processes and measurements. The exercise is designed in a collaborative active learning framework to encourage complex critical thinking and creative problem solving and thus prepares students for the next-generation chemistry laboratories.
Fast and inexpensive characterization of materials properties is a key element to discover novel functional materials. In this work, we suggest an approach employing three classes of Bayesian machine learning (ML) models to correlate electronic absorption spectra of nanoaggregates with the strength of intermolecular electronic couplings in organic conducting and semiconducting materials. As a specific model system, we consider poly(3,4-ethylenedioxythiophene) (PEDOT) polystyrene sulfonate, a cornerstone material for organic electronic applications, and so analyze the couplings between charged dimers of closely packed PEDOT oligomers that are at the heart of the material's unrivaled conductivity. We demonstrate that ML algorithms can identify correlations between the coupling strengths and the electronic absorption spectra. We also show that ML models can be trained to be transferable across a broad range of spectral resolutions and that the electronic couplings can be predicted from the simulated spectra with an 88% accuracy when ML models are used as classifiers. Although the ML models employed in this study were trained on data generated by a multiscale computational workflow, they were able to leverage experimental data.
The discovery of novel materials and functional molecules can help to solve some of society's most urgent challenges, ranging from efficient energy harvesting and storage to uncovering novel pharmaceutical drug candidates. Traditionally matter engineering -- generally denoted as inverse design -- was based massively on human intuition and high-throughput virtual screening. The last few years have seen the emergence of significant interest in computer-inspired designs based on evolutionary or deep learning methods. The major challenge here is that the standard strings molecular representation SMILES shows substantial weaknesses in that task because large fractions of strings do not correspond to valid molecules. Here, we solve this problem at a fundamental level and introduce SELFIES (SELF-referencIng Embedded Strings), a string-based representation of molecules which is 100\% robust. Every SELFIES string corresponds to a valid molecule, and SELFIES can represent every molecule. SELFIES can be directly applied in arbitrary machine learning models without the adaptation of the models; each of the generated molecule candidates is valid. In our experiments, the model's internal memory stores two orders of magnitude more diverse molecules than a similar test with SMILES. Furthermore, as all molecules are valid, it allows for explanation and interpretation of the internal working of the generative models.
Understanding the fundamental processes of light-harvesting is crucial to the development of clean energy materials and devices. Biological organisms have evolved complex metabolic mechanisms to efficiently convert sunlight into chemical energy. Unraveling the secrets of this conversion has inspired the design of clean energy technologies, including solar cells and photocatalytic water splitting. Describing the emergence of macroscopic properties from microscopic processes poses the challenge to bridge length and time scales of several orders of magnitude. Machine learning experiences increased popularity as a tool to bridge the gap between multi-level theoretical models and Edisonian trial-and-error approaches. Machine learning offers opportunities to gain detailed scientific insights into the underlying principles governing light-harvesting phenomena and can accelerate the fabrication of light-harvesting devices.
This chapter provides an overview of established algorithmic strategies for experiment planning for closed-loop experimentation highlighting their strengths and limitations through key examples from academia and industry. It also details the need for a transition from automation to autonomy in materials innovation and process optimization to accelerate discovery across sectors. In this context, we review the early realization of autonomous laboratories, and their associated strategies to optimization, and lay out a roadmap for deploying and orchestrating self-driving laboratories. As a specific tool to enable autonomy in technology innovation, we detail the architecture and suite of applications composing the ChemOS software package. We complete our discussion by highlighting recent demonstrations of ChemOS in chemistry, materials science and process optimization and discuss the specific use of ChemOS to accelerate drug discovery.