Predicting the sensory properties of compounds is challenging due to the subjective nature of the experimental measurements. This testing relies on a panel of human participants and is therefore also expensive and time-consuming. We describe the application of a state-of-the-art deep learning method, Alchemite™, to the imputation of sparse physicochemical and sensory data and compare the results with conventional quantitative structure–activity relationship methods and a multi-target graph convolutional neural network. The imputation model achieved a substantially higher accuracy of prediction, with improvements in R2 between 0.26 and 0.45 over the next best method for each sensory property. We also demonstrate that robust uncertainty estimates generated by the imputation model enable the most accurate predictions to be identified and that imputation also more accurately predicts activity cliffs, where small changes in compound structure result in large changes in sensory properties. In combination, these results demonstrate that the use of imputation, based on data from less expensive, early experiments, enables better selection of compounds for more costly studies, saving experimental time and resources.
In this paper, we describe a combination of structural informatics approaches developed to mine data extracted from existing structure knowledge bases (Protein Data Bank and the GVK database) with a focus on kinase ATP-binding site data. In contrast to existing systems that retrieve and analyze protein structures, our techniques are centered on a database of ligand-bound geometries in relation to residues lining the binding site and transparent access to ligand-based SAR data. We illustrate the systems in the context of the Abelson kinase and related inhibitor structures.
The medicinal chemistry community has become increasingly aware of the value of tracking calculated physical properties such as molecular weight, topological polar surface area, rotatable bonds, and hydrogen bond donors and acceptors. We hypothesized that the shift to high-throughput synthetic practices over the past decade may be another factor that may predispose molecules to fail by steering discovery efforts toward achiral, aromatic compounds. We have proposed two simple and interpretable measures of the complexity of molecules prepared as potential drug candidates. The first is carbon bond saturation as defined by fraction sp(3) (Fsp(3)) where Fsp(3) = (number of sp(3) hybridized carbons/total carbon count). The second is simply whether a chiral carbon exists in the molecule. We demonstrate that both complexity (as measured by Fsp(3)) and the presence of chiral centers correlate with success as compounds transition from discovery, through clinical testing, to drugs. In an attempt to explain these observations, we further demonstrate that saturation correlates with solubility, an experimental physical property important to success in the drug discovery setting.
High throughput microsomal stability assays have been widely implemented in drug discovery and many companies have accumulated experimental measurements for thousands of compounds. Such datasets have been used to develop in silico models to predict metabolic stability and guide the selection of promising candidates for synthesis. This approach has proven most effective when selecting compounds from proposed virtual libraries prior to synthesis. However, these models are not easily interpretable at the structural level, and thus provide little insight to guide traditional synthetic efforts. We have developed global classification models of rat, mouse and human liver microsomal stability using in-house data. These models were built with FCFP_6 fingerprints using a Naïve Bayesian classifier within Pipeline Pilot. The test sets were correctly classified as stable or unstable with satisfying accuracies of 78, 77 and 75% for rat, human and mouse models, respectively. The prediction confidence was assigned using the Bayesian score to assess the applicability of the models. Using the resulting models, we developed a novel data mining strategy to identify structural features associated with good and bad microsomal stability. We also used this approach to identify structural features which are good for one species but bad for another. With these findings, the structure-metabolism relationships are likely to be understood faster and earlier in drug discovery.