The bromodomain and extra terminal (BET) family of bromodomain-containing proteins are important epigenetic regulators that elicit their effect through binding histone tail N-acetyl lysine (KAc) post-translational modifications. Recognition of such markers has been implicated in a range of oncology and immune diseases and, as such, small-molecule inhibition of the BET family bromodomain-KAc protein-protein interaction has received significant interest as a therapeutic strategy, with several potential medicines under clinical evaluation. This work describes the structure- and property-based optimization of a ligand and lipophilic efficient pan-BET bromodomain inhibitor series to deliver candidate I-BET787 (70) that demonstrates efficacy in a mouse model of inflammation and suitable properties for both oral and intravenous (IV) administration. This focused two-phase explore-exploit medicinal chemistry effort delivered the candidate molecule in 3 months with less than 100 final compounds synthesized.
Accurate methods to predict solubility from molecular structure are highly sought after in the chemical sciences. To assess the state of the art, the American Chemical Society organized a "Second Solubility Challenge " in 2019, in which competitors were invited to submit blinded predictions of the solubilities of 132 drug-like molecules. In the first part of this article, we describe the development of two models that were submitted to the Blind Challenge in 2019 but which have not previously been reported. These models were based on computationally inexpensive molecular descriptors and traditional machine learning algorithms and were trained on a relatively small data set of 300 molecules. In the second part of the article, to test the hypothesis that predictions would improve with more advanced algorithms and higher volumes of training data, we compare these original predictions with those made after the deadline using deep learning models trained on larger solubility data sets consisting of 2999 and 5697 molecules. The results show that there are several algorithms that are able to obtain near state-of-the-art performance on the solubility challenge data sets, with the best model, a graph convolutional neural network, resulting in an RMSE of 0.86 log units. Critical analysis of the models reveals systematic differences between the performance of models using certain feature sets and training data sets. The results suggest that careful selection of high quality training data from relevant regions of chemical space is critical for prediction accuracy but that other methodological issues remain problematic for machine learning solubility models, such as the difficulty in modeling complex chemical spaces from sparse training data sets.
Graph Neural Networks (GNNs) have recently gained in popularity, challenging molecular fingerprints or SMILES-based representations as the predominant way to represent molecules for binding affinity prediction. Although simple ligand-based graphs alone are already useful for affinity prediction, better performance on multi-target datasets has been achieved with models that incorporate 3D structural information. Most recent advances utilize complex GNN architectures to capture 3D protein-ligand information by incorporating ligand-interacting protein atoms as additional nodes in the graphs; or by building a second protein-based graph in parallel. This expands the graph considerably while obfuscating the shape of the underlying ligand, diminishing the advantage that GNNs have when encoding molecular structures. There is therefore a need for a simple and elegant molecular graph representation that retains the topology of the ligand while simultaneously encoding 3D protein-ligand interactions. We present Protein-Ligand Interaction Graphs (PLIGs): a simple way of representing atom-atom contacts of 3D protein-ligand complexes as node features for GNNs. PLIGs featurize an atom node in the molecular graph by describing each atom’s properties as well as all atom-atom contacts made with protein atoms within a distance threshold. The edges of the graph are therefore identical to ligand-based graphs, but the nodes encode the 3D protein-ligand contacts. Since PLIGs are applicable to any GNN architecture, we have benchmarked their performance with six different GNN architectures, and compared them to conventional ligand-based graphs and fingerprint-based multi-layer perceptron (MLP) models using the CASF-2016 benchmark set where we found PLIG-based Graph Attention Networks (GATNet) to be the best performing model ( ρ =0.84, RMSE=1.22 pK). In summary, we created a novel graph-based representation that incorporates 3D structural information into the node features of ligand-shaped molecular graphs. The PLIG representation is simple, elegant, flexible and easily customizable, opening up many possibilities of incorporating other 2D and 3D properties into the graph. Access The code and implementation for PLIGs and all models can be found at github.com/MarcMoesser/Protein-Ligand-Interaction-Graphs .
Effective drug discovery relies on the considered deployment of screening assays from the outset of a project. Selected carefully, the data from screening assays can guide series optimization and give insight into the likely performance of compounds in a patient population. Established assays can provide information on a wide variety of compound properties, ranging from target affinity and functional assessment through to pharmacokinetic and toxicological properties, and with capacity on a scale from high-throughput plate formats down to isolated bespoke experiments. Selecting the assays in which to screen a compound at a given stage in a discovery project is a balance between resource cost and information value, and can have a profound impact on the pace and outcomes of the project. This article presents commonly evaluated properties, their physiological relevance, caveats to their interpretation, and the integration of in silico, in vitro and in vivo data to characterize the holistic profiles of discovery compounds and series.
Machine learning approaches promise to accelerate and improve success rates in medicinal chemistry programs by more effectively leveraging available data to guide a molecular design. A key step of an automated computational design algorithm is molecule generation, where the machine is required to design high-quality, drug-like molecules within the appropriate chemical space. Many algorithms have been proposed for molecular generation; however, a challenge is how to assess the validity of the resulting molecules. Here, we report three Turing-inspired tests designed to evaluate the performance of molecular generators. Profound differences were observed between the performance of molecule generators in these tests, highlighting the importance of selection of the appropriate design algorithms for specific circumstances. One molecule generator, based on match molecular pairs, performed excellently against all tests and thus provides a valuable component for machine-driven medicinal chemistry design workflows.
recycling centres will remain open throughout the period of additional covid-19 restrictions
Successful management of classical ballet dancers with overuse injuries requires an understanding of the art form, precise knowledge of anatomy and awareness of certain conditions. Turnout is the single most fundamental physical attribute in classical ballet and ‘forcing turnout’ frequently contributes to overuse injuries. Common presenting conditions arising from the foot and ankle include problems at the first metatarsophalangeal joint, second metatarsal stress fractures, flexor hallucis longus tendinitis and anterior and posterior ankle impingement syndromes. Persistent shin pain in dancers is often due to chronic compartment syndrome, stress fracture of the posteromedial or anterior tibia. Knee pain can arise from patellofemoral syndrome, patellar tendon insertional pathologies, or a comination of both. Hip and back problems are also prevalent in dancers.