Inverse folding models have proven to be highly effective zero-shot predictors of protein stability. Despite this success, the link between the amino acid preferences of an inverse folding model and the free-energy considerations underlying thermodynamic stability remains incompletely understood. A better understanding would be of interest not only from a theoretical perspective, but also potentially provide the basis for stronger zero-shot stability prediction. In this paper, we take steps to clarify the free-energy foundations of inverse folding models. Our derivation reveals the standard practice of likelihood ratios as a simplistic approximation and suggests several paths towards better estimates of the relative stability. We empirically assess these approaches and demonstrate that considerable gains in zero-shot performance can be achieved with fairly simple means.
Bayesian optimization (BO) is an attractive machine learning framework for performing sample-efficient global optimization of black-box functions. The optimization process is guided by an acquisition function that selects points to acquire in each round of BO. In batched BO, when multiple points are acquired in parallel, commonly used acquisition functions are often high-dimensional and intractable, leading to the use of sampling-based alternatives. We propose a statistical physics inspired acquisition function that can natively handle batches. Batched Energy-Entropy acquisition for BO (BEEBO) enables tight control of the explore-exploit trade-off of the optimization process and generalizes to heteroskedastic black-box problems. We demonstrate the applicability of BEEBO on a range of problems, showing competitive performance to existing acquisition functions.
Pre-trained protein language models have demonstrated significant applicability in different protein engineering task. A general usage of these pre-trained transformer models latent representation is to use a mean pool across residue positions to reduce the feature dimensions to further downstream tasks such as predicting bio-physics properties or other functional behaviours. In this paper we provide a two-fold contribution to machine learning (ML) driven drug design. Firstly, we demonstrate the power of sparsity by promoting penalization of pre-trained transformer models to secure more robust and accurate melting temperature (Tm) prediction of single-chain variable fragments with a mean absolute error of 0.23C. Secondly, we demonstrate the power of framing our prediction problem in a probabilistic framework. Specifically, we advocate for the need of adopting probabilistic frameworks especially in the context of ML driven drug design.
Phosphatidylinositol 4,5-bisphosphate (PI(4,5)P2) lipids have been shown to stabilize an active conformation of class A G-protein coupled receptors (GPCRs) through a conserved binding site, not present in class B GPCRs. For class B GPCRs, previous molecular dynamics (MD) simulation studies have shown PI(4,5)P2 interacting with the Glucagon receptor (GCGR), which constitutes an important target for diabetes and obesity therapeutics. In this work, we applied MD simulations supported by native mass spectrometry (nMS) to study lipid interactions with GCGR. We demonstrate how tail composition plays a role in modulating the binding of PI(4,5)P2 lipids to GCGR. Specifically, we find the PI(4,5)P2 lipids to have a higher affinity toward the inactive conformation of GCGR. Interestingly, we find that in contrast to class A GPCRs, PI(4,5)P2 appear to stabilize the inactive conformation of GCGR through a binding site conserved across class B GPCRs but absent in class A GPCRs. This suggests differences in the regulatory function of PI(4,5)P2 between class A and class B GPCRs.
Nanodisc technology is increasingly being applied for structural and biophysical studies of membrane proteins. In this work, we present a general protocol for constructing molecular models of nanodiscs for molecular dynamics simulations. The protocol is written in python and based on geometric equations, making it fast and easy to modify, enabling automation and customization of nanodiscs in silico. The novelty being the ability to construct any membrane scaffold protein (MSP) variant fast and easy given only an input sequence. We validated and tested the protocol by simulating seven different nanodiscs of various sizes and with different membrane scaffold proteins, both circularized and noncircularized. The structural and biophysical properties were analyzed and shown to be in good agreement with previously reported experimental data and simulation studies.
We describe a novel Bayesian method for estimating protein concentration and phosphorylation site occupancy ratios from mass spectrometry experiments. Our variance model assigns standard deviations to all quantitative ratios, even when only a single peptide is observed, increasing the number of quantifiable observations in a sample compared to conventional methods. We further demonstrate the application of this method using a dataset investigating the impact of the PRKAR1A-RET gene fusion in immortalized thyroid cells.
Cell migration is a central biological process that requires fine coordination of molecular events in time and space. A deregulation of the migratory phenotype is also associated with pathological conditions including cancer where cell motility has a causal role in tumor spreading and metastasis formation. Thus cell migration is of critical and strategic importance across the complex disease spectrum as well as for the basic understanding of cell phenotype. Experimental studies of the migration of cells in monolayers are often conducted with 'wound healing' assays. Analysis of these assays has traditionally relied on how the wound area changes over time. However this method does not take into account the shape of the wound. Given the many options for creating a wound healing assay and the fact that wound shape invariably changes as cells migrate this is a significant flaw. Here we present a novel software package for analyzing concerted cell velocity in wound healing assays. Our method encompasses a wound detection algorithm based on cell confluency thresholding and employs a Bayesian approach in order to estimate concerted cell velocity with an associated likelihood. We have applied this method to study the effect of siRNA knockdown on the migration of a breast cancer cell line and demonstrate that cell velocity can track wound healing independently of wound shape and provides a more robust quantification with significantly higher signal to noise ratios than conventional analyses of wound area. The software presented here will enable other researchers in any field of cell biology to quantitatively analyze and track live cell migratory processes and is therefore expected to have a significant impact on the study of cell migration, including cancer relevant processes. Installation instructions, documentation and source code can be found at http://bowhead.lindinglab.science licensed under GPLv3.
It is well known that the development of drug resistance in cancer cells can lead to changes in cell morphology. Here, we describe the use of deep neural networks to analyze this relationship, demonstrating that complex cell morphologies can encode states of signaling networks and unravel cellular mechanisms hidden to conventional approaches. We perform high-content screening of 17 cancer cell lines, generating more than 500 billion data points from similar to 850 million cells. We analyze these data using a deep learning model, resulting in the identification of a continuous 27-dimension space describing all of the observed cell morphologies. From its morphology alone, we could thus predict whether a cell was resistant to ErbB-family drugs, with an accuracy of 74%, and predict the potential mechanism of resistance, subsequently validating the role of MET and insulin-like growth factor 1 receptor (IGF1R) as drivers of cetuximab resistance in in vitro models of lung and head/neck cancer.
What do you do to start reading introduction to vector and tensor analysis? Searching the book that you love to read first or find an interesting book that will make you want to read? Everybody has difference with their reason of reading a book. Actuary, reading habit must be from earlier. Many people may be love to read, but not a book. It's not fault. Someone will be bored to open the thick book with small words to read. In more, this is the real condition. So do happen probably with this introduction to vector and tensor analysis.
Signaling networks are nonlinear and complex, involving a large ensemble of dynamic interaction states that fluctuate in space and time. However, therapeutic strategies, such as combination chemotherapy, rarely consider the timing of drug perturbations. If we are to advance drug discovery for complex diseases, it will be essential to develop methods capable of identifying dynamic cellular responses to clinically relevant perturbations. Here, we present a Bayesian dose-response framework and the screening of an oncological drug matrix, comprising 10,000 drug combinations in melanoma and pancreatic cancer cell lines, from which we predict sequentially effective drug combinations. Approximately 23% of the tested combinations showed high-confidence sequential effects (either synergistic or antagonistic), demonstrating that cellular perturbations of many drug combinations have temporal aspects, which are currently both underutilized and poorly understood.
JF acknowledge funding from the Danish Council for Independent Research | Natural Sciences. ZG acknowledge funding from EPSRC EP/I036575/1 and Google.
Despite the development of powerful computational tools, the full-sequence design of proteins still remains a challenging task. To investigate the limits and capabilities of computational tools, we conducted a study of the ability of the program Rosetta to predict sequences that recreate the authentic fold of thioredoxin. Focusing on the influence of conformational details in the template structures, we based our study on 8 experimentally determined template structures and generated 120 designs from each. For experimental evaluation, we chose six sequences from each of the eight templates by objective criteria. The 48 selected sequences were evaluated based on their progressive ability to (1) produce soluble protein in Escherichia coli and (2) yield stable monomeric protein, and (3) on the ability of the stable, soluble proteins to adopt the target fold. Of the 48 designs, we were able to synthesize 32, 20 of which resulted in soluble protein. Of these, only two were sufficiently stable to be purified. An X-ray crystal structure was solved for one of the designs, revealing a close resemblance to the target structure. We found a significant difference among the eight template structures to realize the above three criteria despite their high structural similarity. Thus, in order to improve the success rate of computational full-sequence design methods, we recommend that multiple template structures are used. Furthermore, this study shows that special care should be taken when optimizing the geometry of a structure prior to computational design when using a method that is based on rigid conformations.
Cancer cells acquire pathological phenotypes through accumulation of mutations that perturb signaling networks. However, global analysis of these events is currently limited. Here, we identify six types of network-attacking mutations (NAMs), including changes in kinase and SH2 modulation, network rewiring, and the genesis and extinction of phosphorylation sites. We developed a computational platform (ReKINect) to identify NAMs and systematically interpreted the exomes and quantitative (phospho-)proteomes of five ovarian cancer cell lines and the global cancer genome repository. We identified and experimentally validated several NAMs, including PKCγ M501I and PKD1 D665N, which encode specificity switches analogous to the appearance of kinases de novo within the kinome. We discover mutant molecular logic gates, a drift toward phospho-threonine signaling, weakening of phosphorylation motifs, and kinase-inactivating hotspots in cancer. Our method pinpoints functional NAMs, scales with the complexity of cancer genomes and cell signaling, and may enhance our capability to therapeutically target tumor-specific networks.
Proteins are biomolecules that are of great importance in science, biotechnology and medicine. Their function relies heavily on their three-dimensional shape, which in turn follows from their amino acid sequence. Therefore, there is great interest in modelling the three-dimensional structure of proteins in silico given their sequence. We discuss the formulation of a tractable probabilistic model of protein structure that features atomic detail and can be used for protein structure prediction. The model unites dynamic Bayesian networks and directional statistics to cover the short-range features of proteins. Long-range features are added by making use of probability kinematics – a little known variant of Bayesian belief updating first proposed by the probability theorist Richard Jeffrey in the 1950s. The method we describe can be generalized to formulate tractable probabilistic models that involve high dimensionality and need to cover multiple scales
Microbial biofilm colonies will in many cases form a smart material capable of responding to external threats dependent on their size and internal state. The microbial community accordingly switches between passive, protective, or attack modes of action. In order to decide which strategy to employ, it is essential for the biofilm community to be able to sense its own size. The sensor designed to perform this task is termed a quorum sensor, since it only permits collective behaviour once a sufficiently large assembly of microbes have been established. The generic quorum sensor construct involves two genes, one coding for the production of a diffusible signal molecule and one coding for a regulator protein dedicated to sensing the signal molecules. A positive feedback in the signal molecule production sets a well-defined condition for switching into the collective mode. The activation of the regulator involves a slow dimerization, which allows low-pass filtering of the activation of the collective mode. Here, we review and combine the model components that form the basic quorum sensor in a number of Gram-negative bacteria, e.g., Pseudomonas aeruginosa.
The multicanonical, or flat-histogram, method is a common technique to improve the sampling efficiency of molecular simulations. The idea is that free-energy barriers in a simulation can be removed by simulating from a distribution where all values of a reaction coordinate are equally likely, and subsequently reweight the obtained statistics to recover the Boltzmann distribution at the temperature of interest. While this method has been successful in practice, the choice of a flat distribution is not necessarily optimal. Recently, it was proposed that additional performance gains could be obtained by taking the position-dependent diffusion coefficient into account, thus placing greater emphasis on regions diffusing slowly. Although some promising examples of applications of this approach exist, the practical usefulness of the method has been hindered by the difficulty in obtaining sufficiently accurate estimates of the diffusion coefficient. Here, we present a simple, yet robust solution to this problem. Compared to current state-of-the-art procedures, the new estimation method requires an order of magnitude fewer data to obtain reliable estimates, thus broadening the potential scope in which this technique can be applied in practice.
Molecular dynamics and Monte Carlo simulation can provide detailed structural insights into biological systems by exploring the distribution of accessible conformational states. However, due to the rough free energy landscapes characterizing many systems, obtaining converged statistics in such simulation remains an important challenge. In order to address this issue, over the past decades, a number of enhanced sampling strategies have been proposed, including methods like simulated annealing, replica exchange, metadynamics and generalized ensemble approaches such as multicanonical (flat histogram) sampling. Most of these methods were designed to soften the landscape, making it easier to overcome the free energy barriers. Recently, it was proposed that additional performance gains could be obtained by taking the position-dependent diffusion coefficient along the reaction coordinate into account, thus placing greater emphasis on regions diffusing slowly so that the round trip rate is maximized (Trebst et al. 2004. 10.1103/PhysRevE.70.046701). Although the diffusion-optimized ensemble approach has shown great promise, the practical application of it has been hindered by the serious challenge of estimating the diffusion coefficient accurately, especially at the preliminary stage of a simulation when only a small part of the phase space has been explored. In this study, a simple yet robust solution to this problem is presented. We propose replacing the standard global estimation procedure with a local estimation technique that was previously introduced in the context of trajectory analysis and reaction coordinate optimization (Krivov et al. 2008. 10.1073/pnas.0800228105). We demonstrate that the precision of the proposed estimation method is dramatically increased, requiring considerable fewer data compared to current state-of-the-art procedures. Finally, we show that with our method, diffusion-optimized sampling arises as a natural generalization of the standard multicanonical ensemble, estimating a simple histogram of transitions rather than the usual histogram of states.
Ensembles of bacteria are able to coordinate their phenotypic behavior in accordance with the size, density, and growth state of the ensemble. This is achieved through production and exchange of diffusible signal molecules in a cell-cell regulatory system termed quorum sensing. In the generic quorum sensor a positive feedback in the production of signal molecules defines the conditions at which the collective behavior switches on. In spite of its conceptual simplicity, a proper measure of biofilm colony "size'' appears to be lacking. We establish that the cell density multiplied by a geometric factor which incorporates the boundary conditions constitutes an appropriate size measure. The geometric factor is the square of the radius for a spherical colony or a hemisphere attached to a reflecting surface. If surrounded by a rapidly exchanged medium, the geometric factor is divided by three. For a disk-shaped biofilm the geometric factor is the horizontal dimension multiplied by the height, and the square of the height of the biofilm if there is significant flow above the biofilm. A remarkably simple factorized expression for the size is obtained, which separates the all-or-none ignition caused by the positive feedback from the smoother activation outside the switching region.