To overcome the challenges of designing novel peptides that bind to targeted sites in protein subdomains with specificity a computational method is presented based on a general design paradigm which generates candidate peptides that exhibit favorable binding specificity within a target protein. Candidate peptides are ranked by performing data mining for sequence and structural complementarity, followed by structure remodeling using molecular dynamics simulation and applying molecular docking protocols in a serial pipeline. Validation results are presented for the generated peptides on two model peptide-protein systems: the WW domain and the PDZ domain. Several runs of the pipeline are contrasted as hyper-parameters that control diversification levels are varied. Different rank-ordered lists of candidate peptides are compared to known binding peptides. Multiple examples of known binding motifs in various stages of the refinement process are found, as well as novel peptides. Candidate peptides are shown to be specific through intensive docking simulations. The biophysical properties of candidates reflect their specific binding partners, and a clear expansion beyond local sequence space for peptide generation is observed. After docking scores are established as a surrogate to binding affinity, binding predictions for the WW domain contain 20% short polyproline peptides that coincide with canonical motifs. Further, 37% of the candidate peptides predicted by the pipeline to bind with the PDZ domain match with known peptide binders validated by previous experiments. Finally, we present results on the tumor suppressing p53 protein. This exploratory protein system is framed with known alternative conformations to deduce candidate peptides that interfere with p53-MDM2 binding. This research was supported by NIGMS of the National Institutes of Health under award number R15GM146200.
The most common class of antibiotics, Beta Lactams, are most often resisted by an enzyme called Beta-Lactamase (BL) which hydrolyses the antibiotic's beta lactam ring, rendering it useless. Extended Spectrum BL (ESBL) expands substrate specificity to include third generation cephalosporins. A machine learning method called Supervised Projective Learning with Orthogonal Completeness (SPLOC) is applied to discriminate the functional dynamics in BL associated with substrate specificity while exploring potential binding targets for inhibitors.
Machine learning (ML) has been an important arsenal in computational biology used to elucidate protein function for decades. With the recent burgeoning of novel ML methods and applications, new ML approaches have been incorporated into many areas of computational biology dealing with protein function. We examine how ML has been integrated into a wide range of computational models to improve prediction accuracy and gain a better understanding of protein function. The applications discussed are protein structure prediction, protein engineering using sequence modifications to achieve stability and druggability characteristics, molecular docking in terms of protein–ligand binding, including allosteric effects, protein–protein interactions and protein-centric drug discovery. To quantify the mechanisms underlying protein function, a holistic approach that takes structure, flexibility, stability, and dynamics into account is required, as these aspects become inseparable through their interdependence. Another key component of protein function is conformational dynamics, which often manifest as protein kinetics. Computational methods that use ML to generate representative conformational ensembles and quantify differences in conformational ensembles important for function are included in this review. Future opportunities are highlighted for each of these topics.
Molecular dynamics data associated with the publication: "Functional Dynamics of Substrate Recognition in TEM Beta-Lactamase" Trajectores were generated in GROMACS, and the carbon alpha coordinates were extracted and aligned with the JEDi analysis software. Details of the simulations and analysis are given in the publication. Data in apo.zip contains trajectories for 32 total trajectories of TEM-1, TEM-2, TEM-10, and TEM-52 beta-lactamase, each starting form different 8 crystal structures Data in holo.zip contains 16 trajectories of TEM-1, TEM-2, TEM-10, and TEM-52 beta-lactamase in complex with ampicillin, amoxicillin, cefotaxime, and ceftazidime each. Trajectories files are in comma delimited format, with rows representing degrees for freedom (789 total), and columns representing samples (10000 per trajectory file). Supervised Projective Learning for Orthogonal Completeness (SPLOC) software for analysis as performed in the publication can be found at: https://github.com/BioMolecularPhysicsGroup-UNCC/MachineLearning/tree/master/SPLOC
The beta-lactamase enzyme provides effective resistance to beta-lactam antibiotics due to substrate recognition controlled by point mutations. Recently, extended-spectrum and inhibitor-resistant mutants have become a global health problem. Here, the functional dynamics that control substrate recognition in TEM beta-lactamase are investigated using all-atom molecular dynamics simulations. Comparisons are made between wild-type TEM-1 and TEM-2 and the extended-spectrum mutants TEM-10 and TEM-52, both in apo form and in complex with four different antibiotics (ampicillin, amoxicillin, cefotaxime and ceftazidime). Dynamic allostery is predicted based on a quasi-harmonic normal mode analysis using a perturbation scan. An allosteric mechanism known to inhibit enzymatic function in TEM beta-lactamase is identified, along with other allosteric binding targets. Mechanisms for substrate recognition are elucidated using multivariate comparative analysis of molecular dynamics trajectories to identify changes in dynamics resulting from point mutations and ligand binding, and the conserved dynamics, which are functionally important, are extracted as well. The results suggest that the H10-H11 loop (residues 214-221) is a secondary anchor for larger extended spectrum ligands, while the H9-H10 loop (residues 194-202) is distal from the active site and stabilizes the protein against structural changes. These secondary non-catalytically-active loops offer attractive targets for novel noncompetitive inhibitors of TEM beta-lactamase.
The effect of bias on hypothesis formation is characterized for an automated data-driven projection pursuit neural network to extract and select features for binary classification of data streams. This intelligent exploratory process partitions a complete vector state space into disjoint subspaces to create working hypotheses quantified by similarities and differences observed between two groups of labeled data streams. Data streams are typically time sequenced, and may exhibit complex spatio-temporal patterns. For example, given atomic trajectories from molecular dynamics simulation, the machine’s task is to quantify dynamical mechanisms that promote function by comparing protein mutants, some known to function while others are nonfunctional. Utilizing synthetic two-dimensional molecules that mimic the dynamics of functional and nonfunctional proteins, biases are identified and controlled in both the machine learning model and selected training data under different contexts. The refinement of a working hypothesis converges to a statistically robust multivariate perception of the data based on a context-dependent perspective. Including diverse perspectives during data exploration enhances interpretability of the multivariate characterization of similarities and differences.
Identifying mechanisms that control molecular function is a significant challenge in pharmaceutical science and molecular engineering. Here, we present a novel projection pursuit recurrent neural network to identify functional mechanisms in the context of iterative supervised machine learning for discovery-based design optimization. Molecular function recognition is achieved by pairing experiments that categorize systems with digital twin molecular dynamics simulations to generate working hypotheses. Feature extraction decomposes emergent properties of a system into a complete set of basis vectors. Feature selection requires signal-to-noise, statistical significance, and clustering quality to concurrently surpass acceptance levels. Formulated as a multivariate description of differences and similarities between systems, the data-driven working hypothesis is refined by analyzing new systems prioritized by a discovery-likelihood. Utility and generality are demonstrated on several benchmarks, including the elucidation of antibiotic resistance in TEM-52 beta-lactamase. The software is freely available, enabling turnkey analysis of massive data streams found in computational biology and material science.
Background Principal component analysis (PCA) is commonly applied to the atomic trajectories of biopolymers to extract essential dynamics that describe biologically relevant motions. Although application of PCA is straightforward, specialized software to facilitate workflows and analysis of molecular dynamics simulation data to fully harness the power of PCA is lacking. The Java Essential Dynamics inspector (JEDi) software is a major upgrade from the previous JED software. Results Employing multi-threading, JEDi features a user-friendly interface to control rapid workflows for interrogating conformational motions of biopolymers at various spatial resolutions and within subregions, including multiple chain proteins. JEDi has options for Cartesian-based coordinates (cPCA) and internal distance pair coordinates (dpPCA) to construct covariance (Q), correlation (R), and partial correlation (P) matrices. Shrinkage and outlier thresholding are implemented for the accurate estimation of covariance. The effect of rare events is quantified using outlier and inlier filters. Applying sparsity thresholds in statistical models identifies latent correlated motions. Within a hierarchical approach, small-scale atomic motion is first calculated with a separate local cPCA calculation per residue to obtain eigenresidues. Then PCA on the eigenresidues yields rapid and accurate description of large-scale motions. Local cPCA on all residue pairs creates a map of all residue-residue dynamical couplings. Additionally, kernel PCA is implemented. JEDi output gives high quality PNG images by default, with options for text files that include aligned coordinates, several metrics that quantify mobility, PCA modes with their eigenvalues, and displacement vector projections onto the top principal modes. JEDi provides PyMol scripts together with PDB files to visualize individual cPCA modes and the essential dynamics occurring within user-selected time scales. Subspace comparisons performed on the most relevant eigenvectors using several statistical metrics quantify similarity/overlap of high dimensional vector spaces. Free energy landscapes are available for both cPCA and dpPCA. Conclusion JEDi is a convenient toolkit that applies best practices in multivariate statistics for comparative studies on the essential dynamics of similar biopolymers. JEDi helps identify functional mechanisms through many integrated tools and visual aids for inspecting and quantifying similarity/differences in mobility and dynamic correlations.
Data needed to reproduce the synthetic molecule analysis from SPLOC paper and the Egghunt also described in that work. Scripts to run this analysis can be found here https://github.com/BioMolecularPhysicsGroup-UNCC/Publications
Beta-lactamases are characterized by their ability to hydrolyze beta-lactam antibiotics. Two common beta-lactamases, TEM-1 and TEM-52, are compared to better understand the role of dynamics in this process. TEM-1 is active against penicillin and first generation cephalosporins, whereas three mutations confer a much broader set of extended spectrum activities to TEM-52. Combined with geometric simulation results, we employ molecular dynamics (MD) simulation to explore motions within these enzymes. We generated eight all-atom 500 ns MD simulations in explicit solvent for both TEM-1 and TEM-52 starting from a variety of different crystal structures. We quantify dynamical differences using various types of Principal Component Analysis (PCA). PCA is commonly used to reduce the high dimensional space describing protein motion to a much smaller subspace in order to characterize essential motions. An important challenge facing protein modelling is the identification of the functional dynamical modes that are often hidden within nonfunctional global dynamics. We investigate the ability of PCA to identify functional modes that discriminate among proteins exhibiting different substrate specificities. We find that traditional Cartesian PCA methods perform well in identifying different dynamics between different crystal structure starting points, but they are mostly unable to distinguish dynamical differences due to sequence variation. We are optimizing methods to uncover functional difference by focusing on subsets of residues and utilizing multiple PCA methods in combination with each other. Moreover, we apply clustering and machine learning methods to better integrate and evaluate the different types of PCA results. We have been able to detect dynamical differences in the TEM-1 and TEM-52 sequence through the projection of top PCA modes onto the conformational ensembles in several instances and we are optimizing the method to obtain a robust measure of statistical significance.