Predicting favorable protein-peptide binding events remains a central challenge in biophysics, with continued uncertainty surrounding how nonlocal effects shape the global energy landscape. Here, we introduce peripheral surface information (PSI) entropy, a quantitative measure of the statistical variability in apolar and charged non-interacting surface (NIS) proportions across conformational ensembles. Using energy-directed molecular docking via HADDOCK3 and explicit-solvent molecular dynamics simulations, it is demonstrated that favorable binding partners exhibit emergent, low-entropy N-states (discrete macrostates in NIS state space) indicative of preferential apolar/charged surface configurations. Across dozens of peptides and multiple receptor systems (WW, PDZ, and MDM2 domains), dominant N-states persisted under varied docking parameters and initial conditions. An experimental meta-ensemble of WW domains from 36 high-resolution structures confirmed the presence of dominant NIS modes independent of in silico methodology, suggesting an evolutionary selection pressure toward specific NIS fingerprints. These findings establish PSI entropy as a thermoinformatic descriptor that encodes favorable binding constraints into unique statistical signatures of the NIS.
To overcome the challenges of designing novel peptides that bind to targeted sites in protein subdomains with specificity a computational method is presented based on a general design paradigm which generates candidate peptides that exhibit favorable binding specificity within a target protein. Candidate peptides are ranked by performing data mining for sequence and structural complementarity, followed by structure remodeling using molecular dynamics simulation and applying molecular docking protocols in a serial pipeline. Validation results are presented for the generated peptides on two model peptide-protein systems: the WW domain and the PDZ domain. Several runs of the pipeline are contrasted as hyper-parameters that control diversification levels are varied. Different rank-ordered lists of candidate peptides are compared to known binding peptides. Multiple examples of known binding motifs in various stages of the refinement process are found, as well as novel peptides. Candidate peptides are shown to be specific through intensive docking simulations. The biophysical properties of candidates reflect their specific binding partners, and a clear expansion beyond local sequence space for peptide generation is observed. After docking scores are established as a surrogate to binding affinity, binding predictions for the WW domain contain 20% short polyproline peptides that coincide with canonical motifs. Further, 37% of the candidate peptides predicted by the pipeline to bind with the PDZ domain match with known peptide binders validated by previous experiments. Finally, we present results on the tumor suppressing p53 protein. This exploratory protein system is framed with known alternative conformations to deduce candidate peptides that interfere with p53-MDM2 binding. This research was supported by NIGMS of the National Institutes of Health under award number R15GM146200.
Motivation Ab initio gene prediction in nonmodel organisms is a difficult task. While many ab initio methods have been developed, their average accuracy over long segments of a genome, and especially when assessed over a wide range of species, generally yields results with sensitivity and specificity levels in the low 60% range. A common weakness of most methods is the tendency to learn patterns that are species-specific to varying degrees. The need exists for methods to extract genetic features that can distinguish coding and noncoding regions that are not sensitive to specific organism characteristics.Results A new method based on a neural network (NN) that uses a collection of sensors to create input features is presented. It is shown that accurate predictions are achieved even when trained on organisms that are significantly different phylogenetically than test organisms. A consensus prediction algorithm for a CoDing Sequence (CDS) is subsequently applied to the first nucleotide level of NN predictions that boosts accuracy through a data-driven procedure that optimizes a CDS/non-CDS threshold. An aggregate accuracy benchmark at the nucleotide level shows that this new approach performs better than existing ab initio methods, while requiring significantly less training data.Availability and implementation https://github.com/BioMolecularPhysicsGroup-UNCC/MachineLearning.
We present a novel nonparametric adaptive partitioning and stitching (NAPS) algorithm to estimate a probability density function (PDF) of a single variable. Sampled data is partitioned into blocks using a branching tree algorithm that minimizes deviations from a uniform density within blocks of various sample sizes arranged in a staggered format. The block sizes are constructed to balance the load in parallel computing as the PDF for each block is independently estimated using the nonparametric maximum entropy method (NMEM) previously developed for automated high throughput analysis. Once all block PDFs are calculated, they are stitched together to provide a smooth estimate throughout the sample range. Each stitch is an averaging process over weight factors based on the estimated cumulative distribution function (CDF) and a complementary CDF that characterize how data from flanking blocks overlap. Benchmarks on synthetic data show that our PDF estimates are fast and accurate for sample sizes ranging from 2(9) to 2(27), across a diverse set of distributions that account for single and multi-modal distributions with heavy tails or singularities. We also generate estimates by replacing NMEM with kernel density estimation (KDE) within blocks. Our results indicate that NAPS(NMEM) is the best-performing method overall, while NAPS(KDE) improves estimates near boundaries compared to standard KDE.
Nonparametric estimation for a probability density function that describes multivariate data has typically been addressed by kernel density estimation (KDE). A novel density estimator recently developed by Farmer and Jacobs offers an alternative high-throughput automated approach to univariate nonparametric density estimation based on maximum entropy and order statistics, improving accuracy over univariate KDE. This article presents an extension of the single variable case to multiple variables. The univariate estimator is used to recursively calculate a product array of one-dimensional conditional probabilities. In combination with interpolation methods, a complete joint probability density estimate is generated for multiple variables. Good accuracy and speed performance in synthetic data are demonstrated by a numerical study using known distributions over a range of sample sizes from 100 to 106 for two to six variables. Performance in terms of speed and accuracy is compared to KDE. The multivariate density estimate developed here tends to perform better as the number of samples and/or variables increases. As an example application, measurements are analyzed over five filters of photometric data from the Sloan Digital Sky Survey Data Release 17. The multivariate estimation is used to form the basis for a binary classifier that distinguishes quasars from galaxies and stars with up to 94% accuracy.
The most common class of antibiotics, Beta Lactams, are most often resisted by an enzyme called Beta-Lactamase (BL) which hydrolyses the antibiotic's beta lactam ring, rendering it useless. Extended Spectrum BL (ESBL) expands substrate specificity to include third generation cephalosporins. A machine learning method called Supervised Projective Learning with Orthogonal Completeness (SPLOC) is applied to discriminate the functional dynamics in BL associated with substrate specificity while exploring potential binding targets for inhibitors.
This article presents PDFEstimator, an R package for nonparametric probability density estimation and analysis, as both a practical enhancement and alternative to kernel-based estimators. PDFEstimator creates fast, highly accurate, data-driven probability density estimates for continuous random data through an intuitive interface. Excellent results are obtained for a diverse set of data distributions ranging from 10 to 106 samples when invoked with default parameter definitions in the absence of user directives. Additionally, the package contains methods for assessing the quality of any estimate, including robust plotting functions for detailed visualization and trouble-shooting. Usage of PDFEstimator is illustrated through a variety of examples, including comparisons to several kernel density methods.
Machine learning (ML) has been an important arsenal in computational biology used to elucidate protein function for decades. With the recent burgeoning of novel ML methods and applications, new ML approaches have been incorporated into many areas of computational biology dealing with protein function. We examine how ML has been integrated into a wide range of computational models to improve prediction accuracy and gain a better understanding of protein function. The applications discussed are protein structure prediction, protein engineering using sequence modifications to achieve stability and druggability characteristics, molecular docking in terms of protein–ligand binding, including allosteric effects, protein–protein interactions and protein-centric drug discovery. To quantify the mechanisms underlying protein function, a holistic approach that takes structure, flexibility, stability, and dynamics into account is required, as these aspects become inseparable through their interdependence. Another key component of protein function is conformational dynamics, which often manifest as protein kinetics. Computational methods that use ML to generate representative conformational ensembles and quantify differences in conformational ensembles important for function are included in this review. Future opportunities are highlighted for each of these topics.
The usual goal of antibody engineering is to affect the binding affinity of antibodies with their antigens or other ligands through directed mutation. Experiments in engineered ScFv antibody fragments have shown that mutations entirely distant from the active site can have very significant affects on binding affinity[1]. A method based on molecular dynamics simulations and normal mode analysis is used to predict and test allosteric mutations in engineered antibodies and antibody fragments. A course grained carbon-alpha covariance matrix is extracted from an all atom MD simulation.
Molecular dynamics data associated with the publication: "Functional Dynamics of Substrate Recognition in TEM Beta-Lactamase" Trajectores were generated in GROMACS, and the carbon alpha coordinates were extracted and aligned with the JEDi analysis software. Details of the simulations and analysis are given in the publication. Data in apo.zip contains trajectories for 32 total trajectories of TEM-1, TEM-2, TEM-10, and TEM-52 beta-lactamase, each starting form different 8 crystal structures Data in holo.zip contains 16 trajectories of TEM-1, TEM-2, TEM-10, and TEM-52 beta-lactamase in complex with ampicillin, amoxicillin, cefotaxime, and ceftazidime each. Trajectories files are in comma delimited format, with rows representing degrees for freedom (789 total), and columns representing samples (10000 per trajectory file). Supervised Projective Learning for Orthogonal Completeness (SPLOC) software for analysis as performed in the publication can be found at: https://github.com/BioMolecularPhysicsGroup-UNCC/MachineLearning/tree/master/SPLOC
A MATLAB function is presented for nonparametric probability density estimation, based on an iterative method that employs the principle of maximum entropy and characteristic properties of single order statistics. Featuring a robust and adaptive design, the method is well-suited for high throughput applications. The implementation comprises a MATLAB interface and underlying C++ code with extensible components that can be easily integrated into third party software. The functionality includes plotting capabilities and model independent diagnostics featuring the scaled quantile residual that is invariant to sample size, distribution, and estimation method.
The beta-lactamase enzyme provides effective resistance to beta-lactam antibiotics due to substrate recognition controlled by point mutations. Recently, extended-spectrum and inhibitor-resistant mutants have become a global health problem. Here, the functional dynamics that control substrate recognition in TEM beta-lactamase are investigated using all-atom molecular dynamics simulations. Comparisons are made between wild-type TEM-1 and TEM-2 and the extended-spectrum mutants TEM-10 and TEM-52, both in apo form and in complex with four different antibiotics (ampicillin, amoxicillin, cefotaxime and ceftazidime). Dynamic allostery is predicted based on a quasi-harmonic normal mode analysis using a perturbation scan. An allosteric mechanism known to inhibit enzymatic function in TEM beta-lactamase is identified, along with other allosteric binding targets. Mechanisms for substrate recognition are elucidated using multivariate comparative analysis of molecular dynamics trajectories to identify changes in dynamics resulting from point mutations and ligand binding, and the conserved dynamics, which are functionally important, are extracted as well. The results suggest that the H10-H11 loop (residues 214-221) is a secondary anchor for larger extended spectrum ligands, while the H9-H10 loop (residues 194-202) is distal from the active site and stabilizes the protein against structural changes. These secondary non-catalytically-active loops offer attractive targets for novel noncompetitive inhibitors of TEM beta-lactamase.
The effect of bias on hypothesis formation is characterized for an automated data-driven projection pursuit neural network to extract and select features for binary classification of data streams. This intelligent exploratory process partitions a complete vector state space into disjoint subspaces to create working hypotheses quantified by similarities and differences observed between two groups of labeled data streams. Data streams are typically time sequenced, and may exhibit complex spatio-temporal patterns. For example, given atomic trajectories from molecular dynamics simulation, the machine’s task is to quantify dynamical mechanisms that promote function by comparing protein mutants, some known to function while others are nonfunctional. Utilizing synthetic two-dimensional molecules that mimic the dynamics of functional and nonfunctional proteins, biases are identified and controlled in both the machine learning model and selected training data under different contexts. The refinement of a working hypothesis converges to a statistically robust multivariate perception of the data based on a context-dependent perspective. Including diverse perspectives during data exploration enhances interpretability of the multivariate characterization of similarities and differences.
Identifying mechanisms that control molecular function is a significant challenge in pharmaceutical science and molecular engineering. Here, we present a novel projection pursuit recurrent neural network to identify functional mechanisms in the context of iterative supervised machine learning for discovery-based design optimization. Molecular function recognition is achieved by pairing experiments that categorize systems with digital twin molecular dynamics simulations to generate working hypotheses. Feature extraction decomposes emergent properties of a system into a complete set of basis vectors. Feature selection requires signal-to-noise, statistical significance, and clustering quality to concurrently surpass acceptance levels. Formulated as a multivariate description of differences and similarities between systems, the data-driven working hypothesis is refined by analyzing new systems prioritized by a discovery-likelihood. Utility and generality are demonstrated on several benchmarks, including the elucidation of antibiotic resistance in TEM-52 beta-lactamase. The software is freely available, enabling turnkey analysis of massive data streams found in computational biology and material science.
Molecular dynamics was used to study the aggregation in water of three versions of polyhedral oligomeric silsesquioxane porphyrin (POSSP) molecules. The POSSP molecules have different functional groups; POSSP-IB has one hepta-isobutyl POSS unit, POSSP-Ph has one hepta-phenyl POSS molecule, and POSSP-TIB has four hepta-isobutyl POSS units. Three control porphyrins (TPP, ATPP, TATPP) were also simulated in this study. The effects of the different substituents on the POSSP aggregation process and final morphology were investigated. It is observed that the isobutyl substituents in the POSS units drives the aggregation mainly through the hydrophobic effect. In the case of the phenyl POSS unit, the self-assembly process is also carried through the hydrophobic effect, but π-π and H-bonding interactions play a role too. The final morphology of the aggregates show that the porphyrins associated with POSSP-IB and POSSP-TIB are far apart from each other contrary to POSSP-Ph, which may have a major implication on the optical properties of these aggregates. This study provides valuable insights on the aggregation in water of POSSP molecules, which are a tunable platform to design novel functional materials.
Background Principal component analysis (PCA) is commonly applied to the atomic trajectories of biopolymers to extract essential dynamics that describe biologically relevant motions. Although application of PCA is straightforward, specialized software to facilitate workflows and analysis of molecular dynamics simulation data to fully harness the power of PCA is lacking. The Java Essential Dynamics inspector (JEDi) software is a major upgrade from the previous JED software. Results Employing multi-threading, JEDi features a user-friendly interface to control rapid workflows for interrogating conformational motions of biopolymers at various spatial resolutions and within subregions, including multiple chain proteins. JEDi has options for Cartesian-based coordinates (cPCA) and internal distance pair coordinates (dpPCA) to construct covariance (Q), correlation (R), and partial correlation (P) matrices. Shrinkage and outlier thresholding are implemented for the accurate estimation of covariance. The effect of rare events is quantified using outlier and inlier filters. Applying sparsity thresholds in statistical models identifies latent correlated motions. Within a hierarchical approach, small-scale atomic motion is first calculated with a separate local cPCA calculation per residue to obtain eigenresidues. Then PCA on the eigenresidues yields rapid and accurate description of large-scale motions. Local cPCA on all residue pairs creates a map of all residue-residue dynamical couplings. Additionally, kernel PCA is implemented. JEDi output gives high quality PNG images by default, with options for text files that include aligned coordinates, several metrics that quantify mobility, PCA modes with their eigenvalues, and displacement vector projections onto the top principal modes. JEDi provides PyMol scripts together with PDB files to visualize individual cPCA modes and the essential dynamics occurring within user-selected time scales. Subspace comparisons performed on the most relevant eigenvectors using several statistical metrics quantify similarity/overlap of high dimensional vector spaces. Free energy landscapes are available for both cPCA and dpPCA. Conclusion JEDi is a convenient toolkit that applies best practices in multivariate statistics for comparative studies on the essential dynamics of similar biopolymers. JEDi helps identify functional mechanisms through many integrated tools and visual aids for inspecting and quantifying similarity/differences in mobility and dynamic correlations.
How can an income tax system be designed to exploit human nature and a free market to create a poverty free society, while balancing budgets without disproportional tax burdens? Such a tax system, with universal character, is deduced from the following guiding principles: (1) a single tax rate applies to all income types and levels; (2) the tax rate adjusts to satisfy budget projections; (3) government transfer only supplements the income of households with self-generated income below the poverty line; (4) deductions for basic living expenses, itemized investments and capital losses are allowed; (5) deductions cannot be applied to government transfer. A general framework emerges with three parameters that determine a minimum allowed tax deduction, a maximum allowed itemized deduction, and a maximum deduction defined by income percentage. An income distribution that mimics the United States, and a series of log-normal distributions are considered to quantitatively compare detailed characteristics of this tax system to progressive and flat tax systems. To minimize government dependency while maximizing after-tax income, the effective tax rate (ETR) as a function of income percentile takes the shape of the letter, V, inspiring the name victory tax, where the middle class has the lowest ETR.
A novel bottom-up approach for EEG signal artefact removal and classification is presented. An eigenbrain is constructed from two eigenchannels on a per subject basis. An eigenchannel provides a characteristic vector space to discriminate motor imagery from resting state signals in both the µ and β frequency bands. Discrimination between left and right hand motor imagery is achieved at the eigenbrain level as a direct sum of the eigenchannel vector spaces. It was found this methodology is viable using only 5% of available motor imagery trials during training, yielding 92.6% classification accuracy on the selected test subject. Furthermore, the utility for real-time applications is promising with a rapid classification that takes less than 100 ms after the initial calibration procedure is completed.