Abstract The human gut, containing 100 trillion microbes, is also considered the “second brain,” having control over the different functions of the physiological system. With advancements in bioinformatics and the development of sequencing technologies, researchers are able to explore the diversity and functional implications of gut microbiota (GM), which have become strongly associated with a variety of diseases. Microbial imbalance, or dysbiosis, acts as a biomarker for early detection and prognosis of a disease. Artificial Intelligence and Machine Learning (AI/ML) methods, although extensively used in predicting GM associated diseases, are seldom translated to having practical real-world outcomes, necessitating the design of robust AI/ML models applicable in real-world scenario. We have therefore come up with designing stacking-based ensemble architectures (EM1 and EM2), developed by integrating multiple ML-based learning algorithms for improving disease prediction accuracy. The GM datasets, after split into training and test sets, were eventually fed into the proposed two-layer ensemble models, which combines the output from standardized base learners via a meta-classifier, strengthening classification robustness as well as ensuring consistency in optimized performance across diverse datasets. Both the proposed hybrid ensemble models have emerged to be superior performers over all baseline and deep learning models, with an average accuracy of 0.87 and 0.84 respectively. By combining multiple learners, the proposed ensemble models outperform traditional single-algorithm-based approaches to attain higher accuracy and robustness on complex GM datasets. Key messages Development of stacking-based hybrid ensemble models (EM), which can be employed to integrate different AI/ML algorithms with better prediction accuracy of gut microbiome (GM)-associated diseases. Use of independent GM datasets with preprocessing methods such as SMOTE and PCA to address class imbalance and high dimensionality. All the proposed EM architectures are mostly superior to the existing state-of-the-art AI/ML methods (highest prediction accuracy: 0.87 and 0.84 with EM1 and EM2 models respectively) for GM diseases predictions. The cross-cohort validation demonstrates high prediction accuracy and robustness, (AUC values close to 0.98 and 0.99, for EM1 and EM2). These therefore demonstrate the effectiveness of EM frameworks for GM associated disease prediction, paving the way for corresponding applications in precision medicine.
Deep learning models are data-hungry, and synthetic (artificial) data has been shown to be invaluable when data availability is low. While this has been demonstrated in certain technology areas, adopting such an approach is new in machine learning (ML) applications in chemistry, except for some pre-training tasks. In drug discovery, predicting binding energy between proteins and ligands is crucial. Many ML-based studies have been proposed to predict protein-ligand binding affinity using existing experimental data. However, these models suffer from inherent biases. Recent efforts have produced PLAS-20k, a synthetic dataset of multiple protein-ligand complex (PLC) conformations generated using molecular dynamics (MD) simulations as a viable option to complement existing experimental data and improve binding affinity prediction. For the binding affinity prediction task, we employ Pafnucy, a deep convolutional neural network, and propose using multiple structures for each PLC from PLAS-20k for training. We compare four different statistical and ML-based result-aggregation techniques. This work demonstrates the utility of dynamic datasets in enhancing binding affinity predictions, laying the foundation for future improvements in predicting similar protein properties using synthetic datasets and more sophisticated models and methods. We propose that physics-based synthetic datasets can significantly help develop more accurate data-driven methods. Scientific contribution This study shows that incorporating synthetic molecular dynamics data improves deep learning models for protein–ligand binding affinity prediction beyond static experimental structures. By systematically evaluating frame selection and prediction aggregation strategies, we demonstrate that training on diverse conformational snapshots significantly enhances generalization and accuracy. Our results highlight that dynamic synthetic datasets can enable deep learning models to outperform conventional methods such as MM-PBSA while remaining computationally efficient.
Antimicrobial peptides (AMPs) are emerging as potent alternatives to conventional antibiotics, yet their diverse nature due to divergent mechanisms of action hinders rational design. Here, we present an electrostatics-stratified computational framework that uncovers key physicochemical principles governing AMP activity. Experimentally validated peptides were grouped by average charge per residue (i.e., the charge/length of the peptide) and analyzed through integrated sequence-, structure-, and chemistry-based descriptors. Distinct molecular signatures emerged across electrostatic regimes: low-charge/length peptides rely on amphipathic organization via structural compactness, whereas the intermediate-charge/length peptides exhibit balanced hydrophobicity and electrostatics. The high-charge peptides couple strong cationic attraction with lipophilicity and tryptophan anchoring to mainly disrupt membranes. Interestingly, hydrophobic moment, which is a measure of the amphipathicity, is found to be important in all three classes of AMPs. This study identifies distinguishing features of AMP sub-groups and suggests design guidelines for developing selective and potent next-generation AMPs.
Recent progress and development of artificial intelligence and machine learning (AI/ML) techniques have enabled addressing complex biomolecular problems. AI/ML models learn the underlying distribution of data they are trained on and when exposed to new inputs, they make predictions based on patterns and relationships previously observed in the training set. Further, generative artificial intelligence (GenAI) can be used to accurately generate protein structure or sequence from specific selected properties. This review specifically focuses on the applications of AI/ML in predicting important functional properties of proteins, and the potential prospects of reverseengineering in depicting the sequence and structure, from available protein-property information.
Intrinsically disordered proteins (IDPs), unlike globular proteins, lack stable secondary structure and exist as dynamic ensembles of conformations in physiological conditions. These conformations allow them to adapt to many roles while interacting with other proteins, forming partially folded soluble oligomers or insoluble plaques rich in β -sheet, and are responsible for different pathological diseases. The dynamic nature of IDPs and the lack of well-defined binding sites make drug discovery and development both challenging. Although computational protocols, predominantly involving molecular dynamics (MD) simulations studies, have been complementing in understanding the structure and dynamics of IDPs, the choice of force fields become important in reproducing experimental observations. We herewith provide a systematic study to investigate the structural propensity of four IDPs (A β 42, Tau43, amylin and α S), using 13 combinations of force field-water models (FF-wm). In addition to validating with NMR observables, we also examine the water dynamics surrounding the chosen IDPs to underscore the generalizability of chosen FF-wms. C36-IDPSFF with amber99sb-dispersion water model (A99SB-disp) is found to perform optimal in simulating IDPs in comparison with other FF-wm combination based on structural propensity and comparison with experimental NMR 3 J HN−Hα -Coupling data as well as dynamical observables. ### Competing Interest Statement The authors have declared no competing interest. DST-SERB, CRG/2021/008036 Ihub-Data, NA
In recent years, the rapid advancement of generative artificial intelligence (GenAI) has revolutionized the landscape of drug design, offering innovative solutions to potentially expedite the discovery of novel therapeutics. GenAI encompasses algorithms and models that autonomously create new data, including text, images, and molecules, often mirroring characteristics of existing datasets. This comprehensive review delves into the realm of GenAI for drug design, emphasizing recent advancements and methodologies that have propelled the field forward. Specifically, we focus on three prominent paradigms: transformers, diffusion models, and reinforcement learning algorithms, which have been exceptionally impactful in the last few years. By synthesizing insights from a myriad of studies and developments, we elucidate the potential of these approaches in accelerating the drug discovery process. Through a detailed analysis, we explore the current state and future directions of GenAI in the context of drug design, highlighting its transformative impact on pharmaceutical research and development.
Computing binding affinities is of great importance in drug discovery pipeline and its prediction using advanced machine learning methods still remains a major challenge as the existing datasets and models do not consider the dynamic features of protein-ligand interactions. To this end, we have developed PLAS-20k dataset, an extension of previously developed PLAS-5k, with 97,500 independent simulations on a total of 19,500 different protein-ligand complexes. Our results show good correlation with the available experimental values, performing better than docking scores. This holds true even for a subset of ligands that follows Lipinski’s rule, and for diverse clusters of complex structures, thereby highlighting the importance of PLAS-20k dataset in developing new ML models. Along with this, our dataset is also beneficial in classifying strong and weak binders compared to docking. Further, OnionNet model has been retrained on PLAS-20k dataset and is provided as a baseline for the prediction of binding affinities. We believe that large-scale MD-based datasets along with trajectories will form new synergy, paving the way for accelerating drug discovery.
A major difference between amyloid precursor protein (APP) isoforms (APP695 and APP751) is the existence of a Kunitz type protease inhibitor (KPI) domain which has a significant impact on the homo- and hetero-dimerization of APP isoforms. However, the exact molecular mechanisms of dimer formation remain elusive. To characterize the role of the KPI domain in APP dimerization, we performed a single molecule pull down (SiMPull) assay where homo-dimerization between tethered APP molecules and soluble APP molecules was highly preferred regardless of the type of APP isoforms, while hetero-dimerization between tethered APP751 molecules and soluble APP695 molecules was limited. We further investigated the domain level APP-APP interactions using coarse-grained models with the Martini force field. Though the model initial ternary complexes (KPI-E1, KPI-KPI, KPI-E2, E1-E1, E2-E2, and E1-E2) generated using HADDOCK (HD) and AlphaFold2 (AF2), the binding free energy profiles and the binding affinities of the domain combinations were investigated via the umbrella sampling with Martini force field. Additionally, membrane-bound microenvironments at the domain level were modeled. As a result, it was revealed that the KPI domain has a stronger attractive interaction with itself than the E1 and E2 domains, as reported elsewhere. Thus, the KPI domain of APP751 may form additional attractive interactions with E1, E2 and the KPI domain itself, whereas it is absent in APP695. In conclusion, we found that the APP751 homo-dimer formation is predominant than the homodimerization in APP695, which is facilitated by the presence of the KPI domain.
Recent studies in Alzheimer's disease (AD) investigated the precise mechanisms responsible for neurofibrillary tangles (NFT) and senile plaques formation. NFTs, the aggregated Tau protein isoforms, are one of the primary factors behind AD. The corresponding smallest variant Tau43, as paired helical filaments (PHFs), self‐assemble into pathological diseased aggregates. However, the molecular details in rationalizing the aggregation propensity of Tau43 remain elusive. Herein, using molecular dynamics simulations on aqueous Tau43, we identify the molecular factors responsible for early behavior of Tau43 aggregation propensity in water. The variant is intrinsically unstructured yet compact in nature, in agreement with previous studies. The PHF6 (11VQIVYK16) segment is relatively less fluctuating, yet most extended and hydrophobic, thereby shielded from aqueous environment by intermolecular polar interactions between terminal residues. We also compared structural propensities of Tau43 and oppositely charged Aβ42 peptide, providing a comparative understanding of early AD pathway, leading to corresponding drug designing avenues.
Amyloid β (Aβ) senile plaques and Tau neurofibrillary tangles (NFTs) are major hallmarks of Alzheimer's disease (AD). However, early stages of Tau aggregation are still limitedly recognized. Here, we present atomistic molecular dynamics simulations and thermodynamics characterizations of heterogeneous Tau43‐Aβ42 and homogeneous Tau43‐Tau43 dimerization processes. Two‐stage approaching‐accommodation mechanism after individual diffusive regime is observed. The approach step involves opposing forces driving two distant monomers to come closer to each other, which are the decrease in protein internal and water‐induced energies, respectively. In the accommodation step, a decrease in protein internal energy is the main driving force for stable compact structure formation. While the charged residues differently initiate and stabilize the dimer structures, the hydrophobic residues (11VQIVYK16 in Tau43 and 39VVIA42 in Aβ42) facilitate the formation of compact dimers, in agreement with experiments. Our results of Tau43‐Aβ42 and Tau43‐Tau43 dimerization will illuminate early onset mechanisms of AD pathology and corresponding therapeutic initiatives.
The advent of nanotechnology has seen a growing interest in the nature of fluid flow and transport under nanoconfinement. The present study leverages fully atomistic molecular dynamics (MD) simulations to study the effect of nanochannel length and intrusion of molecules of the organic solvent, hexafluoro-2-propanol (HFIP), on the dynamical characteristics of water within it. Favorable interactions of HFIP with the nanochannels comprised of single-walled carbon nanotubes traps them over time scales greater than 100 ns, and confinement confers small but distinguishable spatial redistribution between neighboring HFIP pairs. Water molecules within the nanochannels show clear signatures of dynamical slowdown relative to bulk water even for pure systems. The presence of HFIP causes further rotational and translational slowdown in waters when the nanochannel dimension falls below a critical length of 30 angstrom. The enhanced slowdown in the presence of HFIP is quantified from characteristic relaxation parameters and diffusion coefficients in the absence and presence of HFIP. It is finally seen that the net flow of water between the ends of the nanochannel shows a decreasing dependence with nanochannel length only when the number of HFIP molecules is small. These results lend insights into devising ways of modulating solvent properties within nanochannels with cosolvent impurities.
The investigation of intrinsically disordered proteins (IDPs) is a new frontier in structural and molecular biology that requires a new paradigm to connect structural disorder to function. Molecular dynamics simulations and statistical thermodynamics potentially offer ideal tools for atomic-level characterizations and thermodynamic descriptions of this fascinating class of proteins that will complement experimental studies. However, IDPs display sensitivity to inaccuracies in the underlying molecular mechanics force fields. Thus, achieving an accurate structural characterization of IDPs via simulations is a challenge. It is also daunting to perform a configuration-space integration over heterogeneous structural ensembles sampled by IDPs to extract, in particular, protein configurational entropy. In this review, we summarize recent efforts devoted to the development of force fields and the critical evaluations of their performance when applied to IDPs. We also survey recent advances in computational methods for protein configurational entropy that aim to provide a thermodynamic link between structural disorder and protein activity.
We investigate, using atomistic molecular dynamics simulations, the association of surface hydration accompanying local unfolding in the mesophilic protein Yfh1 under a series of thermal conditions spanning its cold and heat denaturation temperatures. The results are benchmarked against the thermally stable protein, Ubq, and behavior at the maximum stability temperature. Local unfolding in Yfh1, predominantly in the beta sheet regions, is in qualitative agreement with recent solution NMR studies; the corresponding Ubq unfolding is not observed. Interestingly, all domains, except for the beta sheet domains of Yfh1, show increased effective surface hydrophobicity with increase in temperature, as reflected by the density fluctuations of the hydration layer. Velocity autocorrelation functions (VACF) of oxygen atoms of water within the hydration layers and the corresponding vibrational density of states (VDOS) are used to characterize alteration in dynamical behavior accompanying the temperature dependent local unfolding. Enhanced caging effects accompanying transverse oscillations of the water molecules are found to occur with the increase in temperature preferentially for the beta sheet domains of Yfh1. Helical domains of both proteins exhibit similar trends in VDOS with changes in temperature. This work demonstrates the existence of key signatures of the local onset of protein thermal denaturation in solvent dynamical behavior.
•Serine protease from Nocardiopsis sp. cloned and expressed.•pET-39b (+) served as better vector than pET-22b (+).•Homology model suggested the protein to be a member of PA clan.•Unfolding cooperativity was studied by novel computational approach.•Docking with substrate showed important binding interactions.
The mechanism of cold denaturation in proteins is often incompletely understood due to limitations in accessing the denatured states at extremely low temperatures. Using atomistic molecular dynamics simulations, we have compared early (nanosecond timescale) structural and solvation properties of yeast frataxin (Yfh1) at its temperature of maximum stability, 292 K (Ts), and the experimentally observed temperature of complete unfolding, 268 K (Tc). Within the simulated timescales, discernible “global” level structural loss at Tc is correlated with a distinct increase in surface hydration. However, the hydration and the unfolding events do not occur uniformly over the entire protein surface, but are sensitive to local structural propensity and hydrophobicity. Calculated infrared absorption spectra in the amide-I region of the whole protein show a distinct red shift at Tc in comparison to Ts. Domain specific calculations of IR spectra indicate that the red shift primarily arises from the beta strands. This is commensurate with a marked increase in solvent accessible surface area per residue for the beta-sheets at Tc. Detailed analyses of structure and dynamics of hydration water around the hydrophobic residues of the beta-sheets show a more bulk water like behavior at Tc due to preferential disruption of the hydrophobic effects around these domains. Our results indicate that in this protein, the surface exposed beta-sheet domains are more susceptible to cold denaturing conditions, in qualitative agreement with solution NMR experimental results.
Self-assembly of the intrinsically unstructured proteins, amyloid beta (Aβ) and alpha synclein (αSyn), are associated with Alzheimer's Disease, and Parkinson's and Lewy Body Diseases, respectively. Importantly, pathological overlaps between these neurodegenerative diseases, and the possibilities of interactions between Aβ and αSyn in biological milieu emerge from several recent clinical reports and in vitro studies. Nevertheless, there are very few molecular level studies that have probed the nature of spontaneous interactions between these two sequentially dissimilar proteins and key characteristics of the resulting cross complexes. In this study, we have used atomistic molecular dynamics simulations to probe the possibility of cross dimerization between αSyn1-95 and Aβ1-42, and thereby gain insights into their plausible early assembly pathways in aqueous environment. Our analyses indicate a strong probability of association between the two sequences, with inter-protein attractive electrostatic interactions playing dominant roles. Principal component analysis revealed significant heterogeneity in the strength and nature of the associations in the key interaction modes. In most, the interactions of repeating Lys residues, mainly in the imperfect repeats 'KTKEGV' present in αSyn1-95 were found to be essential for cross interactions and formation of inter-protein salt bridges. Additionally, a hydrophobicity driven interaction mode devoid of salt bridges, where the non-amyloid component (NAC) region of αSyn1-95 came in contact with the hydrophobic core of Aβ1-42 was observed. The existence of such hetero complexes, and therefore hetero assembly pathways may lead to polymorphic aggregates with variations in pathological attributes. Our results provide a perspective on development of therapeutic strategies for preventing pathogenic interactions between these proteins.
Atomistic molecular dynamics simulation has been used to probe the effect of the A30P mutation on the structural dynamics of micelle-bound, helical αSynuclein when released in an aqueous environment. On the timescales simulated, the effect of the mutation on the secondary structure is restricted to local changes close to the mutation site in the N-terminal helical domain. The changes are transient, and all residues except Lys23 recover their initial structure. The local behavior due to the mutation gives rise to a global difference in the A30P mutant in the form of a permanent kink in the N-terminal helical domain.