Accurate estimation of binding free energy ( Δ Δ G ) remains a critical challenge in drug discovery, directly influencing the development of effective therapeutics. Physics-based computational methods-such as molecular dynamics simulations and implicit solvent modeling-offer rigorous, atomistic insights into molecular recognition but are often constrained by computational costs and limited scalability. In contrast, deep learning models, particularly graph convolutional networks (GCNs), have demonstrated the ability to rapidly predict molecular properties by learning hierarchical representations from large-scale chemical data, yet frequently lack explicit incorporation of physical laws, leading to potential issues with interpretability and generalizability. Hybrid physics-guided neural networks seek to overcome these limitations by embedding physically meaningful features and constraints within deep learning architectures. In this study, we introduce a multi-objective loss function that simultaneously optimizes empirical error, structural similarity, and physical consistency by integrating molecular fingerprints with physics-based features such as electrostatic and van der Waals energies in a unified GCN framework. Applied to 65-75 host-guest systems, this approach yields substantial improvements in entropy ( T Δ S ) prediction accuracy, reducing the mean absolute difference (MAD) from 5.80 to 1.07 kcal/mol, while also enhancing the accuracy and convergence of binding free energy ( Δ Δ G ) predictions (MAD reduced to 1.54 kcal/mol), mitigating overfitting, and improving model transferability to larger complexes. These results demonstrate that principled loss function engineering is pivotal not only for accurate entropy estimation but also for enhancing the reliability and interpretability of binding free energy predictions in molecular machine learning models.
The use of fast in silico prediction methods for protein-ligand binding free energies holds significant promise for the initial phases of drug development. Numerous traditional physics-based models (e.g., implicit solvent models), however, tend to either neglect or heavily approximate entropic contributions to binding due to their computational complexity. Consequently, such methods often yield imprecise assessments of binding strength. Machine learning models provide accurate predictions and can often outperform physics-based models. They, however, are often prone to overfitting, and the interpretation of their results can be difficult. Physics-guided machine learning models combine the consistency of physics-based models with the accuracy of modern data-driven algorithms. This work integrates physics-based model conformational entropies into a graph convolutional network. We introduce a new neural network architecture (a rule-based graph convolutional network) that generates molecular fingerprints according to predefined rules specifically optimized for binding free energy calculations. Our results on 100 small host-guest systems demonstrate significant improvements in convergence and preventing overfitting. We additionally demonstrate the transferability of our proposed hybrid model by training it on the aforementioned host-guest systems and then testing it on six unrelated protein-ligand systems. Our new model shows little difference in training set accuracy compared to a previous model but an order-of-magnitude improvement in test set accuracy. Finally, we show how the results of our hybrid model can be interpreted in a straightforward fashion.
With the investment landscape becoming more competitive, efficiently scaling deal sourcing and improving deal insights have become a dominant strategy for funds. While funds are already spending significant efforts on these two tasks, they cannot be scaled with traditional approaches; hence, there is a surge in automating them. Many third party software providers have emerged recently to address this need with productivity solutions, but they fail due to a lack of personalization for the fund, privacy constraints, and natural limits of software use cases. Therefore, most major funds and many smaller funds have started developing their in-house AI platforms: a game changer for the industry. These platforms grow smarter by direct interactions with the fund and can be used to provide personalized use cases. Recent developments in large language models, e.g. ChatGPT, have provided an opportunity for other funds to also develop their own AI platforms. While not having an AI platform now is not a competitive disadvantage, it will be in two years. Funds require a practical plan and corresponding risk assessments for such AI platforms.
Structure-based drug discovery aims to identify small molecules that can attach to a specific target protein and change its functionality. Recently, deep learning has shown great promise in generating drug-like molecules with specific biochemical features and conditioned with structural features. However, they usually fail to incorporate an essential factor: the underlying physics which guides molecular formation and binding in real-world scenarios. In this work, we describe a physics-guided deep generative model for new ligand discovery, conditioned not only on the binding site but also on physics-based features that describe the binding mechanism between a receptor and a ligand. The proposed hybrid model has been tested on large protein-ligand complexes and small host-guest systems. Using the top-N methodology, on average more than 75% of the generated structures by our hybrid model were stronger binders than the original reference ligand. All of them had higher ΔGbind (affinity) values than the ones generated by the previous state-of-the-art method by an average margin of 1.88 kcal/mol. The visualization of the top-5 ligands generated by the proposed physics-guided model and the reference deep learning model demonstrate more feasible conformations and orientations by the former. The future directions include training and testing the hybrid model on larger datasets, adding more relevant physics-based features, and interpreting the deep learning outcomes from biophysical perspectives.
Molecular fingerprints are essential cheminformatics tools for machine learning with applications in drug discovery. Standard fingerprint software compute fixed-size feature vectors, and employ them as inputs to a deep neural network or other machine learning methods. Fixed-size fingerprint representation of molecules, however, requires extremely large vectors to encode all possible substructures. Limited accuracy and poor interpretation also occur due to the underlying neural network which emphasizes particular and exclusive aspects of the molecular structure. In this study, we develop a novel graph convolutional network (GCN) to predict the binding free energy of protein-ligand complexes. By adding a physics-based layer to the network architecture, the accuracy of the molecular fingerprint has been improved while potential overfitting has been avoided. In addition to standard information about the substuctures, our hyrbid physics-data model encodes atomic bonds features which captures structural features and enables further analysis. It has been shown that machine-optimized fingerprints, compared to fixed-sized fingerprints, can provide more accurate predictions, better performance, and more interpretable results. We show that the proposed GCN fingerprint outperforms the predictive performance of standard fingerprints on binding free energy of host-guest systems and PDBbind database.
AmberTools is a free and open-source collection of programs used to set up, run, and analyze molecular simulations. The newer features contained within AmberTools23 are briefly described in this Application note.
In-silico calculation of binding free energy between protein and ligands has vast applications in the early stages of drug discovery. Most of the classical physics-based models, including implicit solvents, ignore entropy contributions from the system. Instead, a simplified solvent entropy is indirectly considered. This simplification is often done because of an under-sampled conformal space due to physics calculation complexity. Machine learning (ML) methods offer a practical venue to incorporate accurate binding entropy predictions from the experiment. While accurate, there are growing concerns about the over-fitting of ML models to the training set, lack of interpretation due to its “black box” characteristics, and failure to comply with well-known physical models. Recently emerged, physics-guided models are a class of ML models that combine the robust consistency of physics-based models with the accuracy of modern data-driven algorithms. This work presents a method to design two hybrid models by coupling ML with a physics model. Implementing these hybrid models have been done through careful modification of various model learning parameters or hyperparameters. The proposed hybrid models not only outperform purely data-driven models but also show more consistent performance on both training and test sets. We review the basic theory, investigate binding entropy calculation methods, present hybrid models that take advantage of end-point simulation software, and analyze the performance of these models.
This work investigates the non-destructive detection of defects in thermal images of industrial materials based on segmentation of images generated using enhanced truncated-correlation photothermal coherence tomography (eTC-PCT). eTC-PCT is an active infrared thermography modality, which is being applied to the field of non-destructive testing (NDT) and in biomedical & dental thermophotonic imaging. In this report, we combine eTC-PCT with a computer vision algorithm to sharply delineate holes and manufacturing defects (cracks) inside industrial materials. To this end, the eTC-PCT reconstructed image is processed through three consecutive algorithm stages: A threshold selection filter is followed by filtered image segmentation using the K-means algorithm (clustering method) and the outcome is applied to the delineation of (otherwise blurred) discontinuity boundaries by means of the Canny edge detection algorithm. The role of each method is described and it is demonstrated that the combination of these three algorithms is optimal for achieving significant delineation enhancement (sharpness) of blind hole and crack boundaries in industrial materials.
Calculation of protein–ligand binding affinity is a cornerstone of drug discovery. Classic implicit solvent models, which have been widely used to accomplish this task, lack accuracy compared to experimental references. Emerging data-driven models, on the other hand, are often accurate yet not fully interpretable and also likely to be overfitted. In this research, we explore the application of Theory-Guided Data Science in studying protein–ligand binding. A hybrid model is introduced by integrating Graph Convolutional Network (data-driven model) with the GBNSR6 implicit solvent (physics-based model). The proposed physics-data model is tested on a dataset of 368 complexes from the PDBbind refined set and 72 host–guest systems. Results demonstrate that the proposed Physics-Guided Neural Network can successfully improve the “accuracy” of the pure data-driven model. In addition, the “interpretability” and “transferability” of our model have boosted compared to the purely data-driven model. Further analyses include evaluating model robustness and understanding relationships between the physical features.
Calculation of binding affinity of biomolecules is an essential part of drug discovery processes. Mainstream implicit solvent models that are widely used to accomplish this task lack accuracy compared to experiments. Data-driven models, on the other hand, are often accurate yet not fully interpretable and also likely to be overfitted. In this study, we explore the application of “Theory Guided Data Science” in protein-ligand binding prediction. A hybrid model is constructed by combining Graph Convolutional Network (data-driven model) with the GBNSR6 implicit solvent (physics-based model). The proposed physics-data model is tested on a dataset of 72 small and rigid complexes from the host-guest benchmark and SAMPL challenges. Results demonstrate that the Physics-Guided Neural Network was successfully able to improve the accuracy of the physics-based implicit solvent model. In addition, the interpretability and transferability of our hybrid model have been shown to be improved compared to a purely data-driven model. The complete code repository, dataset, and scripts are publicly available at https://github.consaharctech/Binding-Free-Energy-Prediction-Host-Guest-System.
Non-linear mapping is one of the most popular solutions for complex data structures and distinct patterns to cluster data. Auto encoder Networks (AENs) are widely used in clustering as they improve data representation. In this paper, we collect Alexa.com data by crawling popular websites profiles, where dataset has 84 columns with type number and array of words. Next, an AEN architecture is presented to identify specific websites with exceptional patterns and the encoded data expresses new feature space of our original data. (Our) Encoded data is clustered by Affinity Propagation which is a partitioning algorithm without the need for specifying the number of clusters. There are important results based on 194 clusters and exemplars which are filtered and analyzed. One remarkable fact about results is that the first 11 columns as raw data are not clustered by Affinity Propagation a which re considered all as outlier. The results are summarized by selecting the best option w.r.t. the statistics and charts. Some of the obtained results are useful for website owners and provide some suggestions and solutions for Search Engine Optimization (SEO). Finally, we propose our crawler application which crawls and records data over 4 days. It must be added that the proposed web crawler faces challenges and their solutions can be helpful in most commonly used web crawling algorithms and libraries.
In this paper we develop a reliable system for smart irrigation of greenhouses using artificial neural networks, and an IoT architecture. Our solution uses four sensors in different layers of soil to predict future moisture. Using a dataset we collected by running experiments on different soils, we show high performance of neural networks compared to existing alternative method of support vector regression. To reduce the processing power of neural network for the IoT edge devices, we propose using transfer learning. Transfer learning also speeds up training performance with small amount of training data, and allows integrating climate sensors to a pre-trained model, which are the other two challenges of smart irrigation of greenhouses. Our proposed IoT architecture shows a complete solution for smart irrigation.
Ray Luo合作论文数School of Biological Sciences, University of California1