
Microtubules usually form groups at high densities and can glide in the same direction on a kinesin-coated glass surface. They tend to move together (snuggling) to avoid collisions and overlaps. Tracking microtubule groups is difficult due to their intrinsically complex and nonlinear behavior, as well as their sudden appearance and disappearance. We developed a microtubule motion analysis workflow incorporating a U-Net-like fully convolutional neural network (FCN) for noise filtering, a template testing method for endpoint detection, Sparse Optical Flow (SOF) for motion detection, SOF clustering for group detection, and cluster matching for group tracking. We fine-tuned the parameters using videos generated by a microtubule gliding simulation system, and then applied this workflow to real experimental videos. This workflow significantly facilitates microtubule motion analysis.
Accurate prediction of protein-ligand binding affinities is a critical step in accelerating drug discovery by reducing experimental costs and development times. Recently developed co-folding AI models predict how multiple biomolecules fold and interact with each other in threedimensional space. The emergence of a new co-folding model, Boltz-2, has made highly accurate and efficient predictions of protein-ligand binding affinities increasingly feasible. However, the generalization and reliability of these models remain unclear due to the absence of standardized and target-wide benchmark datasets. In this study, we constructed an independent external benchmark dataset derived from ChEMBL version 35 to rigorously evaluate Boltz-2's performance for affinity prediction. The dataset includes 356 unique protein targets and 10,933 compounds, carefully collected to ensure no overlap with the Boltz-2 training data. Binding affinity measurements were standardized into pChEMBL values and linked to the compound SMILES and protein UniProt accessions. Using this benchmark dataset, we compared the performance of the original Boltz-2 model and its NVIDIA Inference Microservice (NIM) implementation. The results showed that the original Boltz-2 and NIM achieved comparable and fair predictive performance across targets (mean absolute error of approximately 0.9), while NIM reduced computational time by approximately 60-90%. The error analysis indicated that no clear correlation existed between the prediction errors and sequence or compound novelty relative to the Boltz-2 training data, underscoring the model's broad coverage. This work provides a transparent and reproducible benchmark for evaluating AI-driven affinity prediction models and offers valuable insights into Boltz-2's applicability, limitations, and potential as a practical tool for data-driven drug discovery.
All seven human mitogen-activated protein kinase kinases (MAP2Ks) are essential for cellular processes such as cell proliferation and apoptosis, and dysfunction of MAP2Ks is associated with cancers and autoimmune diseases. 5Z-7-oxozeaenol (5Z7O) strongly inhibits MAP2K1, 2, 3, and 6, and weakly inhibits MAP2K4, 5, and 7. In this study, the potencies of 5Z7O toward MAP2K2, 3, and 5 were explored by computational methods including docking and molecular dynamics (MD) simulations using homology models. Docking simulations showed that the alpha, beta-unsaturated ketone moiety of 5Z7O bound close to the conserved cysteine residue located in front of the DFG motif (DFG-1) in MAP2K2 and 3 as in previous studies of MAP2K1 and 6; in strong inhibition, 5Z7O binds covalently with this cysteine. MD simulations of MAP2K2 and 3 in the apo state showed that the high flexibility of a conserved "gatekeeper" methionine residue enabled access of 5Z7O to this binding site. However, in docking simulation of MAP2K5, the alpha, beta-unsaturated ketone moiety of 5Z7O was far from the conserved DFG-1 cysteine. A threonine residue at the gatekeeper position in MAP2K5 likely prevents access of the ketone moiety to that cysteine. These findings, together with previous data, provide guidance for the development of inhibitors selective for each type of MAP2K.
The fragment molecular orbital (FMO) method enables quantum chemical evaluation of intra-and intermolecular interactions in biomacromolecules at the fragment level. To facilitate data reuse and advance research in structural biology and drug discovery, we have developed the FMO database (FMODB, URL: https://drugdesign.riken.jp/FMODB/), which currently contains 77,277 entries. In this paper, we summarize the key features added to FMODB since 2021, including advanced search capabilities, enhanced interfragment interaction energy (IFIE) analysis tools, a batch IFIE analysis module, a Web API, and cross-links to PDBj. These updates markedly enhance usability and interoperability, thereby enabling more effective application of FMO data.
To elucidate the molecular recognition mechanism between renin and its inhibitor, we analyze intermolecular interaction energies based on fragment molecular orbital (FMO) calculations for twenty different complexes of various inhibitors. We discuss a relationship between the experimental activity value of inhibitors and the calculated binding energy, and clustering analyses for inter-fragment interaction energies (IFIEs) between the inhibitor and the amino acid residues in renin. We estimated the sum of IFIEs as binding energy between an inhibitor and renin, and found that the calculated binding energies have a relatively strong correlation (R-2 = 0.73) with the experimental IC50 values of each inhibitor. The high-activity correlation between the calculated and experimental values can lead to predicting the effects of drugs and the activity value of new compounds. In addition, we carried out a detailed interaction energy analysis between inhibitors and the amino acid residues in renin, and performed clustering of inhibitors not only by their structure/binding mode but also by the characteristics of interaction, such as energy values, energy patterns, and interacting amino acid residues. As a result, we found that the difference in interaction due to a slight difference in structure, such as the addition/replacement of a single atom/functional group, can be related to the difference in IC50 values. Consequently, the inhibitors in the finally classified clusters tend to show the same order of IC50 value. These results indicate that the structure and the activity value of inhibitor are related to each other through the interaction between an inhibitor and relevant amino acid residues, and suggest that it is certainly possible to predict the IC50 values of inhibitors. Therefore, we consider that our FMO-IFIE analysis would be a useful method to contribute toward innovative drug discovery.
Nirmatrelvir is a severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) main protease (Mpro) inhibitor that exerts its antiviral activity by covalently binding to the catalytic cysteine (Cys145) of Mpro; however, the emergence of drug-resistant variants remains an obstacle to successful antiviral therapy. In this study, we established 23 artificial SARS-CoV-2 Mpro mutants by substituting each active-site residue with alanine in the Mpro-nirmatrelvir complex using a computational approach referred to as a virtual alanine scan. Although the methods were primarily used for non-covalent inhibitor complexes, we conducted a virtual alanine scan for a protein-covalent inhibitor complex. Mutants, in which the structural changes of the main chain and the catalytic dyad were minimal, while the ligand configuration was significantly shifted, were considered to potentially confer drug resistance. The analysis revealed 13 residues: Ser1, His41, Tyr54, Phe140, Leu141, Gly143, Ser144, Cys145, Met165, Glu166, Pro168, Gln189, and Gln192 that are important for the recognition of nirmatrelvir by Mpro, and mutations at these residues may result in drug resistance. The ligand shifts observed in the experimentally reported resistant mutant G143S and the artificially mutant G143A were very similar. These results also indicate that a virtual alanine scan can be applied to covalent inhibitors.
The interplay between the fragment molecular orbital method (FMO) and molecular dynamics (MD) simulations is reviewed. Subsequently, opinions and aspirations related to the further enhancement of this interplay are presented, referring to recent advancements in reactive force fields and machine learning MD, with regard to the simulation of enzymatic reactions. Overall, the interplay between FMO and MD represents a promising frontier in the fields of computational chemistry and quantum life science.
Big data and artificial intelligence (AI) are now creating a whirlwind in the general methodology of both medical research and healthcare practice. It fundamentally transforms the entire paradigm of medical field, which could be called "the third revolution of medicine", where the first one was caused by invention of antibiotics, contributed to eradicate the bacterial infection and the second one was brought by the invention of biopharmaceutical, such as molecular-targeted drug and antibody agent, introducing innovative treatment methods for cancer and several incurable diseases. This article discusses the future form of medicine which the third revolution will bring about. As for the methodological innovation of medical research, this revolution will bring about the data-driven approach for medical science, promoting the reverse science, which will be supported by big data and inductive AI. As for the medical practice and healthcare, this revolution will advance the current mobile health and real world medicine, which, incorporating the cutting edge molecular instrumental methods, will realize PM (precision medicine) mobile health and PM real world medicine. Moreover, on account of current progress in natural language processing as seen in the large language model, AI method will expect to develop to understand the interrelationship among the clinical events described in EMR to comprehend disease progression course, which will realize the "predictive control medicine". Those innovation in both medical research and clinical practice will contribute to reduce the disparity of medical cure level of the clinical practice.
Accurate prediction of drug response in cancer treatment remains a critical challenge due to the complex biological interactions underlying tumor sensitivity and resistance. In this work, we introduce OT-GNN, a novel graph neural network framework that leverages optimal transport theory to integrate prior drug-target interaction knowledge with gene expression profiles for interpretable and robust drug response prediction. By embedding an optimal transport-based alignment mechanism into the GNN architecture, OT-GNN dynamically reweights gene importance tailored to each drug-cell line pair, enhancing both predictive accuracy and biological interpretability. We evaluate OT-GNN on a processed NCI-60 dataset under zero-shot learning settings, demonstrating superior performance compared to traditional machine learning models, recent deep learning methods, and standard GNN variants without our proposed alignment. OT-GNN achieves state-of-the-art ROC-AUC and PR-AUC scores, with improved stability across multiple runs, highlighting its potential as a reliable tool for precision oncology applications. Our approach bridges the gap between data-driven modeling and biological prior knowledge, providing a pathway toward more transparent and effective drug response prediction.
The cost and time required for drug discovery have increased, prompting the need for more efficient methods to predict compound properties using artificial intelligence (AI). In particular, the prediction of absorption, distribution, metabolism, and excretion (ADME) properties is crucial. Intrinsic metabolic clearance (CLint) is an essential property in ADME because it affects the side effects or the dosage schedule. In silico machine learning (ML) models for estimating CLint have been developed, but due to the inherent complexity of AI techniques, it is difficult to obtain ideas for a more effective structure. Therefore, we employed SHAP (SHapley Additive exPlanations), an explainable AI (XAI) method, to elucidate the contributions of molecular substructures to CLint predictions. We constructed a random forest model, classified into low and high clearance categories. SHAP values were calculated to visualize the importance of features, and significant substructures influencing CLint were identified. The model demonstrated a high recall for predicting low-CLint compounds; however, its overall accuracy for high-CLint classification was low. This highlights its potential for filtering out low-CLint compounds rather than accurately identifying those with high-CLint. Visualization of SHAP values provided insights into substructure modifications to improve CLint, thus providing valuable guidance for drug optimization. This approach highlights the effectiveness of integrating explainable AI methods to improve the interpretability of ML models in drug discovery.
Medication-induced hiccups, although uncommon, can significantly reduce patients' quality of life. We had previously identified nicotine as a potential trigger of drug-induced hiccups. However, the mechanisms and risk factors, particularly those related to the route of administration, remain unclear. This study used the U.S. Food and Drug Administration (FDA) Adverse Event Reporting System (FAERS) to investigate nicotine-induced hiccups, particularly focusing on the routes of administration and patient factors. We analyzed 14,836,467 cases reported between January 1, 2004, and March 31, 2022, which were downloaded from the FDA website. Of the 198,556 adverse event reports for nicotine, 970 involved hiccups. We performed univariate analyses on routes of administration and drug formulations to determine their influence on the occurrence of hiccups. Furthermore, we analyzed patient information related to nicotine-induced hiccups. Men were more frequently affected by nicotine-induced hiccups than women, with a higher incidence in older patients. Oral nicotine administration via gum and lozenges was more significantly associated with the occurrence of hiccups than other routes. Nicotine-induced hiccups are influenced by the administration route, particularly oral formulations, such as gum and lozenges. These findings indicate the need for further studies to elucidate the mechanisms of nicotine-induced hiccups and to develop preventive strategies.
Relevant exposure routes need to be taken into account when performing chemical risk assessments for humans. Chemical risks are often assessed by performing route-to-route extrapolations based on oral repeated dose toxicity studies if route-specific toxicity data are unavailable. When performing a route-to-route extrapolation, an extrapolation factor derived from differences in absorption after exposure through different routes needs to be estimated. In this study, we used a machine learning (ML) and regression-based approach to estimate extrapolation-factor-like coefficients for oral-to-inhalation extrapolations using chemical structures and physicochemical properties. We used well-reviewed chemicals with human chronic toxicity values for specific administration routes (oral reference dose and inhalation reference concentration) available. ML regression models for predicting inhalation reference concentrations were developed using oral reference doses and molecular features as descriptors. The ML-based regression models gave better predictions than models using only molecular features or even single constant extrapolation factors, suggesting that the ML-based approach offers advantages over other oral-to-inhalation extrapolation methods.
In structure-based drug design, the fragment molecular orbital (FMO) method, which can quantitatively evaluate interaction energy with high accuracy, helps identify critical amino acid residues and types of inter-and intramolecular interactions for molecular recognition of ligands, peptides, and antigens. In this study, we performed FMO calculations using 679 apoprotein structures, classified them according to their characteristic amino acid residue pair groups (e.g., charged, polar, hydrophobic, and aromatic residues), and statistically analyzed the characteristics of their intramolecular interactions by using inter-fragment interaction energy (IFIE). The average total IFIE between amino acid residues was stronger in the order of pairs of oppositely charged residues, pairs of charged and neutral polar residues, pairs of neutral polar residues, and pairs of hydrophobic residues. Focusing on the dispersion interaction energy obtained using energy decomposition, except for the pairs of both charged residues, pairs of amino acid residues that consisted of aromatic rings had the largest attractive interaction energy compared to other residue pair groups. Furthermore, we compared the intramolecular interaction energy between the FMO and the classic molecular mechanics (MM) methods. The latter tended to estimate the attractive and repulsive interaction energies were weaker than those using the FMO method. As for the dispersion force, the van der Waals interaction energy obtained by the MM method tended to be partially repulsive of amino acid residue pairs with relatively short distances. Such statistical energy interaction analyses of amino acid residue pairs not only depend on understanding the intramolecular interactions of proteins, but is also expected to act as the indicators of IFIEs in molecular recognition (e.g., antigen-antibody and peptide drugs). Furthermore, analysis of critical amino acid residues and their interactions will help understand proteins' molecular bonding and structural stability.
Enzyme kinetics is widely used in many biological studies. Most of enzyme kinetic formulae are derived based on the steady state assumption, which is mathematically expressed by a time- differential equation, namely, d(an enzyme state)/dt = 0 in conventional textbooks. In this opinion article, I would like to propose flux method for handling the steady state assumption.In the proposed flux method, a steady state is defined as the condition wherein the flux entering to the state is equal to the flux leaving that state, implying that the traffic of materials passing through the state is nonzero. How the concept of flux, a type of flow rate, is useful for understanding the nature of enzyme reactions and enzyme kinetic studies will be discussed.
In modeling structures with multiple conformers, the one with the higher occupancy is typically used under implicit understanding. Multiple conformers have rarely been studied to determine whether a structure with a higher ligand occupancy conformer is more stable and has stronger protein-ligand binding than the one with the lower occupancy. We performed fragment molecular orbital (FMO) calculations on complexes with high and low ligand occupancies, comparing their energies, such as total energy and ligand binding energy, to investigate this question. The structures for FMO calculation were the SARS-CoV-2 main protease (M-pro) and the covalent inhibitor, GC376. The initial X-ray crystal structure of the complex (PDB ID: 7CB7) possessed stereo isomers of GC376, namely B1S (R form) and K36 (S form), with different occupancies. Our findings showed that for all structural optimization conditions, the total energy of the high ligand occupancy B1S complex was lower than that of the low ligand occupancy K36 complex. In addition, stronger interfragment interaction energy (IFIE) of B1S than of K36 was confirmed for most structural optimization conditions. These differences between B1S and K36 are caused by the presence and absence of a hydrogen bond between the ligand and His41, as revealed by IFIE and structural geometry analyses between the ligand and each amino acid residue of M-pro. The present study, under several structural optimization conditions, would have the potential to model energetically stable structures by efficiently selecting high occupancy conformers from structures with different ones.
The fragment molecular orbital (FMO) method enables quantum mechanical calculations for macromolecules by dividing the target into fragments. However, most calculations, even for metalloproteins, have been performed by removing metal ions from the structures registered in the Protein Data Bank (PDB). For more realistic and useful calculations, FMO calculations must be performed without removing the metal ions. In this study, we discuss the results obtained from FMO calculations performed using 6-31G* and model core potentials (MCPs) for metal proteins containing Zn and Mg ions. Subsequently, we analyze the differences in atomic charges and interactions.
The cost and time required for drug discovery must be reduced. Recent in silico models have focused on accelerating seed compound discovery based solely on chemical structure. Estimating pharmacokinetic characteristics, including absorption, distribution, metabolism, and excretion (ADME), is essential in the early stage of drug discovery. Therefore, in silico models have used artificial intelligence (AI) techniques to predict the ADME properties of potential compounds. Large experimental data are necessary when constructing in silico models for ADME prediction. However, it remains difficult for one pharmaceutical company or academic laboratory to collect enough data for modeling. Therefore, collecting data from open databases with the assistance of dry scientists is one of the most effective strategies utilized by researchers. However, incorrect values are occasionally included in open databases because of human errors. Furthermore, to construct high-performance ADME in silico models, data curation must include not only chemical structure but also experimental conditions, which requires expert knowledge of pharmacokinetic experiments. Trials to ease the difficulties of data curation have been developed as reported. These tools enable the effective collection and checking of published data. Additionally, they accelerate collaboration between dry and wet scientists, enabling them to collect vast amounts of data to construct high-performance and widespread chemical space ADME in silico models. Collecting much accurate data for constructing ADME in silico models is an expectation of the new era of efficient drug discovery when entirely using AI technology.
Computer scientists have studied artificial intelligence and machine learning since computers were invented. There have been three waves of developments, in this field. First development, in the 1950s, computer chess and reversiTM, were realized. In the second development, “expert systems” attracted attention and started the “5th Generation Computer Project.” in 1980s. However, the results of both the developments were insufficient (for example, 5th Generation Computer was not created), and the enthusiasm for expecting new type of computers such as artificial intelligence quickly diminished. We are now in the third wave of the AI developments. Will it produce satisfactory results? Will it revolutionize life sciences? In this article, I have tried to answer these questions.
The second and subsequent waves of coronavirus disease 2019 (COVID-19) have caused problems worldwide 1. Here, using an objective analytical method 2, we present the changes that occurred in the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2), the causative virus of COVID-19, over time. The virus has mutated in three major directions, resulting in three groups to date. Analysis of the basic structure of the group of viruses was completed by April and shared across all continents. However, the virus continued to mutate independently in each country after the borders were closed. In particular, the virus mutated before the occurrence of the second and subsequent peaks. It seems that the mutations conferred higher infectivity to the virus, because of which the virus overcame previously effective protection and caused second waves of the disease. Currently, each country may possess such a unique, stronger variant. Some of them slowly entered other countries and caused epidemics. These viruses could also serve as sources of further mutations by exchanging parts of the genome, which could create variants with superior infectivity.