
We studied the machine learning models for high-refractive organic compounds, based on literature data. Although, the prediction data by the model were well toward common organic compounds, that of the high-refractive compounds were not accord with experimental data, due to the few literature data of high-refractive compounds. Therefore, we created data for 62 high-refractive compounds using high-throughput quantum chemical calculations and performed the transfer learning based on the literature data model. The obtained transfer learning model showed well results for selected high-refractive organic compounds.
Epoxy resins are one of the functional materials used as paints, adhesives, and the like. The hardeners used in resin synthesis have a great impact on resins in terms of reactivity and physical properties, so the selection of hardener is very important. However, once curing reaction starts, it is difficult to experimentally analyze that reactivity and the like. In this study, we investigated the reaction mechanism of the curing reaction of epoxy-imidazole resin with imidazole as a hardener using density functional theory calculations. From the calculation results, it was clarified that the epoxy-imidazole curing reaction proceeds through a five-step reaction pathway with the reaction substrate and the imidazole formed during reaction the reaction serving as nucleophilic species. The active species in this reaction is the imidazole anion, and by generating that, the ring-opening reaction of the epoxide with low activation free energy proceeds repeatedly
Chemical graphs are utilized to predict various physical properties of molecules. A molecule can be represented as an undirected, labeled graph in which atoms are nodes and bonds are the edges of the graph. In this paper, we defined two indexes for a molecule called HG1 and HG2 based on the variance of elements in the eigenvector corresponding to the highest eigenvalue of the adjacency matrix of the molecular graph. We calculated and examined HG1 and HG2 of a huge number of natural products listed in KNApSAcK database. Heterogeneities of molecules can be assessed based on HG1 and HG2 but HG1 is more suitable for this purpose.
Terpenoids, phenylpropanoids, and polyketides are the majority of the secondary metabolites containing carbon, hydrogen, and oxygen. In this work, 19,769 metabolites accumulated in KNApSAcK Core DB were classified into 71 subgroups comprising three major groups (terpenoids, phenylpropanoids, and polyketides) according to scientific literatures. We represented the metabolites as molecular fingerprint including chemical properties, and used those descriptors for classification by random forest model. We found that both training and test metabolites were well classified into the subgroups, with 94.06 %, and 94.23 % accuracy, respectively. Though classification of metabolites based on metabolic pathways is very time-consuming works, machine learnings with molecular fingerprint made it possible to attain the classification. This work will lead a light for systematical and evolutional understanding of diverged secondary metabolites based on secondary metabolic pathways. Data science is an interdisciplinary and applied field that uses techniques and theories drawn from statistics, mathematics, computer science, and information science. Combining these resources data science enables extracting meaningful and practical insights for secondary metabolites.
Deep generative models can virtually generate chemical structures with desired properties. These models are widely used in de -novo molecular design projects, and are becoming an alternative to conventional approaches to chemical structure generation. Although the usefulness of the generative models has already been proven in retrospective validations: using an already known data set, deployment of the generative models in applications has not been frequently reported. Herein, several research articles are surveyed where deep generative models are employed in de novo molecular design projects to clarify the usage of the generative models for successful de novo design.
Proteolytic cleavage is influenced by the physicochemical properties of amino acids surrounding the cleavage site. Among these properties are 553 amino acid indices, and we considered that combining these indices with machine learning could create QSAR models for protease activity. In this study, we focused on gamma-secretase, an enzyme known to be involved in the pathogenesis of Alzheimer's disease. We created 10,680 regression models for the protease activity of gamma-secretase by using 10 amino acid indices compressed from the 553 amino acid indices through principal component analysis, 12 pocket models of protease binding sites, and 89 machine learning models. We used these regression models to predict cleavage sites for 23 substrates where the cleavage sites were known and examined the amino acid property information used in the model with the highest prediction accuracy (87.0%). We found that the amino acid property information used in this model was related to the secondary structure of proteins, which may imply that it contains important information on the transmembrane cleavage of gamma-secretase.
In recent years, competition in organic photovoltaic cells (OPVs) performance improvement and organic semiconductor development has intensified. In response, there has been an upsurge in the development of predictive models for OPV performance utilizing machine learning. Until now, chemistry researchers have used various approaches when creating OPV cells as well as developing new materials to improve power conversion efficiency (PCE). However, not many of those original approaches have been used for performance prediction due to the small sample size. In this study, we conducted Data-science approach where we collected information from 115 scientific literatures and constructed a dataset with the addition of some new proposed variables to describe the structure and material composition of the active layer. This allows us to use 25 variables to describe OPVs in which the active layer forms a 1 similar to 3-level structure (1-layer, twotiered and three-tiered). Proposed work also includes post-processing and measurement data that have not been addressed in existing studies. Several regression models were constructed with coefficients of determination exceeding 0.9 by supervised learning methods (random forest (RF), monmlp, etc.) using this data.
Solvatochromism of 4-(diethylamino)-4’-nitroazobenzene was observed by visible absorption spectroscopy, and mechanism of the shift of absorption maxima was explained on the basis of semi-empirical molecular orbital calculation results. The wavelength of the absorption maximum measured for methanol solution was longer than that measured for cyclohexane solution; solvatochromic color change was observed. Dipole moments of the ground state and excited state of the azo compound calculated by the CNDO/S method indicated that the excited state was more polar than the ground state. This polarization at the excited state corresponds to so-called charge transfer excitation, and the polarization results in lowering of excitation energy in polar solvent. Similar tendency was also confirmed by the density functional theory calculations. The experiment and molecular orbital calculation carried out in this work were suitable for promoting the understanding of the solvatochromic shift mechanism of absorption wavelength.
The development of synthetic routes for functional chemicals has been heavily depending on experience and intuition of synthetic organic chemists. In case that target molecules have complex structures, there are many possible synthetic routes, and it is often difficult to determine which one should be adopted. In order to decrease synthesis routes for experiments, we introduced “in silico screening” which requires to search TSs for synthesis routes, we have proposed a method to locate the new TS structure of a target reaction by using TS structures in TSDB. However, this method seldom gives the most stable TS structure within possible conformers. That is, the stability of transition states (TS), reactants and products is highly dependent on initial structures used for optimization. Therefore, this method is likely to give inadequate data to compare calculated and measured values of other synthetic reactions. For these purposes, we have to find reaction mechanisms with the most stable TS and molecules involved in the reactions. In this paper, we proposed a method to search the most stable reaction pathway and applied it to the Pinner Pyrimidine reaction of ethyl 3-oxobutanoate and 3-ethoxypropanimidamide.
The extreme gradient boosting regression (XGBR) method was applied to 245 phenol derivatives in order to establish a regression model to predict their toxicity to Tetrahymena pyriformis. The modeling was done by using the set of "electronic-structure informatics" (ESI) descriptors recently suggested by the present authors. It is shown that the XGBR method is successful in predicting the toxicity of phenols in each class of five modes of action (MOA). A feature importance analysis showed that the different ESI descriptors were found to be important depending on the MOA. Through comparisons with the optimized descriptor set previously suggested by M. T. D. Cronin et al. (Chemosphere., 49, 1201-1221, (2002)), it is shown that the ESI descriptor set, which has been applied to different types of target variables, is of similar quality in regression modeling.
In parallel to developments in Next-Generation Sequencing for cancer patient therapy decision making, personalized approaches to chemotherapy selection are also becoming desired. In an ideal situation, an individual's genomic, transcriptomic, and tumor-specific in-vitro response to chemical perturbation would be combined, and the US National Cancer Institute NCI-60 project has systematically screened a large chemical library against a variety of cell lines from various tumor types. Therefore, chemoinformatics approaches to make effective use of this data and identify the chemical and biological factors are of value. In this work, we investigate the impact of both chemical and biological descriptions of tumor response to chemical inhibition, and assess how well modeling approaches can predict tumor inhibition response on external datasets. We find that external datasets in both the classification and regression problems are reasonably well addressed, with the impact of chemical description outweighing the contribution from transcriptome or genome descriptions of tumors.
The in sillico method to predict the ecotoxicity of chemical substances for reducing animal testing has become attracted attention. A most common model for predicting ecotoxicity classify chemical substances empirically based on functional groups and then predict ecotoxicity with a linear regression by using a descriptor of a chemical substance such as Log Kow. But the conventional method outputs duplicate result for one chemical substance when it has multiple functional groups. Moreover, this method is not appropriate for predicting the ecotoxicity of metal compounds. To overcome these challenges, this study developed a new fingerprint as a feature set for machine learning, and a new prediction model with supervised machine learning for chronic ecotoxicity on fish. The new fingerprint extracts feature of a chemical substance by judging the existence of the structure contributing to ecotoxicities such as carbamate insecticide, organophosphorus pesticides, organic halogen, various metal elements, and hexavalent chromium. Moreover, we compared the accuracy for predicting chronic ecotoxicity on fish with various machine learning models by 10-fold cross-validation using this new fingerprint, general fingerprints, and descriptors together as a feature set. As a result, our developed method with the stacking ensemble was the most accurate in this study. This method improved accuracy by using the result of multiple machine learning algorithms as a part of a feature set. The result of the benchmark test show that the prediction accuracy of this method was better than conventional methods.
Multi-target activity (promiscuity) of small molecules provides the basis of drug polypharmacology.Computationally, promiscuity can be explored through systematic analysis of compound activity data.Inhibitors of the human kinome represent an instructive example.
Computer-assisted de novo drug design has been a central research topic in the field of chemoinformatics for approximately 30 years. Professor Kimito Funatsu's research has been a formative component in these developments. His seminal work has contributed inverse quantitative-structure-activity relationship (QSAR) models for small molecule and peptide design. This article highlights a class of recurrent neural networks, so-called long short-term memory (LSTM) networks for generative molecular design, which further the conceptual approach of inverse QSAR. We review the LSTM method for molecular design along with selected practical applications.
Fatty acid synthase (FASN) inhibitors are known to work as anti-cancer drugs. In order to find important factors in their structure-activity relationships and to derive a predictive model for the activity, we herein tried to develop regression models by using descriptors representing chemical reactivities and intermolecular interactions. By employing the descriptors calculated with the electronic-structure theory, regression models for the experimental IC50 values were derived. Good correlations between the predicted and experimental values were obtained for the natural products having inhibitor activity to FASN. The obtained models are expected useful for systematic search for more efficient inhibitors. At the same time, the present results justify the use of the newly suggested descriptors evaluated in electronic-structure calculations.
In polymer material development, we often need to optimize some physical and chemical properties simultaneously. On the other hand, there is no established method to predict some different properties of polymers by the same approach. In this study, property values of various polymers were collected from the literature. Their relevance was considered by hierarchical clustering. PLSR models were constructed which predicted density, glass transition temperature, and dissolution parameter using descriptors obtained from the monomer unit structure information. R2 of the models were 0.88 ~ 0.97. The concept of informatics has shown the possibility to predict different polymer properties in a similar way.
Finding direct correlations between electronic structures of molecules and their properties, which we call “electronic-structure informatics”, is one of the challenging issues in chemoinformatics because the electronic degree of freedom is an essential factor determining the chemical characteristics. Herein we develop computational methods to automatically draw two types of orbital correlation diagrams. They are expected useful to perform machine learning including electronic degrees of freedom. In the present approach, we focus on electronic similarity called orbital similarity whose score is defined as spatial overlap between two molecular orbitals (MOs) enclosed with their iso-value surfaces. The similarity scores are also used to derive another orbital correlation diagram called “orbital interaction diagram”. This diagram is to relate MOs of a target molecule with those of its fragments. Through applications to benzene derivatives, these diagrams are shown to be reasonable, indicating potential usefulness of the present method in machine learning for quantitative predictions of molecular properties and chemical reactivities.
Information of transition states of similar reactions is the key to locating those of unknown reactions. In order to utilize this feature, we are constructing a database, called QMRDB, which gathers results of quantum mechanical calculations for elementary reactions as well as those for related molecules. Another database (TSDB) stores information of name reactions in organic synthesis. Retrieval results from these databases are used for analyzing reaction mechanisms which have not been experimentally examined. We developed a cloud system managing both the two databases and theoretical calculations. The present paper describes the summary of the TSDB cloud system and how to use it to perform in silico screenings for synthesizing drug candidates.
A history of collaboration between French and Japanese chemoinformatics groups, and Professor Funatsu’s establishment of a Japanese chemoinformatics school, is presented.
Prof. Kimito Funatsu received the Honor Award in Division of Chemoinformatics, the Chemical Society of Japan in the 42th Annual Meeting of Chemoinformatics held on Nov. 28th 2019. The awarding recognizes his significant contributions in the development of the cheminformatics discipline in the world as well as in Japan. His research efforts extend over multiple domains such as (i) system development including elucidation of chemical structures and prediction of organic reactions, (ii) quantitative structure activity relationship (QSAR), (iii) quantitative structure property relationship (QSPR), and (iv) international collaborations in chemoinformatics. In the present review, we focus on chemoinformatics in the world as well as in Japan based on “Special issue dedicating to Honor Award: Prof. Kimito Funatsu”, which consists of five invited papers by the world-famous distinguished foreign researchers, and six papers from domestic researchers. Taking these papers into consideration, we try to discuss the meanings of the Honor Award dedicating to Prof. Kimito Funatsu.