Parkinson’s disease (PD) is a serious neurodegenerative disease. Most of the current treatment can only alleviate symptoms, but not stop the progress of the disease. Therefore, it is crucial to find medicines to completely cure PD. Finding new indications of existing drugs through drug repositioning can not only reduce risk and cost, but also improve research and development efficiently. A drug repurposing method was proposed to identify potential Parkinson’s disease-related drugs based on multi-source data integration and convolutional neural network. Multi-source data were used to construct similarity networks, and topology information were utilized to characterize drugs and PD-associated proteins. Then, diffusion component analysis method was employed to reduce the feature dimension. Finally, a convolutional neural network model was constructed to identify potential associations between existing drugs and LProts (PD-associated proteins). Based on 10-fold cross-validation, the developed method achieved an accuracy of 91.57%, specificity of 87.24%, sensitivity of 95.27%, Matthews correlation coefficient of 0.8304, area under the receiver operating characteristic curve of 0.9731 and area under the precision–recall curve of 0.9727, respectively. Compared with the state-of-the-art approaches, the current method demonstrates superiority in some aspects, such as sensitivity, accuracy, robustness, etc. In addition, some of the predicted potential PD therapeutics through molecular docking further proved that they can exert their efficacy by acting on the known targets of PD, and may be potential PD therapeutic drugs for further experimental research. It is anticipated that the current method may be considered as a powerful tool for drug repurposing and pathological mechanism studies.
As an important biomarker in organisms, miRNA is closely related to various small molecules and diseases. Research on small molecule-miRNA-cancer associations is helpful for the development of cancer treatment drugs and the discovery of pathogenesis. It is very urgent to develop theoretical methods for identifying potential small molecular-miRNA-cancer associations, because experimental approaches are usually time-consuming, laborious, and expensive. To overcome this problem, we developed a new computational method, in which features derived from structure, sequence, and symptoms were utilized to characterize small molecule, miRNA, and cancer, respectively. A feature vector was construct to characterize small molecule-miRNA-cancer association by concatenating these features, and a random forest algorithm was utilized to construct a model for recognizing potential association. Based on the 5-fold cross-validation and benchmark data set, the model achieved an accuracy of 93.20 ± 0.52%, a precision of 93.22 ± 0.51%, a recall of 93.20 ± 0.53%, and an F1-measure of 93.20 ± 0.52%. The areas under the receiver operating characteristic curve and precision recall curve were 0.9873 and 0.9870. The real prediction ability and application performance of the developed method have also been further evaluated and verified through an independent data set test and case study. Some potential small molecules and miRNAs related to cancer have been identified and are worthy of further experimental research. It is anticipated that our model could be regarded as a useful high-throughput virtual screening tool for drug research and development. All source codes can be downloaded from https://github.com/LeeKamlong/Multi-class-SMMCA.
As a kind of post-translational modifications, hydroxylation drew less attention than other modifications, such as phosphorylation and acetylation. However, besides protein stability regulation, it has been found that hydroxylation may affect the activity of proteins. Therefore, it is necessary to better understand the biological processes of hydroxylation. Identification of hydroxylated substrates and their corresponding sites is important for the studies of its molecular mechanism. Fast and convenient computational methods for hydroxylation sites identification are much desired, because experimental approaches are time-consuming and labor-intensive. Here, we present HydLoc (Hydroxylation sites Location), a random forest-based hydroxylation sites predictor for human proteins using sequential information and physicochemical properties. The accuracies of leave-one-out cross-validation on the training dataset are 84.25% and 80.61% for residue proline (P) and lysine (K), respectively. Based on the independent test dataset, it achieved an accuracy of 90.74% and 81.25% for P and K hydroxylation sites prediction, respectively. Meanwhile, the sensitivity values of 96.29% and 75.00% were obtained for residue P and K, which outperforms the existing methods. A user-friendly web server of HydLoc is now available at https://www.gdpu-bioinfolab.com/hydloc/
Identifying drug-disease associations is helpful for not only predicting new drug indications and recognizing lead compounds, but also preventing, diagnosing, treating diseases. Traditional experimental methods are time consuming, laborious and expensive. Therefore, it is urgent to develop computational method for predicting potential drug-disease associations on a large scale. Herein, a novel method was proposed to identify drug-disease associations based on the deep learning technique. Molecular structure and clinical symptom information were used to characterize drugs and diseases. Then, a novel two-dimensional matrix was constructed and mapped to a gray-scale image for representing drug-disease association. Finally, deep convolution neural network was introduced to build model for identifying potential drug-disease associations. The performance of current method was evaluated based on the training set and test set, and accuracies of 89.90 and 86.51% were obtained. Prediction ability for recognizing new drug indications, lead compounds and true drug-disease associations was also investigated and verified by performing various experiments. Additionally, 3,620,516 potential drug-disease associations were identified and some of them were further validated through docking modeling. It is anticipated that the proposed method may be a powerful large scale virtual screening tool for drug research and development. The source code of MATLAB is freely available on request from the authors.
Prediction of disease–gene association based on a deep convolutional neural network.
Increasing evidence indicates that miRNAs play a vital role in biological processes and are closely related to various human diseases. Research on miRNA-disease associations is helpful not only for disease prevention, diagnosis and treatment, but also for new drug identification and lead compound discovery. A novel sequence- and symptom-based random forest algorithm model (Seq-SymRF) was developed to identify potential associations between miRNA and disease. Features derived from sequence information and clinical symptoms were utilized to characterize miRNA and disease, respectively. Moreover, the clustering method by calculating the Euclidean distance was adopted to construct reliable negative samples. Based on the fivefold cross-validation, Seq-SymRF achieved the accuracy of 98.00%, specificity of 99.43%, sensitivity of 96.58%, precision of 99.40% and Matthews correlation coefficient of 0.9604, respectively. The areas under the receiver operating characteristic curve and precision recall curve were 0.9967 and 0.9975, respectively. Additionally, case studies were implemented with leukemia, breast neoplasms and hsa-mir-21. Most of the top-25 predicted disease-related miRNAs (19/25 for leukemia; 20/25 for breast neoplasms) and 15 of top-25 predicted miRNA-related diseases were verified by literature and dbDEMC database. It is anticipated that Seq-SymRF could be regarded as a powerful high-throughput virtual screening tool for drug research and development. All source codes can be downloaded from https://github.com/LeeKamlong/Seq-SymRF .
The establishment of water-saving crop planning is an inevitable choice of the water-saving agriculture for the water-deficiency region in the arid and semiarid Loess Plateau of China and the world. The water-saving crop planning refers to the planting structure that centres the adjustment of the crop's adaptation to water, the optimization of temporal and spatial layout for crops, the local natural resources, marketing resources, human resources and financial input to enable region or basin with limited water resources to achieve the maximum economic, social and ecological benefits of planting industry under certain technology and economy. After the analysis on the research progress of optimization theory, optimization goals, optimization methods of water-saving cultivation structure and macro-control measures, it is pointed out that the main deficiencies of the current research of water-saving cultivation pattern optimization are lacking of a strong theoretical basis, and the immaturity of optimization technologies. The future crucial research direction will focus on five aspects such as the special optimization theory system, the division methods by studying the watershed unit and using 3S technology, optimization model based on multi-objective evolutionary algorithm, evaluation of rationality and macro-control measures on the basis of the public participation.
There is renewed interest in the crucifer Camelina sativa (false flax, camelina, gold of pleasure) as an alternative oilseed crop because of its potential value for food, feed, and industrial applications. This species is adapted to canola-growing areas in many regions of the world and is generally considered to be resistant to many diseases. A review of the literature indicates that C. sativa is highly resistant to alternaria black spot and blackleg of crucifers. Genotypes resistant to sclerotinia stem rot, brown girdling root rot, and downy mildew can be found among C. sativa accessions, raising the possibility of developing cultivars resistant to these diseases. However, C. sativa is susceptible to clubroot, white rust, and aster yellows disease. Until resistant cultivars or effective management practices have been developed, the susceptibility of C. sativa to these diseases will limit the cultivation of the crop in areas where these diseases are prevalent.