Predicting interactions between drugs and target proteins is a crucial step to decipher many biological processes, and plays a critical role in drug discovery. In this work, we present an Improved Laplacian Regularized Least Square Method (ILRLS) for drug-target interaction prediction. We predict unknown drug-target interactions from chemical structure information, genomic sequence information simultaneously and drug-protein interaction network space. We obtain the better achievement from enzymes, ion channels, nuclear receptors and GPCRs interaction networks. The result indicates that the method could play a complementary role to the existing prediction methods.
G-protein-coupled receptors (GPCRs) play an important role in physiological processes which are the targets of more than 50% of marketed drugs. In this research, we use a hybrid approach of predicted secondary structural features (PSSF) and approximate entropy (ApEn) as the feature selection method for predicting G-protein-coupled receptors in low homology. The low homology dataset is used to validate the proposed method for its objectivity. The classification model based on the fuzzy K-nearest neighbor classifier has been utilized on the classification of membrane proteins data. In order to enhance the prediction accuracies, here we propose an ensemble classifier as the prediction engine. Compared with the previous best-performing method, the success rate is encouraging. The reliable results also demonstrate the proposed method could contribute more to the characterization of various proteomes and further utilized in neuroscience.
In recent years, more than 300 accidents related to the road transport of dangerous goods (RTDG) occurred in China, arousing the attention of the government. There are various causes of accidents: transport vehicles, dangerous cargo, weather, and road infrastructure (such as bridges/tunnels). Therefore, a monitoring system is needed to dynamically examine potential hazards and to regulate the RTDG. In this paper, according to the data resources for road and transportation in China, a three-layer framework for a monitoring system is proposed to dynamically regulate the RTDG. The three layers are information collection, network, and application. Highway infrastructure and information sharing system databases, constructed by the China Transport Department, also provide necessary information. Unlike the monitoring types that only use GPS, the proposed framework involves more information, and full use of databases resources, making it more suitable for dynamically monitoring RTDG in China.
The number of radio frequency identification (RFID) applications in different fields is increasing rapidly. The test technical for RFID application become important research hot point because of some RFID solutions lost. A novel automatic evaluation system for testing RFID applications is designed which has capability to test RFID solution automaticly. At first, the function necessary of the testing system are analyzed, and the framework is proposed. Then, a RFID application about wine anti-counterfeit is evaluated based on proposed testing system. Test results are presented. It indicates that proposed testing system is useful and effective.
Some of the approaches have been developed to retrieve data automatically from one or multiple remote biological data sources. However, most of them require researchers to remain online and wait for returned results. The latter not only requires highly available network connection, but also may cause the network overload. Moreover, so far none of the existing approaches has been designed to address the following problems when retrieving the remote data in a mobile network environment: (1) the resources of mobile devices are limited; (2) network connection is relatively of low quality; and (3) mobile users are not always online. To address the aforementioned problems, we integrate an agent migration approach with a multi-agent system to overcome the high latency or limited bandwidth problem by moving their computations to the required resources or services. More importantly, the approach is fit for the mobile computing environments. Presented in this paper are also the system architecture, the migration strategy, as well as the security authentication of agent migration. As a demonstration, the remote data retrieval from GenBank was used to illustrate the feasibility of the proposed approach.
Protein function prediction with computational method is becoming an important research field in protein science and bioinformatics. In eukaryotic cells, the knowledge of subnuclear localization is essential for understanding the life function of nucleus. In this study, A novel ensemble classifier is designed incorporating three AdaBoost classifiers to predict protein subnuclear localization. The base classifier algorithms in AdaBoost classifier is fuzzy K nearest neighbors (FKNN). Three parts amino acid pair compositions with different spaces are computed to construct features vector for representing a protein sample. Jackknife cross-validation test are used to evaluate performance of proposed with two benchmark datasets. Compared with prior works, promising results obtained indicate that the proposed method is more effective and practical. Current approach may also be used to improve the prediction quality of other protein attributes. The software written in Matlab are available freely by contacting the corresponding author.
Discovering a protein motif is an important research topic in both bioinformatics and protein sciences. This paper presents a novel motif discovery algorithm which is capable of finding a motif set to represent a protein family. The algorithm involves an abstraction method of important features, a location-sensitive connection approach to link two features, and a repeated connection procedure to generate a motif set. The novel algorithm is applied to discovering motifs in 21 ligase subfamilies. The results show that the obtained motifs are able to represent the characteristics of the subfamilies effectively. The proposed algorithm could become a potential useful tool for protein family prediction.
Apoptosis proteins have a central role in the development and the homeostasis of an organism. These proteins are very important for understanding the mechanism of programmed cell death. The function of an apoptosis protein is closely related to its subcellular location. It is crucial to develop powerful tools to predict apoptosis protein locations for rapidly increasing gap between the number of known structural proteins and the number of known sequences in protein databank. In this study, amino acids pair compositions with different spaces are used to construct feature sets for representing sample of protein feature selection approach based on binary particle swarm optimization, which is applied to extract effective feature. Ensemble classifier is used as prediction engine, of which the basic classifier is the fuzzy K-nearest neighbor. Each basic classifier is trained with different feature sets. Two datasets often used in prior works are selected to validate the performance of proposed approach. The results obtained by jackknife test are quite encouraging, indicating that the proposed method might become a potentially useful tool for subcellular location of apoptosis protein, or at least can play a complimentary role to the existing methods in the relevant areas. The supplement information and software written in Matlab are available by contacting the corresponding author.
We use approximate entropy and hydrophobicity patterns to predict G-protein-coupled receptors. Adaboost classifier is adopted as the prediction engine. A low homology dataset is used to validate the proposed method. Compared with the results reported, the successful rate is encouraging. The source code is written by Matlab.
The Vehicle Routing Problem with Simultaneous Pickup and Delivery (VRPSPD) is a variant of the Capacitated Vehicle Routing Problem (CVRP), in which clients require both pickup and delivery services. This paper proposes an improved particle swarm optimization algorithm based on multiple social structures for solving VRPSPD. The decoding of single particle in swarm consists two parts: the first is m dimensional (m-D) variables for m customers, and the second comprises 2n dimensions (2n-D) for n vehicles which presents vehicle route orientation. The particle is transformed to customers’ list and vehicles matrix. A benchmark dataset is used to validate the performance of proposed algorithm. Comparing with prior works, promising results indicate that the proposed algorithm may hold high potentials for generating powerful tool for solving VRPSPD and other attributes of vehicle routing problem.
In order to enhance the efficiency of supply chain, Radio frequency identification (RFID) technology is used to label every cargo. Logistics center (warehouse) is the important node in supply chain. The RFID application in warehouse is complicated. The environment influence and multi-tags collision would make the reader lost targets. It is desired to raise the detect accuracy and detect speed of RFID reader. In this study, we propose multiple antenna beam former algorithm based on SDMA (Space Division Multiple Access) theory to get rid of the multi-tags interference in warehouse with RFID technology. The simulation results show that proposed algorithm can enhance the detect accuracy of RFID reader. It will become an effective tool in warehouse management. Further more, the RFID reader using multiple antenna beam former algorithm will save energy.
G-protein-coupled receptors (GPCRs), the largest family of cell surface receptors play an important role in production of therapeutic drugs. However, the functions of many of GPCRs are unknown. Hence we develop an new method for classifying the family of GPCRs. It is difficult to predict the classification of GPCRs by means of conventional sequence alignment approaches because of their highly divergent nature. In this study, based on the concept of pseudo amino acid composition (PseAA), approximate entropy (ApEn) of protein sequence as additional characteristics is used to construct PseAA. A 21-D (dimensional) PseAA is formulated to represent the sample of a protein. Fuzzy K nearest neighbors (FKNN) classifier is applied as prediction engine. The datasets in low homology are used to validate the performance of the proposed method. Compared to others' research by now, the prediction accuracies of our research is the highest. The test results indicate that ApEn can play a complimentary role to many of the existing methods, which will be a useful tool for GPCRs function prediction.
It was crucial to develop powerful tools to predict apoptosis protein locations for rapidly increasing gap between the number of known structural proteins and the number of known sequences in protein databank.In this study,based on the concept of pseudo-amino acid(PseAA) composition originally introduced by Chou,novel approximate entropy(ApEn) based PseAA composition was proposed to represent apoptosis protein sequences.An integration classifier was introduced,of which the basic classifier was the FKNN(fuzzy K-nearest neighbor) one,as prediction engine.Each basic classifier was trained in different dimensions of PseAA composition of protein sequences.The immune genetic algorithm(ICA) was used to search the optimal weight factors in generating the PseAA composition for crucial of weight factors in PseAA composition.The results obtained by jackknife test were quite encouraging,indicating that the proposed method might become a potentially useful tool for protein function,or at least can play a complimentary role to the existing methods in relevant areas.
Prediction of protein secondary structure is somewhat reminiscent of the efforts by many previous investigators but yet still worthy of revisiting it owing to its importance in protein science. Several studies indicate that the knowledge of protein structural classes can provide useful information towards the determination of protein secondary structure. Particularly, the performance of prediction algorithms developed recently have been improved rapidly by incorporating homologous multiple sequences alignment information. Unfortunately, this kind of information is not available for a significant amount of proteins. In view of this, it is necessary to develop the method based on the query protein sequence alone, the so-called single-sequence method. Here, we propose a novel single-sequence approach which is featured by that various kinds of contextual information are taken into account, and that a maximum entropy model classifier is used as the prediction engine. As a demonstration, cross-validation tests have been performed by the new method on datasets containing proteins from different structural classes, and the results thus obtained are quite promising, indicating that the new method may become an useful tool in protein science or at least play a complementary role to the existing protein secondary structure prediction methods.
The knowledge of subnuclear localization in eukaryotic cells is essential for understanding the life function of nucleus. Developing prediction methods and tools for proteins subnuclear localization become important research fields in protein science for special characteristics in cell nuclear. In this study, a novel approach has been proposed to predict protein subnuclear localization. Sample of protein is represented by Pseudo Amino Acid (PseAA) composition based on approximate entropy (ApEn) concept, which reflects the complexity of time series. A novel ensemble classifier is designed incorporating three AdaBoost classifiers. The base classifier algorithms in three AdaBoost are decision stumps, fuzzy K nearest neighbors classifier, and radial basis-support vector machines, respectively. Different PseAA compositions are used as input data of different AdaBoost classifier in ensemble. Genetic algorithm is used to optimize the dimension and weight factor of PseAA composition. Two datasets often used in published works are used to validate the performance of the proposed approach. The obtained results of Jackknife cross-validation test are higher and more balance than them of other methods on same datasets. The promising results indicate that the proposed approach is effective and practical. It might become a useful tool in protein subnuclear localization. The software in Matlab and supplementary materials are available freely by contacting the corresponding author.
The knowledge of protein subnuclear locations in eukaryotic cell provides strongly help for annotation of protein function.The gap between the number of known function proteins and the number of known sequence in protein databank is increasing rapidly.Prediction of protein subnuclear locations becomes an important research hot point in protein science.A novel approach based on pseudo amino acid composition(PseAA) was proposed to predict protein subnuclear localization.According to the concept of PseAA originally introduced by Chou,a novel pseudo amino acid(PseAA) composition based on the concept of approximate entropy(ApEn) was pressented.The AdaBoost classifier was used as prediction engine.The quite encouraging obtained results indicate that the current approach is effective and might become potential tools in this area and other protein attributes.
Prediction of protein structural classes is an important search field in bioinformatics and protein science. According to the concept of Pseudo Amino Acid (PseAA) composition originally introduced by Chou, a novel PseAA composition was proposed from protein sequence to represent the sample of protein. Combined with amino acid composition (AAC), ten frequencies of various amino acids combination (10-D) and hydrophobic pattern (6-D), the 36-dimensional (36-D) PseAA was constructed. Fuzzy support vector machine was used as prediction engine. Three benchmark dataset were adopted to validate the performance of the proposed approach. The weight factors, which were crucial in PseAA, were optimized by genetic algorithm. The prediction results by jackknife test denote that the proposed approach is effective. It might be potential tool for prediction of protein function.
Secondary structure prediction plays an important role in function prediction of protein. In this paper, maximum entropy model is used to predict protein secondary structure. We build feature function sets based on the influential factors which are crucial to the states of secondary structure of residues in protein sequence. Multi-factors are taken into account in the model, including charge of amino acids, conformational parameter for the states of secondary structure, short and long ranges of interaction of residues in sequence. As such, multi-source information is integrated into a single probability model by the method. Compared with the reported methods, our method gets a higher accuracy rate in predicting protein secondary structure. The results demonstrate that the proposed method is practical.
It is crucial to develop powerful tools to predict apoptosis protein locations for rapidly increasing gap between the number of known structural proteins and the number of known sequences in protein databank. In this study, based on the concept of pseudo amino acid (PseAA) composition originally introduced by Chou, a novel approximate entropy (ApEn) based PseAA composition is proposed to represent apoptosis protein sequences. An ensemble classifier is introduced, of which the basic classifier is the FKNN (fuzzy K-nearest neighbor) one, as prediction engine. Each basic classifier is trained in different dimensions of PseAA composition of protein sequences. The immune genetic algorithm (IGA) is used to search the optimal weight factors in generating the PseAA composition for crucial of weight factors in PseAA composition. The results obtained by Jackknife test are quite encouraging, indicating that the proposed method might become a potentially useful tool for protein function, or at least can play a complimentary role to the existing methods in the relevant areas.
The function of protein is closely correlated with it subcellular location. Prediction of subcellular location of apoptosis proteins is an important research area in post-genetic era because the knowledge of apoptosis proteins is useful to understand the mechanism of programmed cell death. Compared with the conventional amino acid composition (AAC), the Pseudo Amino Acid composition (PseAA) as originally introduced by Chou can incorporate much more information of a protein sequence so as to remarkably enhance the power of using a discrete model to predict various attributes of a protein. In this study, a novel approach is presented to predict apoptosis protein solely from sequence based on the concept of Chou's PseAA composition. The concept of approximate entropy (ApEn), which is a parameter denoting complexity of time series, is used to construct PseAA composition as additional features. Fuzzy K-nearest neighbor (FKNN) classifier is selected as prediction engine. Particle swarm optimization (PSO) algorithm is adopted for optimizing the weight factors which are important in PseAA composition. Two datasets are used to validate the performance of the proposed approach, which incorporate six subcellular location and four subcellular locations, respectively. The results obtained by jackknife test are quite encouraging. It indicates that the ApEn of protein sequence could represent effectively the information of apoptosis proteins subcellular locations. It can at least play a complimentary role to many of the existing methods, and might become potentially useful tool for protein function prediction. The software in Matlab is available freely by contacting the corresponding author.