The problem of computing the Elementary Flux Modes (EFMs) and Minimal Cut Sets (MCSs) of metabolic network is a fundamental one in metabolic networks. A key insight is that they can be understood as a dual pair of monotone Boolean functions (MBFs). Using this insight, this computation reduces to the question of generating from an oracle a dual pair of MBFs. If one of the two sets (functions) is known, then the other can be computed through a process known as dualization. Fredman and Khachiyan provided two algorithms, which they called simply A and B that can serve as an engine for oracle-based generation or dualization of MBFs. We look at efficiencies available in implementing their algorithm B, which we will refer to as FK-B. Like their algorithm A, FK-B certifies whether two given MBFs in the form of Conjunctive Normal Form and Disjunctive Normal Form are dual or not, and in case of not being dual it returns a conflicting assignment (CA), that is, an assignment that makes one of the given Boolean functions True and the other one False. The FK-B algorithm is a recursive algorithm that searches through the tree of assignments to find a CA. If it does not find any CA, it means that the given Boolean functions are dual. In this article, we propose six techniques applicable to the FK-B and hence to the dualization process. Although these techniques do not reduce the time complexity, they considerably reduce the running time in practice. We evaluate the proposed improvements by applying them to compute the MCSs from the EFMs in the 19 small- and medium-sized models from the BioModels database along with 4 models of biomass synthesis in Escherichia coli that were used in an earlier computational survey Haus et al. (2008).
Prediction of drug resistance and identification of its mechanisms in bacteria such as Mycobacterium tuberculosis, the etiological agent of tuberculosis, is a challenging problem. Solving this problem requires a transparent, accurate, and flexible predictive model. The methods currently used for this purpose rarely satisfy all of these criteria. On the one hand, approaches based on testing strains against a catalogue of previously identified mutations often yield poor predictive performance; on the other hand, machine learning techniques typically have higher predictive accuracy, but often lack interpretability and may learn patterns that produce accurate predictions for the wrong reasons. Current interpretable methods may either exhibit a lower accuracy or lack the flexibility needed to generalize them to previously unseen data. In this paper we propose a novel technique, inspired by group testing and Boolean compressed sensing, which yields highly accurate predictions, interpretable results, and is flexible enough to be optimized for various evaluation metrics at the same time. We test the predictive accuracy of our approach on five first-line and seven second-line antibiotics used for treating tuberculosis. We find that it has a higher or comparable accuracy to that of commonly used machine learning models, and is able to identify variants in genes with previously reported association to drug resistance. Our method is intrinsically interpretable, and can be customized for different evaluation metrics. Our implementation is available at github.com/hoomanzabeti/INGOT_DR and can be installed via The Python Package Index (Pypi) under ingotdr. This package is also compatible with most of the tools in the Scikit-learn machine learning library.
Motivation: Drug resistance in Mycobacterium tuberculosis (MTB) is a growing threat to human health worldwide. One way to mitigate the risk of drug resistance is to enable clinicians to prescribe the right antibiotic drugs to each patient through methods that predict drug resistance in MTB using whole-genome sequencing (WGS) data. Existing machine learning methods for this task typically convert the WGS data from a given bacterial isolate into features corresponding to single-nucleotide polymorphisms (SNPs) or short sequence segments of a fixed length K (K-mers). Here, we introduce a gene burden-based method for predicting drug resistance in TB. We define one numerical feature per gene corresponding to the number of mutations in that gene in a given isolate. This representation greatly reduces the number of model parameters. We further propose a model architecture that considers both gene order and locality structure through a Long-term Recurrent Convolutional Network (LRCN) architecture, which combines convolutional and recurrent layers. Results: We find that using these strategies yields a substantial, statistically significant improvement over state-of-the-art methods on a large dataset of M. tuberculosis isolates, and suggest that this improvement is driven by our method's ability to account for the order of the genes in the genome and their organization into operons. Availability: The implementations of our feature preprocessing pipeline1 and our LRCN model2 are publicly available, as is our complete dataset3. Supplementary information: Additional data are available in the Supplementary Materials document4.
Motivation The prediction of drug resistance and the identification of its mechanisms in bacteria such as Mycobacterium tuberculosis, the etiological agent of tuberculosis, is a challenging problem. Modern methods based on testing against a catalogue of previously identified mutations often yield poor predictive performance. On the other hand, machine learning techniques have demonstrated high predictive accuracy, but many of them lack interpretability to aid in identifying specific mutations which lead to resistance. We propose a novel technique, inspired by the group testing problem and Boolean compressed sensing, which yields highly accurate predictions and interpretable results at the same time. Results We develop a modified version of the Boolean compressed sensing problem for identifying drug resistance, and implement its formulation as an integer linear program. This allows us to characterize the predictive accuracy of the technique and select an appropriate metric to optimize. A simple adaptation of the problem also allows us to quantify the sensitivity-specificity trade-off of our model under different regimes. We test the predictive accuracy of our approach on a variety of commonly used antibiotics in treating tuberculosis and find that it has accuracy comparable to that of standard machine learning models and points to several genes with previously identified association to drug resistance. Availability https://github.com/hoomanzabeti/TB_Resistance_RuleBasedClassifier Contact hooman_zabeti@sfu.ca
BackgroundNutrigenomic has revolutionized our understanding of nutrition. As plants make up a noticeable part of our diet, in the present study we chose microRNAs of edible plants and investigated if they can perfectly match human genes, indicating potential regulatory functionalities.MethodsmiRNAs were obtained using the PNRD database. Edible plants were separated and microRNAs in common in at least four of them entered our analysis. Using vmatchPattern, these 64 miRNAs went through four steps of refinement to improve target prediction: Alignment with the whole genome (2581 results), filtered for those in gene regions (1371 results), filtered for exon regions (66 results) and finally alignment with the human CDS (41 results). The identified genes were further analyzed in-silico to find their functions and relations to human diseases.ResultsFour common plant miRNAs were identified to match perfectly with 22 human transcripts. The identified target genes were involved in a broad range of body functions, from muscle contraction to tumor suppression. We could also indicate some connections between these findings and folk herbology and botanical medicine.ConclusionsThe food that we regularly eat has a great potential in affecting our genome and altering body functions. Plant miRNAs can provide means of designing drugs for a vast range of health problems including obesity and cancer, since they target genes involved in cell cycle (CCNC), digestion (GIPR) and muscular contractions (MYLK). They can also target regions of CDS for which we still have no sufficient information, to help boost our knowledge of the human genome.
MicroRNAs (miRNAs) are short non-coding RNAs which bind to mRNAs and regulate their expression. MiRNAs have been found to be associated with initiation and progression of many complex diseases. Investigating miRNAs and their targets can thus help develop new therapies by designing anti-miRNA oligonucleotides. While existing computational approaches can predict miRNA targets, these predictions have low accuracy. In this paper, we propose a two-step approach to refine the results of sequence-based prediction algorithms. The first step, which is based on our previous work, uses an ensemble learning approach that combines multiple existing methods. The second step utilizes support vector machine (SVM) classifiers in one- and two-class modes to infer miRNA-mRNA interactions based on both binding features, as well as network features extracted from gene regulatory network. Experimental results using two real data sets from TCGA indicate that the use of two-class SVM classification significantly improves the precision of miRNA-mRNA prediction.
RationaleUnderstanding mechanisms of resistance to M. tuberculosis (M. tb) infection in humans could identify novel therapeutic strategies as it has for other infectious diseases, such as HIV.ObjectivesTo compare the early transcriptional response of M. tb-infected monocytes between Ugandan household contacts of tuberculosis patients who demonstrate clinical resistance to M. tb infection (cases) and matched controls with latent tuberculosis infection.MethodsCases (n = 10) and controls (n = 18) were selected from a long-term household contact study in which cases did not convert their tuberculin skin test (TST) or develop tuberculosis over two years of follow up. We obtained genome-wide transcriptional profiles of M. tb-infected peripheral blood monocytes and used Gene Set Enrichment Analysis and interaction networks to identify cellular processes associated with resistance to clinical M. tb infection.Measurements and main resultsWe discovered gene sets associated with histone deacetylases that were differentially expressed when comparing resistant and susceptible subjects. We used small molecule inhibitors to demonstrate that histone deacetylase function is important for the pro-inflammatory response to in-vitro M. tb infection in human monocytes.ConclusionsMonocytes from individuals who appear to resist clinical M. tb infection differentially activate pathways controlled by histone deacetylase in response to in-vitro M. tb infection when compared to those who are susceptible and develop latent tuberculosis. These data identify a potential cellular mechanism underlying the clinical phenomenon of resistance to M. tb infection despite known exposure to an infectious contact.
Genetic networks provide compact representations of interactions between genes, and offer a systems perspective into biological processes and cellular functions. Many algorithms have been developed to estimate such networks based on steady-state gene expression profiles. However, the estimated networks using different methods are often very different from each other. On the other hand, it is not clear whether differences observed between estimated networks in two different biological conditions are truly meaningful, or due to variability in estimation procedures. In this paper, we aim to answer these questions by conducting a comprehensive empirical study to compare networks obtained from different estimation methods and for different subtypes of cancer. We evaluate various network descriptors to assess complex properties of estimated networks, beyond their local structures, and propose a simple permutation test for comparing estimated networks. The results provide new insight into properties of estimated networks using different reconstructionmethods, as well as differences in estimated networks in different biological conditions.
Testicular cancer is the most common cancer in men aged between 15 and 35 and more than 90% of testicular neoplasms are originated at germ cells. Recent research has shown the impact of microRNAs (miRNAs) in different types of cancer, including testicular germ cell tumor (TGCT). MicroRNAs are small non-coding RNAs which affect the development and progression of cancer cells by binding to mRNAs and regulating their expressions. The identification of functional miRNA–mRNA interactions in cancers, i.e. those that alter the expression of genes in cancer cells, can help delineate post-regulatory mechanisms and may lead to new treatments to control the progression of cancer. A number of sequence-based methods have been developed to predict miRNA–mRNA interactions based on the complementarity of sequences. While necessary, sequence complementarity is, however, not sufficient for presence of functional interactions. Alternative methods have thus been developed to refine the sequence-based interactions using concurrent expression profiles of miRNAs and mRNAs. This study aims to find functional cancer-specific miRNA–mRNA interactions in TGCT. To this end, the sequence-based predicted interactions are first refined using an ensemble learning method, based on two well-known methods of learning miRNA–mRNA interactions, namely, TaLasso and GenMiR++. Additional functional analyses were then used to identify a subset of interactions to be most likely functional and specific to TGCT. The final list of 13 miRNA–mRNA interactions can be potential targets for identifying TGCT-specific interactions and future laboratory experiments to develop new therapies.
Traveling salesman problem is of the known and classical problems at Research in Operations.Many scientific activities can be solved as traveling salesman problem.Existing methods for solving hard problems (such as the traveling salesman problem) consists of a large number of variables and constraints which reduces their practical efficiency in solving problems with the original size.In recent decades, the use of heuristic and meta-heuristic algorithms such as genetic algorithms is considered.Due to the simple structure of metaheuristic algorithms that have shown greater ability is more used by researchers in operational research.In this study, the improved genetic algorithm is used to solve TSP that the difference of it with the standard genetic algorithm is in the evaluation function.The new evaluation function is from a common evaluation function and a new idea.
Network reconstruction is an important yet challenging task in systems biology. While many methods have been recently proposed for reconstructing biological networks from diverse data types, properties of estimated networks and differences between reconstruction methods are not well understood. In this paper, we conduct a comprehensive empirical evaluation of seven existing network reconstruction methods, by comparing the estimated networks with different sparsity levels for both normal and tumor samples. The results suggest substantial heterogeneity in networks reconstructed using different reconstruction methods. Our findings also provide evidence for significant differences between networks of normal and tumor samples, even after accounting for the considerable variability in structures of networks estimated using different reconstruction methods. These differences can offer new insight into changes in mechanisms of genetic interaction associated with cancer initiation and progression.
A 2-dimensional Xover-based non-dominated sorting (2DXNS) genetic algorithm is proposed for multi-objective scheduling of static soft real-time tasks on heterogeneous computing systems. The objectives of scheduling problems are generally to minimize makespan, total processor idle time, and the number of processors as well as maximize guarantee ratio. In contrast to earlier approaches that commonly convert the problem to a single objective form; we propose an inherently multi-objective approach that aims to make appropriate tradeoffs among all of the above objectives. A new 2-dimensional structure is also proposed for crossover/mutation operator, in comparison to the more common crossover operators. This operator promotes better recombination due to processor based task decomposition. The simultaneous operation of this crossover operator across these ‘smaller’ problems increases better mating/combination, while trying to maintain more ‘sections’ of the parent chromosomes. This leads to better results with fewer generation passes. Finally, we compare the proposed algorithm with HEFT, a well-known list scheduling algorithm. Experiments on a real world application as well as three sets of random DAGs demonstrate the high efficiency of the proposed 2DXNS in multiprocessor task scheduling. KEYWORDSDirected Acyclic Task Graph (DAG), Heterogeneous Multiprocessor Systems, Load Balancing, Multi-objective Optimization, Task Scheduling, Genetic Algorithm Multiprocessor task scheduling is a challenging problem both in terms of its complexity in decision making and optimization as well as its importance in terms of applicability to modern systems. More specifically, from the perspective of decision making and optimization, this problem can be seen at the crossroads of three important fields of research, i.e. multiprocessor systems, real time decision making, and multi-objective optimization, each of which present a challenging area with its own significant set of complexities [1]. Furthermore, from an application perspective, many real world problems such as aircraft[2], networks [3], electronics and semiconductor manufacturing[4] require tradeoffs to be made among multiple objectives in order to optimize the overall performance of the system. One specific real world application of multiprocessor task scheduling is project management. Each processor corresponds to a resource (human or non-human) and each task corresponds to a sub-project that must be performed withina specific parallel order so that total time of project is minimized. Obviously, the multi-objective scheduling problems are more complex than their traditional single objective problems due to inconsistency, confluence or even contradiction among objectives [5]. Multiprocessor systems present their own set of complexities where inappropriate scheduling of tasks can lead to deteriorated performance of the system in terms of either excessive communication overhead or underutilization of resources [6, 7]. In a heterogeneous computing system, for instance, if one processor executes a given task faster than other processors, it does not necessarily mean that it executes all tasks fastest, due to the differences in internal architecture of different processor. Finally, from a real-time system perspective, additional soft or hard time constraints must be satisfied. In hard real-time systems, the violation of timing constraints of a certain task can be dangerous. Some examples of hard real-time systems are patient monitoring systems and nuclear plant control. While in the soft real time system, the violation of timing constraints of a task results in decreasing usefulness of results over time after the deadline expires without causing any damage to the controlled environment. Some examples of soft real-time systems are telephone switching systems and image processing applications [8]. Task scheduling problem is considered for two different multiprocessor systems: multiprocessor systems with fully connected topology [8-14] and arbitrary topology [6, 15-20]. In its general form, multiprocessor task scheduling is an NP-complete problem [9, 21]. There are two well-known special cases for this problem where optimal polynomial-time solutions are known. These SEDAGHAT ET AL.: A 2-DIMENTIONAL CROSSOVER FOR MULTI-OBJECTIVE EVOLUTIONARY SCHEDULING OF... Indian J.Sci.Res. 4 (3): 318-334, 2014 cases are scheduling tree-structured task graphs with identical computation costs on an arbitrary number of processors and scheduling arbitrary task graphs with identical computation costs on two processors. In these two cases, no communication is assumed among the tasks of the parallel program [9]. Hence, while many researchers have to date proposed various algorithms, their solutions generally simplify the actual problem and fail in one or more ways to consider the true complexity of an actual real time multiprocessor task scheduling problem. Task scheduling problem is inherently multiobjective fashion. Because of the complexity of the scheduling problem, many researchers use evolutionary algorithms in its single objective form [8, 10, 11, 13, 14, 22-24]. Converting multi-objective problem to single objective problem has to make hard tradeoffs among objectives that can lead to reduced performance. Here, we propose a new 2-dimensional Xover-based nondominated sorting (2DXNS) genetic algorithm that takes benefit from the pioneering works of [25] on nondominated sorting genetic algorithm II (NSGA-II) for multi-objective scheduling of static soft real-time tasks on heterogeneous computing systems. Real-world constraints including the precedence relationship between tasks as well as communication costs between the processors are considered. In contrast to the existing literature, the proposed multi-objective approach simultaneously considers several main objectives in scheduling problem such as length of scheduling (makespan), total tardiness (for real time tasks), number of processors as well as guarantee ratio. Furthermore, a new 2-dimensional Xover/mutation operator is proposed that, in comparison to the more common crossover operators such as cycle [10] and one-point [8, 11, 14] crossovers, promotes better recombination in the task scheduling problem. The 2-D representation decomposes the list (vector) of tasks to a task matrix with each row corresponding to a processor; hence it essentially decomposes the large task scheduling problem to several smaller problems. Simultaneous crossover operation across this matrix leads to better mating/combination while larger segments of each parent’s solutions are maintained. As a benchmark for comparison, the proposed 2DXNS is compared with HEFT (heterogeneous earliest finish time)[26], a well-known list scheduling algorithm. In these experiments, the performance of the 2DXNS is examined under different conditions: varying number of tasks, varying communication to computation ratio (CCR) and varying Sparsity of directed acyclic graph (DAG). The rest of the paper is organized as follows. In Section 2, the multiprocessor task scheduling problem is formulated. Section 3 presents a brief review of the prior relevant works. 2DXNS is presented in Section 4. Experimental results are illustrated in Section 5. Finally, Section 6 concludes the paper and suggests several ideas for further development.
Feature selection of the raw data is a fundamental step in the most pattern recognition and machine learning applications. The primary problem of feature selection is the criterion which evaluates a feature set. In the context of classification problems, optimal criterion would be the Bayesian error rate for selected subset of features. The Bayesian error rate bounds to some values that are related to mutual information. This interval shrinks as the mutual information increases. In this paper, we investigated the relationship between dependency and the Naïve Bayes error; dependency of the selected features is calculated as mutual information between the selected features and class. We designed some experiments to examine it about a two classes and two binary features problem. We found that in binary feature selection problem, the Naïve Bayes error increases as dependency increases; however, we showed that there are some states that the Naïve Bayes classifier is optimal while its default assumption is strongly violated (dependency is more than 0.8).
Path planning problem of mobile robot is one that has intrigued and has received much attention throughout the history of robotics, since it is at the essence of what a mobile robot needs to be considered truly “autonomous”. A mobile robot must be able to find collision-free paths to move from one location to another, and in order to truly show a level of intelligence these paths must be optimized under some criteria that are important to the robot, working space and given problem. In this paper we propose a new structured multi-objective genetic algorithm to solve this problem. In our method we explore only valid search space that results in a smaller search space. Also we show the defect of earlier evaluation function and present a new evaluation function. To evaluate our idea we compare our evaluation function with other ones and show the performance of our method. Experiments show the ability of our method in finding best paths with low generation and population.
Task scheduling is an essential aspect of parallel processing system. This problem assumes fully connected processors and ignores contention on the communication links. However, as arbitrary processor network (APN), communication contention has a strong influence on the execution time of a parallel application. In this paper, we propose genetic algorithms with fuzzy routing to face with link contention. In fuzzy routing algorithm, we consider speed of links and also busy time of intermediate links. To evaluate our method, we generate random DAGs with different Sparsity value based on Bernoulli distribution and compare our method with genetic algorithm and classic routing algorithm and also with BSA (bubble scheduling and allocation) method that is a well-known algorithm in this field. Experimental results show our method (GA with fuzzy routing) is able to find a scheduling with lower makespan than GA with classic routing and also BSA.
Edge detection is a basic and important subject in computer vision and image processing. An edge detector is defined as a mathematical operator of small spatial extent that responds in some way to these discontinuities, usually classifying every image pixel as either belonging to an edge or not. Many researchers have been spent attempting to develop effective edge detection algorithms. Despite this extensive research, the task of finding the edges that correspond to true physical boundaries remains a difficult problem. Edge detection algorithms based on the application of human knowledge show their flexibility and suggest that the use of human knowledge is a reasonable alternative. In this paper we propose a fuzzy inference system with two inputs: gradient and wavelet details. First input is calculated by Sobel operator and the second is calculated by wavelet transform of input image and then reconstruction of image only with details subimages by inverse wavelet transform. There are many fuzzy edge detection methods, but none of them utilize wavelet transform as it is used in this paper. For evaluating our method, we detect edges of images with different brightness characteristics and compare results with canny edge detector. The results show the high performance of our method in finding true edges.
Task scheduling is an essential aspect of parallel processing system This problem assumes fully connected processors and ignores contention on the communication links However, as arbitrary processor network (APN), communication contention has a strong influence on the execution time of a parallel application In this paper, we propose multi-objective genetic algorithm to solve task scheduling problem with time constraints in unstructured heterogeneous processors to find the scheduling with minimum makespan and total tardiness To optimize objectives, we use Pareto front based technique, vector based method In this problem, just like tasks, we schedule messages on suitable links during the minimization of the makespan and total tardiness To find a path for transferring a message between processors we use classic routing algorithm We compare our method with BSA method that is a well known algorithm Experimental results show our method is better than BSA and yield better makespan and total tardiness.
Scheduling of real-time tasks on a multi-processor system is an NP-hard problem. This paper aims to propose an algorithm based on multi-objective GA (MOGA) for scheduling of static soft real-time tasks on a heterogeneous multi-processor system when the real-world constraints including the precedence relationship between tasks, different arrival time for each task as well as communication delays between the processors are all considered. The objectives of the proposed scheduling algorithm are maximizing system utilization and minimizing total tardiness. Since these objectives are conflicting, the proposed method applies adaptive weight approach (AWA) where some useful information from the current population is utilized to readjust the weights for obtaining a search pressure toward a positive ideal point. In this paper, we also propose two greedy algorithms in which each algorithm aims to optimize a single objective either idle time or communication delay. The performance of the proposed MOGA is compared with the performance of the greedy algorithms on two types of DAGs, sparse and non-sparse. The results demonstrate the high efficiency of the proposed MOGA in solving real-world task scheduling problems.
Mohammad R. Akbarzadeh-Totonchi合作论文数Department of Electrical Engineering;Faculty of Engineering3