![IEEE International Conference on Computational Advances in Bio and Medical Sciences : [proceedings]](https://originalfileserver.aminer.cn/sys/aminer/magazine.png)
Durbin's positional Burrows-Wheeler transform (PBWT) enables algorithms with the optimal time complexity of O(MN) for reporting all vs all haplotype matches in a population panel with M haplotypes and N variant sites. However, even this efficiency may still be too slow when the number of haplotypes reaches millions. To further reduce the run time, in this paper, a parallel version of the PBWT algorithms is introduced for all versus all haplotype matching, which is called HP-PBWT (haplotype-based parallel PBWT). HP-PBWT parallelly executes the PBWT by splitting a haplotype panel into blocks of haplotypes. HP-PBWT algorithms achieve parallelization for PBWT construction, reporting all versus all L-long matches, and reporting all versus all set -maximal matches while maintaining memory efficiency. HP-PBWT has an O((M/T + T)N) time complexity in PBWT construction, and O((M/T + T + c*)N) time complexity for reporting all versus all L -long matches and reporting all versus all set -maximal matches, where T is the number of threads and c* is the maximum number of matches (of length L or maximum divergence value for L -long matches and set -maximal matches respectively) per haplotype per site. HP-PBWT achieves 4-fold speed-up in UK Biobank genotyping array data with 30 threads in the IO-included benchmarks. When applying HP-PBWT to a dataset of 8 million randomized haplotypes (random binary strings of equal length) in the IO-excluded benchmarks, it can achieve a 22-fold speed-up with 60 cores on the Amazon EC2 server. With further hardware optimization, HP-PBWT is expected to handle billions of haplotypes efficiently.
The ubiquitous approach of transfer learning for feature extraction is harnessed for image based detection of two types of craniofacial abnormalities: pediatric cleft and craniosynostosis. In the current study, using features extracted from pre-trained AlexNet activations, we train a multi class support vector machine (SVM) for cleft lip abnormality and developed a multi-view classifier using max voting for craniosynostosis anomaly detection. We achieved Area under the ROC curve (AUC) value of 0.95 for cleft abnormality and 0.84 for craniosynostosis.
Background Partial Least-Squares Discriminant Analysis (PLS-DA) is a popular machine learning tool that is gaining increasing attention as a useful feature selector and classifier. In an effort to understand its strengths and weaknesses, we performed a series of experiments with synthetic data and compared its performance to its close relative from which it was initially invented, namely Principal Component Analysis (PCA). Results We demonstrate that even though PCA ignores the information regarding the class labels of the samples, this unsupervised tool can be remarkably effective as a feature selector. In some cases, it outperforms PLS-DA, which is made aware of the class labels in its input. Our experiments range from looking at the signal-to-noise ratio in the feature selection task, to considering many practical distributions and models encountered when analyzing bioinformatics and clinical data. Other methods were also evaluated. Finally, we analyzed an interesting data set from 396 vaginal microbiome samples where the ground truth for the feature selection was available. All the 3D figures shown in this paper as well as the supplementary ones can be viewed interactively at http://biorg.cs.fiu.edu/plsda Conclusions Our results highlighted the strengths and weaknesses of PLS-DA in comparison with PCA for different underlying data models.
Network measures that reflect the most salient properties of complex large-scale networks are in high demand in the network research community. In this paper we adapt a combinatorial measure of negative curvature ( also called hyperbolicity) to parametrized finite networks, and show that a variety of biological and social networks are hyperbolic. This hyperbolicity property has strong implications on the higher-order connectivity and other topological properties of these networks. Specifically, we derive and prove bounds on the distance among shortest or approximately shortest paths in hyperbolic networks. We describe two implications of these bounds to crosstalk in biological networks, and to the existence of central, influential neighborhoods in both biological and social networks.
Background Automatic segmentation and localization of lesions in mammogram (MG) images are challenging even with employing advanced methods such as deep learning (DL) methods. We developed a new model based on the architecture of the semantic segmentation U-Net model to precisely segment mass lesions in MG images. The proposed end-to-end convolutional neural network (CNN) based model extracts contextual information by combining low-level and high-level features. We trained the proposed model using huge publicly available databases, (CBIS-DDSM, BCDR-01, and INbreast), and a private database from the University of Connecticut Health Center (UCHC). Results We compared the performance of the proposed model with those of the state-of-the-art DL models including the fully convolutional network (FCN), SegNet, Dilated-Net, original U-Net, and Faster R-CNN models and the conventional region growing (RG) method. The proposed Vanilla U-Net model outperforms the Faster R-CNN model significantly in terms of the runtime and the Intersection over Union metric (IOU). Training with digitized film-based and fully digitized MG images, the proposed Vanilla U-Net model achieves a mean test accuracy of 92.6%. The proposed model achieves a mean Dice coefficient index (DI) of 0.951 and a mean IOU of 0.909 that show how close the output segments are to the corresponding lesions in the ground truth maps. Data augmentation has been very effective in our experiments resulting in an increase in the mean DI and the mean IOU from 0.922 to 0.951 and 0.856 to 0.909, respectively. Conclusions The proposed Vanilla U-Net based model can be used for precise segmentation of masses in MG images. This is because the segmentation process incorporates more multi-scale spatial context, and captures more local and global context to predict a precise pixel-wise segmentation map of an input full MG image. These detected maps can help radiologists in differentiating benign and malignant lesions depend on the lesion shapes. We show that using transfer learning, introducing augmentation, and modifying the architecture of the original model results in better performance in terms of the mean accuracy, the mean DI, and the mean IOU in detecting mass lesion compared to the other DL and the conventional models.
Characterizing intratumor heterogeneity (ITH) is crucial to understanding cancer development, but it is hampered by limits of available data sources. Bulk DNA sequencing is the most common technology to assess ITH, but mixes many genetically distinct cells in each sample, which must then be computationally deconvolved. Single-cell sequencing (SCS) is a promising alternative, but its limitations — e.g., high noise, difficulty scaling to large populations, technical artifacts, and large data sets — have so far made it impractical for studying cohorts of sufficient size to identify statistically robust features of tumor evolution. We have developed strategies for deconvolution and tumor phylogenetics combining limited amounts of bulk and single-cell data to gain some advantages of single-cell resolution with much lower cost, with specific focus on deconvolving genomic copy number data. We developed a mixed membership model for clonal deconvolution via non-negative matrix factorization (NMF) balancing deconvolution quality with similarity to single-cell samples via an associated efficient coordinate descent algorithm. We then improve on that algorithm by integrating deconvolution with clonal phylogeny inference, using a mixed integer linear programming (MILP) model to incorporate a minimum evolution phylogenetic tree cost in the problem objective. We demonstrate the effectiveness of these methods on semi-simulated data of known ground truth, showing improved deconvolution accuracy relative to bulk data alone.
Summary form only given. The human body contains as many as ten times as many microbial cells as human cells, and these are organized into communities that colonize many tissues throughout the body. Each community has a different structure as the tissues/habitats provide different environments. The microbial communities provide many functions such as degradation or biosynthesis of nutrients, protection against infection, and are in constant interaction with the host. Microbial communities show plasticity in their composition, although there are general characteristics to beneficial communities. The microbiome changes after birth in ways that are becoming clearer, and the microbiome in the elderly is likewise different. An unbalanced community structure occasionally occurs and results in diseases, both chronic and acute. Manipulation of the microbiome as a therapeutic, as well as recognition of abnormal microbiomes as a diagnostic, are new frontiers for medicine.
BACKGROUND:Cancer progression reconstruction is an important development stemming from the phylogenetics field. In this context, the reconstruction of the phylogeny representing the evolutionary history presents some peculiar aspects that depend on the technology used to obtain the data to analyze: Single Cell DNA Sequencing data have great specificity, but are affected by moderate false negative and missing value rates. Moreover, there has been some recent evidence of back mutations in cancer: this phenomenon is currently widely ignored.RESULTS:We present a new tool, gpps, that reconstructs a tumor phylogeny from Single Cell Sequencing data, allowing each mutation to be lost at most a fixed number of times. The General Parsimony Phylogeny from Single cell (gpps) tool is open source and available at https://github.com/AlgoLab/gpps .CONCLUSIONS:gpps provides new insights to the analysis of intra-tumor heterogeneity by proposing a new progression model to the field of cancer phylogeny reconstruction on Single Cell data.
Stochastic rule-based models serve as natural and compact representations for biochemical reactions. The Gillespie stochastic simulation algorithm and its variants are employed to predict the behavior of biochemical systems modeled by such stochastic rule-based models. However, it is often not feasible to create a complete stochastic rule-based model from first principles. Instead, our knowledge of the biochemical system is used to obtain the set of chemical reactions of the stochastic rule-based model. The lack of knowledge about the rate constants of biochemical reactions is readily modeled by using unknown parameters in stochastic rule-based models.A primary challenge in the use of such a parameterized stochastic rule-based model for predicting the behavior of a biological system is the determination of the parameters of the model from multiple experimental observations. However, the focus of many earlier efforts has been on discovering parameter values of a parameterized stochastic biological model from a single specification written down in a computer-readable language such as probabilistic temporal logic.In practice, a biological model must satisfy multiple experimental observations made on the biological system being modeled. Hence, it is important to synthesize a single set of parameter values that cause a parameterized stochastic model to satisfy multiple probabilistic temporal logic specifications simultaneously.We present a new approach for estimating parameter values of stochastic biochemical models so that a single parameterized model satisfies all the given probabilistic temporal logic behavioral specifications simultaneously. Our approach first computes a quantitative metric describing how well a stochastic biochemical model satisfies a given specification. It then utilizes a multiple hypothesis testing based statistical model checking method to simultaneously validate the model against multiple probabilistic temporal logic behavioral specifications.We demonstrate the usefulness of our method by estimating the parameters of two stochastic rule-based models of biochemical receptors with 26 and 29 parameters against three probabilistic temporal logic behavioral specifications each. Our computational experiments are performed on an AMD Ryzen Threadripper 1900X 8-Core 3.8 GHz processor with 16 GB of RAM operating under Ubuntu 16.04, and obtained a set of parameter values for each model within one day.
Intrinsically disordered proteins or regions (IDPs or IDRs) lack stable structures in solution yet often fold upon binding with partners. IDPs or IDRs are highly abundant in all proteomes and represent a significant modification of the sequence → structure → function paradigm. In this study, we first present the results of various disorder predictions on a non-redundant set of PDB complexes. Further, structural formation by these predicted-to-be-disordered proteins was observed to depend upon various factors (including hetero group binding, protein/DNA/RNA binding, disulfide bonds, and ion binding). This study provides additional support for the newly proposed paradigm of sequence → IDP/IDR ensemble → function.
Trauma is one of the main causes of hospitalization. Time is of the essence in diagnosis and treatment of trauma patients with severe injuries. To assist in decision-making, we propose a hidden Markov model for identification of disease states through which patients progress. An important property of our model is that it is based on features which are routinely collected in hospital trauma centers. Using a hidden Markov model based on fifteen features, six different patient states are identified. The resulting Markov model can be useful in identifying patients' states to assist in diagnosis and treatment.
Background: In the string correction problem, we are to transform one string into another using a set of prescribed edit operations. In string correction using the Damerau-Levenshtein (DL) distance, the permissible edit operations are: substitution, insertion, deletion and transposition. Several algorithms for string correction using the DL distance have been proposed. The fastest and most space efficient of these algorithms is due to Lowrance and Wagner. It computes the DL distance between strings of length m and n, respectively, in O(mn) time and O(mn) space. In this paper, we focus on the development of algorithms whose asymptotic space complexity is less and whose actual runtime and energy consumption are less than those of the algorithm of Lowrance and Wagner. Results: We develop space-and cache-efficient algorithms to compute the Damerau-Levenshtein (DL) distance between two strings as well as to find a sequence of edit operations of length equal to the DL distance. Our algorithms require O(s min{m, n}+ m+ n) space, where s is the size of the alphabet and m and n are, respectively, the lengths of the two strings. Previously known algorithms require O(mn) space. The space-and cache-efficient algorithms of this paper are demonstrated, experimentally, to be superior to earlier algorithms for the DL distance problem on time, space, and enery metrics using three different computational platforms. Conclusion: Our benchmarking shows that, our algorithms are able to handle much larger sequences than earlier algorithms due to the reduction in space requirements. On a single core, we are able to compute the DL distance and an optimal edit sequence faster than known algorithms by as much as 73.1% and 63.5%, respectively. Further, we reduce energy consumption by as much as 68.5%. Multicore versions of our algorithms achieve a speedup of 23.2 on 24 cores.
The following topics are dealt with: bioinformatics; genomics; molecular biophysics; cellular biophysics; genetics; RNA; biology computing; diseases; proteins; learning (artificial intelligence).
Discovering patterns in biological sequences is a crucial step to extract useful information from them. Motifs can be viewed as patterns that occur exactly or with minor changes across some or all of the biological sequences. Motif search has numerous applications including the identification of transcription factors and their binding sites, composite regulatory patterns, similarity among families of proteins, etc. The general problem of motif search is intractable. One of the most studied models of motif search proposed in literature is Edit-distance based Motif Search (EMS). In EMS, the goal is to find all the patterns of length l that occur with an edit-distance of at most d in each of the input sequences. EMS algorithms existing in the literature do not scale well on challenging instances and large datasets. In this paper, the current state-of-the-art EMS solver is advanced by exploiting the idea of dimension reduction. A novel idea to reduce the cardinality of the alphabet is proposed. The algorithm we propose, EMS3, is an exact algorithm. I.e., it finds all the motifs present in the input sequences. EMS3 can be also viewed as a divide and conquer algorithm. In this paper, we provide theoretical analyses to establish the efficiency of EMS3. Extensive experiments on standard benchmark datasets (synthetic and real-world) show that the proposed algorithm outperforms the existing state-of-the-art algorithm (EMS2).
The variation in gene expression profiles of cells captured in different phases of the cell cycle can interfere with cell type identification and functional analysis of single cell RNA-Seq (scRNA-Seq) data. In this paper, we introduce SC1CC ( SC1 C ell C ycle analysis tool), a computational approach for clustering and ordering single cell transcriptional profiles according to their progression along cell cycle phases. We also introduce a new robust metric, Gene Smoothness Score (GSS) for assessing the cell cycle based order of the cells. SC1CC is available as part of the SC1 web-based scRNA-Seq analysis pipeline, publicly accessible at https://sc1.engr.uconn.edu/ .
Detrusor smooth muscle (DSM) instability is a common symptom of over active bladder (OAB) which effects quality of millions of lives. The enhanced spontaneous contraction due to generation of more frequent spontaneous action potentials (AP) is the fundamental cause of DSM instability. Ca <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2+</sup> dependent potassium channels are abundantly distributed in DSM cells and regulate the membrane excitability by modulating AP shape and the resting membrane potential. Quantitative studies of these channels will help to predict the underlying mechanism of electrical signaling in DSM cells. Here, we present a large conductance Ca <sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2+</sup> dependent potassium (BK) channel based upon Hodgkin-Huxley (HH) formalism which is incorporated into a published single DSM cell model to investigate several modulating effects. This simple BK ion channel model is constructed with biophysical details from documented electrophysiological studies. As it reflects several modulating properties found in DSM cell electrophysiology, further investigation can shed light in genesis of OB.
Microbe-microbe and host-microbe interactions in a microbiome play a vital role in both health and disease. However, the structure of the microbial community and the colonization patterns are highly complex to infer even under controlled wet laboratory conditions. In this study, we investigate what information, if any, can be provided by a Bayesian Network (BN) about a microbial community. Unlike the previously proposed Co-occurrence Networks (CoNs), BNs are based on conditional dependencies and can help in revealing complex associations. In this paper, we report a surprising association between directed edges in BNs and known colonization order. Furthermore, when combined with the sign of the correlations from CoNs, BNs allow for many useful conclusions.In this paper, we ask the pertinent question: what can BNs reveal about the relationships between microbial taxa in a microbiome? In particular, we argue that BNs are able to capture temporal order (colonization order) when combined with the sign of correlation. The main contribution of this paper is to show how to carefully use BNs in conjunction with CoNs to make potential inferences about colonization order.We carried out our experiments on two datasets. The first is oral microbiome data from the Human Microbiome Project (HMP) [1], which included eight different sites all from within the oral cavity from 242 healthy adults (129 males, 113 females). The second is data from preterm infant gut microbiome samples as described in the paper by La Rosa et al. [2]. A total of 922 stool samples from 58 premature babies (each weighing ≤ 1500 g at birth) collected on different days were used for our analysis.
During protein synthesis, the speed of ribosomes as they move along mRNAs can vary based on multiple biological factors. It has been reported that pausing of ribosomes can influence gene expression through multiple mechanisms, such as facilitating protein folding and triggering translational abandonment. However, a deeper understanding of ribosomal pausing and its' regulation is still missing. Ribosome profiling is a rapidly emerging technique for capturing a global "snap-shot" of ribosome position and activity. The method shows great potential for investigating many aspects of translating ribosomes. In particular, pause sites can be found by identifying peaks in ribosome profiling data. Here we present riboSmoothR, a novel method for identifying such peaks. Our algorithm utilizes the smoothed z-score method to identify anomalous ribosome footprint counts along a transcript. We show that riboSmoothR identifies a larger number of conserved peak sites than other commonly used peak finding methods. Additionally, we describe sequence-based features associated with peak sites identified by our method, such as codon and amino acid biases.
Single cell RNA-Seq (scRNA-Seq) is critical for studying cellular heterogeneity and development of tissues and tumors. There are currently only few tools that allow researchers to analyze scRNA-seq data without considerable coding efforts. Here, we present an interactive web-based pipeline for single-cell RNA-Seq data analysis that allows researchers to analyze data in an efficient work-flow. The tool also implements a novel and unique method of selecting informative genes before clustering and differential expression analysis.
The advance of high-throughput sequencing has made it possible to perform large-scale mutation analysis by altering codons on a sequence and investigating the effect of the changes. While sequence alignment algorithms can be applied to compare the reads to the original unaltered sequence, neither nucleotide-based nor protein-based alignment is completely suitable for analyzing changes at the codon level.We develop a codon-based sequence alignment algorithm by modifying the dynamic programming equations so that an exact match of three letters in a codon is assigned a positive score, a mismatch of at least one letter in a codon is assigned a negative score, and an indel of either one, two or three letters is assigned a constant negative score. This strategy models what could happen within each codon directly. It has the same time complexity as the nucleotide-based dynamic programming algorithm.We apply our algorithm to analyze the effect of mutations within the RNA Polymerase II trigger loop from high-throughput sequencing libraries that we have generated in our lab. We compare our results to the ones obtained with nucleotide-based alignment. We show that our algorithm is able to avoid systematic errors that could be made with nucleotide-based alignment.