Enumerative studies of RNA secondary structures were initiated four decades ago by Waterman and his coworkers. Since then, RNA secondary structures have been explored according to many different structural characteristics, for instance, helices, components and loops by Hofacker, Schuster and Stadler, orders by Nebel, saturated structures by Clote, the 5^'-3^' end distance by Clote, Ponty and Steyaert, and the rainbow spectrum by Li and Reidys. However, the majority of the contributions are asymptotic results, and it is harder to derive explicit formulas. In this paper, we obtain exact formulas counting RNA secondary structures with a given number of helices as well as a given joint size distribution of helices and loops, while some related asymptotic results due to Hofacker, Schuster and Stadler have been known for about twenty years. Our approach is combinatorial, analyzing a recent bijection between RNA secondary structures and plane trees discovered by the first author and proposing a variation of Chen's bijective approach of counting trees by forests of simple trees.
The study of native motifs of RNA secondary structures helps us better understand the formation and eventually the functions of these molecules. Commonly known structural motifs include helices, hairpin loops, bulges, interior loops, exterior loops and multiloops. However, enumerative results and generating algorithms taking into account the joint distribution of these motifs are sparse. In this paper, we present progress on deriving such distributions employing a tree-bijection of RNA secondary structures obtained by Schmitt and Waterman and a novel rake decomposition of plane trees. The key feature of the latter is that the derived components encode motifs of the RNA secondary structures without pseudoknots associated with the plane trees very well. As an application, we present an algorithm (RakeSamp) generating uniformly random secondary structures without pseudoknots that satisfy fine motif specifications on the length and degree of various types of loops as well as helices.
Transcriptional interactions in a cell are modulated by a variety of mechanisms that prevent their representation as pure pairwise interactions between a transcription factor and its target(s). These include, among others, transcription factor activation by phosphorylation and acetylation, formation of active complexes with one or more co-factors, and mRNA/protein degradation and stabilization processes. This paper presents a first step towards the systematic, genome-wide computational inference of genes that modulate the interactions of specific transcription factors at the post-transcriptional level. The method uses a statistical test based on changes in the mutual information between a transcription factor and each of its candidate targets, conditional on the expression of a third gene. The approach was first validated on a synthetic network model, and then tested in the context of a mammalian cellular system. By analyzing 254 microarray expression profiles of normal and tumor related human B lymphocytes, we investigated the post transcriptional modulators of the MYC proto-oncogene, an important transcription factor involved in tumorigenesis. Our method discovered a set of 100 putative modulator genes, responsible for modulating 205 regulatory relationships between MYC and its targets. The set is significantly enriched in molecules with function consistent with their activities as modulators of cellular interactions, recapitulates established MYC regulation pathways, and provides a notable repertoire of novel regulators of MYC function. The approach has broad applicability and can be used to discover modulators of any other transcription factor, provided that adequate expression profile data are available.
COVID-19 is an infectious disease caused by the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2). The viral genome is considered to be relatively stable and the mutations that have been observed and reported thus far are mainly focused on the coding region. This article provides evidence that macrolevel pandemic dynamics, such as social distancing, modulate the genomic evolution of SARS-CoV-2. This view complements the prevalent paradigm that microlevel observables control macrolevel parameters such as death rates and infection patterns. First, we observe differences in mutational signals for geospatially separated populations such as the prevalence of A23404G in CA versus NY and WA. We show that the feedback between macrolevel dynamics and the viral population can be captured employing a transfer entropy framework. Second, we observe complex interactions within mutational clades. Namely, when C14408T first appeared in the viral population, the frequency of A23404G spiked in the subsequent week. Third, we identify a noncoding mutation, G29540A, within the segment between the coding gene of the N protein and the ORF10 gene, which is largely confined to NY (>95%). These observations indicate that macrolevel sociobehavioral measures have an impact on the viral genomics and may be useful for the dashboard-like tracking of its evolution. Finally, despite the fact that SARS-CoV-2 is a genetically robust organism, our findings suggest that we are dealing with a high degree of adaptability. Owing to its ample spread, mutations of unusual form are observed and a high complexity of mutational interaction is exhibited.
In May 1985 there was at University of California Santa Cruz an influential meeting that was the first serious discussion of sequencing the entire human genome. The author was one of the participants and described the meeting and related issues.
Sterol biosynthesis, primarily associated with eukaryotic kingdoms of life, occurs as an abbreviated pathway in the bacterium Methylococcus capsulatus. Sterol 14 alpha-demethylation is an essential step in this pathway and is catalyzed by cytochrome P450 51 (CYP51). In M. capsulatus, the enzyme consists of the P450 domain naturally fused to a ferredoxin domain at the C-terminus (CYP51fx). The structure of M. capsulatus CYP51fx was solved to 2.7 angstrom resolution and is the first structure of a bacterial sterol biosynthetic enzyme. The structure contained one P450 molecule per asymmetric unit with no electron density seen for ferredoxin. We connect this with the requirement of P450 substrate binding in order to activate productive ferredoxin binding. Further, the structure of the P450 domain with bound detergent (which replaced the substrate upon crystallization) was solved to 2.4 angstrom resolution. Comparison of these two structures to the CYP51s from human, fungi, and protozoa reveals strict conservation of the overall protein architecture. However, the structure of an "orphan" P450 from nonsterol-producing Mycobacterium tuberculosis that also has CYP51 activity reveals marked differences, suggesting that loss of function in vivo might have led to alterations in the structural constraints. Our results are consistent with the idea that eukaryotic and bacterial CYP51s evolved from a common cenancestor and that early eukaryotes may have recruited CYP51 from a bacterial source. The idea is supported by bioinformatic analysis, revealing the presence of CYP51 genes in >1,000 bacteria from nine different phyla, >50 of them being natural CYP51fx fusion proteins.
Levenshtein edit distance has played a central role—both past and present—in sequence alignment in particular and biological database similarity search in general. We start our review with a history of dynamic programming algorithms for computing Levenshtein distance and sequence alignments. Following, we describe how those algorithms led to heuristics employed in the most widely used software in bioinformatics, BLAST, a program to search DNA and protein databases for evolutionarily relevant similarities. More recently, the advent of modern genomic sequencing and the volume of data it generates has resulted in a return to the problem of local alignment. We conclude with how the mathematical formulation of Levenshtein distance as a metric made possible additional optimizations to similarity search in biological contexts. These modern optimizations are built around the low metric entropy and fractional dimensionality of biological databases, enabling orders of magnitude acceleration of biological similarity search.
Sterol biosynthesis, primarily associated with eukaryotic kingdoms of life, occurs as an abbreviated pathway in the bacterium Methylococcus capsulatus. Sterol 14a-demethylation is an essential step in this pathway and is catalyzed by cytochrome P450 51 (CYP51). In M. capsulatus, the enzyme consists of the P450 domain naturally fused to a ferredoxin domain at the C-terminus (CYP51fx). The structure of M. capsulatus CYP51fx was solved to 2.7 Å resolution and is the first structure of a bacterial sterol biosynthetic enzyme. The structure contained one P450 molecule per asymmetric unit with no electron density seen for ferredoxin. We connect this with the requirement of P450 substrate binding in order to activate productive ferredoxin binding. Further, the structure of the P450 domain with bound detergent (which replaced the substrate upon crystallization) was solved to 2.4 Å resolution. Comparison of these two structures to the CYP51s from human, fungi, and protozoa reveals strict conservation of the overall protein architecture. However, the structure of an “orphan” P450 from nonsterol-producing Mycobacterium tuberculosis that also has CYP51 activity reveals marked differences, suggesting that loss of function in vivo might have led to alterations in the structural constraints. Our results are consistent with the idea that eukaryotic and bacterial CYP51s evolved from a common cenancestor and that early eukaryotes may have recruited CYP51 from a bacterial source. The idea is supported by bioinformatic analysis, revealing the presence of CYP51 genes in >1,000 bacteria from nine different phyla, >50 of them being natural CYP51fx fusion proteins.
BackgroundAlignment-free (AF) sequence comparison is attracting persistent interest driven by data-intensive applications. Hence, many AF procedures have been proposed in recent years, but a lack of a clearly defined benchmarking consensus hampers their performance assessment.ResultsHere, we present a community resource (http://afproject.org) to establish standards for comparing alignment-free approaches across different areas of sequence-based research. We characterize 74 AF methods available in 24 software tools for five research applications, namely, protein sequence classification, gene tree inference, regulatory element detection, genome-based phylogenetic inference, and reconstruction of species trees under horizontal gene transfer and recombination events.ConclusionThe interactive web service allows researchers to explore the performance of alignment-free tools relevant to their data types and analytical goals. It also allows method developers to assess their own algorithms and compare them with current state-of-the-art tools, accelerating the development of new, more accurate AF solutions.
MOTIVATION:Detecting sequences containing repetitive regions is a basic bioinformatics task with many applications. Several methods have been developed for various types of repeat detection tasks. An efficient generic method for detecting most types of repetitive sequences is still desirable. Inspired by the excellent properties and successful applications of the D2 family of statistics in comparative analyses of genomic sequences, we developed a new statistic D2R that can efficiently discriminate sequences with or without repetitive regions.RESULTS:Using the statistic, we developed an algorithm of linear time and space complexity for detecting most types of repetitive sequences in multiple scenarios, including finding candidate clustered regularly interspaced short palindromic repeats regions from bacterial genomic or metagenomics sequences. Simulation and real data experiments show that the method works well on both assembled sequences and unassembled short reads.AVAILABILITY AND IMPLEMENTATION:The codes are available at https://github.com/XuegongLab/D2R_codes under GPL 3.0 license.SUPPLEMENTARY INFORMATION:Supplementary data are available at Bioinformatics online.
Motivation:Capturing association patterns in gene expression levels under different conditions or time points is important for inferring gene regulatory interactions. In practice, temporal changes in gene expression may result in complex association patterns that require more sophisticated detection methods than simple correlation measures. For instance, the effect of regulation may lead to time-lagged associations and interactions local to a subset of samples. Furthermore, expression profiles of interest may not be aligned or directly comparable (e.g. gene expression profiles from two species).Results:We propose a count statistic for measuring association between pairs of gene expression profiles consisting of ordered samples (e.g. time-course), where correlation may only exist locally in subsequences separated by a position shift. The statistic is simple and fast to compute, and we illustrate its use in two applications. In a cross-species comparison of developmental gene expression levels, we show our method not only measures association of gene expressions between the two species, but also provides alignment between different developmental stages. In the second application, we applied our statistic to expression profiles from two distinct phenotypic conditions, where the samples in each profile are ordered by the associated phenotypic values. The detected associations can be useful in building correspondence between gene association networks under different phenotypes. On the theoretical side, we provide asymptotic distributions of the statistic for different regions of the parameter space and test its power on simulated data.Availability and implementation:The code used to perform the analysis is available as part of the Supplementary Material.Contact:msw@usc.edu or hhuang@stat.berkeley.edu.Supplementary information:Supplementary data are available at Bioinformatics online.
The efficiency of treatment of human infections with the unicellular eukaryotic pathogens such as fungi and protozoa remains deeply unsatisfactory. For example, the mortality rates from nosocomial fungemia in critically ill, immunosuppressed or post-cancer patients often exceed 50%. A set of six systemic clinical azoles [sterol 14α-demethylase (CYP51) inhibitors] represents the first-line antifungal treatment. All these drugs were discovered empirically, by monitoring their effects on fungal cell growth, though it had been proven that they kill fungal cells by blocking the biosynthesis of ergosterol in fungi at the stage of 14α-demethylation of the sterol nucleus. This review briefs the history of antifungal azoles, outlines the situation with the current clinical azole-based drugs, describes the attempts of their repurposing for treatment of human infections with the protozoan parasites that, similar to fungi, also produce endogenous sterols, and discusses the most recently acquired knowledge on the CYP51 structure/function and inhibition. It is our belief that this information should be helpful in shifting from the traditional phenotypic screening to the actual target-driven drug discovery paradigm, which will rationalize and substantially accelerate the development of new, more efficient and pathogen-oriented CYP51 inhibitors.
Alignment-free genome and metagenome comparisons are increasingly important with the development of next generation sequencing (NGS) technologies. Recently developed state-of-the-art k-mer based alignment-free dissimilarity measures including CVTree, |$d_2^*$| and |$d_2^S$| are more computationally expensive than measures based solely on the k-mer frequencies. Here, we report a standalone software, aCcelerated Alignment-FrEe sequence analysis (CAFE), for efficient calculation of 28 alignment-free dissimilarity measures. CAFE allows for both assembled genome sequences and unassembled NGS shotgun reads as input, and wraps the output in a standard PHYLIP format. In downstream analyses, CAFE can also be used to visualize the pairwise dissimilarity measures, including dendrograms, heatmap, principal coordinate analysis and network display. CAFE serves as a general k-mer based alignment-free analysis platform for studying the relationships among genomes and metagenomes, and is freely available at https://github.com/younglululu/CAFE.
BACKGROUND:Alignment-free sequence comparison using counts of word patterns (grams, k-tuples) has become an active research topic due to the large amount of sequence data from the new sequencing technologies. Genome sequences are frequently modelled by Markov chains and the likelihood ratio test or the corresponding approximate χ 2-statistic has been suggested to compare two sequences. However, it is not known how to best choose the word length k in such studies.RESULTS:We develop an optimal strategy to choose k by maximizing the statistical power of detecting differences between two sequences. Let the orders of the Markov chains for the two sequences be r 1 and r 2, respectively. We show through both simulations and theoretical studies that the optimal k= max(r 1,r 2)+1 for both long sequences and next generation sequencing (NGS) read data. The orders of the Markov chains may be unknown and several methods have been developed to estimate the orders of Markov chains based on both long sequences and NGS reads. We study the power loss of the statistics when the estimated orders are used. It is shown that the power loss is minimal for some of the estimators of the orders of Markov chains.CONCLUSION:Our studies provide guidelines on choosing the optimal word length for the comparison of Markov sequences.
Yang Young Lu,Kujin Tang,Jie Ren,Jed A. Fuhrman,Michael S. Waterman and Fengzhu Sun1,3∗ Molecular and Computational Biology Program, Department of Biological Sciences, University of Southern California, CA, USA Department of Biological Sciences and Wrigley Institute for Environmental Studies, University of Southern California, Los Angeles, California, USA Centre for Computational Systems Biology, School of Mathematical Sciences, Fudan University, Shanghai, China
*Correspondence: fsun@usc.edu 1Centre for Computational Systems Biology, School of Mathematical Sciences, Fudan University , Shanghai, China 2Molecular and Computational Biology Program, University of Southern California, Los Angeles, California, USA Full list of author information is available at the end of the article 1 Appendix 1.1 Proof of Theorem 1 Let r1 and r2 (r1 ≤ r2) be the orders of sequences A1 and A2, respectively. If k < r2 + 1, even under the null hypothesis that the two sequences are from the same MC, the statistic Sk is large and does not follow a χ-distribution. For a given type I error α, the threshold for rejecting the null hypothesis should be very big. Therefore, the power of Sk under the alternative hypothesis should be low. When k ≥ r2 + 1, Sk will be χ-distributed with Ck−1(C − 1) degrees of freedom under the null hypothesis, where C is the size of the alphabet. Under the alternative hypothesis, Sk will have a non-central χ-distribution with non-centrality parameter that can be computed as follows. Without loss of generality, we assume that the sequences have the same length L. Note from the law of large numbers for MCs [1, 2],
Cytochrome P450 (P450) enzymes are important in the metabolism of drugs, steroids, fat-soluble vitamins, carcinogens, pesticides, and many other types of chemicals. Their catalytic activities are important issues in areas such as drug-drug interactions and endocrine function. During the past 30 years, structures of P450s have been very helpful in understanding function, particularly the mammalian P450 structures available in the past 15 years. We review recent activity in this area, focusing on the past 2 years (2014-2015). Structural work with microbial P450s includes studies related to the biosynthesis of natural products and the use of parasitic and fungal P450 structures as targets for drug discovery. Studies on mammalian P450s include the utilization of information about 'drug-metabolizing' P450s to improve drug development and also to understand the molecular bases of endocrine dysfunction.
Biosynthesis of steroid hormones in vertebrates involves three cytochrome P450 hydroxylases, CYP11A1, CYP17A1 and CYP19A1, which catalyze sequential steps in steroidogenesis. These enzymes are conserved in the vertebrates, but their origin and existence in other chordate subphyla (Tunicata and Cephalochordata) have not been clearly established. In this study, selected protein sequences of CYP11A1, CYP17A1 and CYP19A1 were compiled and analyzed using multiple sequence alignment and phylogenetic analysis. Our analyses show that cephalochordates have sequences orthologous to vertebrate CYP11A1, CYP17A1 or CYP19A1, and that echinoderms and hemichordates possess CYP11-like but not CYP19 genes. While the cephalochordate sequences have low identity with the vertebrate sequences, reflecting evolutionary distance, the data show apparent origin of CYP11 prior to the evolution of CYP19 and possibly CYP17, thus indicating a sequential origin of these functionally related steroidogenic CYPs. Co-occurrence of the three CYPs in early chordates suggests that the three genes may have coevolved thereafter, and that functional conservation should be reflected in functionally important residues in the proteins. CYP19A1 has the largest number of conserved residues while CYP11A1 sequences are less conserved. Structural analyses of human CYP11A1, CYP17A1 and CYP19A1 show that critical substrate binding site residues are highly conserved in each enzyme family. The results emphasize that the steroidogenic pathways producing glucocorticoids and reproductive steroids are several hundred million years old and that the catalytic structural elements of the enzymes have been conserved over the same period of time. Analysis of these elements may help to identify when precursor functions linked to these enzymes first arose.