The spliceosome is the complex molecular machinery that sequentially assembles on eukaryotic messenger RNA precursors to remove introns (pre-mRNA splicing), a physiologically regulated process altered in numerous pathologies. We report transcriptome-wide analyses upon systematic knock down of 305 spliceosome components and regulators in human cancer cells and the reconstruction of functional splicing factor networks that govern different classes of alternative splicing decisions. The results disentangle intricate circuits of splicing factor cross-regulation, reveal that the precise architecture of late-assembling U4/U6.U5 tri-small nuclear ribonucleoprotein (snRNP) complexes regulates splice site pairing, and discover an unprecedented division of labor among protein components of U1 snRNP for regulating exon definition and alternative 5' splice site selection. Thus, we provide a resource to explore physiological and pathological mechanisms of splicing regulation.
Understanding the pathogenic mechanisms of SF3B1 mutation can help unravel their contribution in patients’ worse prognosis. The cytotoxic effects and delayed leukemic infiltration induced by H3B-8800 support the potential use of SF3B1 inhibitors as a novel therapy in CLL. Splicing factor 3B subunit 1 (SF3B1) is involved in pre-mRNA branch site recognition and is the target of antitumor-splicing inhibitors. Mutations in SF3B1 are observed in 15% of patients with chronic lymphocytic leukemia (CLL) and are associated with poor prognosis, but their pathogenic mechanisms remain poorly understood. Using deep RNA-sequencing data from 298 CLL tumor samples and isogenic SF3B1 WT and K700E-mutated CLL cell lines, we characterize targets and pre-mRNA sequence features associated with the selection of cryptic 3′ splice sites upon SF3B1 mutation, including an event in the MAP3K7 gene relevant for activation of NF-κB signaling. Using the H3B-8800 splicing modulator, we show, for the first time in CLL, cytotoxic effects in vitro in primary CLL samples and in SF3B1 -mutated isogenic CLL cell lines, accompanied by major splicing changes and delayed leukemic infiltration in a CLL xenotransplant mouse model. H3B-8800 displayed preferential lethality towards SF3B1 -mutated cells and synergism with the BCL2 inhibitor venetoclax, supporting the potential use of SF3B1 inhibitors as a novel therapeutic strategy in CLL.
Although splicing occurs largely co-transcriptionally, the order by which introns are removed does not necessarily follow the order in which they are transcribed. Whereas several genomic features are known to influence whether or not an intron is spliced before its downstream neighbor, multiple questions related to adjacent introns' splicing order (AISO) remain unanswered. Here, we present Insplico, the first standalone software for quantifying AISO that works with both short and long read sequencing technologies. We first demonstrate its applicability and effectiveness using simulated reads and by recapitulating previously reported AISO patterns, which unveiled overlooked biases associated with long read sequencing. We next show that AISO around individual exons is remarkably constant across cell and tissue types and even upon major spliceosomal disruption, and it is evolutionarily conserved between human and mouse brains. We also establish a set of universal features associated with AISO patterns across various animal and plant species. Finally, we used Insplico to investigate AISO in the context of tissue-specific exons, particularly focusing on SRRM4-dependent microexons. We found that the majority of such microexons have non-canonical AISO, in which the downstream intron is spliced first, and we suggest two potential modes of SRRM4 regulation of microexons related to their AISO and various splicing-related features. Insplico is available on gitlab.com/aghr/insplico.
Transition from maternal to embryonic transcriptional control is crucial for embryogenesis. However, alternative splicing regulation during this process remains understudied. Using transcriptomic data from human, mouse, and cow preimplantation development, we show that the stage of zygotic genome activation (ZGA) exhibits the highest levels of exon skipping diversity reported for any cell or tissue type. Much of this exon skipping is temporary, leads to disruptive noncanonical isoforms, and occurs in genes enriched for DNA damage response in the three species. Two core spliceosomal components, Snrpb and Snrpd2, regulate these patterns. These genes have low maternal expression at ZGA and increase sharply thereafter. Microinjection of Snrpb/d2 messenger RNA into mouse zygotes reduces the levels of exon skipping at ZGA and leads to increased p53-mediated DNA damage response. We propose that mammalian embryos undergo an evolutionarily conserved, developmentally programmed splicing failure at ZGA that contributes to the attenuation of cellular responses to DNA damage.
Understanding the regulatory interactions that control gene expression during the development of novel tissues is a key goal of evolutionary developmental biology. Here, we show that Mbnl3 has undergone a striking process of evolutionary specialization in eutherian mammals resulting in the emergence of a novel placental function for the gene. Mbnl3 belongs to a family of RNA-binding proteins whose members regulate multiple aspects of RNA metabolism. We find that, in eutherians, while both Mbnl3 and its paralog Mbnl2 are strongly expressed in placenta, Mbnl3 expression has been lost from nonplacental tissues in association with the evolution of a novel promoter. Moreover, Mbnl3 has undergone accelerated protein sequence evolution leading to changes in its RNA-binding specificities and cellular localization. While Mbnl2 and Mbnl3 share partially redundant roles in regulating alternative splicing, polyadenylation site usage and, in turn, placenta maturation, Mbnl3 has also acquired novel biological functions. Specifically, Mbnl3 knockout (M3KO) alone results in increased placental growth associated with higher Myc expression. Furthermore, Mbnl3 loss increases fetal resource allocation during limiting conditions, suggesting that location of Mbnl3 on the X chromosome has led to its role in limiting placental growth, favoring the maternal side of the parental genetic conflict.
Alternative splicing (AS) can vastly expand animal transcriptomes and proteomes. Two main open questions in the field are how AS is regulated across cell/tissue types and disease, and what roles different AS events play. To facilitate AS research, we have created the computational VastDB framework, which comprises a series of complementary software and resources that we describe in this chapter. The VastDB framework is especially designed to aid biomedical researchers without a strong computational background. It offers tools and resources to: (a) quantify AS and identify differentially spliced AS events using RNA-seq data (vast-tools), (b) perform multiple genomic and sequence analyses for investigating AS events (Matt), (c) identify AS events with genomic and regulatory conservation among species (ExOrthist), and (d) help with the biological interpretation of the results, and, ultimately, with the identification of interesting AS events to design wet-lab experiments (VastDB and PastDB).
Summary:Tracking thousands of alternative splicing (AS) events genome-wide makes their downstream analysis computationally challenging and laborious. Here, we present Matt, the first UNIX command-line toolkit with focus on high-level AS analyses. With 50 commands it facilitates computational AS analyses by (i) expediting repetitive data-preparation tasks, (ii) offering routine high-level analyses, including the extraction of exon/intron features, discriminative feature detection, motif enrichment analysis, and the generation of motif RNA-maps, (iii) improving reproducibility by documenting all analysis steps and (iv) accelerating the implementation of own analysis pipelines by offering users to exploit its modular functionality.Availability and implementation:matt.crg.eu under GNU LGPLv3, together with comprehensive documentation and application examples. Matt is implemented in Perl and R, invokes pdfLATEX and depends only on Perl Core modules/the R Base package simplifying its installation.Supplementary information:Supplementary data are available at Bioinformatics online.
Pre-mRNA splicing is a critical step of gene expression in eukaryotes. Transcriptome-wide splicing patterns are complex and primarily regulated by a diverse set of recognition elements and associated RNA-binding proteins. The retention and splicing (RES) complex is formed by three different proteins (Bud13p, Pml1p and Snu17p) and is involved in splicing in yeast. However, the importance of the RES complex for vertebrate splicing, the intronic features associated with its activity, and its role in development are unknown. In this study, we have generated loss-of-function mutants for the three components of the RES complex in zebrafish and showed that they are required during early development. The mutants showed a marked neural phenotype with increased cell death in the brain and a decrease in differentiated neurons. Transcriptomic analysis of bud13, snip1 (pml1) and rbmx2 (snu17) mutants revealed a global defect in intron splicing, with strong mis-splicing of a subset of introns. We found these RES-dependent introns were short, rich in GC and flanked by GC depleted exons, all of which are features associated with intron definition. Using these features, we developed and validated a predictive model that classifies RES dependent introns. Altogether, our study uncovers the essential role of the RES complex during vertebrate development and provides new insights into its function during splicing.
Recursive splicing (RS) starts by defining an "RS-exon,'' which is then spliced to the preceding exon, thus creating a recursive 5' splice site (RS-5ss). Previous studies focused on cryptic RS-exons, and now we find that the exon junction complex (EJC) represses RS of hundreds of annotated, mainly constitutive RS-exons. The core EJC factors, and the peripheral factors PNN and RNPS1, maintain RS-exon inclusion by repressing spliceosomal assembly on RS-5ss. The EJC also blocks 5ss located near exon-exon junctions, thus repressing inclusion of cryptic microexons. The prevalence of annotated RS-exons is high in deuterostomes, while the cryptic RS-exons are more prevalent in Drosophila, where EJC appears less capable of repressing RS. Notably, incomplete repression of RS also contributes to physiological alternative splicing of several human RS-exons. Finally, haploinsufficiency of the EJC factor Magoh in mice is associated with skipping of RS-exons in the brain, with relevance to the microcephaly phenotype and human diseases.
Alterations in gene regulation are considered major driving forces in divergent evolution. This is reflected in different species by the variable architecture of regulatory networks controlling highly conserved metabolic pathways. While many regulatory proteins are surprisingly conserved their wiring has evolved more rapidly. This project focuses on the adaptation to nutrient limitation, which requires the activation of the conserved AMP-activated protein kinase (AMPK alias Snf1 in yeast) and its downstream effectors. The goal is to uncover basic principles of adaptation and steps in the evolutionary process associated with regulatory network rearrangement. This requires improving the prediction of gene regulation based experimental data, DNA sequence information and information theory. In this project Context Tree (CT) models and Parsimonious Context Tree (PCT) models and the corresponding algorithms for extended Context Tree Maximization (CTM) and extended Parsimonious Context Tree Maximization (PCTM) are derived, implemented, and applied. Computational predictions and experimental validation will establish an iterative cycle to improve algorithms in each cycle leading to a growing set of experimentally verified and falsified predictions, finally allowing a deeper understanding of the evolution of the transcriptional regulatory network controlling energy metabolism, one of the most fundamental processes conserved across all kingdoms of life.
Epithelial-mesenchymal interactions are crucial for the development of numerous animal structures. Thus, unraveling how molecular tools are recruited in different lineages to control interplays between these tissues is key to understanding morphogenetic evolution. Here, we study Esrp genes, which regulate extensive splicing programs and are essential for mammalian organogenesis. We find that Esrp homologs have been independently recruited for the development of multiple structures across deuterostomes. Although Esrp is involved in a wide variety of ontogenetic processes, our results suggest ancient roles in non-neural ectoderm and regulating specific mesenchymal-to-epithelial transitions in deuterostome ancestors. However, consistent with the extensive rewiring of Esrp -dependent splicing programs between phyla, most developmental defects observed in vertebrate mutants are related to other types of morphogenetic processes. This is likely connected to the origin of an event in Fgfr , which was recruited as an Esrp target in stem chordates and subsequently co-opted into the development of many novel traits in vertebrates.
Several splicing-modulating compounds, including Sudemycins and Spliceostatin A, display anti-tumor properties. Combining transcriptome, bioinformatic and mutagenesis analyses, we delineate sequence determinants of the differential sensitivity of 3' splice sites to these drugs. Sequences 5' from the branch point (BP) region strongly influence drug sensitivity, with additional functional BPs reducing, and BP-like sequences allowing, drug responses. Drug-induced retained introns are typically shorter, displaying higher GC content and weaker polypyrimidine-tracts and BPs. Drug-induced exon skipping preferentially affects shorter alternatively spliced regions with weaker BPs. Remarkably, structurally similar drugs display both common and differential effects on splicing regulation, SSA generally displaying stronger effects on intron retention, and Sudemycins more acute effects on exon skipping. Collectively, our results illustrate how splicing modulation is exquisitely sensitive to the sequence context of 3' splice sites and to small structural differences between drugs.
Alternative splicing (AS) generates remarkable regulatory and proteomic complexity in metazoans. However, the functions of most AS events are not known, and programs of regulated splicing remain to be identified. To address these challenges, we describe the Vertebrate Alternative Splicing and Transcription Database (VastDB), the largest resource of genome-wide, quantitative profiles of AS events assembled to date. VastDB provides readily accessible quantitative information on the inclusion levels and functional associations of AS events detected in RNA-seq data from diverse vertebrate cell and tissue types, as well as developmental stages. The VastDB profiles reveal extensive new intergenic and intragenic regulatory relationships among different classes of AS and previously unknown and conserved landscapes of tissue-regulated exons. Contrary to recent reports concluding that nearly all human genes express a single major isoform, VastDB provides evidence that at least 48% of multiexonic protein-coding genes express multiple splice variants that are highly regulated in a cell/tissue-specific manner, and that >18% of genes simultaneously express multiple major isoforms across diverse cell and tissue types. Isoforms encoded by the latter set of genes are generally coexpressed in the same cells and are often engaged by translating ribosomes. Moreover, they are encoded by genes that are significantly enriched in functions associated with transcriptional control, implying they may have an important and wide-ranging role in controlling cellular activities. VastDB thus provides an unprecedented resource for investigations of AS function and regulation.
Cellular responses to starvation are of ancient origin since nutrient limitation has always been a common challenge to the stability of living systems. Hence, signaling molecules involved in sensing or transducing information about limiting metabolites are highly conserved, whereas transcription factors and the genes they regulate have diverged. In eukaryotes the AMP-activated protein kinase (AMPK) functions as a central regulator of cellular energy homeostasis. The yeast AMPK ortholog SNF1 controls the transcriptional network that counteracts carbon starvation conditions by regulating a set of transcription factors. Among those Cat8 and Sip4 have overlapping DNA-binding specificity for so-called carbon source responsive elements and induce target genes upon SNF1 activation. To analyze the evolution of the Cat8-Sip4 controlled transcriptional network we have compared the response to carbon limitation of Saccharomyces cerevisiae to that of Kluyveromyces lactis. In high glucose, S. cerevisiae displays tumor cell-like aerobic fermentation and repression of respiration (Crabtree-positive) while K. lactis has a respiratory-fermentative life-style, respiration being regulated by oxygen availability (Crabtree-negative), which is typical for many yeasts and for differentiated higher cells. We demonstrate divergent evolution of the Cat8-Sip4 network and present evidence that a role of Sip4 in controlling anabolic metabolism has been lost in the Saccharomyces lineage. We find that in K. lactis, but not in S. cerevisiae, the Sip4 protein plays an essential role in C2 carbon assimilation including induction of the glyoxylate cycle and the carnitine shuttle genes. Induction of KlSIP4 gene expression by KlCat8 is essential under these growth conditions and a primary function of KlCat8. Both KlCat8 and KlSip4 are involved in the regulation of lactose metabolism in K. lactis. In chromatin-immunoprecipitation experiments we demonstrate binding of both, KlSip4 and KlCat8, to selected CSREs and provide evidence that KlSip4 counteracts KlCat8-mediated transcription activation by competing for binding to some but not all CSREs. The finding that the hierarchical relationship of these transcription factors differs between K. lactis and S. cerevisiae and that the sets of target genes have diverged contributes to explaining the phenotypic differences in metabolic life-style.
The meaningful correlation of sensory data with analytical data is one of the most challenging tasks in flavor research. In beef stocks in particular, due to the presence of low levels of aroma-active compounds and the taste contribution of non-volatile molecules to the typical “juiciness” character, the consumer encounters a complex matrix situation. The goal of our study was to carry out a comprehensive analysis of all relevant flavor molecules and the correlation to human sensory data. A technique recently developed at the IPB and termed “reverse metabolomics” was used to link biological activity (i.e., sensory data) with variations in the metabolic profile (i.e., analytical data). We used this methodology for the first time to correlate sensorial attributes and GC-MS, LC-MS, and NMR data in culinary beef stocks. Reverse metabolomics was applied to study the link between sensory and chemical composition in a series of freshly prepared culinary beef stocks. A set of 10 different beef stocks was prepared. The degree of liking of the samples was recorded on a hedonic 1–9 scale. Analysis of the stocks was performed by LC-MS, GC-MS, and NMR. 1H-NMR data directly obtained from the meat stock were very complex. Analysis of this dataset by reverse metabolomics revealed some basic structural elements of the key taste compounds, such as carnosine or anserine. The reverse metabolomics correlation of quantitative data with partiality revealed the importance of a set of compounds. This relevance of these compounds has been confirmed by additional sensory experiments which showed an increase in perceived juiciness.
The binding affinity of DNA-binding proteins such as transcription factors is mainly determined by the base composition of the corresponding binding site on the DNA strand. Most proteins do not bind only a single sequence, but rather a set of sequences, which may be modeled by a sequence motif. Algorithms for de novo motif discovery differ in their promoter models, learning approaches, and other aspects, but typically use the statistically simple position weight matrix model for the motif, which assumes statistical independence among all nucleotides. However, there is no clear justification for that assumption, leading to an ongoing debate about the importance of modeling dependencies between nucleotides within binding sites. In the past, modeling statistical dependencies within binding sites has been hampered by the problem of limited data. With the rise of high-throughput technologies such as ChIP-seq, this situation has now changed, making it possible to make use of statistical dependencies effectively. In this work, we investigate the presence of statistical dependencies in binding sites of the human enhancer-blocking insulator protein CTCF by using the recently developed model class of inhomogeneous parsimonious Markov models, which is capable of modeling complex dependencies while avoiding overfitting. These findings lead to a more detailed characterization of the CTCF binding motif, which is only poorly represented by independent nucleotide frequencies at several positions, predominantly at the 3 ' end.
We introduce inhomogeneous parsimonious Markov models for modeling statistical patterns in discrete sequences. These models are based on parsimonious context trees, which are a generalization of context trees, and thus generalize variable order Markov models. We follow a Bayesian approach, consisting of structure and parameter learning. Structure learning is a challenging problem due to an overexponential number of possible tree structures, so we describe an exact and efficient dynamic programming algorithm for finding the optimal tree structures. We apply model and learning algorithm to the problem of modeling binding sites of the human transcription factor C/EBP, and find an increased prediction performance compared to fixed order and variable order Markov models. We investigate the reason for this improvement and find several instances of context-specific dependences that can be captured by parsimonious context trees but not by traditional context trees.
We propose a visualization technique for summarizing contents of document streams, such as news or scientific archives. The content of streaming documents change over time and so do themes the documents are about. Topic evolution is a relatively new research subject that encompasses the unsupervised discovery of thematic subjects in a document collection and the adaptation of these subjects as new documents arrive. While many powerful topic evolution methods exist, the combination of learning and visualization of the evolving topics has been less explored, although it is indispensable for understanding a dynamic document collection. We propose Topic Table, a visualization technique that builds upon topic modeling for deriving a condensed representation of a document collection. Topic Table captures important and intuitively comprehensible aspects of a topic over time: the importance of the topic within the collection, the words characterizing this topic, the semantic changes of a topic from one timepoint to the next. As an example, we visualize content of the NIPS proceedings from 1987 to 1999.
Dvir Aran Rolf Backofen Sandra Baldauf Mukul Bansal Robert Beiko Asa Ben-Hur Bonnie Berger Christoph Bock Mikael Boden Yana Bromberg Sharon Browning Emidio Capriotti Minh Cao Ho-Ryun Chung Komalapriya Chandrasekaran Charlotte Dean Frank Dudbridge Danny Durand Nadia El-Mabrouk Arne Elofsson Arcadio-Rubio Garcia Andre Gohr Ivo Grosse Asaf Hellman Paul Horton Daniel Huson Sorin Istrail Shalev Itzkovitz Asif Javed Noam Kaplan Tommy Kaplan Szymon Kielbasa Roy Kishony Mehmet Koyuturk Jens Lagergren Wei Li Mangus Liden Jun Ma Veli Mäkinen Serghei Mangul Gabriel Margarido Debora Marks Yosi Maruvka Alice McHardy Bernard Moret Alexandre Morozov Kenta Nakai Zoran Nikoloski Lior Pachter Osnat Penn Martin Porsch Dariusz Przybylski Jörg Rahnenführer
Stefan Posch合作论文数Institut f?r Informatik;Martin-Luther-Universit?t Halle Wittenberg6
Marc Strickert合作论文数Institute of Plant Genetics and Crop Plant Research4