The UCSC Genome Browser (https://genome.ucsc.edu) is a widely utilized web-based tool for visualization and analysis of genomic data, encompassing over 4000 assemblies from diverse organisms. Since its release in 2001, it has become an essential resource for genomics and bioinformatics research. Annotation data available on Genome Browser includes both internally created and maintained tracks as well as custom tracks and track hubs provided by the research community. This last year's updates include over 25 new annotation tracks such as the gnomAD 4.1 track on the human GRCh38/hg38 assembly, the addition of three new public hubs, and significant expansions to the Genome Archive[GenArk) system for interacting with the enormous variety of assemblies. We have also made improvements to our interface, including updates to the browser graphic page, such as a new popup dialog feature that now displays item details without requiring navigation away from the main Genome Browser page. GenePred tracks have been upgraded with right-click options for zooming and precise navigation, along with enhanced mouseOver functions. Additional improvements include a new grouping feature for track hubs and hub description info links. A new tutorial focusing on Clinical Genetics has also been added to the UCSC Genome Browser.
The accurate identification and quantitation of RNA isoforms present in the cancer transcriptome is key for analyses ranging from the inference of the impacts of somatic variants to pathway analysis to biomarker development and subtype discovery. The ICGC-TCGA DREAM Somatic Mutation Calling in RNA (SMC-RNA) challenge was a crowd-sourced effort to benchmark methods for RNA isoform quantification and fusion detection from bulk cancer RNA sequencing (RNA-seq) data. It concluded in 2018 with a comparison of 77 fusion detection entries and 65 isoform quantification entries on 51 synthetic tumors and 32 cell lines with spiked-in fusion constructs. We report the entries used to build this benchmark, the leaderboard results, and the experimental features associated with the accurate prediction of RNA species. This challenge required submissions to be in the form of containerized workflows, meaning each of the entries described is easily reusable through CWL and Docker containers at https://github.com/SMC-RNA-challenge. A record of this paper's transparent peer review process is included in the supplemental information.
While splicing changes caused by somatic mutations in SF3B1 are known, identifying full-length isoform changes may better elucidate the functional consequences of these mutations. We report nanopore sequencing of full-length cDNA from CLL samples with and without SF3B1 mutation, as well as normal B cell samples, giving a total of 149 million pass reads. We present FLAIR (Full-Length Alternative Isoform analysis of RNA), a computational workflow to identify high-confidence transcripts, perform differential splicing event analysis, and differential isoform analysis. Using nanopore reads, we demonstrate differential 3' splice site changes associated with SF3B1 mutation, agreeing with previous studies. We also observe a strong downregulation of intron retention events associated with SF3B1 mutation. Full-length transcript analysis links multiple alternative splicing events together and allows for better estimates of the abundance of productive versus unproductive isoforms. Our work demonstrates the potential utility of nanopore sequencing for cancer and splicing research.
Micromonas is a unicellular marine green alga that thrives from tropical to polar ecosystems. We investigated the growth and cellular characteristics of acclimated mid-exponential phase Micromonas commoda RCC299 over multiple light levels and over the diel cycle (14:10 hour light:dark). We also exposed the light:dark acclimated M. commoda to experimental shifts from moderate to high light (HL), and to HL plus ultraviolet radiation (HL+UV), 4.5 hours into the light period. Cellular responses of this prasinophyte were quantified by flow cytometry and changes in gene expression by qPCR and RNA-seq. While proxies for chlorophyll a content and cell size exhibited similar diel variations in HL and controls, with progressive increases during day and decreases at night, both parameters sharply decreased after the HL+UV shift. Two distinct transcriptional responses were observed among chloroplast genes in the light shift experiments: i) expression of transcription and translation-related genes decreased over the time course, and this transition occurred earlier in treatments than controls; ii) expression of several photosystem I and II genes increased in HL relative to controls, as did the growth rate within the same diel period. However, expression of these genes decreased in HL+UV, likely as a photoprotective mechanism. RNA-seq also revealed two genes in the chloroplast genome, ycf2-like and ycf1-like, that had not previously been reported. The latter encodes the second largest chloroplast protein in Micromonas and has weak homology to plant Ycf1, an essential component of the plant protein translocon. Analysis of several nuclear genes showed that the expression of LHCSR2, which is involved in non-photochemical quenching, and five light-harvesting-like genes, increased 30 to >50-fold in HL+UV, but was largely unchanged in HL and controls. Under HL alone, a gene encoding a novel nitrite reductase fusion protein (NIRFU) increased, possibly reflecting enhanced N-assimilation under the 625 μmol photons m-2 s-1 supplied in the HL treatment. NIRFU’s domain structure suggests it may have more efficient electron transfer than plant NIR proteins. Our analyses indicate that Micromonas can readily respond to abrupt environmental changes, such that strong photoinhibition was provoked by combined exposure to HL and UV, but a ca. 6-fold increase in light was stimulatory.
BACKGROUND:Prasinophytes are widespread marine green algae that are related to plants. Cellular abundance of the prasinophyte Micromonas has reportedly increased in the Arctic due to climate-induced changes. Thus, studies of these unicellular eukaryotes are important for marine ecology and for understanding Viridiplantae evolution and diversification.RESULTS:We generated evidence-based Micromonas gene models using proteomics and RNA-Seq to improve prasinophyte genomic resources. First, sequences of four chromosomes in the 22 Mb Micromonas pusilla (CCMP1545) genome were finished. Comparison with the finished 21 Mb genome of Micromonas commoda (RCC299; named herein) shows they share ≤8,141 of ~10,000 protein-encoding genes, depending on the analysis method. Unlike RCC299 and other sequenced eukaryotes, CCMP1545 has two abundant repetitive intron types and a high percent (26 %) GC splice donors. Micromonas has more genus-specific protein families (19 %) than other genome sequenced prasinophytes (11 %). Comparative analyses using predicted proteomes from other prasinophytes reveal proteins likely related to scale formation and ancestral photosynthesis. Our studies also indicate that peptidoglycan (PG) biosynthesis enzymes have been lost in multiple independent events in select prasinophytes and plants. However, CCMP1545, polar Micromonas CCMP2099 and prasinophytes from other classes retain the entire PG pathway, like moss and glaucophyte algae. Surprisingly, multiple vascular plants also have the PG pathway, except the Penicillin-Binding Protein, and share a unique bi-domain protein potentially associated with the pathway. Alongside Micromonas experiments using antibiotics that halt bacterial PG biosynthesis, the findings highlight unrecognized phylogenetic complexity in PG-pathway retention and implicate a role in chloroplast structure or division in several extant Viridiplantae lineages.CONCLUSIONS:Extensive differences in gene loss and architecture between related prasinophytes underscore their divergence. PG biosynthesis genes from the cyanobacterial endosymbiont that became the plastid, have been selectively retained in multiple plants and algae, implying a biological function. Our studies provide robust genomic resources for emerging model algae, advancing knowledge of marine phytoplankton and plant evolution.
Micromonas is a unicellular motile alga within the Prasinophyceae, a green algal group that is related to land plants. This picoeukaryote (<2 μm diameter) is widespread in the marine environment but is not well understood at the cellular level. Here, we examine shifts in mRNA and protein expression over the course of the day-night cycle using triplicated mid-exponential, nutrient replete cultures of Micromonas pusilla CCMP1545. Samples were collected at key transition points during the diel cycle for evaluation using high-throughput LC-MS proteomics. In conjunction, matched mRNA samples from the same time points were sequenced using pair-ended directional Illumina RNA-Seq to investigate the dynamics and relationship between the mRNA and protein expression programs of M. pusilla. Similar to a prior study of the marine cyanobacterium Prochlorococcus, we found significant divergence in the mRNA and proteomics expression dynamics in response to the light:dark cycle. Additionally, expressional responses of genes and the proteins they encoded could also be variable within the same metabolic pathway, such as we observed in the oxygenic photosynthesis pathway. A regression framework was used to predict protein levels from both mRNA expression and gene-specific sequence-based features. Several features in the genome sequence were found to influence protein abundance including codon usage as well as 3’ UTR length and structure. Collectively, our studies provide insights into the regulation of the proteome over a diel cycle as well as the relationships between transcriptional and translational programs in the widespread marine green alga Micromonas.
Spliceosomal introns are a hallmark of eukaryotic genes that are hypothesized to play important roles in genome evolution but have poorly understood origins. Although most introns lack sequence homology to each other, new families of spliceosomal introns that are repeated hundreds of times in individual genomes have recently been discovered in a few organisms. The prevalence and conservation of these introner elements (IEs) or introner-like elements in other taxa, as well as their evolutionary relationships to regular spliceosomal introns, are still unknown. Here, we systematically investigate introns in the widespread marine green alga Micromonas and report new families of IEs, numerous intron presence-absence polymorphisms, and potential intron insertion hot-spots. The new families enabled identification of conserved IE secondary structure features and establishment of a novel general model for repetitive intron proliferation across genomes. Despite shared secondary structure, the IE families from each Micromonas lineage bear no obvious sequence similarity to those in the other lineages, suggesting that their appearance is intimately linked with the process of speciation. Two of the new IE families come from an Arctic culture (Micromonas Clade E2) isolated from a polar region where abundance of this alga is increasing due to climate induced changes. The same two families were detected in metagenomic data from Antarctica-a system where Micromonas has never before been reported. Strikingly high identity between the Arctic isolate and Antarctic coding sequences that flank the IEs suggests connectivity between populations in the two polar systems that we postulate occurs through deep-sea currents. Recovery of Clade E2 sequences in North Atlantic Deep Waters beneath the Gulf Stream supports this hypothesis. Our research illuminates the dynamic relationships between an unusual class of repetitive introns, genome evolution, speciation, and global distribution of this sentinel marine alga.
Significance Phytochromes are photosensory signaling proteins widely distributed in unicellular organisms and multicellular land plants. Best known for their global regulatory roles in photomorphogenesis, plant phytochromes are often assumed to have arisen via gene transfer from the cyanobacterial endosymbiont that gave rise to photosynthetic chloroplast organelles. Our analyses support the scenario that phytochromes were acquired prior to diversification of the Archaeplastida, possibly before the endosymbiosis event. We show that plant phytochromes are structurally and functionally related to those discovered in prasinophytes, an ecologically important group of marine green algae. Based on our studies, we propose that these phytochromes share light-mediated signaling mechanisms with those of plants. Phytochromes presumably perform critical acclimative roles for unicellular marine algae living in fluctuating light environments.
Premise of the study: We developed and tested primers for 218 nuclear loci for studying population genetics, phylogeography, and genome evolution in bryophytes. Methods and Results: We aligned expressed sequence tags (ESTs) from Ceratodon purpureus to the Physcomitrella patens genome sequence, and designed primers that are homologous to conserved exons but span introns in the P. patens genome. We tested these primers on four isolates from New York, USA; Otavalo, Ecuador; and two laboratory isolates from Austria (WT4 and GG1). The median genome-wide nucleotide diversity was 0.008 substitutions/site, but the range was large (0–0.14), illustrating the among-locus heterogeneity in the species. Conclusions: These loci provide a valuable resource for finely resolved, genome-wide population genetic and species-level phylogenetic analyses of C. purpureus and its relatives.
To gain insight into how genomic information is translated into cellular and developmental programs, the Drosophila model organism Encyclopedia of DNA Elements (modENCODE) project is comprehensively mapping transcripts, histone modifications, chromosomal proteins, transcription factors, replication proteins and intermediates, and nucleosome properties across a developmental time course and in multiple cell lines. We have generated more than 700 data sets and discovered protein-coding, noncoding, RNA regulatory, replication, and chromatin elements, more than tripling the annotated portion of the Drosophila genome. Correlated activity patterns of these elements reveal a functional regulatory network, which predicts putative new functions for genes, reveals stage- and tissue-specific regulators, and enables gene-expression prediction. Our results provide a foundation for directed experimental and computational studies in Drosophila and related species and also a model for systematic data integration toward comprehensive genomic and functional annotation.
Drosophila melanogaster cell lines are important resources for cell biologists. Here, we catalog the expression of exons, genes, and unannotated transcriptional signals for 25 lines. Unannotated transcription is substantial (typically 19% of euchromatic signal). Conservatively, we identify 1405 novel transcribed regions; 684 of these appear to be new exons of neighboring, often distant, genes. Sixty-four percent of genes are expressed detectably in at least one line, but only 21% are detected in all lines. Each cell line expresses, on average, 5885 genes, including a common set of 3109. Expression levels vary over several orders of magnitude. Major signaling pathways are well represented: most differentiation pathways are "off" and survival/growth pathways "on." Roughly 50% of the genes expressed by each line are not part of the common set, and these show considerable individuality. Thirty-one percent are expressed at a higher level in at least one cell line than in any single developmental stage, suggesting that each line is enriched for genes characteristic of small sets of cells. Most remarkable is that imaginal disc-derived lines can generally be assigned, on the basis of expression, to small territories within developing discs. These mappings reveal unexpected stability of even fine-grained spatial determination. No two cell lines show identical transcription factor expression. We conclude that each line has retained features of an individual founder cell superimposed on a common "cell line" gene expression pattern.
RNA-Seq enables rapid sequencing of total cellular RNA and should allow the reconstruction of spliced transcripts in a cell population. Trapnell et al. achieve this and transcript quantification using only paired-end RNA-Seq data and an unannotated genome sequence, and apply the approach to characterize isoform switching over a developmental time course. High-throughput mRNA sequencing (RNA-Seq) promises simultaneous transcript discovery and abundance estimation1,2,3. However, this would require algorithms that are not restricted by prior gene annotations and that account for alternative transcription and splicing. Here we introduce such algorithms in an open-source software program called Cufflinks. To test Cufflinks, we sequenced and analyzed >430 million paired 75-bp RNA-Seq reads from a mouse myoblast cell line over a differentiation time series. We detected 13,692 known transcripts and 3,724 previously unannotated ones, 62% of which are supported by independent expression data or by homologous genes in other species. Over the time series, 330 genes showed complete switches in the dominant transcription start site (TSS) or splice isoform, and we observed more subtle shifts in 1,304 other genes. These results suggest that Cufflinks can illuminate the substantial regulatory flexibility and complexity in even this well-studied model of muscle development and that it can improve transcriptome-based genome annotation.
Access the most recent version at doi: published online December 22, 2010 Genome Res. Lucy Cherbas, Aarron Willingham, Dayu Zhang, et al. cell lines Drosophila The transcriptional diversity of 25 Material Supplemental http://genome.cshlp.org/content/suppl/2010/12/02/gr.112961.110.DC1.html P<P Published online December 22, 2010 in advance of the print journal. Open Access Freely available online through the Genome Research Open Access option.
Since its start, the Mammalian Gene Collection (MGC) has sought to provide at least one full-protein-coding sequence cDNA clone for every human and mouse gene with a RefSeq transcript, and at least 6200 rat genes. The MGC cloning effort initially relied on random expressed sequence tag screening of cDNA libraries. Here, we summarize our recent progress using directed RT-PCR cloning and DNA synthesis. The MGC now contains clones with the entire protein-coding sequence for 92% of human and 89% of mouse genes with curated RefSeq (NM-accession) transcripts, and for 97% of human and 96% of mouse genes with curated RefSeq transcripts that have one or more PubMed publications, in addition to clones for more than 6300 rat genes. These high-quality MGC clones and their sequences are accessible without restriction to researchers worldwide.
N-SCAN is a gene-prediction system that combines the methods of ab initio predictors like GENSCAN with information derived from genome comparison. It is the latest in the TWINSCAN series of programs. This unit describes the use of N-SCAN to identify gene structures in eukaryotic genomic sequences. Protocols for using N-SCAN through its Web interface and from the command line in a Linux environment are provided. Detailed discussion about the appropriate parameter settings, input-sequence processing, and choice of genome for comparison are included.