When arthropod-borne viruses (arboviruses) are delivered to vector mosquitoes in infectious bloodmeals, viral components interact with host proteins to hijack cells and initiate replication. The extent to which arbovirus infection alters mosquito host transcriptional regulatory processes is currently unknown. We hypothesized that histone modifications would be altered in mosquitoes exposed to Rift Valley fever virus (RVFV MP12). H3K27ac and H3K9me3 marks were interrogated using CUT&RUN in a mosquito species that has a predicted dissemination barrier, Aedes aegypti. Global H3K27ac peaks showed progressive depletion over time compared to bloodfed controls. Gene set enrichment analysis revealed that immune response transcripts were enriched at 1 and 3 days post-feeding (dpf). For virus-exposed samples, the highest proportion of DEGs proximal to histone marks occurred with depletion of repressive H3K9me3 peaks at 3 dpf. Associated DEGs included transcription factors, secondary messengers and processes affecting cell polarization. Analysis of midguts after a non-infectious bloodmeal versus sugar-fed controls revealed global changes to H3K27ac and H3K9me3 marks, as well. Differential H3K27ac marks were proximal to one quarter of all DEGs at 1 dpf, consistent with an important role of H3K27ac in bloodmeal digestion. Together, these results demonstrate that H3K27ac and H3K9me3 patterns are altered upon virus exposure in a complex interplay that could be due to viral manipulation or host defense.
ELT-2 is the major transcription factor (TF) required for Caenorhabditis elegans intestinal development. ELT-2 expression initiates in embryos to promote development and then persists after hatching through the larval and adult stages. Though the sites of ELT-2 binding are characterized and the transcriptional changes that result from ELT-2 depletion are known, an intestine-specific transcriptome profile spanning developmental time has been missing. We generated this dataset by performing Fluorescence Activated Cell Sorting on intestine cells at distinct developmental stages. We analyzed this dataset in conjunction with previously conducted ELT-2 studies to evaluate the role of ELT-2 in directing the intestinal gene regulatory network through development. We found that only 33% of intestine-enriched genes in the embryo were direct targets of ELT-2 but that number increased to 75% by the L3 stage. This suggests additional TFs promote intestinal transcription especially in the embryo. Furthermore, only half of ELT-2's direct target genes were dependent on ELT-2 for their proper expression levels, and an equal proportion of those responded to elt-2 depletion with over-expression as with under-expression. That is, ELT-2 can either activate or repress direct target genes. Additionally, we observed that ELT-2 repressed its own promoter, implicating new models for its autoregulation. Together, our results illustrate that ELT-2 impacts roughly 20-50% of intestine-specific genes, that ELT-2 both positively and negatively controls its direct targets, and that the current model of the intestinal regulatory network is incomplete as the factors responsible for directing the expression of many intestinal genes remain unknown.
Aneuploidy, the incorrect number of whole chromosomes, is a common feature of tumors that contributes to their initiation and evolution. Preventing aneuploidy requires properly functioning kinetochores, which are large protein complexes assembled on centromeric DNA that link mitotic chromosomes to dynamic spindle microtubules and facilitate chromosome segregation. The kinetochore leverages at least two mechanisms to prevent aneuploidy: error correction and the spindle assembly checkpoint (SAC). BubR1, a factor involved in both processes, was identified as a cancer dependency and therapeutic target in multiple tumor types; however, it remains unclear what specific oncogenic pressures drive this enhanced dependency on BubR1 and whether it arises from BubR1's regulation of the SAC or error-correction pathways. Here, we use a genetically controlled transformation model and glioblastoma tumor isolates to show that constitutive signaling by RAS or MAPK is necessary for cancer-specific BubR1 vulnerability. TheMAPK pathway enzymatically hyperstimulates a network of kinetochore kinases that compromises chromosome segregation, rendering cells more dependent on two BubR1 activities: counteracting excessive kinetochore-microtubule turnover for error correction and maintaining the SAC. This work expands our understanding of how chromosome segregation adapts to different cellular states and reveals an oncogenic trigger of a cancer-specific defect.
Low complexity domains (LCDs) in proteins are regions predominantly composed of a small subset of the possible amino acids. LCDs are involved in a variety of normal and pathological processes across all domains of life. Existing methods define LCDs using information-theoretical complexity thresholds, sequence alignment with repetitive regions, or statistical overrepresentation of amino acids relative to whole-proteome frequencies. While these methods have proven valuable, they are all indirectly quantifying amino acid composition, which is the fundamental and biologically-relevant feature related to protein sequence complexity. Here, we present a new computational tool, LCD-Composer, that directly identifies LCDs based on amino acid composition and linear amino acid dispersion. Using LCD-Composer's default parameters, we identified simple LCDs across all organisms available through UniProt and provide the resulting data in an accessible form as a resource. Furthermore, we describe large-scale differences between organisms from different domains of life and explore organisms with extreme LCD content for different LCD classes. Finally, we illustrate the versatility and specificity achievable with LCD-Composer by identifying diverse classes of LCDs using both simple and multifaceted composition criteria. We demonstrate that the ability to dissect LCDs based on these multifaceted criteria enhances the functional mapping and classification of LCDs.
Bunyaviruses (Negarnaviricota: Bunyavirales) are a large and diverse group of viruses that include important human, veterinary, and plant pathogens. The rapid characterization of known and new emerging pathogens depends on the availability of comprehensive reference sequence databases that can be used to match unknowns, infer evolutionary relationships and pathogenic potential, and make response decisions in an evidence-based manner. In this study, we determined the coding-complete genome sequences of 99 bunyaviruses in the Centers for Disease Control and Prevention’s Arbovirus Reference Collection, focusing on orthonairoviruses (family Nairoviridae), orthobunyaviruses (Peribunyaviridae), and phleboviruses (Phenuiviridae) that either completely or partially lacked genome sequences. These viruses had been collected over 66 years from 27 countries from vertebrates and arthropods representing 37 genera. Many of the viruses had been characterized serologically and through experimental infection of animals but were isolated in the pre-sequencing era. We took advantage of our unusually large sample size to systematically evaluate genomic characteristics of these viruses, including reassortment, and co-infection. We corroborated our findings using several independent molecular and virologic approaches, including Sanger sequencing of 197 genome segments, and plaque isolation of viruses from putative co-infected virus stocks. This study contributes to the described genetic diversity of bunyaviruses and will enhance the capacity to characterize emerging human pathogenic bunyaviruses.
*Department of Biochemistry and Molecular Biology, Colorado State University, Fort Collins, CO, USA 80523 † Authors contributed equally ‡Current affiliation – Experimental Immunology Branch, National Cancer Institute, US National Institutes of Health, Bethesda, MD, USA § Current affiliation – Center for Vector-Borne Infectious Diseases, Colorado State University, Fort Collins, CO, USA ** Corresponding author: erin.nishimura@colostate.edu
Part 2 of 2. Each file contains the proteome (in FASTA format) corresponding to a single organism. Files ending in "_additional" represent additional isoforms and translation products associated with that organism. All bacterial proteomes were downloaded from the UniProt FTP site (ftp://ftp.uniprot.org/pub/databases/uniprot/) on 8/21/2020 and are included here to ensure that LCDs correctly map to sequence locations. Please see the LICENSE_UniProt file for license information regarding UniProt sequence data.
Each file contains the proteome (in FASTA format) corresponding to a single organism. Files ending in "_additional" represent additional isoforms and translation products associated with that organism. All archaeal proteomes were downloaded from the UniProt FTP site (ftp://ftp.uniprot.org/pub/databases/uniprot/) on 8/21/2020 and are included here to ensure that LCDs correctly map to sequence locations. Please see the LICENSE_UniProt file for license information regarding UniProt sequence data.
LCD-Composer results are stored in a separate file for each organism. Columns are ordered as follows:1) The protein identifier (header in the FASTA proteome file),2) the LCD sequence,3) the location of the LCD within the protein,4) the LCD class (i.e. the amino acid of interest used in the LCD-Composer search criteria),5) the percent composition of the amino acid of interest within the LCD,6) the linear dispersion of the amino acid of interest within the LCD
This lightning talk will introduce biodiversity researchers to the sub-species of humans known as 'computer scientists'. This is a particularly fruitful area of study owing to the rapid speciation evident in the population. From the original off-shoot of mathematicians to today's bewildering variety of specialists, the talk aims to introduce the legacy population of Homo sapiens sapiens to the subtleties of interacting with their recently evolved cousins. Along the way, the talk will include other nuggets such as the difference between data mining (it is lucrative) and text mining (it is not). This talk should be of interest to someone. Note, the organisers of TDWG will probably want to make it clear that all views expressed are purely those of the presenter and yes the presenter's mother was a computer.
November 2011 - January 2013 marks the 200th anniversary of the Luddite uprisings in England: a great opportunity to celebrate their struggle and to redress the wrongs done to them.
The transcription factor GATA1 regulates an extensive program of gene activation and repression during erythroid development. However, the associated mechanisms, including the contributions of distal versus proximal cis-regulatory modules, co-occupancy with other transcription factors, and the effects of histone modifications, are poorly understood. We studied these problems genome-wide in a Gata1 knockout erythroblast cell line that undergoes GATA1-dependent terminal maturation, identifying 2616 GATA1-responsive genes and 15,360 GATA1-occupied DNA segments after restoration of GATA1. Virtually all occupied DNA segments have high levels of H3K4 monomethylation and low levels of H3K27me3 around the canonical GATA binding motif, regardless of whether the nearby gene is induced or repressed. Induced genes tend to be bound by GATA1 close to the transcription start site (most frequently in the first intron), have multiple GATA1-occupied segments that are also bound by TAL1, and show evolutionary constraint on the GATA1-binding site motif. In contrast, repressed genes are further away from GATA1-occupied segments, and a subset shows reduced TAL1 occupancy and increased H3K27me3 at the transcription start site. Our data expand the repertoire of GATA1 action in erythropoiesis by defining a new cohort of target genes and determining the spatial distribution of cis-regulatory modules throughout the genome. In addition, we begin to establish functional criteria and mechanisms that distinguish GATA1 activation from repression at specific target genes. More broadly, these studies illustrate how a "master regulator" transcription factor coordinates tissue differentiation through a panoply of DNA and protein interactions.
DNA sequence motifs and epigenetic modifications contribute to specific binding by a transcription factor, but the extent to which each feature determines occupancy in vivo is poorly understood. We addressed this question in erythroid cells by identifying DNA segments occupied by GATA1 and measuring the level of trimethylation of histone H3 lysine 27 (H3K27me3) and monomethylation of H3 lysine 4 (H3K4me1) along a 66 Mb region of mouse chromosome 7. While 91% of the GATA1-occupied segments contain the consensus binding-site motif WGATAR, only approximately 0.7% of DNA segments with such a motif are occupied. Using a discriminative motif enumeration method, we identified additional motifs predictive of occupancy given the presence of WGATAR. The specific motif variant AGATAA and occurrence of multiple WGATAR motifs are both strong discriminators. Combining motifs to pair a WGATAR motif with a binding site motif for GATA1, EKLF or SP1 improves discriminative power. Epigenetic modifications are also strong determinants, with the factor-bound segments highly enriched for H3K4me1 and depleted of H3K27me3. Combining primary sequence and epigenetic determinants captures 52% of the GATA1-occupied DNA segments and substantially increases the specificity, to one out of seven segments with the required motif combination and epigenetic signals being bound.
Material Supplemental http://genome.cshlp.org/content/suppl/2008/11/06/gr.083089.108.DC1.html References http://genome.cshlp.org/content/18/12/1896.full.html#ref-list-1 This article cites 63 articles, 35 of which can be accessed free at: Open Access Freely available online through the Genome Research Open Access option. service Email alerting click here top right corner of the article or Receive free email alerts when new articles cite this article sign up in the box at the
Tissue development and function are exquisitely dependent on proper regulation of gene expression, but it remains controversial whether the genomic signals controlling this process are subject to strong selective constraint. While some studies show that highly constrained noncoding regions act to enhance transcription, other studies show that DNA segments with biochemical signatures of regulatory regions, such as occupancy by a transcription factor, are seemingly unconstrained across mammalian evolution. To test the possible correlation of selective constraint with enhancer activity, we used chromatin immunoprecipitation as an approach unbiased by either evolutionary constraint or prior knowledge of regulatory activity to identify DNA segments within a 66-Mb region of mouse chromosome 7 that are occupied by the erythroid transcription factor GATA1. DNA segments bound by GATA1 were identified by hybridization to high-density tiling arrays, validated by quantitative PCR, and tested for gene regulatory activity in erythroid cells. Whereas almost all of the occupied segments contain canonical WGATAR binding site motifs for GATA1, in only 45% of the cases is the motif deeply preserved (found at the orthologous position in placental mammals or more distant species). However, GATA1-bound segments with high enhancer activity tend to be the ones with an evolutionarily preserved WGATAR motif, and this relationship was confirmed by a loss-of-function assay. Thus, GATA1 binding sites that regulate gene expression during erythroid maturation are under strong selective constraint, while nonconstrained binding may have only a limited or indirect role in regulation.
This article describes a set of alignments of 28 vertebrate genome sequences that is provided by the UCSC Genome Browser. The alignments can be viewed on the Human Genome Browser (March 2006 assembly) at http://genome.ucsc.edu, downloaded in bulk by anonymous FTP from http://hgdownload.cse.ucsc.edu/goldenPath/hg18/multiz28way, or analyzed with the Galaxy server at http://g2.bx.psu.edu. This article illustrates the power of this resource for exploring vertebrate and mammalian evolution, using three examples. First, we present several vignettes involving insertions and deletions within protein-coding regions, including a look at some human-specific indels. Then we study the extent to which start codons and stop codons in the human sequence are conserved in other species, showing that start codons are in general more poorly conserved than stop codons. Finally, an investigation of the phylogenetic depth of conservation for several classes of functional elements in the human genome reveals striking differences in the rates and modes of decay in alignability. Each functional class has a distinctive period of stringent constraint, followed by decays that allow (for the case of regulatory regions) or reject (for coding regions and ultraconserved elements) insertions and deletions.