Methylation of cytosine residues in nucleic acids plays a critical role in a range of biological activities in eukaryotes, including regulation of transcription, organization of chromatin structure, modulation of translation, cellular differentiation, and development. While much of the scientific focus in this field was centered on DNA methylation over the past few decades, it has also become clear that methylation of RNA is a crucial modification. A group of homologous DNMT2 methyltransferase enzymes in different model organisms are now known to catalyze the transfer of a methyl group to the cytosine at position 38 in tRNAAspGUC molecules. The important biological role for tRNA methyltransferases is highlighted by the fact that the genomes of some model eukaryotes, including Dictyostelium discoideum, Drosophila melanogaster, Entamoeba histolytica, and Schizosaccharomyces pombe, possess a DNMT2 homolog but do not encode any other enzymes of the DNMT family. In this study, we explore the function of the DNMT2 homolog (DNMA) in D. discoideum by examining the phenotypic effects resulting from deletion of this enzyme. Pleiotropic impacts on cell growth, morphology and motility, nuclear organization, and disruption to the developmental program are detected. We also analyze global gene expression in the dnmA knock-out cells and develop a homology-based structural model of DNMA, allowing us to perform docking simulations of the molecular interaction with tRNAAspGUC. Our findings demonstrate that DNMA, as a tRNA methyltransferase, is critical to normal cellular activity and development in Dictyostelium.
The retinal determination genetic network controls the development of the visual system in all seeing animals through the molecular regulation of cells to adopt an eye tissue fate. The compound eye of the fruit fly, Drosophila melanogaster, is an excellent model system to study the complex mechanisms within the network that regulate specification of cellular identity during embryogenesis. In Drosophila, the two Pax6 paralogues, eyeless and twin of eyeless, sit at the very pinnacle of the network and their expression early in development activates critical downstream components of the retinal determination pathway. In this study, we investigate the expression of 21 known components of the network in two established embryonic cell lines, Kc167 and S2 cells, that show reciprocal expression patterns for the two Pax6 paralogues. Network mapping reveals that many of the components of the network demonstrate extensive interactions with additional factors. Integrating the transcriptional profile of the cell lines, interaction maps and embryonic expression patterns enables us to identify 16 potential novel components of the genetic network, 11 of which are transcription factors. We confirm the regulatory potential for a subset of the novel transcription factors through the identification of predicted binding sites in previously characterized enhancers for the core genes in the network.
The Drosophila PAX6 homolog twin of eyeless (toy) sits at the pinnacle of the genetic pathway controlling eye development, the retinal determination network. Expression of toy in the embryo is first detectable at cellular blastoderm stage 5 in an anterior-dorsal band in the presumptive procephalic neuroectoderm, which gives rise to the primordia of the visual system and brain. Although several maternal and gap transcription factors that generate positional information in the embryo have been implicated in controlling toy, the regulation of toy expression in the early embryo is currently not well characterized. In this study, we adopt an integrated experimental approach utilizing bioinformatics, molecular genetic testing of putative enhancers in transgenic reporter gene assays and quantitative analysis of expression patterns in the early embryo, to identify 2 novel coacting enhancers at the toy gene. In addition, we apply mathematical modeling to dissect the regulatory landscape for toy. We demonstrate that relatively simple thermodynamic-based models, incorporating only 5 TF binding sites, can accurately predict gene expression from the 2 coacting enhancers and that the HUNCHBACK TF plays a critical regulatory role through a dual-modality function as an activator and repressor. Our analysis also reveals that the molecular architecture of the 2 enhancers is very different, indicating that the underlying regulatory logic they employ is distinct.
Homeodomain transcription factors (TFs) bind to specific DNA sequences to regulate the expression of target genes. Structural work has provided insight into molecular identities and aided in unraveling structural features of these TFs. However, the detailed affinity and specificity by which these TFs bind to DNA sequences is still largely unknown. Qualitative methods, such as DNA footprinting, Electrophoretic Mobility Shift Assays (EMSAs), Systematic Evolution of Ligands by Exponential Enrichment (SELEX), Bacterial One Hybrid (B1H) systems, Surface Plasmon Resonance (SPR), and Protein Binding Microarrays (PBMs) have been widely used to investigate the biochemical characteristics of TF-DNA binding events. In addition to these qualitative methods, bioinformatic approaches have also assisted in TF binding site discovery. Here we discuss the advantages and limitations of these different approaches, as well as the benefits of utilizing more quantitative approaches, such as Mechanically Induced Trapping of Molecular Interactions (MITOMI), Microscale Thermophoresis (MST) and Isothermal Titration Calorimetry (ITC), in determining the biophysical basis of binding specificity of TF-DNA complexes and improving upon existing computational approaches aimed at affinity predictions.
Background Transcription factor (TF) proteins are a key component of the gene regulatory networks that control cellular fates and function. TFs bind DNA regulatory elements in a sequence-specific manner and modulate target gene expression through combinatorial interactions with each other, cofactors, and chromatin-modifying proteins. Large-scale studies over the last two decades have helped shed light on the complex network of TFs that regulate development in Drosophila melanogaster. Results Here, we present a detailed characterization of expression of all known and predicted Drosophila TFs in two well-established embryonic cell lines, Kc167 and S2 cells. Using deep coverage RNA sequencing approaches we investigate the transcriptional profile of all 707 TF coding genes in both cell types. Only 103 TFs have no detectable expression in either cell line and 493 TFs have a read count of 5 or greater in at least one of the cell lines. The 493 TFs belong to 54 different DNA-binding domain families, with significant enrichment of those in the zf-C2H2 family. We identified 123 differentially expressed genes, with 57 expressed at significantly higher levels in Kc167 cells than S2 cells, and 66 expressed at significantly lower levels in Kc167 cells than S2 cells. Network mapping reveals that many of these TFs are crucial components of regulatory networks involved in cell proliferation, cell-cell signaling pathways, and eye development. Conclusions We produced a reference TF coding gene expression dataset in the extensively studied Drosophila Kc167 and S2 embryonic cell lines, and gained insight into the TF regulatory networks that control the activity of these cells.
Abstract DNA methylation, the addition of a methyl (CH3) group to a cytosine residue, is an evolutionarily conserved epigenetic mark involved in a number of different biological functions in eukaryotes, including transcriptional regulation, chromatin structural organization, cellular differentiation and development. In the social amoeba Dictyostelium, previous studies have shown the existence of a DNA methyltransferase (DNMA) belonging to the DNMT2 family, but the extent and function of 5-methylcytosine in the genome are unclear. Here, we present the whole genome DNA methylation profile of Dictyostelium discoideum using deep coverage replicate sequencing of bisulfite-converted gDNA extracted from post-starvation cells. We find an overall very low number of sites with any detectable level of DNA methylation, occurring at significant levels in only 303–3432 cytosines out of the ∼7.5 million total cytosines in the genome depending on the replicate. Furthermore, a knockout of the DNMA enzyme leads to no overall decrease in DNA methylation. Of the identified sites, significant methylation is only detected at 11 sites in all four of the methylomes analyzed. Targeted bisulfite PCR sequencing and computational analysis demonstrate that the methylation profile does not change during development and that these 11 cytosines are most likely false positives generated by protection from bisulfite conversion due to their location in hairpin-forming palindromic DNA sequences. Our data therefore provide evidence that there is no significant DNA methylation in Dictyostelium before fruiting body formation and identify a reproducible experimental artifact from bisulfite sequencing.
The core promoter elements are important DNA sequences for the regulation of RNA polymerase II transcription in eukaryotic cells. Despite the broad evolutionary conservation of these elements, there is extensive variation in the nucleotide composition of the actual sequences. In this study, we aim to improve our understanding of the complexity of this sequence variation in the TATA box and initiator core promoter elements in Drosophila melanogaster. Using computational approaches, including an enhanced version of our previously developed MARZ algorithm that utilizes gapped nucleotide matrices, several sequence landscape features are uncovered, including an interdependency between the nucleotides in position 2 and 5 in the initiator. Incorporating this information in an expanded MARZ algorithm improves predictive performance for the identification of the initiator element. Overall our results demonstrate the need to carefully consider detailed sequence composition features in core promoter elements in order to make more robust and accurate bioinformatic predictions.
Abstract Drosophila melanogaster cell lines are an important resource for a range of studies spanning genomics, molecular genetics, and cell biology. Amongst these valuable lines are Kc167 (Kc) and Schneider 2 (S2) cells, which were originally isolated in the late 1960s from embryonic sources and have been used extensively to investigate a broad spectrum of biological activities including cell–cell signaling and immune system function. Whole-genome tiling microarray analysis of total RNA from these two cell types was performed as part of the modENCODE project over a decade ago and revealed that they share a number of gene expression features. Here, we expand on these earlier studies by using deep-coverage RNA-sequencing approaches to investigate the transcriptional profile in Kc and S2 cells in detail. Comparison of the transcriptomes reveals that ∼75% of the 13,919 annotated genes are expressed at a detectable level in at least one of the cell lines, with the majority of these genes expressed at high levels in both cell lines. Despite the overall similarity of the transcriptional landscape in the two cell types, 2,588 differentially expressed genes are identified. Many of the genes with the largest fold change are known only by their “CG” designations, indicating that the molecular control of Kc and S2 cell identity may be regulated in part by a cohort of relatively uncharacterized genes. Our data also indicate that both cell lines have distinct hemocyte-like identities, but share active signaling pathways and express a number of genes in the network responsible for dorsal–ventral patterning of the early embryo.
A detailed comprehension of transcriptional regulation is critical to understanding the genetic control of development and disease across many different organisms. To more fully investigate the complex molecular interactions controlling the precise expression of genes, many groups have constructed mathematical models to complement their experimental approaches. A critical step in such studies is choosing the most appropriate parameter estimation algorithm to enable detailed analysis of the parameters that contribute to the models.In this study, we develop a novel set of evolutionary algorithms that use a pseudo-random Sobol Set to construct the initial population and incorporate parameter sensitivities into the adaptation of mutation rates, using local, global, and hybrid strategies. Comparison of the performance of these new algorithms to a number of current state-of-the-art global parameter estimation algorithms on a range of continuous test functions, as well as synthetic biological data representing models of gene regulatory systems, reveals improved performance of the new algorithms in terms of runtime, error and reproducibility. In addition, by analyzing the ability of these algorithms to fit datasets of varying quality, we provide the experimentalist with a guide to how the algorithms perform across a range of noisy data.These results demonstrate the improved performance of the new set of parameter estimation algorithms and facilitate meaningful integration of model parameters and predictions in our understanding of the molecular mechanisms of gene regulation.
Homeodomain transcription factors (HD TFs) are a large class of evolutionarily conserved DNA binding proteins that contain a basic 60-amino acid region required for binding to specific DNA sites. In Drosophila melanogaster, many of these HD TFs are expressed in the early embryo and control transcription of target genes in development through their interaction with cis-regulatory modules. Previous studies where some of the Drosophila HD TFs were purified required the use of strong denaturants (i.e. 6 M urea) and multiple chromatography columns, making the downstream biochemical examination of the isolated protein difficult. To circumvent these obstacles, we have developed a streamlined expression and purification protocol to produce large yields of Drosophila HD TFs. Using the HD TFs FUSHI-TARAZU (FTZ), ANTENNAPEDIA (ANTP), ABDOMINAL-A (ABD-A), ABDOMINAL-B (ABD-B), and ULTRABITHORAX (UBX) as examples, we demonstrate that our 3-day protocol involving the overexpression of His(6)-SUMO fusion constructs in E. coli followed by a Ni2+-IMAC, SUMO-tag cleavage with the SUMO protease Ulp1, and a heparin column purification produces pure, soluble protein in biological buffers around pH 7 in the absence of denaturants. Electrophoretic mobility shift assays (EMSA) confirm that the purified HD proteins are functional and nuclear magnetic resonance (NMR) spectra confirm that the purified HDs are well-folded. These purified HD TFs can be used in future biophysical experiments to structurally and biochemically characterize how and why these HD TFs bind to different DNA sequences and further probe how nucleotide differences contribute to TF-DNA specificity in the HD family.
In all complex organisms, the precise levels and timing of gene expression controls vital biological processes. In higher eukaryotes, including the fruit fly Drosophila melanogaster, the complex molecular control of transcription (the synthesis of RNA from DNA) and translation (the synthesis of proteins from RNA) events driving this gene expression are not fully understood. In particular, for Drosophila melanogaster, there is a plethora of experimental data, including quantitative measurements of both RNA and protein concentrations, but the precise mechanisms that control the dynamics of gene expression during early development and the processes which lead to steady-state levels of certain proteins remain elusive. This study analyzes a current mathematical modeling approach in an attempt to better understand the long-term behavior of gene regulation. The model is a modified reaction-diffusion equation which has been previously employed in predicting gene expression levels and studying the relative contributions of transcription and translation events to protein abundance [10,11,24]. Here, we use Matrix Algebra and Analysis techniques to study the stability of the gene expression system and analyze equilibria, using very general assumptions regarding the parameter values incorporated into the model. We prove that, given realistic biological parameter values, the system will result in a unique, stable equilibrium solution. Additionally, we give an example of this long-term behavior using the model alongside actual experimental data obtained from Drosophila embryos.
DNA methylation, the addition of a methyl (CH 3 ) group to a cytosine residue, is an evolutionarily conserved epigenetic mark involved in a number of different biological functions in eukaryotes, including transcriptional regulation, chromatin structural organization, cellular differentiation and development. In the slime mold Dictyostelium , previous studies have shown the existence of a DNA methyltransferase (DNMA) belonging to the DNMT2 family, but the extent and function of 5-methyl-cytosine in the genome is unclear. Here we present the whole genome DNA methylation profile of Dictyostelium discoideum using deep coverage, replicate sequencing of bisulfite converted gDNA extracted from post-starvation cells. We find an overall very low level of DNA methylation, occurring at only 462 out of the ~7.5 million (0.006%) cytosines in the genome. Despite this sparse profile, significant methylation can be detected at 51 of these sites in replicate experiments, suggesting they are robust targets for DNA methylation. These 5-methyl-cytosines are associated with a broad range of protein-coding genes, tRNA-encoding genes and retrotransposable elements. Our data provides evidence of a minimal, but functional, methylome in Dictyostelium , thereby making Dictyostelium a candidate model organism to further investigate the evolutionary function of DNA methylation.
Understanding the molecular machinery involved in transcriptional regulation is central to improving our knowledge of an organism's development, disease, and evolution. The building blocks of this complex molecular machinery are an organism's genomic DNA sequence and transcription factor proteins. Despite the vast amount of sequence data now available for many model organisms, predicting where transcription factors bind, often referred to as 'motif detection' is still incredibly challenging. In this study, we develop a novel bioinformatic approach to binding site prediction. We do this by extending pre-existing SVM approaches in an unbiased way to include all possible gapped k-mers, representing different combinations of complex nucleotide dependencies within binding sites. We show the advantages of this new approach when compared to existing SVM approaches, through a rigorous set of cross-validation experiments. We also demonstrate the effectiveness of our new approach by reporting on its improved performance on a set of 127 genomic regions known to regulate gene expression along the anterio-posterior axis in early Drosophila embryos.
A long-standing objective in modern biology is to characterize the molecular components that drive the development of an organism. At the heart of eukaryotic development lies gene regulation. On the molecular level, much of the research in this field has focused on the binding of transcription factors (TFs) to regulatory regions in the genome known as cis-regulatory modules (CRMs). However, relatively little is known about the sequence-specific binding preferences of many TFs, especially with respect to the possible interdependencies between the nucleotides that make up binding sites. A particular limitation of many existing algorithms that aim to predict binding site sequences is that they do not allow for dependencies between nonadjacent nucleotides. In this study, we use a recently developed computational algorithm, MARZ, to compare binding site sequences using 32 distinct models in a systematic and unbiased approach to explore nucleotide dependencies within binding sites for 15 distinct TFs known to be critical to Drosophila development. Our results indicate that many of these proteins have varying levels of nucleotide interdependencies within their DNA recognition sequences, and that, in some cases, models that account for these dependencies greatly outperform traditional models that are used to predict binding sites. We also directly compare the ability of different models to identify the known KRUPPEL TF binding sites in CRMs and demonstrate that a more complex model that accounts for nucleotide interdependencies performs better when compared with simple models. This ability to identify TFs with critical nucleotide interdependencies in their binding sites will lead to a deeper understanding of how these molecular characteristics contribute to the architecture of CRMs and the precise regulation of transcription during organismal development.
Enhancers constitute one of the major components of regulatory machinery of metazoans. Although several genome-wide studies have focused on finding and locating enhancers in the genomes, the fundamental principles governing their internal architecture and cis-regulatory grammar remain elusive. Here, we describe an extensive, quantitative perturbation analysis targeting the dorsal-ventral patterning gene regulatory network (GRN) controlled by Drosophila NF-κB homolog Dorsal. To understand transcription factor interactions on enhancers, we employed an ensemble of mathematical models, testing effects of cooperativity, repression, and factor potency. Models trained on the dataset correctly predict activity of evolutionarily divergent regulatory regions, providing insights into spatial relationships between repressor and activator binding sites. Importantly, the collective predictions of sets of models were effective at novel enhancer identification and characterization. Our study demonstrates how experimental dataset and modeling can be effectively combined to provide quantitative insights into cis-regulatory information on a genome-wide scale.
In the development of the Drosophila embryo, gene expression is directed by the sequence-specific interactions of a large network of protein transcription factors (TFs) and DNA cis-regulatory binding sites. Once the identity of the typically 8-10bp binding sites for any given TF has been determined by one of several experimental procedures, the sequences can be represented in a position weight matrix (PWM) and used to predict the location of additional TF binding sites elsewhere in the genome. Often, alignments of large (>200bp) genomic fragments that have been experimentally determined to bind the TF of interest in Chromatin Immunoprecipitation (ChIP) studies are trimmed under the assumption that the majority of the binding sites are located near the center of all the aligned fragments. In this study, ChIP/chip datasets are analyzed using the corresponding PWMs for the well-studied TFs; CAUDAL, HUNCHBACK, KNIRPS and KRUPPEL, to determine the distribution of predicted binding sites. All four TFs are critical regulators of gene expression along the anterio-posterior axis in early Drosophila development. For all four TFs, the ChIP peaks contain multiple binding sites that are broadly distributed across the genomic region represented by the peak, regardless of the prediction stringency criteria used. This result suggests that ChIP peak trimming may exclude functional binding sites from subsequent analyses.