Recent evidence indicates that the bacterial Rho helicase regulates Bacillus subtilis gene expression in a growth-dependent manner. This regulation, along with extensive in vivo trimming of Rho-dependent transcript 3'-ends, complicates the identification of Rho-dependent transcription terminators using standard transcriptomic approaches. To overcome this challenge, we applied Helicase-SELEX to precisely map Rho utilization (Rut) sites genome-wide. Using B. subtilis Rho (BsRho), we identified 600 putative Rut sites, while the more permissive Escherichia coli Rho (EcRho) revealed 4189 sites, including specimens known to regulate B. subtilis genes. Comparative analysis showed that both enzymes recognize similar pyrimidine-rich sequences, though BsRho favors short unstructured Rut motifs whereas EcRho can act on presumably more structured RNAs without requiring accessory factors. In vivo validation of selected Rut sites confirmed Rho-dependent regulation and extensive PNPase-mediated processing of Rho-terminated transcripts. Collectively, our results reveal a rich and complex Rho-dependent regulatory network in B. subtilis, encompassing the widespread control of antisense transcription and genes/operons of both primary and secondary metabolism. Although nonessential under standard laboratory conditions, Rho thus likely contributes to B. subtilis fitness and survival in more demanding environments. Our comprehensive compendium of Rut sites offers a valuable resource for exploring this adaptive regulatory landscape.
The bacterial transcription termination factor Rho is a rare example of an RNA helicase that functions as a ring-shaped ATP-powered six-subunit motor. Recent studies have linked Rho's distinctive architecture to a variety of regulatory mechanisms that shape the bacterial transcriptome at the global scale and control the transcription of individual genes in a context-dependent manner. In this review, we provide a comprehensive overview of the molecular mechanisms by which Rho triggers transcription termination. We examine the two prevailing modes of Rho's action: the "catch-up" mode, where Rho actively translocates along RNA and collides with the RNA polymerase to terminate transcription, and the "stand-by" mode where Rho, recruited by transcription elongation factor NusG, remains poised to engage RNA polymerase at specific sites or under particular constraints. Additionally, we highlight Rho's interplay with nucleoid-structuring protein H-NS in the regulation of bacterial chromatin transcription, as well as the crucial role played by Rho in the conditional regulation of specific genomic loci. We discuss how these mechanisms contribute to the fine-tuning of gene activity and integrate into broader regulatory networks, supporting bacterial adaptation to environmental changes and resilience to external challenges.
Binding of the bacterial Rho helicase to nascent transcripts triggers Rho-dependent transcription termination (RDTT) in response to cellular signals that modulate mRNA structure and accessibility of Rho utilization (Rut) sites. Despite the impact of temperature on RNA structure, RDTT was never linked to the bacterial response to temperature shifts. We show that Rho is a central player in the cold-shock response (CSR), challenging the current view that CSR is primarily a posttranscriptional program. We identify Rut sites in 5'-untranslated regions of key CSR genes/operons (cspA, cspB, cspG, and nsrR-rnr-yjfHI) that trigger premature RDTT at 37°C but not at 15°C. High concentrations of RNA chaperone CspA or nucleotide changes in the cspA mRNA leader reduce RDTT efficiency, revealing how RNA restructuring directs Rho to activate CSR genes during the cold shock and to silence them during cold acclimation. These findings establish a paradigm for how RNA thermosensors can modulate gene expression.
Helicases are ubiquitous motor enzymes that remodel nucleic acids (NA) and NA-protein complexes in key cellular processes. To explore the functional repertoire and specificity landscape of helicases, we devised a screening scheme-Helicase-SELEX (Systematic Evolution of Ligands by EXponential enrichment)-that enzymatically probes substrate and cofactor requirements at global scale. Using the transcription termination Rho helicase of Escherichia coli as a prototype for Helicase-SELEX, we generated a genome-wide map of Rho utilization (Rut) sites. The map reveals many features, including promoter- and intrinsic terminator-associated Rut sites, bidirectional Rut tandems, and cofactor-dependent Rut sites with inverted G > C skewed compositions. We also implemented an H-SELEX variant where we used a model ligand, serotonin, to evolve synthetic Rut sites operating in vitro and in vivo in a ligand-dependent manner. Altogether, our data illustrate the power and flexibility of Helicase-SELEX to seek constitutive or conditional helicase substrates in natural or synthetic NA libraries for fundamental or synthetic biology discovery.
Evolutionarily conserved NusG protein enhances bacterial RNA polymerase processivity but can also promote transcription termination by binding to, and stimulating the activity of, Rho factor. Rho terminates transcription upon anchoring to cytidine-rich motifs, the so-called Rho utilization sites (Rut) in nascent RNA. Both NusG and Rho have been implicated in the silencing of horizontally-acquired A/T-rich DNA by nucleoid structuring protein H-NS. However, the relative roles of the two proteins in H-NS-mediated gene silencing remain incompletely defined. In the present study, a Salmonella strain carrying the nusG gene under the control of an arabinose-inducible repressor was used to assess the genome-wide response to NusG depletion. Results from two complementary approaches, i) screening lacZ protein fusions generated by random transposition and ii) transcriptomic analysis, converged to show that loss of NusG causes massive upregulation of Salmonella pathogenicity islands (SPIs) and other H-NS-silenced loci. A similar, although not identical, SPI-upregulated profile was observed in a strain with a mutation in the rho gene, Rho K130Q. Surprisingly, Rho mutation Y80C, which affects Rho’s primary RNA binding domain, had either no effect or made H-NS-mediated silencing of SPIs even tighter. Thus, while corroborating the notion that bound H-NS can trigger Rho-dependent transcription termination in vivo, these data suggest that H-NS-elicited termination occurs entirely through a NusG-dependent pathway and is less dependent on Rut site binding by Rho. We provide evidence that through Rho recruitment, and possibly through other still unidentified mechanisms, NusG prevents pervasive transcripts from elongating into H-NS-silenced regions. Failure to perform this function causes the feedforward activation of the entire Salmonella virulence program. These findings provide further insight into NusG/Rho contribution in H-NS-mediated gene silencing and underscore the importance of this contribution for the proper functioning of a global regulatory response in growing bacteria. The complete set of transcriptomic data is freely available for viewing through a user-friendly genome browser interface.
Bacterial transcription termination proceeds via two main mechanisms triggered either by simple, well-conserved (intrinsic) nucleic acid motifs or by the motor protein Rho. Although bacterial genomes can harbor hundreds of termination signals of either type, only intrinsic terminators are reliably predicted. Computational tools to detect the more complex and diversiform Rho-dependent terminators are lacking. To tackle this issue, we devised a prediction method based on Orthogonal Projections to Latent Structures Discriminant Analysis [OPLS-DA] of a large set of in vitro termination data. Using previously uncharacterized genomic sequences for biochemical evaluation and OPLS-DA, we identified new Rho-dependent signals and quantitative sequence descriptors with significant predictive value. Most relevant descriptors specify features of transcript C>G skewness, secondary structure, and richness in regularly-spaced 5'CC/UC dinucleotides that are consistent with known principles for Rho-RNA interaction. Descriptors collectively warrant OPLS-DA predictions of Rho-dependent termination with a ∼85% success rate. Scanning of the Escherichia coli genome with the OPLS-DA model identifies significantly more termination-competent regions than anticipated from transcriptomics and predicts that regions intrinsically refractory to Rho are primarily located in open reading frames. Altogether, this work delineates features important for Rho activity and describes the first method able to predict Rho-dependent terminators in bacterial genomes.
Mitochondrial dysfunction due to nuclear or mitochondrial DNA alterations contributes to multiple diseases such as metabolic myopathies, neurodegenerative disorders, diabetes and cancer. Nevertheless, to date, only half of the estimated 1,500 mitochondrial proteins has been identified, and the function of most of these proteins remains to be determined. Here, we characterize the function of M19, a novel mitochondrial nucleoid protein, in muscle and pancreatic β-cells. We have identified a 13-long amino acid sequence located at the N-terminus of M19 that targets the protein to mitochondria. Furthermore, using RNA interference and over-expression strategies, we demonstrate that M19 modulates mitochondrial oxygen consumption and ATP production, and could therefore regulate the respiratory chain activity. In an effort to determine whether M19 could play a role in the regulation of various cell activities, we show that this nucleoid protein, probably through its modulation of mitochondrial ATP production, acts on late muscle differentiation in myogenic C2C12 cells, and plays a permissive role on insulin secretion under basal glucose conditions in INS-1 pancreatic β-cells. Our results are therefore establishing a functional link between a mitochondrial nucleoid protein and the modulation of respiratory chain activities leading to the regulation of major cellular processes such as myogenesis and insulin secretion.
To cite this article: Bousquet J, Anto J, Auffray C, Akdis M, Cambon-Thomsen A, Keil T, Haahtela T, Lambrecht BN, Postma DS, Sunyer J, Valenta R, Akdis CA, Annesi-Maesano I, Arno A, Bachert C, Ballester F, Basagana X, Baumgartner U, Bindslev-Jensen C, Brunekreef B, Carlsen KH, Chatzi L, Crameri R, Eveno E, Forastiere F, Garcia-Aymerich J, Guerra S, Hammad H, Heinrich J, Hirsch D, Jacquemin B, Kauffmann F, Kerkhof M, Kogevinas M, Koppelman GH, Kowalski ML, Lau S, Lodrup-Carlsen KC, Lopez-Botet M, Lotvall J, Lupinek C, Maier D, Makela MJ, Martinez FD, Mestres J, Momas I, Nawijn MC, Neubauer A, Oddie S, Palkonen S, Pin I, Pison C, Rancé F, Reitamo S, Rial-Sebbag E, Salapatas M, Siroux V, Smagghe D, Torrent M, Toskala E, van Cauwenberge P, van Oosterhout AJM, Varraso R, von Hertzen L, Wickman M, Wijmenga C, Worm M, Wright J, Zuberbier T. MeDALL (Mechanisms of the Development of ALLergy): an integrated approach from phenotypes to systems medicine. Allergy 2011; 66: 596–604. The origin of the epidemic of IgE-associated (allergic) diseases is unclear. MeDALL (Mechanisms of the Development of ALLergy), an FP7 European Union project (No. 264357), aims to generate novel knowledge on the mechanisms of initiation of allergy and to propose early diagnosis, prevention, and targets for therapy. A novel phenotype definition and an integrative translational approach are needed to understand how a network of molecular and environmental factors can lead to complex allergic diseases. A novel, stepwise, large-scale, and integrative approach will be led by a network of complementary experts in allergy, epidemiology, allergen biochemistry, immunology, molecular biology, epigenetics, functional genomics, bioinformatics, computational and systems biology. The following steps are proposed: (i) Identification of 'classical' and 'novel' phenotypes in existing birth cohorts; (ii) Building discovery of the relevant mechanisms in IgE-associated allergic diseases in existing longitudinal birth cohorts and Karelian children; (iii) Validation and redefinition of classical and novel phenotypes of IgE-associated allergic diseases; and (iv) Translational integration of systems biology outcomes into health care, including societal aspects. MeDALL will lead to: (i) A better understanding of allergic phenotypes, thus expanding current knowledge of the genomic and environmental determinants of allergic diseases in an integrative way; (ii) Novel diagnostic tools for the early diagnosis of allergy, targets for the development of novel treatment modalities, and prevention of allergic diseases; (iii) Improving the health of European citizens as well as increasing the competitiveness and boosting the innovative capacity of Europe, while addressing global health issues and ethical issues.
BackgroundThe PIP (prolactin-inducible protein) gene has been shown to be expressed in breast cancers, with contradictory results concerning its implication. As both the physiological role and the molecular pathways in which PIP is involved are poorly understood, we conducted combined gene expression profiling and network analysis studies on selected breast cancer cell lines presenting distinct PIP expression levels and hormonal receptor status, to explore the functional and regulatory network of PIP co-modulated genes.Principal findingsMicroarray analysis allowed identification of genes co-modulated with PIP independently of modulations resulting from hormonal treatment or cell line heterogeneity. Relevant clusters of genes that can discriminate between [PIP+] and [PIP-] cells were identified. Functional and regulatory network analyses based on a knowledge database revealed a master network of PIP co-modulated genes, including many interconnecting oncogenes and tumor suppressor genes, half of which were detected as differentially expressed through high-precision measurements. The network identified appears associated with an inhibition of proliferation coupled with an increase of apoptosis and an enhancement of cell adhesion in breast cancer cell lines, and contains many genes with a STAT5 regulatory motif in their promoters.ConclusionsOur global exploratory approach identified biological pathways modulated along with PIP expression, providing further support for its good prognostic value of disease-free survival in breast cancer. Moreover, our data pointed to the importance of a regulatory subnetwork associated with PIP expression in which STAT5 appears as a potential transcriptional regulator.
Background: The molecular mechanisms underlying innate tumor drug resistance, a major obstacle to successful cancer therapy, remain poorly understood. In colorectal cancer (CRC), molecular studies have focused on drug-selected tumor cell lines or individual candidate genes using samples derived from patients already treated with drugs, so that very little data are available prior to drug treatment.Results: Transcriptional profiles of clinical samples collected from CRC patients prior to their exposure to a combined chemotherapy of folinic acid, 5-fluorouracil and irinotecan were established using microarrays. Vigilant experimental design, power simulations and robust statistics were used to restrain the rates of false negative and false positive hybridizations, allowing successful discrimination between drug resistance and sensitivity states with restricted sampling. A list of 679 genes was established that intrinsically differentiates, for the first time prior to drug exposure, subsequently diagnosed chemo-sensitive and resistant patients. Independent biological validation performed through quantitative PCR confirmed the expression pattern on two additional patients. Careful annotation of interconnected functional networks provided a unique representation of the cellular states underlying drug responses.Conclusion: Molecular interaction networks are described that provide a solid foundation on which to anchor working hypotheses about mechanisms underlying in vivo innate tumor drug responses. These broad-spectrum cellular signatures represent a starting point from which by-pass chemotherapy schemes, targeting simultaneously several of the molecular mechanisms involved, may be developed for critical therapeutic intervention in CRC patients. The demonstrated power of this research strategy makes it generally applicable to other physiological and pathological situations.
Understanding the complexity and dynamics of cancer cells in response to effective therapy requires hypothesis-driven, quantitative, and high-throughput measurement of genes and proteins at both spatial and temporal levels. This study was designed to gain insights into molecular networks underlying the clinical synergy between retinoic acid (RA) and arsenic trioxide (ATO) in acute promyelocytic leukemia (APL), which results in a high-quality disease-free survival in most patients after consolidation with conventional chemotherapy. We have applied an approach integrating cDNA microarray, 2D gel electrophoresis with MS, and methods of computational biology to study the effects on APL cell line NB4 treated with RA, ATO, and the combination of the two agents and collected in a time series. Numerous features were revealed that indicated the coordinated regulation of molecular networks from various aspects of granulocytic differentiation and apoptosis at the transcriptome and proteome levels. These features include an array of transcription factors and cofactors, activation of calcium signaling, stimulation of the IFN pathway, activation of the proteasome system, degradation of the PML-RARalpha oncoprotein, restoration of the nuclear body, cell-cycle arrest, and gain of apoptotic potential. Hence, this investigation has provided not only a detailed understanding of the combined therapeutic effects of RA/ATO in APL but also a road map to approach hematopoietic malignancies at the systems level.
In this study, we have used high density cDNA arrays to assess age-related changes in gene expression in the myogenic program of human satellite cells and to elucidate modifications in differentiation capacity that could occur throughout in vitro cellular aging. We have screened a collection of 2016 clones from a human skeletal muscle 3′-end cDNA library in order to investigate variations in the myogenic program of myotubes formed by the differentiation of myoblasts of individuals with different ages (5 days old, 52 years old and 79 years old) and induced to differentiate at different stages of their lifespan (early proliferation, presenescence and senescence). Although our analysis has not been able to underline specific changes in the expression of genes encoding proteins involved in muscle structure and/or function, we have demonstrated an age-related induction of genes involved in stress response and a down-regulation of genes involved both in mitochondrial electron transport/ATP synthase and in glycolysis/TCA cycle. From this global approach of post-mitotic cell aging, we have identified 2 potential new markers of presenescence for human myotubes, both strongly linked to carbohydrate metabolism, which could be useful in developing therapeutic strategies.
While it is universally accepted that intact RNA constitutes the best representation of the steady-state of transcription, there is no gold standard to define RNA quality prior to gene expression analysis. In this report, we evaluated the reliability of conventional methods for RNA quality assessment including UV spectroscopy and 28S:18S area ratios, and demonstrated their inconsistency. We then used two new freely available classifiers, the Degradometer and RIN systems, to produce user-independent RNA quality metrics, based on analysis of microcapillary electrophoresis traces. Both provided highly informative and valuable data and the results were found highly correlated, while the RIN system gave more reliable data. The relevance of the RNA quality metrics for assessment of gene expression differences was tested by Q-PCR, revealing a significant decline of the relative expression of genes in RNA samples of disparate quality, while samples of similar, even poor integrity were found highly comparable. We discuss the consequences of these observations to minimize artifactual detection of false positive and negative differential expression due to RNA integrity differences, and propose a scheme for the development of a standard operational procedure, with optional registration of RNA integrity metrics in public repositories of gene expression data.
The Human Anatomic Gene Expression Library (H-ANGEL) is a resource for information concerning the anatomical distribution and expression of human gene transcripts. The tool contains protein expression data from multiple platforms that has been associated with both manually annotated full-length cDNAs from H-InvDB and RefSeq sequences. Of the H-Inv predicted genes, 18 897 have associated expression data generated by at least one platform. H-ANGEL utilizes categorized mRNA expression data from both publicly available and proprietary sources. It incorporates data generated by three types of methods from seven different platforms. The data are provided to the user in the form of a web-based viewer with numerous query options. H-ANGEL is updated with each new release of cDNA and genome sequence build. In future editions, we will incorporate the capability for expression data updates from existing and new platforms. H-ANGEL is accessible at http://www.jbirc.aist.go.jp/hinv/h-angel/.
The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology.
Defects in nucleotide excision repair have been shown to be associated with the photosensitive form of the disorder trichothiodystrophy (TTD). Most repair-deficient TTD patients are mutated in the XPD gene, a subunit of the transcription factor TFIIH. Knowledge of the kinetics and efficiency of repair of the two major UV-induced photolesions in TTD is critical to understand the role of unrepaired lesions in the process of carcinogenesis and explain the absence of enhanced skin cancer incidence in TTD patients contrarily to the xeroderma pigmentosum D patients. In this study, we used different approaches to quantify repair of UV-induced cyclobutane pyrimidine dimers (CPD) and pyrimidine (6–4) pyrimidone photoproducts (6–4PP) at the gene and the genome overall level. In cells of two TTD patients, repair of CPD and 6–4PP was reduced compared with normal human cells, but the reduction was more severe in confluent cells than in exponentially growing cells. Moreover, the impairment of repair was more drastic for CPD than 6–4PP. Most notably, exponentially growing TTD cells displayed complete repair 6–4PP over a broad dose range, albeit at a reduced rate compared with normal cells. Strand-specific analysis of CPD repair in a transcriptional active gene revealed that TTD cells were capable to perform transcription-coupled repair. Taken together, the data suggest that efficient repair of 6–4PP in dividing TTD cells in concert with transcription-coupled repair might account for the absence of increased skin carcinogenesis in TTD patients.
It is well established that biological aging is associated with functional deficits at the cellular, tissue, organ and system levels, but the molecular mechanisms that control lifespan and age-related phenotypes are still not well understood. In order to investigate the molecular mechanisms underlying myoblast aging, we have used quantitative hybridization of a cDNA array of 2016 clones from a human skeletal muscle 3′-end cDNA library to monitor gene expression patterns of myoblasts of individuals with different ages (5 days old, 52 years old and 79 years old) and at different stages of proliferation (early, presenescent and senescent). We have shown that expression profiles in satellite cells vary with donor age, with an up-regulation of genes involved in muscle structure, muscle differentiation and in metabolism in the newborn, and a down-regulation of genes involved in protein renewal in adults. We have also observed that myoblasts isolated from subjects of different ages have typical expression profiles at the beginning of their proliferative lifespan. However, this phenomenon progressively disappears as the cells approach senescence. In addition, even though some of the modifications are similar to those observed in other cell types, we have observed that many changes in gene expression are characteristic of the myoblasts, confirming the hypothesis that the program of replicative senescence is specific for each cell type. Finally, we have identified four potential new markers of presenescence for human myoblasts, which could be useful in developing therapeutic strategies.