
Despite a recent explosion in the production of vertebrate genome sequence data and large-scale efforts to completely annotate the human genome, we still have scant knowledge of the principles that built vertebrate genomes in evolution, and of genome architecture and its functional significance. We review approaches using bioinformatics, zebrafish transgenesis, and recent findings in the molecular basis of gene regulation and tie these in with mechanisms for the maintenance of long-range conserved synteny across all vertebrate genomes. Specifically, we discuss the recently discovered genomic regulatory blocks which we argue are principal units of vertebrate genome evolution and serve as the foundations onto which evolutionary innovations are built through sequence evolution and insertion of new cis-regulatory elements. We subsequently discuss how these arrangements relate to common human heritable diseases and their significance in disease causality.
Retrotransposons constitute a significant fraction of mammalian genomes. Considering the finding of widespread transcriptional activity across entire genomes, it is not surprising that retrotransposons contribute to the collective RNA pool. However, the transcriptional output from retrotransposons does not merely represent spurious transcription. We review examples of functional RNAs transcribed from retrotransposons, and address the collection of non-protein coding RNAs derived from transposable element sequences, including numerous human microRNAs and the neuronal BC RNAs. Finally, we review the emerging understanding of how retrotransposons themselves are regulated by small RNAs.
The last decade has seen an explosion of interest in new classes of non-coding RNA. While some are now firmly established as new categories of legitimate functional RNAs, the purpose and even existence of others remain to be solidified. Here, we discuss the challenges associated with discovery and characterization of non-traditional categories of non-coding RNA.
Considering the heterogeneity of leukaemias and the widening spectrum of therapeutic strategies, novel diagnostic methods are urgently needed for haematological malignancies. For a decade, gene expression profiling (GEP) has been applied in leukaemia research. Thus, various studies demonstrated worldwide that the majority of genetically defined leukaemia subtypes are accurately predictable by GEP, for example, with respect to reciprocal rearrangements in acute myeloid leukaemia (AML). Moreover, novel prognostically relevant gene classifiers were developed as, for example, in normal karyotype AML. Considering the lymphatic malignancies, GEP studies defined novel clinically relevant subtypes in diffuse large B cell lymphoma (DLBCL), and improved the discrimination of Burkitt lymphoma and DLBCL cases, overcoming considerable overlaps of these entities that exist from morphological and genetic perspectives. Treatment-specific sensitivity assays are being developed for targeted drugs such as farnesyl transferase inhibitors in AML or imatinib in BCR-ABL1 positive acute lymphoblastic leukaemia (ALL). Irrespectively of these proceedings, an introduction of the microarray technology in haematological practice requires diagnostic algorithms and strategies for interaction with currently established diagnostic techniques. Large multicentre studies such as the MILE Study (Microarray Innovations in LEukemia) aim at translating this methodology into clinical routine workflows and to catalyze this process.
Reliable structure prediction is a prerequisite for most types of bioinformatical analysis of RNA. Since the accuracy of structure prediction from single sequences is limited, one often resorts to computing the consensus structure for a set of related RNA sequences. Since functionally important RNA structures are expected to evolve much more slowly than the underlying sequences, the pattern of sequence (co-)variation can be exploited to dramatically improve structure prediction. Since a conserved common structure is only expected when the RNA structure is under selective pressure, consensus structure prediction also provides an ideal starting point for the de novo detection of structured non-coding RNAs. Here, we review different strategies for the prediction of consensus secondary structures, and show how these approaches can be used to predict non-coding RNA genes.
RNA silencing is a complex and highly conserved regulatory mechanism that is now known to be involved in such diverse processes as development, pathogen control, genome maintenance and response to environmental changes. Since its recent discovery, RNA silencing has become a fast moving key area of research in plant and animal molecular biology. Research in this field has greatly profited from recent developments in novel sequencing technologies that allow massive parallel sequencing of small RNA (sRNA) molecules, the key players of all RNA silencing phenomena. As researchers are beginning to decipher the complexity of RNA silencing, novel methodologies have to be developed to make sense of the large amounts of data that are currently being generated. In this review we present an overview of RNA silencing pathways in plants and the current challenges in analysing sRNA data, with a special focus on computational approaches.
Regulatory elements can affect specific genes from megabase distances, often from within or beyond unrelated neighbouring genes. The task of computational charting of regulatory inputs in the genome can be approached from several directions. Typically, computational identification of putative regulatory elements for a gene of interest requires tools that will aid in estimating the extent of the (potentially vast) genomic region around the gene that is likely to contain regulatory elements, as well as tools for the identification and characterization of individual elements. Conversely, starting from a putative regulatory element or a regulatory variation in a non-coding region, one often wants to associate the regulatory element with the correct target gene(s). The design of tools for these purposes relies on the remarkably high level of sequence conservation of thousands of regulatory enhancers, their strong tendency to cluster around their target genes, as well as a constrained range of functional categories of the corresponding target genes, many of which are developmental regulators. Additional evolutionary information, such as conservation of synteny, and a growing body of functional genomic and epigenomic data are being rapidly added to established and emerging tools for studying developmental regulation and cross-species conservation to provide new functional insights into the roles of these regions. In this article, we give an overview of the functionality available in general purpose and new/specialized web tools for the above tasks, and discuss current and future developments in the field.
South Asia is home to more than 1.5 billion humans representing many diverse ethnicities, linguistic and religious groups and representing almost one-quarter of humanity. Modern humans arrived here soon after their departure from Africa approximately 50,000-70,000 years before present (YBP) and several subsequent human migrations and invasions, as well as the unique social structure of the region, have helped shape the pattern of genetic diversity currently observed in these populations. Over the last few decades population geneticists and molecular anthropologists have analyzed DNA variation in indigenous populations from this region in order to catalog their genetic relationships and histories. The emphasis is gradually shifting from the study of population origins to high resolution surveys of DNA variation to address issues of population stratification and genetic susceptibility or resistance to diseases in genome-wide association surveys. We present a historical overview of the genetic studies carried out on populations from this region in order to understand the influence of geographic, linguistic and religious factors on population diversity in this region, and discuss future prospects in light of developments in high throughput genotyping and next generation sequencing technologies.
Proper development and functioning of an organism depends on precise spatial and temporal expression of all its genes. These coordinated expression-patterns are maintained primarily through the process of transcriptional regulation. Transcriptional regulation is mediated by proteins binding to regulatory elements on the DNA in a combinatorial manner, where particular combinations of transcription factor binding sites establish specific regulatory codes. In this review, we survey experimental and computational approaches geared towards the identification of proximal and distal gene regulatory elements in the genomes of complex eukaryotes. Available approaches that decipher the genetic structure and function of regulatory elements by exploiting various sources of information like gene expression data, chromatin structure, DNA-binding specificities of transcription factors, cooperativity of transcription factors, etc. are highlighted. We also discuss the relevance of regulatory elements in the context of human health through examples of mutations in some of these regions having serious implications in misregulation of genes and being strongly associated with human disorders.
The spatiotemporally and quantitatively correct activity of a gene requires the presence of intact coding sequence as well as properly functioning regulatory control. One of the great challenges of the post-genome era is to gain a better understanding of the mechanisms of gene control. Proper gene regulation depends not only on the required transcription factors and associated complexes being present (in the correct dosage), but also on the integrity, chromatin conformation and nuclear positioning of the gene's chromosomal segment. Thus, when either the cis-trans regulatory system of a gene or the normal context of its chromatin structure are disrupted, gene expression may be adversely affected, potentially leading to disease. As transcriptional regulation is a highly complex process depending on many factors, there are many different mechanisms that can cause aberrant gene expression. Traditionally, the term 'position effect' was used to refer to situations where the level of expression of a gene is deleteriously affected by an alteration in its chromosomal environment, while maintaining an intact transcription unit. Over the past years, an ever increasing number of such disease-related position effect cases have come to light, and detailed studies have revealed insight into the variety of causes, which can be categorized into a number of different mechanistic groups. We suggest replacing the outdated term of 'position effect disease' with the new generic name of 'cis-ruption disorder' to describe genetic disease cases that are caused by disruption of the normal cis-regulatory architecture of the disease gene locus. Here, we review these various cis-ruption mechanisms and discuss how their studies have contributed to our understanding of long- range gene regulation.
Several recent studies from the field of epigenetics have combined chromatin-immunoprecipitation (ChIP) with next-generation high-throughput sequencing technologies to describe the locations of histone post-translational modifications (PTM) and DNA methylation genome-wide. While these reports begin to quench the chromatin biologists thirst for visualizing where in the genome epigenetic marks are placed, they also illustrate several advantages of sequencing based genomics compared to microarray analysis. Accordingly, next-generation sequencing (NGS) technologies are now challenging microarrays as the tool of choice for genome analysis. The increased affordability of comprehensive sequence-based genomic analysis will enable new questions to be addressed in many areas of biology. It is inevitable that massively-parallel sequencing platforms will supercede the microarray for many applications, however, there are niches for microarrays to fill and interestingly we may very well witness a symbiotic relationship between microarrays and high-throughput sequencing in the future.
We describe various types of outliers seen in Affymetrix GeneChip data. We have been able to utilise the data in the Gene Expression Omnibus to screen GeneChips across a range of scales, from single probes, to spatially adjacent fractions of arrays, to whole arrays, to whole experiments. In this review we describe a number of causes for why some reported intensities might be misleading on GeneChips.
A large fraction of non-coding RNAs is short and/or poorly conserved in sequence. Most of the longer examples, furthermore, consist of a collection of conserved structural motifs rather than a coherent globally conserved secondary structure. As a consequence, the conceptually simple problem of homology search becomes a complex and technically demanding task. Despite the best efforts of databases such as Rfam, the situation is complicated further by the sparsity of information in many—in particular prokaryotic—RNA families. In this contribution, we review recent efforts to customize sequence-based search tools for ncRNA applications. In particular, semi-global alignments and the development of methods for fragmented pattern search have brought significant practical advances. Current developments in this area focus on the integration of fragmented sequence pattern search with search algorithms for secondary structure patterns. We focus here, in particular, on strategies that can be successful in the ‘twilight zone’ where generic approaches from blast to infernal to start to fail.
Genome-wide analyses of the eukaryotic transcriptome have revealed that the majority of the genome is transcribed, producing large numbers of non-protein-coding RNAs (ncRNAs). This surprising observation challenges many assumptions about the genetic programming of higher organisms and how information is stored and organized within the genome. Moreover, the rapid advances in genomics have given little opportunity for biologists to integrate these emerging findings into their intellectual and experimental frameworks. This problem has been compounded by the perception that genome-wide studies often generate more questions than answers, which in turn has led to confusion and controversy. In this article, we address common questions associated with the phenomenon of pervasive transcription and consider the indices that can be used to evaluate the function (or lack thereof) of the resulting ncRNAs. We suggest that many lines of evidence, including expression profiles, conservation signatures, chromatin modification patterns and examination of increasing numbers of individual cases, argue in favour of the widespread functionality of non-coding transcription. We also discuss how informatic and experimental approaches used to analyse protein-coding genes may not be applicable to ncRNAs and how the general perception that protein-coding genes form the main informational output of the genome has resulted in much of the misunderstanding surrounding pervasive transcription and its potential significance. Finally, we present the conceptual implications of the majority of the eukaryotic genome being functional and describe how appreciating this perspective will provide considerable opportunity to further understand the molecular basis of development and complex diseases.
Premature termination of transcription, or attenuation, is an efficient RNA-based regulatory strategy that is commonly used in bacterial organisms. Attenuators are generally located in the 5' untranslated regions of genes or operons and combine a Rho-independent terminator, controlling transcription, with an RNA element that senses specific environmental signals. A striking diversity of sensing elements enable regulation of gene expression in response to multiple environmental conditions, including temperature changes, availability of small metabolites (such as ions, amino acids, nucleobases or vitamins), or availability of macromolecules such as tRNAs and regulatory proteins. The wide distribution of attenuators suggests an early emergence among bacteria. However, attenuators also display a great mobility and lability, illustrated by a multiplicity of recent horizontal transfers and duplications. For these reasons, attenuation systems are of high interest both from a fundamental evolutionary perspective and for possible biotechnological applications.
The past era of proteomics has been dominated by global studies whose primary aim was to characterize as much of the proteome as possible. The progress was remarkable. Early in this decade, the ability to identify a thousand proteins from a single proteome sample using mass spectrometry (MS) was limited to a handful of laboratories. Presently, the ability to identify upwards of 3000 proteins is pretty much commonplace. This ability to gain comprehensive proteomic coverage was used to conduct a variety of different studies ranging from simply determining the protein content of a complex sample to comparing the phosphoproteomes of different cell types. Unfortunately, the ability to identify large numbers of proteins comes at a price: the amount of information obtained per protein is not very high. The hope of these global studies was that by surveying as much of the proteome as possible, at least identification of a few proteins of biological interest would facilitate important discoveries and expand boundaries of systems biology. Global proteomic studies have had a number of successes with a large number of important findings being made. Unfortunately, most basic researchers are not equipped with bioinformatics tools capable of thorough meta-analysis of these data sets and are interested in more specific information on a small number of proteins. Fortunately, MS has many more attributes than simply being able to identify large numbers of molecules in complex mixtures. For a long time, MS has been the technology of choice for quantitating metabolites because of its sensitivity and specificity. The last decade has witnessed remarkable progress in adopting targeted approaches to quantitate specific proteins within complex mixtures such as serum and plasma. The conventional thinking used to be that MS would play a major role in the discovery of potential disease-specific biomarkers for example; validation would be conducted using antibody-based methods. However, as demonstrated by the many articles present within this special edition of Briefings inFunctional Genomics andProteomics, MS will be a major contributor for quantitating these potential markers across the thousands of samples necessary for validation. Beyond validation of potential biomarkers, translation of targeted quantitative methods to other types of proteins, such as glycosylated or phosphorylated proteins will be key for using MS to specifically define the roles proteins play in cellular systems. As co-editors, we wish to thank each author who contributed a manuscript to the special edition. We understand that your time is valuable and we appreciate you sacrificing some of it to make this an outstanding issue. To the reader, we hope you find this collection of articles enjoyable to read and that they stimulate ideas that you can test and explore within your own facilities.
Gene expression domains are normally not arranged in vertebrate genomes according to their expression patterns. Instead, it is not unusual to find genes expressed in different cell types, or in different developmental stages, sharing a particular region of a chromosome. Therefore, the existence of boundaries, or insulators, as non-coding gene regulatory elements, is instrumental for the adequate organization and function of vertebrate genomes. Through the evolution and natural selection at the molecular level, and according to available DNA sequences surrounding a locus, previously existing or recently mobilized, different elements have been recruited to serve as boundaries, depending on their suitability to properly insulate gene expression domains. In this regard, several gene regulatory elements, including scaffold/matrix-attachment regions, members of families of DNA repetitive elements (such as LINEs or SINEs), target sites for the zinc-finger multipurpose nuclear factor CTCF, enhancers and locus control regions, have been reported to show functional activities as insulators. In this review, we will address how such a variety of apparently different genomic sequences converge in a similar function, namely, to adequately insulate a gene expression domain, thereby allowing the locus to be expressed according to their own gene regulatory elements without interfering itself and being interfered by surrounding loci. The identification and characterization of genomic boundaries is not only interesting as a theoretical exercise for better understanding how vertebrate genomes are organized, but also allows devising new and improved gene transfer strategies to ensure the expression of heterologous DNA constructs in ectopic genomic locations.
Squamous cell carcinoma of head and neck (SCCHN) is the sixth most common malignancy and is a major cause of cancer morbidity and mortality worldwide. As with most solid cancers, the cure rate for SCCHN is excellent if tumors are diagnosed early in the course of the disease. Early diagnosis of cancer remains difficult because of the lack of specific symptoms in early disease as well as the limited understanding of etiology and oncogenesis. Advances in proteomics and genomics contribute to the understanding of the pathophysiology of neoplasia, cancer diagnosis and anticancer drug discovery. The powerful 'omics' technologies have opened new avenues towards biomarker discovery, identification of signaling molecules associated with cell growth, cell death, cellular metabolism and early detection of cancer. Analysis of tumor-specific omics profiles provided a unique opportunity to diagnose, classify, and detect malignant disease; to better understand and define the behavior of specific tumors; and to provide direct and targeted therapy. These technologies however still require integration and standardization of techniques and validation against accepted clinical and pathologic parameters. This article provides a summary of technologies, potential clinical applications, and challenges of omics in head and neck cancer.
Chromatin insulators are DNA-protein complexes with broad functions in nuclear biology. Drosophila has at least five different types of insulators; recent results suggest that these different insulators share some components that may allow them to function through common mechanisms. Data from genome-wide localization studies of insulator proteins indicate a possible functional specialization, with different insulators playing distinct roles in nuclear biology. Cells have developed mechanisms to control insulator activity by recruiting specialized proteins or by covalent modification of core components. Current results suggest that insulators set up cell-specific blueprints of nuclear organization that may contribute to the establishment of different patterns of gene expression during cell differentiation and development.