Type 2 diabetes (T2D) represents a multifactorial metabolic disease with a strong genetic predisposition. Despite elaborate efforts in identifying the genetic variants determining individual susceptibility towards T2D, the majority of genetic factors driving disease development remain poorly understood. With the aim to identify novel T2D risk genes we previously generated an N2 outcross population using the two inbred mouse strains New Zealand obese (NZO) and C3HeB/FeJ (C3H). A linkage study performed in this population led to the identification of the novel T2D-associated quantitative trait locus (QTL) Nbg15 (NZO blood glucose on chromosome 15, Logarithm of odds (LOD) 6.6). In this study we used a combined approach of positional cloning, gene expression analyses and in silico predictions of DNA polymorphism on gene/protein function to dissect the genetic variants linking Nbg15 to the development of T2D. Moreover, we have generated congenic strains that associated the distal sublocus of Nbg15 to mechanisms altering pancreatic beta cell function. In this sublocus, Cbx6, Fam135b and Kdelr3 were nominated as potential causative genes associated with the Nbg15 driven effects. Moreover, a putative mutation in the Kdelr3 gene from NZO was identified, negatively influencing adaptive responses associated with pancreatic beta cell death and induction of endoplasmic reticulum stress. Importantly, knockdown of Kdelr3 in cultured Min6 beta cells altered insulin granules maturation and pro-insulin levels, pointing towards a crucial role of this gene in islets function and T2D susceptibility.
To nominate novel disease genes for obesity and type 2 diabetes (T2D), we recently generated two mouse backcross populations of the T2D-susceptible New Zealand Obese (NZO/HI) mouse strain and two genetically different, lean and T2D-resistant strains, 129P2/OlaHsd and C3HeB/FeJ. Comparative linkage analysis of our two female backcross populations identified seven novel body fat-associated quantitative trait loci (QTL). Only the locus Nbw14 (NZO body weight on chromosome 14) showed linkage to obesity-related traits in both backcross populations, indicating that the causal gene variant is likely specific for the NZO strain as NZO allele carriers in both crosses displayed elevated body weight and fat mass. To identify candidate genes for Nbw14, we used a combined approach of gene expression and haplotype analysis to filter for NZO-specific gene variants in gonadal white adipose tissue, defined as the main QTL-target tissue. Only two genes, Arl11 and Sgcg, fulfilled our candidate criteria. In addition, expression QTL analysis revealed cis-signals for both genes within the Nbw14 locus. Moreover, retroviral overexpression of Sgcg in 3T3-L1 adipocytes resulted in increased insulin-stimulated glucose uptake. In humans, mRNA levels of SGCG correlated with body mass index and body fat mass exclusively in diabetic subjects, suggesting that SGCG may present a novel marker for metabolically unhealthy obesity. In conclusion, our comparative-cross analysis could substantially improve the mapping resolution of the obesity locus Nbw14. Future studies will throw light on the mechanism by which Sgcg may protect from the development of obesity.
Type 2 diabetes (T2D) has a strong genetic component. Most of the gene variants driving the pathogenesis of T2D seem to target pancreatic β-cell function. To identify novel gene variants acting at early stage of the disease, we analyzed whole transcriptome data to identify differential expression (DE) and alternative exon splicing (AS) transcripts in pancreatic islets collected from two metabolically diverse mouse strains at 6 weeks of age after three weeks of high-fat-diet intervention. Our analysis revealed 1218 DE and 436 AS genes in islets from NZO/Hl vs C3HeB/FeJ. Whereas some of the revealed genes present well-established markers for β-cell failure, such as Cd36 or Aldh1a3 , we identified numerous DE/AS genes that have not been described in context with β-cell function before. The gene Lgals2 , previously associated with human T2D development, was DE as well as AS and localizes in a quantitative trait locus (QTL) for blood glucose on Chr.15 that we reported recently in our N 2 (NZOxC3H) population. In addition, pathway enrichment analysis of DE and AS genes showed an overlap of only half of the revealed pathways, indicating that DE and AS in large parts influence different pathways in T2D development. PPARG and adipogenesis pathways, two well-established metabolic pathways, were overrepresented for both DE and AS genes, probably as an adaptive mechanism to cope for increased cellular stress. Our results provide guidance for the identification of novel T2D candidate genes and demonstrate the presence of numerous AS transcripts possibly involved in islet function and maintenance of glucose homeostasis.
To identify novel disease genes for type 2 diabetes (T2D) we generated two backcross populations of obese and diabetes-susceptible New Zealand Obese (NZO/HI) mice with the two lean mouse strains 129P2/OlaHsd and C3HeB/FeJ. Subsequent whole-genome linkage scans revealed 30 novel quantitative trait loci (QTL) for T2D-associated traits. The strongest association with blood glucose [12 cM, logarithm of the odds (LOD) 13.3] and plasma insulin (17 cM, LOD 4.8) was detected on proximal chromosome 7 (designated Nbg7p, NZO blood glucose on proximal chromosome 7) exclusively in the NZOxC3H crossbreeding, suggesting that the causal gene is contributed by the C3H genome. Introgression of the critical C3H fragment into the genetic NZO background by generating recombinant congenic strains and metabolic phenotyping validated the phenotype. For the detection of candidate genes in the critical region (30-46 Mb), we used a combined approach of haplotype and gene expression analysis to search for C3H-specific gene variants in the pancreatic islets, which appeared to be the most likely target tissue for the QTL. Two genes, Atp4a and Pop4, fulfilled the criteria from our candidate gene approaches. The knockdown of both genes in MIN6 cells led to decreased glucose-stimulated insulin secretion, indicating a regulatory role of both genes in insulin secretion, thereby possibly contributing to the phenotype linked to Nbg7p. In conclusion, our combined- and comparative-cross analysis approach has successfully led to the identification of two novel diabetes susceptibility candidate genes, and thus has been proven to be a valuable tool for the discovery of novel disease genes.
Motivation: Extensive drug treatment gene expression data have been generated in order to identify biomarkers that are predictive for toxicity or to classify compounds. However, such patterns are often highly variable across compounds and lack robustness. We and others have previously shown that supervised expression patterns based on pathway concepts rather than unsupervised patterns are more robust and can be used to assess toxicity for entire classes of drugs more reliably. Results: We have developed a database, ToxDB, for the analysis of the functional consequences of drug treatment at the pathway level. We have collected 2694 pathway concepts and computed numerical response scores of these pathways for 437 drugs and chemicals and 7464 different experimental conditions. ToxDB provides functionalities for exploring these pathway responses by offering tools for visualization and differential analysis allowing for comparisons of different treatment parameters and for linking this data with toxicity annotation and chemical information. Database URL: http://toxdb.molgen.mpg.de
When evaluating compound similarity, addressing multiple sources of information to reach conclusions about common pharmaceutical and/or toxicological mechanisms of action is a crucial strategy. In this chapter, we describe a systems biology approach that incorporates analyses of hepatotoxicant data for 33 compounds from three different sources: a chemical structure similarity analysis based on the 3D Tanimoto coefficient, a chemical structure-based protein target prediction analysis, and a cross-study/cross-platform meta-analysis of in vitro and in vivo human and rat transcriptomics data derived from public resources (i.e., the diXa data warehouse). Hierarchical clustering of the outcome scores of the separate analyses did not result in a satisfactory grouping of compounds considering their known toxic mechanism as described in literature. However, a combined analysis of multiple data types may hypothetically compensate for missing or unreliable information in any of the single data types. We therefore performed an integrated clustering analysis of all three data sets using the R-based tool iClusterPlus. This indeed improved the grouping results. The compound clusters that were formed by means of iClusterPlus represent groups that show similar gene expression while simultaneously integrating a similarity in structure and protein targets, which corresponds much better with the known mechanism of action of these toxicants. Using an integrative systems biology approach may thus overcome the limitations of the separate analyses when grouping liver toxicants sharing a similar mechanism of toxicity.
The computational prediction of alternative splicing from high-throughput sequencing data is inherently difficult and necessitates robust statistical measures because the differential splicing signal is overlaid by influencing factors such as gene expression differences and simultaneous expression of multiple isoforms among others. In this work we describe ARH-seq, a discovery tool for differential splicing in case-control studies that is based on the information-theoretic concept of entropy.
Background: The analysis of differential splicing (DS) is crucial for understanding physiological processes in cells and organs. In particular, aberrant transcripts are known to be involved in various diseases including cancer. A widely used technique for studying DS are exon arrays. Over the last decade a variety of algorithms for the detection of DS events from exon arrays has been developed. However, no comprehensive, comparative evaluation including sensitivity to the most important data features has been conducted so far. To this end, we created multiple data sets based on simulated data to assess strengths and weaknesses of seven published methods as well as a newly developed method, KLAS. Additionally, we evaluated all methods on two cancer data sets that comprised RT-PCR validated results.Results: Our studies indicated ARH as the most robust methods when integrating the results over all scenarios and data sets. Nevertheless, special cases or requirements favor other methods. While FIRMA was highly sensitive according to experimental data, SplicingCompass, MIDAS and ANOSVA showed high specificity throughout the scenarios. On experimental data ARH, FIRMA, MIDAS, and KLAS performed best.Conclusions: Each method shows different characteristics regarding sensitivity, specificity, interference to certain data settings and robustness over multiple data sets. While some methods can be considered as generally good choices over all data sets and scenarios, other methods show heterogeneous prediction quality on the different data sets. The adequate method has to be chosen carefully and with a defined study aim in mind.
The computational prediction of alternative splicing from high-throughput sequencing data is inherently difficult and necessitates robust statistical measures because the differential splicing signal is overlaid by influencing factors such as gene expression differences and simultaneous expression of multiple isoforms amongst others. In this work we describe ARH-seq, a discovery tool for differential splicing in case-control studies that is based on the information-theoretic concept of entropy. ARH-seq works on high-throughput sequencing data and is an extension of the ARH method that was originally developed for exon microarrays. We show that the method has inherent features, such as independence of transcript exon number and independence of differential expression, what makes it particularly suited for detecting alternative splicing events from sequencing data. In order to test and validate our workflow we challenged it with publicly available sequencing data derived from human tissues and conducted a comparison with eight alternative computational methods. In order to judge the performance of the different methods we constructed a benchmark data set of true positive splicing events across different tissues agglomerated from public databases and show that ARH-seq is an accurate, computationally fast and high-performing method for detecting differential splicing events.
Background: Leishmania (L.) are intracellular protozoan parasites able to survive and replicate in the hostile phagolysosomal environment of infected macrophages. They cause leishmaniasis, a heterogeneous group of worldwide-distributed affections, representing a paradigm of neglected diseases that are mainly embedded in impoverished populations. To establish successful infection and ensure their own survival, Leishmania have developed sophisticated strategies to subvert the host macrophage responses. Despite a wealth of gained crucial information, these strategies still remain poorly understood. MicroRNAs (miRNAs), an evolutionarily conserved class of endogenous 22-nucleotide non-coding RNAs, are described to participate in the regulation of almost every cellular process investigated so far. They regulate the expression of target genes both at the levels of mRNA stability and translation; changes in their expression have a profound effect on their target transcripts. Methodology/Principal Findings: We report in this study a comprehensive analysis of miRNA expression profiles in L. major-infected human primary macrophages of three healthy donors assessed at different time-points post-infection (three to 24 h). We show that expression of 64 out of 365 analyzed miRNAs was consistently deregulated upon infection with the same trends in all donors. Among these, several are known to be induced by TLR-dependent responses. GO enrichment analysis of experimentally validated miRNA-targeted genes revealed that several pathways and molecular functions were disturbed upon parasite infection. Finally, following parasite infection, miR-210 abundance was enhanced in HIF-1adependent manner, though it did not contribute to inhibiting anti-apoptotic pathways through pro-apoptotic caspase-3 regulation. Conclusions/Significance: Our data suggest that alteration in miRNA levels likely plays an important role in regulating macrophage functions following L. major infection. These results could contribute to better understanding of the dynamics of gene expression in host cells during leishmaniasis. Citation: Lemaire J, Mkannez G, Guerfali FZ, Gustin C, Attia H, et al. (2013) MicroRNA Expression Profile in Human Macrophages in Response to Leishmania major Infection. PLoS Negl Trop Dis 7(10): e2478. doi:10.1371/journal.pntd.0002478 Editor: Jesus G. Valenzuela, National Institute of Allergy and Infectious Diseases, United States of America Received January 7, 2013; Accepted August 30, 2013; Published October 3, 2013 Copyright: 2013 Lemaire et al. This is an open-access article distributed under the terms of the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are credited. Funding: This work was funded by the European Union under its 6th Framework Programme (LSHG-CT-2006-037231) and partially supported by an NIH/NIAID/ DMID Grant Number 5P50AI074178 for DL. The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Competing Interests: The authors have declared that no competing interests exist. * E-mail: dhafer_l@yahoo.ca, dhafer.laouini@pasteur.rns.tn (DL); patsy.renard@fundp.ac.be (PR) . These authors contributed equally to this work. " These authors co-directed this study. ` Membership of the Sysco-Consortium is provided in the Acknowledgments.
We analyzed the transcriptional signatures of mouse bone marrow-derived macrophages at different times after infection with promastigotes of the protozoan parasite Leishmania major. Ingenuity Pathway Analysis revealed that the macrophage metabolic pathways including carbohydrate and lipid metabolisms were among the most altered pathways at later time points of infection. Indeed, L. major promastiogtes induced increased mRNA levels of the glucose transporter and almost all of the genes associated with glycolysis and lactate dehydrogenase, suggesting a shift to anaerobic glycolysis. On the other hand, L. major promastigotes enhanced the expression of scavenger receptors involved in the uptake of Low-Density Lipoprotein (LDL), inhibited the expression of genes coding for proteins regulating cholesterol efflux, and induced the synthesis of triacylglycerides. These data suggested that Leishmania infection disturbs cholesterol and triglycerides homeostasis and may lead to cholesterol accumulation and foam cell formation. Using Filipin and Bodipy staining, we showed cholesterol and triglycerides accumulation in infected macrophages. Moreover, Bodipy-positive lipid droplets accumulated in close proximity to parasitophorous vacuoles, suggesting that intracellular L. major may take advantage of these organelles as high-energy substrate sources. While the effect of infection on cholesterol accumulation and lipid droplet formation was independent on parasite development, our data indicate that anaerobic glycolysis is actively induced by L. major during the establishment of infection.
The yeast two-hybrid (Y2H) system is the most widely applied methodology for systematic protein–protein interaction (PPI) screening and the generation of comprehensive interaction networks. We developed a novel Y2H interaction screening procedure using DNA microarrays for high-throughput quantitative PPI detection. Applying a global pooling and selection scheme to a large collection of human open reading frames, proof-of-principle Y2H interaction screens were performed for the human neurodegenerative disease proteins huntingtin and ataxin-1. Using systematic controls for unspecific Y2H results and quantitative benchmarking, we identified and scored a large number of known and novel partner proteins for both huntingtin and ataxin-1. Moreover, we show that this parallelized screening procedure and the global inspection of Y2H interaction data are uniquely suited to define specific PPI patterns and their alteration by disease-causing mutations in huntingtin and ataxin-1. This approach takes advantage of the specificity and flexibility of DNA microarrays and of the existence of solid-related statistical methods for the analysis of DNA microarray data, and allows a quantitative approach toward interaction screens in human and in model organisms.
Background: Down syndrome (DS; trisomy 21) is the most common genetic cause of mental retardation in the human population and key molecular networks dysregulated in DS are still unknown. Many different experimental techniques have been applied to analyse the effects of dosage imbalance at the molecular and phenotypical level, however, currently no integrative approach exists that attempts to extract the common information.Results: We have performed a statistical meta-analysis from 45 heterogeneous publicly available DS data sets in order to identify consistent dosage effects from these studies. We identified 324 genes with significant genome-wide dosage effects, including well investigated genes like SOD1, APP, RUNX1 and DYRK1A as well as a large proportion of novel genes (N = 62). Furthermore, we characterized these genes using gene ontology, molecular interactions and promoter sequence analysis. In order to judge relevance of the 324 genes for more general cerebral pathologies we used independent publicly available microarry data from brain studies not related with DS and identified a subset of 79 genes with potential impact for neurocognitive processes. All results have been made available through a web server under http://ds-geneminer.molgen.mpg.de/.Conclusions: Our study represents a comprehensive integrative analysis of heterogeneous data including genome-wide transcript levels in the domain of trisomy 21. The detected dosage effects build a resource for further studies of DS pathology and the development of new therapies.
Type-2 diabetes mellitus (T2DM) is a complex disease with multiple causes covering several functional entities of the metabolism. Environmental factors contribute to the pathogenesis of the disease – most notably nutrition and weight of the organism. The identification of disease genes is the driving power of many research projects. In a previous paper (Rasche et al. 2008) we presented a method that integrates results from different T2DM related studies and identifies candidate genes with high disease relevance. This chapter is designated to elaborate on our work from a network based perspective. Network biology is a promising field that can shed light on interrelations between disease genes and from disease genes to their functional neighborhood. We use network-based tools to advance from a single-gene analysis towards a subnet, a functional module, of disease genes. Proteins are gene products that are associated with particular molecular functions. Molecular functions are interpreted as activities that can be performed by individual proteins following the definitions introduced by the Gene Ontology Consortium (Ashburner et al. 2000). Examples of molecular functions are catalytic activity, transporter activity or binding. Additionally, a biological process is accomplished by one or more ordered assemblies of molecular functions (Ashburner et al. 2000). Proteins physically interact with each other in order to carry out a biological function. A biological function is related to the term biological process. A signal transduction cascade whose biological function is to transmit information from a receptor to a transcription factor is a succession of protein-protein interactions (PPIs). Both the molecular function of a protein and the biological function in which it is involved are best deduced by studying the environment where it operates in. To this end, scientists pursue the ambitious goal of assembling all PPIs in an organism – the interactome – to elucidate how proteins work together and promote individual biological processes and eventually the complete cellular machinery. Today, mainly two methods are used to detect PPIs: Yeast two-hybrid screens (Fields & Sternglanz 1994) and affinity purification (Pandey & Mann 2000). These large-scale technologies provide vast numbers of interactions but have high false positive rates. Additionally, such experiments only reflect one environmental condition and not the dynamics of interactions between different phyiological states leading to high false negative rates. Regarding the current size of the human interactome, we have only a draft of the complete set of interactions. However, looking at the course of construction (fig. 1) so far and bearing in mind new quality standards we are continuously moving towards the completion of a
MOTIVATION:Exon arrays allow the quantitative study of alternative splicing (AS) on a genome-wide scale. A variety of splicing prediction methods has been proposed for Affymetrix exon arrays mainly focusing on geometric correlation measures or analysis of variance. In this article, we introduce an information theoretic concept that is based on modification of the well-known entropy function.RESULTS:We have developed an AS robust prediction method based on entropy (ARH). We can show that this measure copes with bias inherent in the analysis of AS such as the dependency of prediction performance on the number of exons or variable exon expression. In order to judge the performance of ARH, we have compared it with eight existing splicing prediction methods using experimental benchmark data and demonstrate that ARH is a well-performing new method for the prediction of splice variants.AVAILABILITY AND IMPLEMENTATION:ARH is implemented in R and provided in the Supplementary Material.
Aims/hypothesis Numerous new genes have recently been identified in genome-wide association studies for type 2 diabetes. Most are highly expressed in beta cells and presumably play important roles in their function. However, these genes account for only a small proportion of total risk and there are likely to be additional candidate genes not detected by current methodology. We therefore investigated islets from the polygenic New Zealand mouse (NZL) model of diet-induced beta cell dysfunction to identify novel genes and pathways that may play a role in the pathogenesis of diabetes. Methods NZL mice were fed a diabetogenic high-fat diet (HF) or a diabetes-protective carbohydrate-free HF diet (CHF). Pancreatic islets were isolated by laser capture microdissection (LCM) and subjected to genome-wide transcriptome analyses. Results In the prediabetic state, 2,109 islet transcripts were differentially regulated (>1.5-fold) between HF and CHF diets. Of the genes identified, 39 (e.g. Cacna1d , Chd2 , Clip2 , Igf2bp2 , Dach1 , Tspan8 ) correlated with data from the Diabetes Genetics Initiative and Wellcome Trust Case Control Consortium genome-wide scans for type 2 diabetes, thus validating our approach. HF diet induced early changes in gene expression associated with increased cell-cycle progression, proliferation and differentiation of islet cells, and oxidative stress (e.g. Cdkn1b , Tmem27 , Pax6 , Cat , Prdx4 and Txnip ). In addition, pathway analysis identified oxidative phosphorylation as the predominant gene-set that was significantly upregulated in response to the diabetogenic HF diet. Conclusions/interpretation We demonstrated that LCM of pancreatic islet cells in combination with transcriptional profiling can be successfully used to identify novel candidate genes for diabetes. Our data strongly implicate glucose-induced oxidative stress in disease progression.