Summary The breast tumor microenvironment of primary and metastatic sites is a complex milieu of differing cell populations, consisting of tumor cells and the surrounding stroma. Despite recent progress in delineating the immune component of the stroma, the genomic expression landscape of the non-immune stroma (NIS) population and their role in mediating cancer progression and informing effective therapies are not well understood. Here we obtained 52 cell-sorted NIS and epithelial tissue samples across 37 patients from i) normal breast, ii) normal breast adjacent to primary tumor, iii) primary tumor, and iv) metastatic tumor sites. Deep RNA-seq revealed diverging gene expression profiles as the NIS evolves from normal to metastatic tumor tissue, with intra-patient normal-primary variation comparable to inter-patient variation. Significant expression changes between normal and adjacent normal tissue support the notion of a cancer field effect, but extended out to the NIS. Most differentially expressed protein-coding genes and lncRNAs were found to be associated with pattern formation, embryogenesis, and the epithelial-mesenchymal transition. We validated the protein expression changes of a novel candidate gene, C2orf88, by immunohistochemistry staining of representative tissues. Significant mutual information between epithelial ligand and NIS receptor gene expression, across primary and metastatic tissue, suggests a unidirectional model of molecular signaling between the two tissues. Furthermore, survival analyses of 827 luminal breast tumor samples demonstrated the predictive power of the NIS gene expression to inform clinical outcomes. Together, these results highlight the evolution of NIS gene expression in breast tumors and suggest novel therapeutic strategies targeting the microenvironment.
The SK-BR-3 cell line is one of the most important models for HER2+ breast cancers, which affect one in five breast cancer patients. SK-BR-3 is known to be highly rearranged, although much of the variation is in complex and repetitive regions that may be underreported. Addressing this, we sequenced SK-BR-3 using long-read single molecule sequencing from Pacific Biosciences and develop one of the most detailed maps of structural variations (SVs) in a cancer genome available, with nearly 20,000 variants present, most of which were missed by short-read sequencing. Surrounding the important ERBB2 oncogene (also known as HER2), we discover a complex sequence of nested duplications and translocations, suggesting a punctuated progression. Full-length transcriptome sequencing further revealed several novel gene fusions within the nested genomic variants. Combining long-read genome and transcriptome sequencing enables an in-depth analysis of how SVs disrupt the genome and sheds new light on the complex mechanisms involved in cancer genome evolution.
Abstract The breast tumor microenvironment of primary and metastatic sites is a complex milieu of differing cell populations. However, the genomic expression landscape of the tumor stroma and its role in mediating cancer progression and informing effective therapies is not well understood. Here we obtained and cultured 52 cell-sorted stromal tissue samples across 37 patients from normal, primary tumor, and metastatic tumor sites. Deep RNA-seq was performed on poly(A)+ transcripts using a multiplexed barcoding sequencing strategy to increase yield and obviate batch effects. A conservative linear model of expression covariates revealed significant (q<0.05) differential expression of protein-coding genes and lncRNAs in the stroma (80 transcripts in normal versus primary and 108 transcripts in primary versus metastatic). The majority of these genes were found to be associated with pattern formation, embryogenesis, and the epithelial-mesenchymal transition. Among the genes that were silenced in the epithelial cells, IHC staining was initially done for C2orf88, ALOX5AP, and UNC5C which are important in PKA binding, immune response, apoptosis, and neural development. Furthermore, survival analyses of 827 luminal breast tumor samples from TCGA demonstrated the predictive power of the stromal gene expression in informing clinical outcome. Together these results highlight the evolution of stromal gene expression in breast tumors and suggest novel therapeutic strategies targeting the microenvironment. Citation Format: Raditya Utama, Anja Bastian, Narayanan Sadagopan, Ying Jin, Eric Antoniou, Yin Huang, Peter Lee, Gurinder Atwal. Transcriptional landscape of stroma progression in the breast tumor microenvironment [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr LB-357.
An improved reference genome for maize, using single-molecule sequencing and high-resolution optical mapping, enables characterization of structural variation and repetitive regions, and identifies lineage expansions of transposable elements that are unique to maize. The maize genome was initially reported in 2009 but with some accuracy limitations. Doreen Ware and colleagues report a new reference genome for maize using single-molecule sequencing and high-resolution optical mapping. The technique shows improvements in the gene space including resolution of gaps and misassemblies and correction of order and orientation of genes. The authors characterize structural variation and repetitive regions, and identify transposable element lineage expansions unique to maize. Complete and accurate reference genomes and annotations provide fundamental tools for characterization of genetic and functional variation1. These resources facilitate the determination of biological processes and support translation of research findings into improved and sustainable agricultural technologies. Many reference genomes for crop plants have been generated over the past decade, but these genomes are often fragmented and missing complex repeat regions2. Here we report the assembly and annotation of a reference genome of maize, a genetic and agricultural model species, using single-molecule real-time sequencing and high-resolution optical mapping. Relative to the previous reference genome3, our assembly features a 52-fold increase in contig length and notable improvements in the assembly of intergenic spaces and centromeres. Characterization of the repetitive portion of the genome revealed more than 130,000 intact transposable elements, allowing us to identify transposable element lineage expansions that are unique to maize. Gene annotations were updated using 111,000 full-length transcripts obtained by single-molecule real-time sequencing4. In addition, comparative optical mapping of two other inbred maize lines revealed a prevalence of deletions in regions of low gene density and maize lineage-specific genes.
UNLABELLED:Currently available bisulfite sequencing tools frequently suffer from low mapping rates and low methylation calls, especially for data generated from the Illumina sequencer, NextSeq. Here, we introduce a sequential trimming-and-retrieving alignment approach for investigating DNA methylation patterns, which significantly improves the number of mapped reads and covered CpG sites. The method is implemented in an automated analysis toolkit for processing bisulfite sequencing reads.AVAILABILITY AND IMPLEMENTATION:http://mysbfiles.stonybrook.edu/~xuefenwang/software.html and https://github.com/xfwang/BStools.
Background Colorectal cancer (CRC) has the third highest mortality rates among the US population. According to the most recent concept of carcinogenesis, human tumors are organized hierarchically, and the top of it is occupied by malignant stem cells (cancer stem cells, CSCs, or cancer-initiating cells, CICs), which possess unlimited self-renewal and tumor-initiating capacities and high resistance to conventional therapies. To reflect the complexity and diversity of human tumors and to provide clinically and physiologically relevant cancer models, large banks of characterized patient-derived low-passage cell lines, and especially CIC-enriched cell lines, are urgently needed. Principal Findings Here we report the establishment of a novel CIC-enriched, highly tumorigenic and clonogenic colon cancer cell line, CR4, derived from liver metastasis. This stable cell line was established by combining 3D culturing and 2D culturing in stem cell media, subcloning of cells with particular morphology, co-culture with carcinoma associated fibroblasts (CAFs) and serial transplantation to NOD/SCID mice. Using RNA-Seq complete transcriptome profiling of the tumorigenic fraction of the CR4 cells in comparison to the bulk tumor cells, we have identified about 360 differentially expressed transcripts, many of which represent stemness, pluripotency and resistance to treatment. Majority of the established CR4 cells express common markers of stemness, including CD133, CD44, CD166, EpCAM, CD24 and Lgr5. Using immunocytochemical, FACS and western blot analyses, we have shown that a significant ratio of the CR4 cells express key markers of pluripotency markers, including Sox-2, Oct3/4 and c-Myc. Constitutive overactivation of ABC transporters and NF-kB and absence of tumor suppressors p53 and p21 may partially explain exceptional drug resistance of the CR4 cells. Conclusions The highly tumorigenic and clonogenic CIC-enriched CR4 cell line may provide an important new tool to support the discovery of novel diagnostic and/or prognostic biomarkers as well as the development of more effective therapeutic strategies.
BACKGROUND:The use of high throughput genome-sequencing technologies has uncovered a large extent of structural variation in eukaryotic genomes that makes important contributions to genomic diversity and phenotypic variation. When the genomes of different strains of a given organism are compared, whole genome resequencing data are typically aligned to an established reference sequence. However, when the reference differs in significant structural ways from the individuals under study, the analysis is often incomplete or inaccurate.RESULTS:Here, we use rice as a model to demonstrate how improvements in sequencing and assembly technology allow rapid and inexpensive de novo assembly of next generation sequence data into high-quality assemblies that can be directly compared using whole genome alignment to provide an unbiased assessment. Using this approach, we are able to accurately assess the "pan-genome" of three divergent rice varieties and document several megabases of each genome absent in the other two.CONCLUSIONS:Many of the genome-specific loci are annotated to contain genes, reflecting the potential for new biological properties that would be missed by standard reference-mapping approaches. We further provide a detailed analysis of several loci associated with agriculturally important traits, including the S5 hybrid sterility locus, the Sub1 submergence tolerance locus, the LRK gene cluster associated with improved yield, and the Pup1 cluster associated with phosphorus deficiency, illustrating the utility of our approach for biological discovery. All of the data and software are openly available to support further breeding and functional studies of rice and other species.
BACKGROUND:High throughput parallel sequencing, RNA-Seq, has recently emerged as an appealing alternative to microarray in identifying differentially expressed genes (DEG) between biological groups. However, there still exists considerable discrepancy on gene expression measurements and DEG results between the two platforms. The objective of this study was to compare parallel paired-end RNA-Seq and microarray data generated on 5-azadeoxy-cytidine (5-Aza) treated HT-29 colon cancer cells with an additional simulation study.METHODS:We first performed general correlation analysis comparing gene expression profiles on both platforms. An Errors-In-Variables (EIV) regression model was subsequently applied to assess proportional and fixed biases between the two technologies. Then several existing algorithms, designed for DEG identification in RNA-Seq and microarray data, were applied to compare the cross-platform overlaps with respect to DEG lists, which were further validated using qRT-PCR assays on selected genes. Functional analyses were subsequently conducted using Ingenuity Pathway Analysis (IPA).RESULTS:Pearson and Spearman correlation coefficients between the RNA-Seq and microarray data each exceeded 0.80, with 66%~68% overlap of genes on both platforms. The EIV regression model indicated the existence of both fixed and proportional biases between the two platforms. The DESeq and baySeq algorithms (RNA-Seq) and the SAM and eBayes algorithms (microarray) achieved the highest cross-platform overlap rate in DEG results from both experimental and simulated datasets. DESeq method exhibited a better control on the false discovery rate than baySeq on the simulated dataset although it performed slightly inferior to baySeq in the sensitivity test. RNA-Seq and qRT-PCR, but not microarray data, confirmed the expected reversal of SPARC gene suppression after treating HT-29 cells with 5-Aza. Thirty-three IPA canonical pathways were identified by both microarray and RNA-Seq data, 152 pathways by RNA-Seq data only, and none by microarray data only.CONCLUSIONS:These results suggest that RNA-Seq has advantages over microarray in identification of DEGs with the most consistent results generated from DESeq and SAM methods. The EIV regression model reveals both fixed and proportional biases between RNA-Seq and microarray. This may explain in part the lower cross-platform overlap in DEG lists compared to those in detectable genes.
Second-generation sequencers such as the Illumina GAIIx make possible the study of all transcribed loci in a genome across an almost endless dynamic range. Although typical protocols call for starting from at least 1 mu g of total RNA, this is not possible when studying small tissues or rare cell types. This chapter explains how to prepare Illumina sequencing libraries from mouse oocytes. The protocol is also suitable for mural and cumulus cells, flow sorted or laser captured cells.
The design of oligonucleotide sequences for the detection of gene expression in species with disparate volumes of genome and EST sequence information has been broadly studied. However, a congruous strategy has yet to emerge to allow the design of sensitive and specific gene expression detection probes. This study explores the use of a phylogenomic approach to align transcribed sequences to vertebrate protein sequences for the detection of gene families to design genomewide 70-mer oligonucleotide probe sequences for bovine and porcine. The bovine array contains 23,580 probes that target the transcripts of 16,341 genes, about 72% of the total number of bovine genes. The porcine array contains 19,980 probes targeting 15,204 genes, about 76% of the genes in the Ensembl annotation of the pig genome. An initial experiment using the bovine array demonstrates the specificity and sensitivity of the array.
Transcription profiling of ovarian follicles. Understanding the mechanisms by which a single follicle is selected for further ovulation is important to control fertility in mammals. However, development of new treatments is limited by our poor understanding of molecular mechanisms regulating follicular selection. Our hypothesis is that genes involved in the control of cell proliferation and apoptosis are differentially regulated during follicular selection. Our objective was to identify these new genes. Bovine follicles were collected and gene expression levels were measured using microarrays. First, follicles were allocated to three groups, according to the time spent from the initiation of follicular wave to surgery (24 H, 36 H, and 48-60 H). Fifty-seven genes are differentially expressed at a false discovery rate of 5%. These genes are involved in the control of lipid metabolism (P-value = 0.0005), cell proliferation (0.007), cell death (0.003), cell morphology (0.003), and immune response (0.003). Follicles were also grouped into four categories, according to the expected time of deviation (early deviation; 8 mm, mid-deviation; 8.5 mm, late deviation; 9 mm, dominant follicles; 10 mm). One hundred and twenty eight genes are differentially expressed between these four groups, including genes involved in cell proliferation (0.00002), cell death (0.0006), cell-to-cell signaling (0.003) cell, morphology (0.003), lipid metabolism (0.0004), and immune response (0.00007). The expression levels of 10 genes were confirmed using quantitative real time PCR. As expected, we identified new differentially regulated genes involved in the control of cell growth and apoptosis. We also discovered a potential role for immune cells, and in particular macrophages, in follicular selection.
Exposure to ergot alkaloids in endophyte-infected fescue (E+) is associated with impaired animal productivity, especially during heat stress, which is commonly referred to as fescue toxicosis. To elucidate the pathogenesis of this condition, the effects of short-term heat stress (HS) on hepatic gene expression in rats exposed to endophytic ergot alkaloids were evaluated. Rats implanted with telemetric transmitters to continuously measure core temperature were fed an E+ diet and maintained under thermoneutral (TN) conditions (21 degrees C) for 5 d, followed by TN or 31 degrees C (HS) conditions for 3 d. Feed intake (FI) and BW were monitored daily. The E+ and HS-induced alterations in hepatic genes were evaluated using DNA microarrays and PCR analyses. Hepatic antioxidant enzyme activities, as well as the incidence of apoptosis, were determined. As expected, intake of E+ reduced FI and BW from pretreatment levels under TN conditions, with greater reductions during short-term HS. Genes involved in gluconeogenesis and apoptosis were upregulated, whereas genes associated with oxidative phosphorylation, xenobiotic metabolism, antioxidative mechanisms, immune function, cellular proliferation, and chaperone activity were all downregulated with short-term HS. Hepatocytic apoptosis was increased and antioxidant enzyme activity decreased in the livers of rats exposed to HS. The hypothesized, exacerbating effects of HS on the direct, endophytic toxin-related and indirect, reduced caloric intake-associated alterations in hepatic gene expression were clearly demonstrated in rats and may help to elucidate the pathogenesis of fescue toxicosis in various animal species.
The objective of the present study was to evaluate the efficacy of curcumin, an antioxidant found in turmeric (Curcuma longa) powder (TMP), to ameliorate changes in gene expression in the livers of broiler chicks fed aflatoxin B1 (AFB1). Four pen replicates of 5 chicks each were assigned to each of 4 dietary treatments, which included the following: A) basal diet containing no AFB1 or TMP (control), B) basal diet supplemented with TMP (0.5%) that supplied 74 mg/kg of curcumin, C) basal diet supplemented with 1.0 mg of AFB1/kg of diet, and D) basal diet supplemented with TMP that supplied 74 mg/kg of curcumin and 1.0 mg of AFB1/kg of diet. Aflatoxin reduced (P < 0.05) feed intake and BW gain and increased (P < 0.05) relative liver weight. Addition of TMP to the AFB1 diet ameliorated (P < 0.05) the negative effects of AFB1 on growth performance and liver weight. At the end of the 3-wk treatment period, livers were collected (6 per treatment) to evaluate changes in the expression of genes involved in antioxidant function [catalase (CAT), superoxide dismutase (SOD), glutathione peroxidase (GPx), glutathione S-transferase (GST)], biotransformation [epoxide hydrolase (EH), cytochrome P450 1A1 and 2H1 (CYP1A1 and CYP2H1)], and the immune system [interleukins 6 and 2 (IL-6 and IL-2)]. Changes in gene expression were determined using the quantitative real-time PCR technique. There was no statistical difference in gene expression among the 4 treatment groups for CAT and IL-2 genes. Decreased expression of SOD, GST, and EH genes due to AFB1 was alleviated by inclusion of TMP in the diet. Increased expression of IL-6, CYP1A1 and CYP2H1 genes due to AFB1 was also alleviated by TMP. The current study demonstrates partial protective effects of TMP on changes in expression of antioxidant, biotransformation, and immune system genes in livers of chicks fed AFB1. Practical application of the research is supplementation of TMP in diets to prevent or reduce the effects of aflatoxin in chicks fed aflatoxin-contaminated diets.
The objective of this study was to determine the effects of dietary aflatoxin B 1 (AFB 1) on he-patic gene expression in male broiler chicks. Seventy-five 1-d-old male broiler chicks were assigned to 3 dietary treatments (5 replicates of 5 chicks each) from hatch to d 21. The diets contained 0, 1 and 2 mg of AFB 1 /kg of feed. Aflatoxin B 1 reduced (P < 0.05) feed intake, BW gain, serum total proteins, and serum Ca and P, but increased (P < 0.01) liver weights in a dose-dependent manner. Microarray analysis was used to identify shifts in genetic expression associated with the affected physiological processes in chicks fed 0 and 2 mg of AFB 1 / kg of feed to identify potential targets for pharmaco-logical/toxicological intervention. A loop design was used for microarray experiments with 3 technical and 4 biological replicates per treatment group. Ribonucleic acid was extracted from liver tissue, and its quality was determined using gel electrophoresis and spectro-photometry. High-quality RNA was purified from DNA contamination, reverse transcribed, and hybridized to an oligonucleotide microarray chip. Microarray data were analyzed using a 2-step ANOVA model and validated by quantitative real-time PCR of selected genes. Genes with false discovery rates less than 13% and fold change greater than 1.4 were considered differentially expressed. Compared with controls (0 mg of AFB 1 /kg), various genes associated with energy production and fatty acid metabolism (carnitine palmitoyl transferase), growth and development (insulin-like growth factor 1), antioxidant protection (glutathione S transferase), de-toxification (epoxide hydrolase), coagulation (coagula-tion factors IX and X), and immune protection (inter-leukins) were downregulated, whereas genes associated with cell proliferation (ornithine decarboxylase) were upregulated in birds fed 2 mg of AFB 1 /kg. This study demonstrates that AFB 1 exposure at a concentration of 2 mg/kg results in physiological responses associated with altered gene expression in chick livers.
Mice were used to study the effects of chronic heat stress on hepatic gene expression. Twenty-five mice were allocated to either chronic heat stress (34°C) or control (24°C) conditions for a period of 2 weeks from 47 to 60d of age. Nineteen genes differentially expressed in liver were identified using DNA microarrays. Genes involved in the anti-oxidant pathway and metabolism were up-regulated. Genes involved in generation of reactive oxygen radicals and mitochondrial expressed genes were down-regulated. Enzyme activity measurements confirmed the array results. Mice exposed to chronic heat stress showed signs of increased oxidative stress in liver cells.
Fertility losses in male mice occur approximately 18-28 d after heat stress. The objective of this study was to identify gene expression differences in males highly versus lowly fertile after heat stress. Mature male mice were exposed to heat stress (35+/-1 degrees C; n=50) or thermoneutral (21+/-1 degrees C; n=10) conditions for 24 h (Day 0) and hemicastrated (Day 1) to collect tissue for gene expression analyses. Males were subjected to a mating test from Days 18 to 26 when variation in fertility was anticipated. A fertility index was used to rank heat-stressed males and identify those males resistant and susceptible to heat stress, respectively. Microarray analyses were conducted on testis tissues from control (n=5), heat stress resistant (n=5), and heat stress susceptible (n=5) males, and 225 genes were observed to be differentially expressed (P<0.05), including genes involved in chaperone (Canx, Hspcb1, and Tcp1) and catalytic (Fkpb6, Psma7, and Idh1) activity. Expression patterns of these genes were confirmed using real-time RT-PCR. Male progeny from selected sires were similarly divergent in fertility after heat stress. Testicular expression levels of Canx, Hspcb, and Tcp1 genes were determined in these progeny. Hspcb expression was moderately heritable (0.31+/-0.25); however, expression patterns of Canx and Tcp1 were not heritable.
Early embryonic development in the pig requires DNA methylation remodeling of the maternal and paternal genomes. Aberrant remodeling, which can be exasperated by in vitro technologies, is detrimental to development and can result in physiological and anatomic abnormalities in the developing fetus and offspring. Here, we developed and validated a microarray based approach to characterize on a global scale the CpG methylation profiles of porcine gametes and blastocyst stage embryos. The relative methylation in the gamete and blastocyst samples showed that 18.5% (921/4,992) of the DNA clones were found to be significantly different (P<0.01) in at least one of the samples. Furthermore, for the different blastocyst groups, the methylation profile of the in vitro-produced blastocysts was less similar to the in vivo-produced blastocysts as compared to the parthenogenetic- and somatic cell nuclear transfer (SCNT)-produced blastocysts. The microarray results were validated by using bisulfite sequencing for 12 of the genomic regions in liver, sperm, and in vivo-produced blastocysts. These results suggest that a generalized change in global methylation is not responsible for the low developmental potential of blastocysts produced by using in vitro techniques. Instead, the appropriate methylation of a relatively small number of genomic regions in the early embryo may enable early development to occur.