ABSTRACT Despite the increasing prevalence of neurodegenerative diseases, the molecular characterization of the brain remains challenging due to limited access to the tissue. Cerebrospinal fluid (CSF) contains a significant proportion of molecular contents originating from the brain, and characterizing these molecules has served as a surrogate to evaluate molecular dysregulation in the brain. Here we performed cell-free messenger RNA (cf-mRNA) RNA-sequencing on 52 human CSF samples, and further compared their transcriptomic profiles to matched plasma samples. In addition, we evaluated the molecular dysregulation of cf-mRNA in CSF between individuals with Alzheimer’s disease (AD) and non-cognitively impaired (NCI) controls. The molecular content of CSF cf-mRNA was distinct from plasma cf-mRNA, with a substantially higher number of brain-associated genes identified in CSF. We identified a large set of dysregulated gene transcripts in the CSF cf-mRNA population of individuals with AD, and these gene transcripts were used to establish a diagnostic classifier to discriminate AD from NCI subjects. Notably, the gene transcripts were enriched in biological processes closely associated with AD, such as brain development and synaptic signaling. We also discovered a subset of gene transcripts within AD subjects that exhibit a strong correlation between CSF and plasma cf-mRNA. This study not only reveals the novel cf-mRNA content of CSF but also highlights the potential of CSF cf-mRNA profiling as a tool to garner pathophysiological insights into AD.
Background: Primary sclerosing cholangitis (PSC) is a rare chronic cholestatic liver disease characterized by multifocal bile duct strictures. To date, underlying molecular mechanisms of PSC remain unclear, and therapeutic options are limited. Methods: We performed cell-free messenger RNA (cf-mRNA) sequencing to characterize the circulating transcriptome of PSC and noninvasively investigate potentially bioactive signals that are associated with PSC. Serum cf-mRNA profiles were compared among 50 individuals with PSC, 20 healthy controls, and 235 individuals with NAFLD. Tissue and cell type-of-origin genes that are dysregulated in subjects with PSC were evaluated. Subsequently, diagnostic classifiers were developed using PSC dysregulated cf-mRNA genes. Results: Differential expression analysis of the cf-mRNA transcriptomes of PSC and healthy controls resulted in identification of 1407 dysregulated genes. Furthermore, differentially expressed genes between PSC and healthy controls or NAFLD shared common genes known to be involved in liver pathophysiology. In particular, genes from liver- and specific cell type-origin, including hepatocyte, HSCs, and KCs, were highly abundant in cf-mRNA of subjects with PSC. Gene cluster analysis revealed that liver-specific genes dysregulated in PSC form a distinct cluster, which corresponded to a subset of the PSC subject population. Finally, we developed a cf-mRNA diagnostic classifier using liver-specific genes that discriminated PSC from healthy control subjects using gene transcripts of liver origin. Conclusions: Blood-based whole-transcriptome cf-mRNA profiling revealed high abundance of liver-specific genes in sera of subjects with PSC, which may be used to diagnose patients with PSC. We identified several unique cf-mRNA profiles of subjects with PSC. These findings may also have utility for noninvasive molecular stratification of subjects with PSC for pharmacotherapy safety and response studies.
BackgroundInflammatory and immune responses are essential and dynamic biological processes that protect the body against acute and chronic adverse stimuli. While conventional protein markers have been used to evaluate systemic inflammatory response, the immunological response to stimulation is complex and involves modulation of a large set of genes and interacting signalling pathways of innate and adaptive immune systems. There is a need for a non-invasive tool that can comprehensively evaluate and monitor molecular dysregulations associated with inflammatory and immune responses in circulation and in inaccessible solid organs.MethodsHere we utilized cell-free messenger RNA (cf-mRNA) RNA-Seq whole transcriptome profiling and computational biology to temporally assess lipopolysaccharide (LPS) induced and JAK inhibitor modulated inflammatory and immune responses in mouse plasma samples.FindingsCf-mRNA profiling displayed a pattern of systemic immune responses elicited by LPS and dysregulation of associated pathways. Moreover, attenuation of several inflammatory pathways, including STAT and interferon pathways, were observed following the treatment of JAK inhibitor. We further identified the dysregulation of liver-specific transcripts in cf-mRNA which reflected changes in the gene-expression pattern in this generally inaccessible biological compartment.InterpretationUsing a preclinical mouse model, we demonstrated the potential of plasma cf-mRNA profiling for systemic and organ-specific characterization of drug-induced molecular alterations that are associated with inflammatory and immune responses.FundingMolecular Stethoscope.
Microglia is one of the major immune cell types in the human brain and plays pivotal roles in regulating inflammatory and immune response in healthy as well as disease states. By analyzing whole transcriptomic data derived from a large cohort of postmortem cortex tissues, we identified two distinct microglial subtypes within the population. The main difference between the two subtypes lies in the differential expression levels of the C1q complex components, Fc γ receptor (CD16) components and CD14. We validated our discovery in independent cohorts of brain autopsy tissues as well as in RNA-seq data generated from isolated microglia. Future investigations into the causes and physiological implications of these subtypes may shed more light on the homeostasis and regulation of the immune related processes in the brain.
Hepatic fibrosis stage is the most important determinant of outcomes in patients with nonalcoholic fatty liver disease (NAFLD). There is an urgent need for noninvasive tests that can accurately stage fibrosis and determine efficacy of interventions. Here, we describe a novel cell-free (cf)-mRNA sequencing approach that can accurately and reproducibly profile low levels of circulating mRNAs and evaluate the feasibility of developing a cf-mRNA-based NAFLD fibrosis classifier. Using separate discovery and validation cohorts with biopsy-confirmed NAFLD (n = 176 and 59, respectively) and healthy subjects (n = 23), we performed serum cf-mRNA RNA-Seq profiling. Differential expression analysis identified 2,498 dysregulated genes between patients with NAFLD and healthy subjects and 134 fibrosis-associated genes in patients with NAFLD. Comparison between cf-mRNA and liver tissue transcripts revealed significant overlap of fibrosis-associated genes and pathways indicating that the circulating cf-mRNA transcriptome reflects molecular changes in the livers of patients with NAFLD. In particular, metabolic and immune pathways reflective of known underlying steatosis and inflammation were highly dysregulated in the cf-mRNA profile of patients with advanced fibrosis. Finally, we used an elastic net ordinal logistic model to develop a classifier that predicts clinically significant fibrosis (F2-F4). In an independent cohort, the cf-mRNA classifier was able to identify 50% of patients with at least 90% probability of clinically significant fibrosis. We demonstrate a novel and robust cf-mRNA-based RNA-Seq platform for noninvasive identification of diverse hepatic molecular disruptions and for fibrosis staging with promising potential for clinical trials and clinical practice. NEW & NOTEWORTHY This work is the first study, to our knowledge, to utilize circulating cell-free mRNA sequencing to develop an NAFLD diagnostic classifier.
Circulating cell-free mRNA (cf-mRNA) holds great promise as a non-invasive diagnostic biomarker. However, cf-mRNA composition and its potential clinical applications remain largely unexplored. Here we show, using Next Generation Sequencing-based profiling, that cf-mRNA is enriched in transcripts derived from the bone marrow compared to circulating cells. Further, longitudinal studies involving bone marrow ablation followed by hematopoietic stem cell transplantation in multiple myeloma and acute myeloid leukemia patients indicate that cf-mRNA levels reflect the transcriptional activity of bone marrow-resident hematopoietic lineages during bone marrow reconstitution. Mechanistically, stimulation of specific bone marrow cell populations in vivo using growth factor pharmacotherapy show that cf-mRNA reflects dynamic functional changes over time associated with cellular activity. Our results shed light on the biology of the circulating transcriptome and highlight the potential utility of cf-mRNA to non-invasively monitor bone marrow involved pathologies.
The lack of accessible noninvasive tools to examine the molecular alterations occurring in the brain limits our understanding of the causes and progression of Alzheimer's disease (AD), as well as the identification of effective therapeutic strategies. Here, we conducted a comprehensive profiling of circulating, cell-free messenger RNA (cf-mRNA) in plasma of 126 patients with AD and 116 healthy controls of similar age. We identified 2591 dysregulated genes in the cf-mRNA of patients with AD, which are enriched in biological processes well known to be associated with AD. Dysregulated genes included brain-specific genes and resembled those identified to be dysregulated in postmortem AD brain tissue. Furthermore, we identified disease-relevant circulating gene transcripts that correlated with the severity of cognitive impairment. These data highlight the potential of high-throughput cf-mRNA sequencing to evaluate AD-related pathophysiological alterations in the brain, leading to precision healthcare solutions that could improve AD patient management.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Circulating cell free mRNA (cf-mRNA) holds great promise as a non-invasive diagnostic biomarker. However, the biological origin of cf-mRNA is still not well understood, limiting the clinical applications of this technology. Here, we use the bone marrow (BM) and pharmacologic manipulation of its resident cells as a window to study the origin of cf-mRNA. Using NGS-based profiling, we show that cf-mRNA is enriched in transcripts derived from the BM compared to circulating cells. Further, BM ablation experiments followed by hematopoietic stem cell transplants in cancer patients show that cf-mRNA levels reflect the transcriptional activity of BM resident hematopoietic lineages during marrow reconstitution. Finally, by stimulating specific BM cell populations in vivo using growth factor therapeutics (i.e. EPO, G-CSF), we show that cf-mRNA reveals dynamic functional changes in growing cell types, suggesting that, unlike other cell-free nucleic acids, cf-mRNA is secreted from living cells, rather than exclusively from apoptotic cells. Our results shed new light on the biology of cf-mRNA and demonstrate its potential applications in clinical practice.
Genomic rearrangements are a hallmark of human cancers. Here, we identify the piggyBac transposable element derived 5 (PGBD5) gene as encoding an active DNA transposase expressed in the majority of childhood solid tumors, including lethal rhabdoid tumors. Using assembly-based whole-genome DNA sequencing, we found previously undefined genomic rearrangements in human rhabdoid tumors. These rearrangements involved PGBD5-specific signal (PSS) sequences at their breakpoints and recurrently inactivated tumor-suppressor genes. PGBD5 was physically associated with genomic PSS sequences that were also sufficient to mediate PGBD5-induced DNA rearrangements in rhabdoid tumor cells. Ectopic expression of PGBD5 in primary immortalized human cells was sufficient to promote cell transformation in vivo. This activity required specific catalytic residues in the PGBD5 transposase domain as well as end-joining DNA repair and induced structural rearrangements with PSS breakpoints. These results define PGBD5 as an oncogenic mutator and provide a plausible mechanism for site-specific DNA rearrangements in childhood and adult solid tumors.
Anton G. Henssen, Richard Koche, Jiali Zhuang, Eileen Jiang, Casie Reed, Amy 4 Eisenberg, Eric Still, Ian C. MacArthur, Elias Rodríguez-Fos, Santiago Gonzalez, Montserrat 5 Puiggròs, Andrew N. Blackford, Christopher E. Mason, Elisa de Stanchina, Mithat Gönen, 6 Anne-Katrin Emde, Minita Shah, Kanika Arora, Catherine Reeves, Nicholas D. Socci, 7 Elizabeth Perlman, Cristina R. Antonescu, Charles W. M. Roberts, Hanno Steen, 8 Elizabeth Mullen, Stephen P. Jackson, David Torrents, Zhiping Weng, Scott A. 9 Armstrong, and Alex Kentsis * 10
Background Numerous human genes encode potentially active DNA transposases or recombinases, but our understanding of their functions remains limited due to shortage of methods to profile their activities on endogenous genomic substrates. Results To enable functional analysis of human transposase-derived genes, we combined forward chemical genetic hypoxanthine-guanine phosphoribosyltransferase 1 ( HPRT1 ) screening with massively parallel paired-end DNA sequencing and structural variant genome assembly and analysis. Here, we report the HPRT1 mutational spectrum induced by the human transposase PGBD5, including PGBD5-specific signal sequences (PSS) that serve as potential genomic rearrangement substrates. Conclusions The discovered PSS motifs and high-throughput forward chemical genomic screening approach should prove useful for the elucidation of endogenous genome remodeling activities of PGBD5 and other domesticated human DNA transposases and recombinases.
Genomic structural variations (SVs) are pervasive in many types of cancers. Characterizing their underlying mechanisms and potential molecular consequences is crucial for understanding the basic biology of tumorigenesis. Here, we engineered a local assembly-based algorithm (laSV) that detects SVs with high accuracy from paired-end high-throughput genomic sequencing data and pinpoints their breakpoints at single base-pair resolution. By applying laSV to 97 tumor-normal paired genomic sequencing datasets across six cancer types produced by The Cancer Genome Atlas Research Network, we discovered that non-allelic homologous recombination is the primary mechanism for generating somatic SVs in acute myeloid leukemia. This finding contrasts with results for the other five types of solid tumors, in which non-homologous end joining and microhomology end joining are the predominant mechanisms. We also found that the genes recursively mutated by single nucleotide alterations differed from the genes recursively mutated by SVs, suggesting that these two types of genetic alterations play different roles during cancer progression. We further characterized how the gene structures of the oncogene JAK1 and the tumor suppressors KDM6A and RB1 are affected by somatic SVs and discussed the potential functional implications of intergenic SVs.
Summary: High-throughput sequencing technologies such as ChIP-seq have deepened our understanding in many biological processes. De novo motif search is one of the key downstream computational analysis following the ChIP-seq experiments and several algorithms have been proposed for this purpose. However, most web-based systems do not perform independent filtering or enrichment analyses to ensure the quality of the discovered motifs. Here, we developed a web server Factorbook Motif Pipeline based on an algorithm used in analyzing ENCODE consortium ChIP-seq datasets. It performs comprehensive analysis on the set of peaks detected from a ChIP-seq experiments: (i) de novo motif discovery; (ii) independent composition and bias analyses and (iii) matching to the annotated motifs. The statistical tests employed in our pipeline provide a reliable measure of confidence as to how significant are the motifs reported in the discovery step. Availability: Factorbook Motif Pipeline source code is accessible through the following URL. https://github.com/joshuabhk/factorbook-motif-pipeline
Insertions and excisions of transposable elements (TEs) affect both the stability and variability of the genome. Studying the dynamics of transposition at the population level can provide crucial insights into the processes and mechanisms of genome evolution. Pooling genomic materials from multiple individuals followed by high-throughput sequencing is an efficient way of characterizing genomic polymorphisms in a population. Here we describe a novel method named TEMP, specifically designed to detect TE movements present with a wide range of frequencies in a population. By combining the information provided by pair-end reads and split reads, TEMP is able to identify both the presence and absence of TE insertions in genomic DNA sequences derived from heterogeneous samples; accurately estimate the frequencies of transposition events in the population and pinpoint junctions of high frequency transposition events at nucleotide resolution. Simulation data indicate that TEMP outperforms other algorithms such as PoPoolationTE, RetroSeq, VariationHunter and GASVPro. TEMP also performs well on whole-genome human data derived from the 1000 Genomes Project. We applied TEMP to characterize the TE frequencies in a wild Drosophila melanogaster population and study the inheritance patterns of TEs during hybrid dysgenesis. We also identified sequence signatures of TE insertion and possible molecular effects of TE movements, such as altered gene expression and piRNA production. TEMP is freely available at github: https://github.com/JialiUMassWengLab/TEMP.git.
The human genome encodes the blueprint of life, but the function of the vast majority of its nearly three billion bases is unknown. The Encyclopedia of DNA Elements (ENCODE) project has systematically mapped regions of transcription, transcription factor association, chromatin structure and histone modification. These data enabled us to assign biochemical functions for 80% of the genome, in particular outside of the well-studied protein-coding regions. Many discovered candidate regulatory elements are physically associated with one another and with expressed genes, providing new insights into the mechanisms of gene regulation. The newly identified elements also show a statistical correspondence to sequence variants linked to human disease, and can thereby guide interpretation of this variation. Overall, the project provides new insights into the organization and regulation of our genes and genome, and is an expansive resource of functional annotations for biomedical research. The human genome sequence provides the underlying code for human biology. Despite intensive study, especially in identifying protein-coding genes, our understanding of the genome is far from complete, particularly with regard to non-coding RNAs, alternatively spliced transcripts and regulatory sequences. Systematic analyses of transcripts and regulatory information are essential for the identification of genes and regulatory regions, and are an important resource for the study of human biology and disease. Such analyses can also provide comprehensive views of the organization and variability of genes and regulatory information across cellular contexts, species and individuals. The Encyclopedia of DNA Elements (ENCODE) project aims to delineate all functional elements encoded in the human genome 1–3. Operationally, we define a functional element as a discrete genome segment that encodes a defined product (for example, protein or non-coding RNA) or displays a reproducible biochemical signature (for example, protein binding, or a specific chromatin structure). Comparative genomic studies suggest that 3–8% of bases are under purifying (negative) selection 4–8 and therefore may be functional, although other analyses have suggested much higher estimates 9–11. In a pilot phase covering 1% of the genome, the ENCODE project annotated 60% of mammalian evolutionarily constrained bases, but also identified many additional putative functional elements without evidence of constraint 2. The advent of more powerful DNA sequencing technologies now enables whole-genome and more precise analyses with a broad repertoire of functional assays. Here we describe the production and initial analysis of 1,640 data sets designed to annotate functional elements in the entire human genome. We integrate results from diverse experiments within …
The Encyclopedia of DNA Elements (ENCODE) consortium aims to identify all functional elements in the human genome including transcripts, transcriptional regulatory regions, along with their chromatin states and DNA methylation patterns. The ENCODE project generates data utilizing a variety of techniques that can enrich for regulatory regions, such as chromatin immunoprecipitation (ChIP), micrococcal nuclease (MNase) digestion and DNase I digestion, followed by deeply sequencing the resulting DNA. As part of the ENCODE project, we have developed a Web-accessible repository accessible at http://factorbook.org. In Wiki format, factorbook is a transcription factor (TF)-centric repository of all ENCODE ChIP-seq datasets on TF-binding regions, as well as the rich analysis results of these data. In the first release, factorbook contains 457 ChIP-seq datasets on 119 TFs in a number of human cell lines, the average profiles of histone modifications and nucleosome positioning around the TF-binding regions, sequence motifs enriched in the regions and the distance and orientation preferences between motif sites.