Background: Fusion genes are typically identified by RNA sequencing (RNA-seq) without elucidating the causal genomic breakpoints. However, non-poly(A)-enriched RNA-seq contains large proportions of intronic reads that also span genomic breakpoints. Results: We have developed an algorithm, Dr. Disco, that searches for fusion transcripts by taking an entire reference genome into account as search space. This includes exons but also introns, intergenic regions, and sequences that do not meet splice junction motifs. Using 1,275 RNA-seq samples, we investigated to what extent genomic breakpoints can be extracted from RNA-seq data and their implications regarding poly(A)-enriched and ribosomal RNA-minus RNA-seq data. Comparison with whole-genome sequencing data revealed that most genomic breakpoints are not, or minimally, transcribed while, in contrast, the genomic breakpoints of all 32 TMPRSS2-ERG-positive tumours were present at RNA level. We also revealed tumours in which the ERG breakpoint was located before ERG, which co-existed with additional deletions and messenger RNA that incorporated intergenic cryptic exons. In breast cancer we identified rearrangement hot spots near CCND1 and in glioma near CDK4 and MDM2 and could directly associate this with increased expression. Furthermore, in all datasets we find fusions to intergenic regions, often spanning multiple cryptic exons that potentially encode neo-antigens. Thus, fusion transcripts other than classical gene-to-gene fusions are prominently present and can be identified using RNA-seq. Conclusion: By using the full potential of non-poly(A)-enriched RNA-seq data, sophisticated analysis can reliably identify expressed genomic breakpoints and their transcriptional effects.
Spliced fusion-transcripts are typically identified by RNA-seq without elucidating the causal genomic breakpoints. However, non poly(A)-enriched RNA-seq contains large proportions of intronic reads spanning also genomic breakpoints. Using 1.274 RNA-seq samples, we investigated what additional information is embedded in non poly(A)-enriched RNA-seq data. Here, we present our novel, graph-based, Dr. Disco algorithm that makes use of both intronic and exonic RNA-seq reads to identify not only fusion transcripts but also genomic breakpoints in gene but also in intergenic regions. Dr. Disco identified TMPRSS2-ERG fusions with genomic breakpoints and other transcribed rearrangements from multiple RNA-sequencing cohorts. In breast cancer and glioma samples Dr. Disco identified rearrangement hotspots near CCND1 and MDM2 and could directly associate this with increased expression. A comparison with matched DNA-sequencing revealed that most genomic breakpoints are not, or minimally, transcribed while also revealing highly expressed translocations missed by DNA-seq. By using the full potential of non poly(A)-enriched RNA-seq data, Dr. Disco can reliably identify expressed genomic breakpoints and their transcriptional effects.
Pulmonary hypertension in interstitial lung diseases is associated with increased mortality and hospitalizations and reduced exercise capacity. Interstitial pneumonia with autoimmune features (IPAF) is a recently described interstitial lung disease. The characteristics of pulmonary hypertension in IPAF patients are unknown. We sought to characterize patients with IPAF based on their echocardiographic probability of pulmonary hypertension and compare patients with and without pulmonary hypertension identified by right heart catheterization. We conducted a retrospective study of patients seen in the interstitial lung disease clinic from 2015 to 2018. Forty-seven patients with IPAF were identified. Patients were classified into low, intermediate and high echocardiographic pulmonary hypertension probabilities. A sub-group analysis of patients with pulmonary hypertension and without pulmonary hypertension (IPAF-PH vs. IPAF-no PH) identified by right heart catheterization was also performed. Linear regression analysis was performed to study the association between 6-min-walk-distance (6MWD) and pulmonary vascular resistance (PVR) while adjusting for age and body mass index. Right ventricular hypertrophy (>5 mm), right ventricular enlargement (>41 mm) and right ventricular systolic dysfunction defined as fractional area change% ≤35 was present in 76%, 24%, and 39% of patients, respectively. Pulmonary hypertension was identified in 12.7% of patients. IPAF-PH patients had higher mean pulmonary artery pressure and lower cardiac output compared to the IPAF-no PH group (34 mmHg vs. 19 mmHg, p = 0.002 and 4.0 vs. 5.7 L/min, p = 0.023, respectively). Lower 6MWD was associated with higher PVR on regression analysis (p = 0.002). Pulmonologists should be aware that a significant number of IPAF patients may develop pulmonary hypertension. Reduced 6MWD may suggest the presence of pulmonary hypertension in IPAF patients.
Autism spectrum disorder (ASD) is highly prevalent in subjects with Tuberous Sclerosis Complex (TSC), but we are not still able to reliably predict which infants will develop ASD. This study aimed to identify the early clinical markers of ASD and/or developmental delay (DD) in infants with an early diagnosis of TSC. We prospectively evaluated 82 infants with TSC (6–24 months of age), using a detailed neuropsychological assessment (Bayley Scales of Infant Development—BSID, and Autism Diagnostic Observation Schedule—ADOS), in the context of the EPISTOP (Long-term, prospective study evaluating clinical and molecular biomarkers of EPIleptogenesiS in a genetic model of epilepsy—Tuberous SclerOsis ComPlex) project (NCT02098759). Normal cognitive developmental quotient at 12 months excluded subsequent ASD (negative predictive value 100%). The total score of ADOS at 12 months clearly differentiated children with a future diagnosis of ASD from children without (p = 0.012). Atypical socio-communication behaviors (p < 0.001) were more frequently observed than stereotyped/repetitive behaviors in children with ASD at 24 months. The combined use of BSID and ADOS can reliably identify infants with TSC with a higher risk for ASD at age 6–12 months, allowing for clinicians to target the earliest symptoms of abnormal neurodevelopment with tailored intervention strategies.
Removal of colorectal adenomas is an effective strategy to reduce colorectal cancer (CRC) mortality rates. However, as only a minority of adenomas progress to cancer, such strategies may lead to overtreatment. The present study aimed to characterize adenomas by in-depth molecular profiling, to obtain insights into altered biology associated with the colorectal adenoma-to-carcinoma progression. We obtained low-coverage whole genome sequencing, RNA sequencing and tandem mass spectrometry data for 30 CRCs, 30 adenomas and 18 normal adjacent colon samples. These data were used for DNA copy number aberrations profiling, differential expression, gene set enrichment and gene-dosage effect analysis. Protein expression was independently validated by immunohistochemistry on tissue microarrays and in patient-derived colorectal adenoma organoids. Stroma percentage was determined by digital image analysis of tissue sections. Twenty-four out of 30 adenomas could be unambiguously classified as high risk (n = 9) or low risk (n = 15) of progressing to cancer, based on DNA copy number profiles. Biological processes more prevalent in high-risk than low-risk adenomas were related to proliferation, tumor microenvironment and Notch, Wnt, PI3K/AKT/mTOR and Hedgehog signaling, while metabolic processes and protein secretion were enriched in low-risk adenomas. DNA copy number driven gene-dosage effect in high-risk adenomas and cancers was observed for POFUT1, RPRD1B and EIF6. Increased POFUT1 expression in high-risk adenomas was validated in tissue samples and organoids. High POFUT1 expression was also associated with Notch signaling enrichment and with decreased goblet cells differentiation. In-depth molecular characterization of colorectal adenomas revealed POFUT1 and Notch signaling as potential drivers of tumor progression.
Consensus molecular subtyping is an RNA expression-based classification system for colorectal cancer (CRC). Genomic alterations accumulate during CRC pathogenesis, including the premalignant adenoma stage, leading to changes in RNA expression. Only a minority of adenomas progress to malignancies, a transition that is associated with specific DNA copy number aberrations or microsatellite instability (MSI). We aimed to investigate whether colorectal adenomas can already be stratified into consensus molecular subtype (CMS) classes, and whether specific CMS classes are related to the presence of specific DNA copy number aberrations associated with progression to malignancy. RNA sequencing was performed on 62 adenomas and 59 CRCs. MSI status was determined with polymerase chain reaction-based methodology. DNA copy number was assessed by low-coverage DNA sequencing (n=30) or array-comparative genomic hybridisation (n=32). Adenomas were classified into CMS classes together with CRCs from the study cohort and from The Cancer Genome Atlas (n=556), by use of the established CMS classifier. As a result, 54 of 62 (87%) adenomas were classified according to the CMS. The CMS3 metabolic subtype', which was least common among CRCs, was most prevalent among adenomas (n=45; 73%). One of the two adenomas showing MSI was classified as CMS1 (2%), the MSI immune' subtype. Eight adenomas (13%) were classified as the canonical' CMS2. No adenomas were classified as the mesenchymal' CMS4, consistent with the fact that adenomas lack invasion-associated stroma. The distribution of the CMS classes among adenomas was confirmed in an independent series. CMS3 was enriched with adenomas at low risk of progressing to CRC, whereas relatively more high-risk adenomas were observed in CMS2. We conclude that adenomas can be stratified into the CMS classes. Considering that CMS1 and CMS2 expression signatures may mark adenomas at increased risk of progression, the distribution of the CMS classes among adenomas is consistent with the proportion of adenomas expected to progress to CRC. (c) 2018 The Authors. The Journal of Pathology published by John Wiley & Sons Ltd on behalf of Pathological Society of Great Britain and Ireland.
Abstract Introduction Acute myeloid leukemia (AML) is characterized by uncontrolled proliferation of malignant myeloid progenitor cells in the bone marrow that are arrested in differentiation. In AML, genetic aberrations often involve the same genes and play an important role in risk assessment and treatment of AML. In the WHO classification 2016 (Arber et al., Blood 2016), nine AML subtypes of clinical and prognostic importance are distinguished by distinct and practically mutually exclusive mutations, covering 50-60% of AML cases. By analysing an extended panel of genes, Papaemmanuil et al. (NEJM 2016) developed a purely genomic classification of AML. In this system, 11 groups are defined including 6 entities characterized by chromosomal translocations. Similar as in WHO 2016, these 6 entities each account for less than 5% of AML and are identified by metaphase cytogenetics. Of these 6 entities, 5 groups are defined by fusion genes and one group by inv(3)/t(3;3) leading to overexpression of EVI1. The 5 remaining classes include 4 entities with cytogenetically normal AML defined by mutations in NPM1 (27%), bi-allelic CEBPA (4%), genes regulating RNA splicing, chromatin or transcription (18%) and IDH2 R172 mutations (1%) and one entity characterized by mutations in TP53, a complex karyotype or specific aneuploidies (13%). Although the majority of patients can be classified by this new system, 15% of patients still lack class-defining lesions and expression levels of structurally normal genes, which can also have a decisive prognostic impact, are not considered. We propose that whole transcriptome messenger RNA sequencing provides a single and flexible platform to identify the diversity of genetic aberrations relevant for classification of AML. Methods A panel of hundred AML were analysed and HAMLET (Human AMLExpedited Transcriptomics) was developed as bioinformatics pipeline to detect fusion genes, small variants in thirteen genes, long tandem duplications in FLT3 and KMT2A and overexpression of EVI1. In HAMLET, a new algorithm based on soft clipped reads was developed to detect long tandem repeats in FLT3 and KMT2A. All genetic aberrations called by HAMLET were validated by diagnostic data and targeted re-sequencing. Results The data showed that HAMLET correctly called all genetic aberrations relevant for current classification of AML with high sensitivity and specificity. Moreover, the new soft clipped approach that has been integrated in HAMLET proved to be useful not only to detect long tandem duplications in FLT3 and KMT2A, but also to determine the allelic ratio of mutant-to-wild type FLT3, which is predictive for overall survival. By filtering small variants for predicted importance according to large AML sequencing data sets (Jaiswal et al., NEJM 2017), we classified the 100 AML according to genomic classification and showed that 87 cases were classified in single entities, 4 cases in two subgroups and 9 cases had no class-defining lesions. Of the 9 cases without class-defining lesions, 8 cases had detectable driver mutations and one case had no detectable driver mutation. These numbers perfectly match percentages reported by Papaemmanuil et al. (NEJM 2016). Apart from genetic aberrations that are relevant for current classification of AML, HAMLET also identified additional abnormalities. Of particular interest is NUP98-NSD1 (Hollink et al., Blood 2011), a cryptic fusion gene that is missed by metaphase cytogenetics in three AML with no class-defining lesions, and EVI1 overexpression in 5 cases without inv(3)/t(3;3) including three KMT2A-rearranged AML with extremely poor prognosis (Groschel et al., JCO 2013). Conclusions HAMLET correctly called all genetic aberrations relevant for current classification of AML and provides a wealth of additional information with potential consequences for patient management. In conclusion, HAMLET is a comprehensive and reliable pipeline for RNA sequence analysis that may contribute to better risk assessment and personalized treatment of AML. Disclosures Borras: GenomeScan B.V.: Employment. Janssen:GenomeScan B.V.: Employment.
The ability to form teratomas in vivo containing multiple somatic cell types is regarded as functional evidence of pluripotency for human pluripotent stem cells (hPSCs). Since the Teratoma assay is animal dependent, laborious, and only qualitative, the PluriTest and the hPSC ScoreCard assay have been developed as in vitro alternatives. Here we compared normal hPSCs, induced hPSCs (hiPSCs) with reactivated reprogramming transgenes, and human embryonal carcinoma cells (hECs) in these assays. While normal hPSCs gave rise to typical teratomas, the xenografts of the hECs and the hiPSCs with reactivated reprogramming transgenes were largely undifferentiated and malignant. The hPSC ScoreCard assay confirmed the line-specific differentiation propensities in vitro. However, when undifferentiated cells were analyzed by the PluriTest, only hECs were identified as abnormal whereas all other cell lines were indistinguishable and resembled normal hPSCs. Our results indicate that pluripotency assays are best selected on the basis of intended downstream applications.
Next-generation sequencing is radically changing how DNA diagnostic laboratories operate. What started as a single-gene profession is now developing into gene panel sequencing and whole-exome and whole-genome sequencing (WES/WGS) analyses. With further advances in sequencing technology and concomitant price reductions, WGS will soon become the standard and be routinely offered. Here, we focus on the critical steps involved in performing WGS, with a particular emphasis on points where WGS differs from WES, the important variables that should be taken into account, and the quality control measures that can be taken to monitor the process. The points discussed here, combined with recent publications on guidelines for reporting variants, will facilitate the routine implementation of WGS into a diagnostic setting.