Background: Fusion genes are typically identified by RNA sequencing (RNA-seq) without elucidating the causal genomic breakpoints. However, non-poly(A)-enriched RNA-seq contains large proportions of intronic reads that also span genomic breakpoints. Results: We have developed an algorithm, Dr. Disco, that searches for fusion transcripts by taking an entire reference genome into account as search space. This includes exons but also introns, intergenic regions, and sequences that do not meet splice junction motifs. Using 1,275 RNA-seq samples, we investigated to what extent genomic breakpoints can be extracted from RNA-seq data and their implications regarding poly(A)-enriched and ribosomal RNA-minus RNA-seq data. Comparison with whole-genome sequencing data revealed that most genomic breakpoints are not, or minimally, transcribed while, in contrast, the genomic breakpoints of all 32 TMPRSS2-ERG-positive tumours were present at RNA level. We also revealed tumours in which the ERG breakpoint was located before ERG, which co-existed with additional deletions and messenger RNA that incorporated intergenic cryptic exons. In breast cancer we identified rearrangement hot spots near CCND1 and in glioma near CDK4 and MDM2 and could directly associate this with increased expression. Furthermore, in all datasets we find fusions to intergenic regions, often spanning multiple cryptic exons that potentially encode neo-antigens. Thus, fusion transcripts other than classical gene-to-gene fusions are prominently present and can be identified using RNA-seq. Conclusion: By using the full potential of non-poly(A)-enriched RNA-seq data, sophisticated analysis can reliably identify expressed genomic breakpoints and their transcriptional effects.
Spliced fusion-transcripts are typically identified by RNA-seq without elucidating the causal genomic breakpoints. However, non poly(A)-enriched RNA-seq contains large proportions of intronic reads spanning also genomic breakpoints. Using 1.274 RNA-seq samples, we investigated what additional information is embedded in non poly(A)-enriched RNA-seq data. Here, we present our novel, graph-based, Dr. Disco algorithm that makes use of both intronic and exonic RNA-seq reads to identify not only fusion transcripts but also genomic breakpoints in gene but also in intergenic regions. Dr. Disco identified TMPRSS2-ERG fusions with genomic breakpoints and other transcribed rearrangements from multiple RNA-sequencing cohorts. In breast cancer and glioma samples Dr. Disco identified rearrangement hotspots near CCND1 and MDM2 and could directly associate this with increased expression. A comparison with matched DNA-sequencing revealed that most genomic breakpoints are not, or minimally, transcribed while also revealing highly expressed translocations missed by DNA-seq. By using the full potential of non poly(A)-enriched RNA-seq data, Dr. Disco can reliably identify expressed genomic breakpoints and their transcriptional effects.
Consensus molecular subtyping is an RNA expression-based classification system for colorectal cancer (CRC). Genomic alterations accumulate during CRC pathogenesis, including the premalignant adenoma stage, leading to changes in RNA expression. Only a minority of adenomas progress to malignancies, a transition that is associated with specific DNA copy number aberrations or microsatellite instability (MSI). We aimed to investigate whether colorectal adenomas can already be stratified into consensus molecular subtype (CMS) classes, and whether specific CMS classes are related to the presence of specific DNA copy number aberrations associated with progression to malignancy. RNA sequencing was performed on 62 adenomas and 59 CRCs. MSI status was determined with polymerase chain reaction-based methodology. DNA copy number was assessed by low-coverage DNA sequencing (n=30) or array-comparative genomic hybridisation (n=32). Adenomas were classified into CMS classes together with CRCs from the study cohort and from The Cancer Genome Atlas (n=556), by use of the established CMS classifier. As a result, 54 of 62 (87%) adenomas were classified according to the CMS. The CMS3 metabolic subtype', which was least common among CRCs, was most prevalent among adenomas (n=45; 73%). One of the two adenomas showing MSI was classified as CMS1 (2%), the MSI immune' subtype. Eight adenomas (13%) were classified as the canonical' CMS2. No adenomas were classified as the mesenchymal' CMS4, consistent with the fact that adenomas lack invasion-associated stroma. The distribution of the CMS classes among adenomas was confirmed in an independent series. CMS3 was enriched with adenomas at low risk of progressing to CRC, whereas relatively more high-risk adenomas were observed in CMS2. We conclude that adenomas can be stratified into the CMS classes. Considering that CMS1 and CMS2 expression signatures may mark adenomas at increased risk of progression, the distribution of the CMS classes among adenomas is consistent with the proportion of adenomas expected to progress to CRC. (c) 2018 The Authors. The Journal of Pathology published by John Wiley & Sons Ltd on behalf of Pathological Society of Great Britain and Ireland.
BACKGROUND:In rectal cancer, total mesorectal excision surgery combined with preoperative (chemo)radiotherapy reduces local recurrence rates but does not improve overall patient survival, a result that may be due to the harmful side effects and/or co-morbidity of preoperative treatment. New biomarkers are needed to facilitate identification of rectal cancer patients at high risk for local recurrent disease. This would allow for preoperative (chemo)radiotherapy to be restricted to high-risk patients, thereby reducing overtreatment and allowing personalized treatment protocols. We analyzed genome-wide DNA copy number (CN) and allelic alterations in 112 tumors from preoperatively untreated rectal cancer patients. Sixty-six patients with local and/or distant recurrent disease were compared to matched controls without recurrence. Results were validated in a second cohort of tumors from 95 matched rectal cancer patients. Additionally, we performed a meta-analysis that included 42 studies reporting on CN alterations in colorectal cancer and compared results to our own data.RESULTS:The genomic profiles in our study were comparable to other rectal cancer studies. Results of the meta-analysis supported the hypothesis that colon cancer and rectal cancer may be distinct disease entities. In our discovery patient study cohort, allelic retention of chromosome 7 was significantly associated with local recurrent disease. Data from the validation cohort were supportive, albeit not statistically significant, of this finding.CONCLUSIONS:We showed that retention of heterozygosity on chromosome 7 may be associated with local recurrence in rectal cancer. Further research is warranted to elucidate the mechanisms and effect of retention of chromosome 7 on the development of local recurrent disease in rectal cancer.
Implementation of next-generation DNA sequencing (NGS) technology into routine diagnostic genome care requires strategic choices. Instead of theoretical discussions on the consequences of such choices, we compared NGS-based diagnostic practices in eight clinical genetic centers in the Netherlands, based on genetic testing of nine pre-selected patients with cardiomyopathy. We highlight critical implementation choices, including the specific contributions of laboratory and medical specialists, bioinformaticians and researchers to diagnostic genome care, and how these affect interpretation and reporting of variants. Reported pathogenic mutations were consistent for all but one patient. Of the two centers that were inconsistent in their diagnosis, one reported to have found 'no causal variant', thereby underdiagnosing this patient. The other provided an alternative diagnosis, identifying another variant as causal than the other centers. Ethical and legal analysis showed that informed consent procedures in all centers were generally adequate for diagnostic NGS applications that target a limited set of genes, but not for exome- and genome-based diagnosis. We propose changes to further improve and align these procedures, taking into account the blurring boundary between diagnostics and research, and specific counseling options for exome- and genome-based diagnostics. We conclude that alternative diagnoses may infer a certain level of 'greediness' to come to a positive diagnosis in interpreting sequencing results. Moreover, there is an increasing interdependence of clinic, diagnostics and research departments for comprehensive diagnostic genome care. Therefore, we invite clinical geneticists, physicians, researchers, bioinformatics experts and patients to reconsider their role and position in future diagnostic genome care.
At least 17 genomic regions are established as harboring melanoma susceptibility variants, in most instances with genome‐wide levels of significance and replication in independent samples. Based on genome‐wide single nucleotide polymorphism (SNP) data augmented by imputation to the 1,000 Genomes reference panel, we have fine mapped these regions in over 5,000 individuals with melanoma (mainly from the GenoMEL consortium) and over 7,000 ethnically matched controls. A penalized regression approach was used to discover those SNP markers that most parsimoniously explain the observed association in each genomic region. For the majority of the regions, the signal is best explained by a single SNP, which sometimes, as in the tyrosinase region, is a known functional variant. However in five regions the explanation is more complex. At the CDKN2A locus, for example, there is strong evidence that not only multiple SNPs but also multiple genes are involved. Our results illustrate the variability in the biology underlying genome‐wide susceptibility loci and make steps toward accounting for some of the “missing heritability.”
In humans, mutations in IGF1 or IGF1R cause intrauterine and postnatal growth restriction; however, data on mutations in IGF2, encoding insulin-like growth factor (IGF) II, are lacking. We report an IGF2 variant (c.191C→A, p.Ser64Ter) with evidence of pathogenicity in a multigenerational family with four members who have growth restriction. The phenotype affects only family members who have inherited the variant through paternal transmission, a finding that is consistent with the maternal imprinting status of IGF2. The severe growth restriction in affected family members suggests that IGF-II affects postnatal growth in addition to prenatal growth. Furthermore, the dysmorphic features of affected family members are consistent with a role of deficient IGF-II levels in the cause of the Silver-Russell syndrome. (Funded by Bundesministerium für Bildung und Forschung and the European Union.).
Dear Editor, Forwardgenetic screens are commonly usedas unbiased tools to isolate genes responsible for a phenotype of interest. In Arabidopsis thaliana, especially T-DNA activation tagging populations are frequently employed. These populations are generated using vectors containingmultiple copies of the constitutive 35Spromoters derived fromcauliflowermosaic virus (35S CaMV) and often result in isolation of dominant gain-of-function alleles (Weigel et al., 2000; Nakazawa et al., 2003). This allows the study ofmembers of large gene families that are often functionally redundant and, therefore, hard to identify in loss-offunction screens. Moreover, due to the dominant nature, the phenotypes can usually be recognized in T1 generations (Ostergaard and Yanofsky, 2004). Plasmid rescue and thermal asymmetric interlaced PCR (TAIL-PCR) have been effectively employed to recover plant-specific sequences flanking the T-DNA insertions (Weigel et al., 2000; Singer and Burke, 2003). However, in some instances, these two techniques do not yield expected results, probably due to potential sequence complexities following integration events (Laufs et al., 1999). Today, Illumina sequencing is the most widely applied nextgeneration sequencing technology. It has been used to identify, for example, transposon insertions in Zeamays (Williams-Carrier et al., 2010), to obtain large transcript sequences in Sesamum indicum (Wei et al., 2011) and for sequencing of chloroplast genomes (Cronn et al., 2008). Here, we present how Illumina paired-end sequencing (sequences of both the beginning and end of a randomly generated DNA fragment) can be employed to identify T-DNA loci in activation-taggedArabidopsis plants. To identify molecular components involved in controlling hyponastic growth in Arabidopsis thaliana (stress-induced upward leaf movement; reviewed in Van Zanten et al., 2010), we conducted a forward genetic screen on a population of plants tagged with a tetramer of 35S CaMV promoters (Weigel et al., 2000). Plants were screened for altered petiole angles at the start of the experiment (initial angle), after 6 h of ethylene exposure and after 6 h in low light intensity conditions. A set of candidate lines with aberrant petiole angle phenotypes was isolated and designated SDS2, SDS4, DDD1, and EDD1 (Figure 1A and 1B; for designation code and all experimental procedures, see ‘Supplemental Methods’). We confirmed that each line harbored only one T-DNA insertion by segregation analysis (Supplemental Figure 1). To identify T-DNA insertion loci in selected candidates, we first employed plasmid rescue and TAIL-PCR. Both methods repeatedly failed for these lines and, therefore, we adopted a novel approach using Illumina next-generation sequencing. Genomic DNA of the four lines was pooled and subjected to sequencing. 50-bp paired reads with 204 6 63 bases insert size were obtained with a total of 20 419 624 reads. The data generated by the Genome Analyzer IIx was a set of two files; one file contained the forward reads (+) and the second file contained the reverse reads (–). The files were in a proprietary format called SCARF (Solexa compact ASCII read format). First, the files were converted to the standard/Sanger FASTQ format. Subsequently, the forward and reverse reads were aligned separately to the reference sequence of the T-DNA (pSKI015; Weigel et al., 2000) and to the Arabidopsis genome, using the Bowtie aligner (Langmead et al., 2009). For each sequence, the position on the chromosome and the orientation were added. The aligned reads were then paired based on their ID using R programming scripts (RDevelopment Core Team.). As shown in Figure 1C, three types of pairing can occur: Arabidopsis/Arabidopsis, T-DNA/T-DNA, and Arabidopsis/T-DNA. The detection of T-DNA insertion loci is possible based on the latter pairing (Arabidopsis/T-DNA). Unaligned reads were also of interest, as some reads could map over the breakpoint and give its exact location. In order todetect such sequences, unaligned readswere blasted against reference genomes and sub-sequences of the same read that mapped both to Arabidopsis and T-DNAwere selected and analyzed. The pipeline used for searching the Arabidopsis/T-DNA paired reads detected three major breakpoints,
We report a genome-wide association study of melanoma conducted by the GenoMEL consortium based on 317K tagging SNPs for 1,650 selected cases and 4,336 controls, with replication in an additional two cohorts (1,149 selected cases and 964 controls from GenoMEL, and a population-based case-control study in Leeds of 1,163 cases and 903 controls). The genome-wide screen identified five loci with genotyped or imputed SNPs reaching P < 5 x 10(-7). Three of these loci were replicated: 16q24 encompassing MC1R (combined P = 2.54 x 10(-27) for rs258322), 11q14-q21 encompassing TYR (P = 2.41 x 10(-14) for rs1393350) and 9p21 adjacent to MTAP and flanking CDKN2A (P = 4.03 x 10(-7) for rs7023329). MC1R and TYR are associated with pigmentation, freckling and cutaneous sun sensitivity, well-recognized melanoma risk factors. Common variants within the 9p21 locus have not previously been associated with melanoma. Despite wide variation in allele frequency, these genetic variants show notable homogeneity of effect across populations of European ancestry living at different latitudes and show independent association to disease risk.
Background: Kinome profiling aims at the parallel analysis of kinase activities in a cell. Novel developed arrays containing consensus substrates for kinases are used to assess those kinase activities. The arrays described in this paper were already used to determine kinase activities in mammalian systems, but since substrates from many organisms are present we decided to test these arrays for the determination of kinase activities in the model plant species Arabidopsis thaliana.Results: Kinome profiling using Arabidopsis cell extracts resulted in the labelling of many consensus peptides by kinases from the plant, indicating the usefulness of this kinome profiling tool for plants. Method development showed that fresh and frozen plant material could be used to make cell lysates containing active kinases. Dilution of the plant extract increased the signal to noise ratio and non-radioactive ATP enhances full development of spot intensities.Upon infection of Arabidopsis with an avirulent strain of the bacterial pathogen Pseudomonas syringae pv. tomato, we could detect differential kinase activities by measuring phosphorylation of consensus peptides.Conclusion: We show that kinome profiling on arrays with consensus substrates can be used to monitor kinase activities in plants. In a case study we show that upon infection with avirulent P. syringae differential kinase activities can be found. The PepChip can for example be used to purify (unknown) kinases that play a role in P. syringae infection.This paper shows that kinome profiling using arrays of consensus peptides is a valuable new tool to study signal-transduction in plants. It complements the available methods for genomics and proteomics research.