Background: Fusion genes are typically identified by RNA sequencing (RNA-seq) without elucidating the causal genomic breakpoints. However, non-poly(A)-enriched RNA-seq contains large proportions of intronic reads that also span genomic breakpoints. Results: We have developed an algorithm, Dr. Disco, that searches for fusion transcripts by taking an entire reference genome into account as search space. This includes exons but also introns, intergenic regions, and sequences that do not meet splice junction motifs. Using 1,275 RNA-seq samples, we investigated to what extent genomic breakpoints can be extracted from RNA-seq data and their implications regarding poly(A)-enriched and ribosomal RNA-minus RNA-seq data. Comparison with whole-genome sequencing data revealed that most genomic breakpoints are not, or minimally, transcribed while, in contrast, the genomic breakpoints of all 32 TMPRSS2-ERG-positive tumours were present at RNA level. We also revealed tumours in which the ERG breakpoint was located before ERG, which co-existed with additional deletions and messenger RNA that incorporated intergenic cryptic exons. In breast cancer we identified rearrangement hot spots near CCND1 and in glioma near CDK4 and MDM2 and could directly associate this with increased expression. Furthermore, in all datasets we find fusions to intergenic regions, often spanning multiple cryptic exons that potentially encode neo-antigens. Thus, fusion transcripts other than classical gene-to-gene fusions are prominently present and can be identified using RNA-seq. Conclusion: By using the full potential of non-poly(A)-enriched RNA-seq data, sophisticated analysis can reliably identify expressed genomic breakpoints and their transcriptional effects.
Spliced fusion-transcripts are typically identified by RNA-seq without elucidating the causal genomic breakpoints. However, non poly(A)-enriched RNA-seq contains large proportions of intronic reads spanning also genomic breakpoints. Using 1.274 RNA-seq samples, we investigated what additional information is embedded in non poly(A)-enriched RNA-seq data. Here, we present our novel, graph-based, Dr. Disco algorithm that makes use of both intronic and exonic RNA-seq reads to identify not only fusion transcripts but also genomic breakpoints in gene but also in intergenic regions. Dr. Disco identified TMPRSS2-ERG fusions with genomic breakpoints and other transcribed rearrangements from multiple RNA-sequencing cohorts. In breast cancer and glioma samples Dr. Disco identified rearrangement hotspots near CCND1 and MDM2 and could directly associate this with increased expression. A comparison with matched DNA-sequencing revealed that most genomic breakpoints are not, or minimally, transcribed while also revealing highly expressed translocations missed by DNA-seq. By using the full potential of non poly(A)-enriched RNA-seq data, Dr. Disco can reliably identify expressed genomic breakpoints and their transcriptional effects.
The cancer transcriptome is remarkably complex, including low-abundance transcripts, many not polyadenylated. To fully characterize the transcriptome of localized prostate cancer, we performed ultra-deep total RNA-seq on 144 tumors with rich clinical annotation. This revealed a linear transcriptomic subtype associated with the aggressive intraductal carcinoma sub-histology and a fusion profile that differentiates localized from metastatic disease. Analysis of back-splicing events showed widespread RNA circularization, with the average tumor expressing 7,232 circular RNAs (circRNAs). The degree of circRNA production was correlated to disease progression in multiple patient cohorts. Loss-of-function screening identified 11.3% of highly abundant circRNAs as essential for cell proliferation; for ∼90% of these, their parental linear transcripts were not essential. Individual circRNAs can have distinct functions, with circCSNK1G3 promoting cell growth by interacting with miR-181. These data advocate for adoption of ultra-deep RNA-seq without poly-A selection to interrogate both linear and circular transcriptomes.
Removal of colorectal adenomas is an effective strategy to reduce colorectal cancer (CRC) mortality rates. However, as only a minority of adenomas progress to cancer, such strategies may lead to overtreatment. The present study aimed to characterize adenomas by in-depth molecular profiling, to obtain insights into altered biology associated with the colorectal adenoma-to-carcinoma progression. We obtained low-coverage whole genome sequencing, RNA sequencing and tandem mass spectrometry data for 30 CRCs, 30 adenomas and 18 normal adjacent colon samples. These data were used for DNA copy number aberrations profiling, differential expression, gene set enrichment and gene-dosage effect analysis. Protein expression was independently validated by immunohistochemistry on tissue microarrays and in patient-derived colorectal adenoma organoids. Stroma percentage was determined by digital image analysis of tissue sections. Twenty-four out of 30 adenomas could be unambiguously classified as high risk (n = 9) or low risk (n = 15) of progressing to cancer, based on DNA copy number profiles. Biological processes more prevalent in high-risk than low-risk adenomas were related to proliferation, tumor microenvironment and Notch, Wnt, PI3K/AKT/mTOR and Hedgehog signaling, while metabolic processes and protein secretion were enriched in low-risk adenomas. DNA copy number driven gene-dosage effect in high-risk adenomas and cancers was observed for POFUT1, RPRD1B and EIF6. Increased POFUT1 expression in high-risk adenomas was validated in tissue samples and organoids. High POFUT1 expression was also associated with Notch signaling enrichment and with decreased goblet cells differentiation. In-depth molecular characterization of colorectal adenomas revealed POFUT1 and Notch signaling as potential drivers of tumor progression.
Abstract Background Consensus molecular subtyping (CMS) is an RNA-expression-based classification of colorectal cancers (CRC). Genomic alterations, resulting in specific RNA-expression patterns, accumulate during CRC pathogenesis, including the premalignant adenoma stage. Aim This study aimed to investigate whether differentiation of colorectal neoplasia into CMS classes can already be recognized at the adenoma stage, and whether specific CMS classes could be associated with DNA copy number aberrations that mark adenomas at high-risk of progressing to CRC. Materials and Methods RNA-sequencing was performed on 62 advanced adenomas and 59 CRCs. DNA copy number analysis in adenomas was performed by low-coverage DNA-sequencing (n=30) or array-Comparative Genomic Hybridization (n=32). Microsatellite instability (MSI) status was determined by PCR methods. Adenomas and CRCs were classified into CMS subtypes together with CRCs (n=556) from The Cancer Genome Atlas, using the Random Forest CMS classifier. Results The majority of the adenomas were classified as CMS3 (n=45; 72%), the 'metabolic subtype', recognized as least common among CRCs. No adenomas were classified as the 'mesenchymal' CMS4 subtype. One adenoma was classified as the 'MSI immune' CMS1 (2%), and 8 adenomas as the 'canonical' CMS2 (13%) type. The remaining 8 (13%) could not be classified. The CMS3 class was enriched with adenomas at low-risk of progression. Conclusion Most adenomas were successfully classified. The lack of CMS4 adenomas is consistent with the fact that adenomas lack invasion-associated stroma. Adenomas showing cancer-associated chromosomal instability (CIN) or MSI (24%) were mostly classified as CMS2 and CMS1, respectively. The CMS3 subtype appeared to be the predominant adenoma signature. Citation Format: Malgorzata A. Komor, Linda J. Bosch, Gergana Bounova, Anne S. Bolijn, Pien Delis-van Diemen, Christian Rausch, Youri Hoogstrate, Andrew Stubbs, Mark de Jong, Guido Jenster, Nicole C. an Grieken, Beatriz Carvalho, Lodewyk Wessels, Connie R. Jimenez, Remond J. Fijneman, Gerrit A. Meijer, NGS-ProToCol Consortium. CMS classification of colorectal adenomas [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 3676.
Consensus molecular subtyping is an RNA expression-based classification system for colorectal cancer (CRC). Genomic alterations accumulate during CRC pathogenesis, including the premalignant adenoma stage, leading to changes in RNA expression. Only a minority of adenomas progress to malignancies, a transition that is associated with specific DNA copy number aberrations or microsatellite instability (MSI). We aimed to investigate whether colorectal adenomas can already be stratified into consensus molecular subtype (CMS) classes, and whether specific CMS classes are related to the presence of specific DNA copy number aberrations associated with progression to malignancy. RNA sequencing was performed on 62 adenomas and 59 CRCs. MSI status was determined with polymerase chain reaction-based methodology. DNA copy number was assessed by low-coverage DNA sequencing (n=30) or array-comparative genomic hybridisation (n=32). Adenomas were classified into CMS classes together with CRCs from the study cohort and from The Cancer Genome Atlas (n=556), by use of the established CMS classifier. As a result, 54 of 62 (87%) adenomas were classified according to the CMS. The CMS3 metabolic subtype', which was least common among CRCs, was most prevalent among adenomas (n=45; 73%). One of the two adenomas showing MSI was classified as CMS1 (2%), the MSI immune' subtype. Eight adenomas (13%) were classified as the canonical' CMS2. No adenomas were classified as the mesenchymal' CMS4, consistent with the fact that adenomas lack invasion-associated stroma. The distribution of the CMS classes among adenomas was confirmed in an independent series. CMS3 was enriched with adenomas at low risk of progressing to CRC, whereas relatively more high-risk adenomas were observed in CMS2. We conclude that adenomas can be stratified into the CMS classes. Considering that CMS1 and CMS2 expression signatures may mark adenomas at increased risk of progression, the distribution of the CMS classes among adenomas is consistent with the proportion of adenomas expected to progress to CRC. (c) 2018 The Authors. The Journal of Pathology published by John Wiley & Sons Ltd on behalf of Pathological Society of Great Britain and Ireland.
BACKGROUND:Recently, much progress has been made in the field of gene-expression in early embryogenesis. However, the dynamic behaviour of transcriptomes in individual embryos has hardly been studied yet and the time points at which pools of embryos are collected are usually still quite far apart. Here, we present a high-resolution gene-expression time series with 180 individual zebrafish embryos, obtained from nine different spawns, developmentally ordered and profiled from late blastula to mid-gastrula stage. On average one embryo per minute was analysed. The focus was on identification and description of the transcriptome dynamics of the expressed genes in this embryonic stage, rather than to biologically interpret profiles in cellular processes and pathways.RESULTS:In the late blastula to mid-gastrula stage, we found 6,734 genes being expressed with low variability and rather gradual changes. Ten types of dynamic behaviour were defined, such as genes with continuously increasing or decreasing expression, and all expressed genes were grouped into these types. Also, the exact expression starting and stopping points of several hundred genes during this developmental period could be pinpointed. Although the resolution of the experiment was so high, that we were able to clearly identify four known oscillating genes, no genes were observed with a peaking expression. Additionally, several genes showed expression at two or three distinct levels that strongly related to the spawn an embryo originated from.CONCLUSION:Our unique experimental set-up of whole-transcriptome analysis of 180 individual embryos, provided an unparalleled in-depth insight into the dynamics of early zebrafish embryogenesis. The existence of a tightly regulated embryonic transcriptome program, even between individuals from different spawns is shown. We have made the expression profile of all genes available for domain experts. The fact that we were able to separate the different spawns by their gene-expression variance over all expressed genes, underlines the importance of spawn specificity, as well as the unexpectedly tight gene-expression regulation in early zebrafish embryogenesis.
Number of expressed and non-expressed Ensembl genes and Vega, Refseq and Unigene genes. (XLSX 9Â kb)
Objective: To develop a reliable, reproducible, and sensitive method for investigating gene-expression profiles from individual human oocytes.Design: Five commercially available protocols were investigated for their efficiency to amplify messenger RNA (mRNA) from 54 single human oocytes. Protocols resulting in sufficient yields were further validated using microarray technology. For the validation, mRNA was isolated from 25 human oocytes. To eliminate biological variation, RNA from 13 human oocytes was pooled together and split into 12 identical samples for further mRNA amplification. From 12 oocytes, mRNA was individually isolated.Setting: University medical center and university microarray laboratory.Patient(s): Couples undergoing intracytoplasmic sperm injection treatment were asked to donate their immature oocytes for research, and written informed consent was obtained in all cases. Seventy-nine human oocytes were used in total.Intervention(s): None.Main Outcome Measure(s): Amplification efficiency and microarray profiles.Result(s): Two of the five protocols (WT-Ovation One-Direct and Arcturus RiboAMp HS Plus) resulted in sufficient yields and high success rates and were further validated for their performance in obtaining reliable, reproducible, and sensitive expression profiles from individual human oocytes. Evaluation of these two protocols demonstrated that they both displayed low technical variation and produced highly reproducible profiles (r >= 0.95). One of them identified significantly more transcripts but also had a higher number of false discoveries.Conclusion(s): Two protocols generated ample amounts of mRNA for (quantitative) polymerase chain reaction, microarray, and sequencing techniques. Further validation using a design that discriminates between biological and technical variation showed that both protocols can be used for gene-expression profiling of individual human oocytes. (C) 2016 by American Society for Reproductive Medicine.
Maternal mRNA present in mature oocytes plays an important role in the proper development of the early embryo. As the composition of the maternal transcriptome in general has been studied with pooled mature eggs, potential differences between individual eggs are unknown. Here we present a transcriptome study on individual zebrafish eggs from clutches of five mothers in which we focus on the differences in maternal mRNA abundance per gene between and within clutches. To minimize technical interference, we used mature, unfertilized eggs from siblings. About half of the number of analyzed genes was found to be expressed as maternal RNA. The expressed and non-expressed genes showed that maternal mRNA accumulation is a non-random process, as it is related to specific biological pathways and processes relevant in early embryogenesis. Moreover, it turned out that overall the composition of the maternal transcriptome is tightly regulated as about half of the expressed genes display a less than twofold expression range between the observed minimum and maximum expression values of a gene in the experiment. Even more, the maximum gene-expression difference within clutches is for 88% of the expressed genes lower than twofold. This means that expression differences observed in maternally expressed genes are primarily caused by differences between mothers, with only limited variability between eggs from the same mother. This was underlined by the fact that 99% of the expressed genes were found to be differentially expressed between any of the mothers in an ANOVA test. Furthermore, linking chromosome location, transcription factor binding sites, and miRNA target sites of the genes in clusters of distinct and unique mother-specific gene-expression, suggest biological relevance of the mother-specific signatures in the maternal transcriptome composition. Altogether, the maternal transcriptome composition of mature zebrafish oocytes seems to be tightly regulated with a distinct mother-specific signature.
Maternal mRNA that is present in the mature oocyte plays an important role in the proper development of the early embryo. To elucidate the role of the maternal transcriptome we recently reported a microarray study on individual zebrafish eggs from five different clutches from sibling mothers and showed differences in maternal RNA abundance between and within clutches, “Mother-specific signature in the maternal transcriptome composition of mature, unfertilized Eggs” [1]. Here we provide in detail the applied preprocessing method as well as the R-code to identify expressed and non-expressed genes in the associated transcriptome dataset. Additionally, we provide a website that allows a researcher to search for the expression of their gene of interest in this experiment.
The stability of the photon beam position on synchrotron beamlines is critical for most if not all synchrotron radiation experiments. The position of the beam at the experiment or optical element location is set by the position and angle of the electron beam source as it traverses the magnetic field of the bend-magnet or insertion device. Thus an ideal photon beam monitor would be able to simultaneously measure the photon beam's position and angle, and thus infer the electron beam's position in phase space. X-ray diffraction is commonly used to prepare monochromatic beams on X-ray beamlines usually in the form of a double-crystal monochromator. Diffraction couples the photon wavelength or energy to the incident angle on the lattice planes within the crystal. The beam from such a monochromator will contain a spread of energies due to the vertical divergence of the photon beam from the source. This range of energies can easily cover the absorption edge of a filter element such as iodine at 33.17 ke V. A vertical profile measurement of the photon beam footprint with and without the filter can be used to determine the vertical centroid position and angle of the photon beam. In the measurements described here an imaging detector is used to measure these vertical profiles with an iodine filter that horizontally covers part of the monochromatic beam. The goal was to investigate the use of a combined monochromator, filter and detector as a phase-space beam position monitor. The system was tested for sensitivity to position and angle under a number of synchrotron operating conditions, such as normal operations and special operating modes where the photon beam is intentionally altered in position and angle at the source point. The results are comparable with other methods of beam position measurement and indicate that such a system is feasible in situations where part of the synchrotron beam can be used for the phase-space measurement.
We are developing electron linear accelerator 100Mo(γ,n)99Mo technology as a replacement to nuclear reactor 235U(n,f)99Mo production. We report irradiation of natural molybdenum disks (25 MeV, 10 kW) and 100Mo-enriched disks (35 MeV, 2 kW), their dissolution and the extraction of 99mTc-pertechnetate. Up to 6.2 GBq 99Mo was produced, solvent extraction was performed at >90 % yields of 99mTc, and quality control showed that a product with high radionuclidic and radiochemical purity could be obtained. Irradiated natural molybdenum products showed more impurities (91mNb, 92mNb, 95mNb and 95Nb) than enriched target material. Linear accelerator technology is feasible for production of quality 99Mo/99mTc, particularly when paired with 100Mo-enriched targets.
Structural variations in genomes are commonly studied by (micro)array-based comparative genomic hybridization.The data analysis methods to infer copy number variation in model organisms (human, mouse) are established.In principle, the procedures are based on signal ratios between test and reference samples and the order of the probe targets in the genome.These procedures are less applicable to experiments with non-model organisms, which frequently comprise non-sequenced genomes with an unknown order of probe targets.We therefore present an additional analysis approach, which does not depend on the structural information of a reference genome, and quantifies the presence or absence of a probe target in an unknown genome.The principle is that intensity values of target probes are compared with the intensities of negative-control probes and positive-control probes from a control hybridization, to determine if a probe target is absent or present.In a test, analyzing the genome content of a known bacterial strain: Staphylococcus aureus MRSA252, this approach proved to be success- ful, demonstrated by receiver operating characteristic area under the curve values larger than 0.9995.We show its usability in various applications, such as comparing genome content and validating nextgeneration sequencing reads from eukaryotic nonmodel organisms.
In transcriptomics research, design for experimentation by carefully considering biological, technological, practical and statistical aspects is very important, because the experimental design space is essentially limitless. Usually, the ranges of variable biological parameters of the design space are based on common practices and in turn on phenotypic endpoints. However, specific sub-cellular processes might only be partially reflected by phenotypic endpoints or outside the associated parameter range. Here, we provide a generic protocol for range finding in design for transcriptomics experimentation based on small-scale gene-expression experiments to help in the search for the right location in the design space by analyzing the activity of already known genes of relevant molecular mechanisms. Two examples illustrate the applicability: in-vitro UV-C exposure of mouse embryonic fibroblasts and in-vivo UV-B exposure of mouse skin. Our pragmatic approach is based on: framing a specific biological question and associated gene-set, performing a wide-ranged experiment without replication, eliminating potentially non-relevant genes, and determining the experimental 'sweet spot' by gene-set enrichment plus dose-response correlation analysis. Examination of many cellular processes that are related to UV response, such as DNA repair and cell-cycle arrest, revealed that basically each cellular (sub-) process is active at its own specific spot(s) in the experimental design space. Hence, the use of range finding, based on an affordable protocol like this, enables researchers to conveniently identify the 'sweet spot' for their cellular process of interest in an experimental design space and might have far-reaching implications for experimental standardization.
High bunch charge, femtosecond, electron pulses were generated using a 95 kV electron gun with an S-band RF rebunching cavity. Laser ponderomotive scattering in a counter-propagating beam geometry is shown to provide high sensitivity with the prerequisite spatial and temporal resolution to fully characterize, in situ, both the temporal profile of the electron pulses and RF time timing jitter. With the current beam parameters, we determined a temporal Instrument Response Function (IRF) of 430 fs FWHM. The overall performance of our system is illustrated through the high-quality diffraction data obtained for the measurement of the electron-phonon relaxation dynamics for Si (001).
Grating enhanced ponderomotive scattering is used to directly characterize arrival time jitter and pulse duration for electron bunches longitudinally compressed with a 3GHz RF cavity. We report a long term arrival jitter over 2 hours of 65 fs and full characterization of femtosecond electron pulses, essential to determining the true operational time resolution for atomically resolved dynamics.