As integral parts of the transcription super elongation complex, the ENL and AF9 proteins are two master regulators of gene expression involved in the development and homeostasis of various tissues. They are encoded by MLLT1 and MLLT3, two genes that frequently undergo oncogenic mutations or chromosomal translocations in cancers. In addition, we discovered an internal promoter driving the production of a YEATS (named after the five proteins first shown to contain this domain: Yaf9, ENL, AF9, Taf14, and Sas5)-domain-devoid AF9 isoform massively expressed in healthy hematopoietic stem cells (HSCs), as well as in the blasts of patients with extremely poor-prognosis acute myeloid leukemia (AML). Here, we show that such internal promoter is also active in a subset of digestive tissues, including the small intestine, colon and stomach. We also reveal the existence of a similar internal promoter in the MLLT1 locus that is inactive in the hematopoietic lineage and AML but functional in other healthy tissues, including the pituitary gland, brain, thyroid, uterus and lung. Furthermore, the YEATS-domain-devoid MLLT1 isoform may exhibit tumor-type-specific biomarker values, whereas full-length MLLT1 does not. It is highly expressed in the gonadotroph subtype of pituitary tumors, and its expression in non-small cell lung cancer inversely correlates with the epidermal growth factor mutation and a transcriptomic signature associated with a good prognosis. Thus, both MLLT1 and MLLT3 have tissue- and tumor-type-specific internal promoters driving expression of similar YEATS domain-devoid isoforms. This suggests that the functions of both MLLT1 and MLLT3 in tissue homeostasis and cancers, as well as their potentials as biomarkers should be re-evaluated taking into account expression of these YEATS-domain-devoid ENL and AF9 transcription factors.
The AF9 (protein AF9) transcription factor, encoded by MLLT3 (mixed-lineage leukemia translocated to 3) on chromosome 9, functions as a chromatin reader. Through its N-terminal YEATS (Yaf9, ENL, AF9, Taf14, and Sas5) protein domain, it interacts with acetylated [1] or crotonylated [2] histone H3, as well as with the PAF1 (RNA polymerase II-associated factor 1 homolog) and P-TEFb (positive transcription elongation factor b) components of the super elongation complex (SEC). AF9 also interacts through its poly-serine domain (Poly-Ser) with the TFIID (Transcription factor II D) subunit of the RNA polymerase II (RNApol II) complex. In addition, its C-terminal transactivation domain, AHD (nuclear anchorage protein1 homology domain), binds other SEC components, such as AFF1 and AFF4 (ALF transcription elongation factor 1 or 4), as well as transcription regulators CBX8 (chromobox 8), DOT1L (disruptor of telomeric silencing 1 like), and BCOR (B cell lymphoma 6 corepressor), as reviewed by Kabra & Bushweller [3] (Figure 1A). Thus, MLLT3 is an integral part of the SEC, which is essential for optimizing the catalytic activity of RNApol II transcription at specific genome loci. Several studies have indicated that MLLT3 is highly and specifically expressed in hematopoietic stem cells (HSCs), but it is rapidly and significantly downregulated during normal differentiation or immediately after HSCs are placed in ex vivo culture. In both scenarios, this shutdown parallels the rapid loss of stemness. Consistently, ectopic expression of MLLT3 significantly prolongs self-renewal capacity of HSCs, suggesting that MLLT3 is a crucial factor for HSC maintenance [4]. Based on standard quantification of RNA-sequencing reads mapping to the MLLT3 locus, we first confirmed that, compared to the MLLT1 paralogue used as an internal control, MLLT3 expression was significantly higher in CD34+ cells than in mature lymphocytes, granulocytes, or monocytes from healthy samples of the Leucegene dataset (Leucegene-NH, detailed in Supplementary Information) (Figure 1B, left panel). To refine this observation, made in CD34+ cells containing a mixture of progenitors but only a few HSCs, we repeated the analysis in HSCs and various stages of progenitor cells sorted form healthy donors (IUCT-NH, detailed in Supplementary Information). The data clearly confirmed that MLLT3 is highly expressed in HSCs but rapidly declines as differentiation proceeds (Figure 1B, right panel). However, closer examination using a k-mer approach (described in Materials and Methods in Supplementary Information), which visualized RNA-sequencing read alignment along the 11 exons (E1-E11) of the reference MLLT3 transcript, revealed an unexpected profile. Strikingly, the substantial MLLT3 expression detected in HSCs was driven by a sharp and pronounced increase in reads starting precisely at the first nucleotide of exon E6 (Figure 1C). This unexpected profile was absent when examining MLLT1 expression in the hematopoietic lineage (Supplementary Figure S1). These findings suggest the expression of one or more 5' end shortened MLLT3 transcripts arising from a hematologic lineage-specific internal promoter located in MLLT3 intron 5. A new set of specific and successive k-mers covering the entire MLLT3 intron 5 revealed the presence of two novel segments retained in poly(A)+ RNAs. Apart from the 5' end of the first segment, which lacked a clearly defined starting point consistent with a probable transcription start site, consensus donor and acceptor splice sites flanked these two segments. These unexpected exons were designated as exon E6a and exon E6b/b' (Figure 1D). Additional k-mer analyses confirmed that these exons were spliced to exon E6, resulting in three possible splice variants: E6a-E6b-E6, E6a-E6b'-E6, and E6a-E6, with the E6a-E6 variant being the most predominant (Figure 1E and Supplementary Figure S2). This assortment of novel CD34+-specific alternative exons was further validated using standard Sanger sequencing following RT-PCR amplification with specific primers (Supplementary Figure S3). Next, exploration of public CHIP-Seq (chromatin immunoprecipitation followed by sequencing) and CAGE (mRNA 5' cap analysis of gene expression) datasets revealed the existence of an alternative P2 promoter in addition to the canonical P1 promoter. Associated with an active promoter H3K4me3 mark and a CAGE peak, this P2 promoter, found exclusively in immature CD34+ cells but absent in CD14+ monocytes, was predicted to drive the expression of transcripts beginning with exon E6a (Supplementary Figure S4). Transcript-specific k-mer quantifications confirmed that the exceptionally high overall level of MLLT3 observed in HSCs was primarily due to these shorter transcripts starting with exon E6a. We have collectively named these shorter transcripts s-MLLT3, in contrast to the reference full-length MLLT3 transcript, referred to as l-MLLT3 (Figure 1F). These findings revealed the existence of an HSC-specific internal promoter (P2) that drives the expression of shorter MLLT3 transcripts (s-MLLT3 mRNAs, Figure 1G). To evaluate the translational potential of these s-MLLT3 transcripts, we cloned a C-terminal myc-tagged version of the predominant E6a-E6 variant into an expression vector (Figure 1H, top), and assessed its protein expression capacity in the HEK (human embryonic kidney) cell line by western blotting. Compared with a similar vector encoding l-MLLT3, the E6a-E6 s-MLLT3 transcript produced shorter s-MLLT3 proteins, (Figure 1H, bottom left). A western blotting analysis of endogenous proteins confirmed the existence of these shorter forms in CD34+ cells but not in the K-562 leukemic cell line, which served as a low-expressing control (Figure 1H, bottom middle). Isoform-specific RT-qPCR quantification further corroborated that, compared to CD34+ cells and full-length l-MLLT3, expression of the three shorter forms (E6a-E6b, E6a-E6b', and E6a-E6) was very low in K-562 cells (Figure 1H, bottom right). These western blot results identified at least two distinct s-MLLT3 proteins, likely arising from alternative translation initiation codons (AUG2 and AUG3) located in exon 7, producing proteins that retain the C-terminal AHD transactivation domain but lack the YEATS and Poly-Ser domains (Supplementary Figure S5). Interestingly, the shorter MLLT3 alternative transcripts initiate within intron 5, which is also a frequent site of chromosomal translocation in acute myeloid leukemia (AML). Specifically, MLLT3 intron 5 is the primary site of the t(9;11) chromosomal translocation, leading to fusion with KMT2A (lysine (K) methyl transferase 2A) [5]. KMT2A is a chromatin writer that deposits epigenetic marks indicating active transcription at specific loci, particularly the HOX genes required for hematopoiesis [6]. The resulting KMT2A-MLLT3 fusion transcript encodes a chimeric protein composing the N-terminal third of KMT2A fused to the C-terminal portion of MLLT3, which lacks the YEATS domain but retains the AHD transactivation domain (Supplementary Figure S6). To investigate global MLLT3 expression in AML, we analyzed two RNA-Seq datasets: IUCT-AML [7] and Beat-AML [8] (detailed in Supplementary Information). Compared with MLLT1, the expression range of MLLT3 was broader, showing a > 120-fold amplitude across samples (Figure 1I). Isoform- specific k-mers targeting alternative (E6a-E6 + E6b/b'-E6 for s-MLLT3) and canonical (E5-E6 for l-MLLT3) exon-exon junctions revealed no correlation between s-MLLT3 and l-MLLT3 expressions. Approximately 20% of samples exhibited high s-MLLT3 levels, whereas in other samples, s-MLLT3 was undetectable despite l-MLLT3 expression (Supplementary Figure S7). This finding suggests differential regulation of the two promoters and/or the resulting transcripts in AML. We next investigated whether the ∼20% of AML samples with high s-MLLT3 expression represented a distinct clinical entity. Samples which were either MECOM+ (myelodysplasia syndrome 1 and EVI1 complex locus) or GATA2-MECOM (Supplementary Figure S8), and those with mutant RUNX1 (runt-related transcription factor 1) and/or TP53 (tumor protein p53) showed higher levels of l-MLLT3, with even greater levels of s-MLLT3 (Supplementary Figure S9, left and middle). No correlation with the KMT2A-MLLT3 translocation was observed. Conversely, NPM1 (nucleophosmin 1)-mutated samples exhibited very low levels of both transcripts (Supplementary Figure S9, left and middle). Notably, elevated s-MLLT3 or l-MLLT3 levels were associated with an adverse ELN2017 (European LeukemiaNet 2017) score (Supplementary Figure S9, right). Median expression-based group separation revealed that patients with the highest overall MLLT3 expression had worse survival outcomes (Figure 1J, left). However, given the lack of correlation between s-MLLT3 and l-MLLT3 expression in AML samples, we assessed their independent impacts on survival. Isoform-specific k-mers showed that s-MLLT3 (but not l-MLLT3) expression significantly influenced poor patient survival (Figure 1J, middle and right panels). In conclusion, these findings demonstrate the existence of an internal promoter within the MLLT3 locus, driving expression of 5'-end-shortened transcripts encoding an AF9 protein lacking the YEAST chromatin reader domain. These alternative transcripts are highly expressed in HSCs and in ∼20% of AML patients with the worst survival outcomes. These results suggest that the role of the MLLT3 locus in HSCs and in AML should be re-evaluated, considering the expression of this YEATS-domain-devoid AF9 transcription factor. Stéphane Pyronnet wrote the manuscript. Chloé Bessière, Ahmed Zamani, Sandra Dailhau, Christian Récher, Marina Bousquet, and Stéphane Pyronnet contributed to the study design, conception, and data analysis. Ahmed Zamani, Romain Pfeifer, and Marina Bousquet designed and performed biological experiments. Chloé Bessière, Sandra Dailhau, Camille Marchet, Benoit Guibert, Anthony Boureux, Raïssa Silva Da Silva, Nicolas Gilbert, and Thérèse Commes developed the k-mer-based bioinformatics tools. Fabienne Meggetto, Christian Touriol, and Marina Bousquet provided comments on and contributed to editing the manuscript. Some of the results presented in this publication are based on data generated by the Leucegene group, primarily based at IRIC in Montreal, Canada, and supported by Genome Canada and Genome Québec. This data was made possible through human AML specimens provided by the BCLQ in Montreal, Canada. Christian Récher declares a consulting or advisory role with Abbvie, Amgen, Astellas, BMS, Boehringer, Jazz Pharmaceuticals, and Servier, and has received research funding from Abbvie, Amgen, Astellas, BMS, Iqvia, and Jazz Pharmaceuticals. All other authors declare no conflict of interest. This work was funded by INSERM, Institut Universitaire du Cancer-Toulouse (IUCT), Labex Toucan, Fondation Leucémie Espoir, Ligue Régionale Contre le Cancer, Fondation ARC, Association Laurette Fugain, Agence Nationale de la Recherche (ANR-18-CE45-0020 Transipedia) and (ANR-22-CE45-0007 full-RNA). Chloé Bessière was supported by Fondation de France, Ahmed Zamani and Raïssa Silva Da Silva by Ligue Nationale Contre le Cancer, Romain Pfeifer by Ministère de l'Enseignement Supérieur et de la Recherche and Fondation ARC, Sandra Dailhau by Ministère de la Santé and Institut National du Cancer (INCA, PRT-K-2022-184CircOma). In accordance with French law, each anonymous volunteer donor or patient was informed, and the HIMIP collection has been declared to the Ministère de l'Enseignement Supérieur et de la Recherche (DC 2008-307). A transfer agreement was obtained (AC 2008-129) after approval by the local ethical committee, Comité de Protection des Personnes Sud-Ouest et Outremer II, and the local Research Ethics Committee of the Etablissement Français du Sang (Toulouse, France, agreement #21PLER2021-007). Clinical and biological annotations have also been declared to the Comité National Informatique et Libertés (CNIL). This study was conducted in accordance with the Declaration of Helsinki. The raw and processed RNA-sequencing data generated in this study have been deposited at the National Center for Biotechnology Information Gene Expression Omnibus (repository number GSE62852). Please note: The publisher is not responsible for the content or functionality of any supporting information supplied by the authors. Any queries (other than missing content) should be directed to the corresponding author for the article.
Acute Myeloid Leukemia (AML) is a genetically and clinically heterogeneous disease that can develop at any age. While AML incidence increases with age and distinct genetic alterations are observed in younger versus older patients, current classification systems do not incorporate age as a defining factor. In this study, we analyzed RNA-seq data from 404 AML patients at initial diagnosis, leveraging a k-mer-based machine learning approach to uncover age-related transcriptomic differences in favorable and adverse risk groups. Our model achieved over 90% accuracy in risk prediction and identified key gene signatures distinguishing ELN2017 favorable and adverse groups. From these signatures, we selected prognostic biomarkers with significant impacts on survival. Additionally, we explored the biological context underlying transcriptomic complexity across age groups, revealing distinct tumor profiles and differences in immune and stromal cell populations, particularly in older patients. These findings underscore the importance of age-related molecular features in AML and provide new insights for risk stratification and therapeutic targeting.
Osteoarthritis (OA) is the most common age-induced degenerative joint disease associated with synovial inflammation, subchondral bone remodeling, and cartilage degradation. One of the significant emerging causes of OA progression is senescent cell accumulation within the joint compartment during lifespan. Currently, there are no therapeutic approaches nor stratification tools that rely on the senescence burden in OA. In this study, we identified the b-series ganglioside 3 (GD3) as new senescent cell surface marker associated with OA. Joint RNA sequencing analysis revealed an increase expression of the GD3 synthase, ST8SIA1 in cartilage, synovial tissue, and subchondral bone marrow from OA patients compared to healthy donors. Moreover, we revealed a strong correlative association between the expression of ST8SIA1 and GD3 production with senescence hallmarks in an in vitro-induced 3D organotypic OA cartilage model but also with cartilage histological grading scores in human and preclinical murine OA joints. Anti-GD3 cell sorting showed that GD3-positive human OA chondrocytes or human OA synoviocytes are enriched in senescence and SASP markers compared to GD3-negative counterparts confirming that GD3 is a cell surface marker linked to the senescence stage. Intra-articular anti-GD3 antibody delivery in experimental OA model reduced local expression of senescence and OA markers in association with a protection against OA-induced subchondral bone remodeling. Our research demonstrates a compelling linkage between ST8SIA1 gene, GD3, and senescence in OA pathology, revealing knowledge and perspectives for a better understanding and anti-senescence treatment of OA pathogenesis.
Acute Myeloid Leukemia (AML) is a highly heterogeneous disease. The current AML classifications are based mainly on molecular markers, including cytogenetics features, fusion genes, and the presence or absence of mutations. In this study, we investigated mutation status in AML patients through RNA-seq data in link with differential gene expression. We applied seven machine learning algorithms to identify the presence or absence of NPM1, IDH1/IDH2, and FLT3-ITD mutations, reaching 95%, 93%, and 87% accuracy, respectively. In each case, the best performing models were complex models, suggesting highly complex biological processes at work behind AML. ### Competing Interest Statement The authors have declared no competing interest.
Analyzing the immense diversity of RNA isoforms in large RNA-seq repositories requires laborious data processing using specialized tools. Indexing techniques based on k-mers have previously been effective at searching for RNA sequences across thousands of RNA-seq libraries but falling short of enabling direct RNA quantification. We show here that RNAs queried in the form of k-mer sets can be quantified in seconds, with a precision akin to that of conventional RNA quantification methods. We showcase several applications by exploring an index of the Cancer Cell Line Encyclopedia (CCLE) collection consisting of 1019 RNA-seq samples. Non-reference RNA sequences such as RNAs harboring driver mutations and fusions, splicing isoforms or RNAs derived from repetitive elements, can be retrieved with high accuracy. Moreover, we show that k-mer indexing offers a powerful means to reveal variant RNAs induced by specific gene alterations, for instance in splicing factors. A web server allows public queries in CCLE and other indexes: . Code is provided to allow users to set up their own server from any RNA-seq dataset. ### Competing Interest Statement The authors have declared no competing interest.
RNA sequencing technology combining short read and long read analysis can be used to detect chimeric RNAs in malignant cells. Here, we propose an integrated approach that uses k-mers to analyze indexed datasets. This approach is used to identify chimeric RNA in chronic myelomonocytic leukemia (CMML) cells, a myeloid malignancy that associates features of myelodysplastic and myeloproliferative neoplasms. In virtually every CMML patient, new generation sequencing identifies one or several somatic driver mutations, typically affecting epigenetic, splicing and signaling genes. In contrast, cytogenetic aberrations are currently detected in only one third of the cases. Nevertheless, chromosomal abnormalities contribute to patient stratification, some of them being associated with higher risk of poor outcome, e.g. through transformation into acute myeloid leukemia (AML). Our approach selects four chimeric RNAs that have been detected and validated in CMML cells. We further focus on NRIP1-MIR99AHG, as this fusion has also recently been detected in AML cells. We show that this fusion encodes three isoforms, including a novel one. Further studies will decipher the biological significance of such a fusion and its potential to improve disease stratification. Taken together, this report demonstrates the ability of a large-scale approach to detect chimeric RNAs in cancer cells. Graphical Abstract
Indexing techniques relying on k-mers have proven effective in searching for RNA sequences across thousands of RNA-seq libraries, but without enabling direct RNA quantification. We show here that arbitrary RNA sequences can be quantified in seconds through their decomposition into k-mers, with a precision akin to that of conventional RNA quantification methods. Using an index of the Cancer Cell Line Encyclopedia (CCLE) collection consisting of 1019 RNA-seq samples, we show that k-mer indexing offers a powerful means to reveal non-reference sequences, and variant RNAs induced by specific gene alterations, for instance in splicing factors.
Motivation Acute Myeloid Leukemia is a highly heterogeneous disease. Although current classifications are well-known and widely adopted, many patients experience drug resistance and disease relapse. New biomarkers are needed to make classifications more reliable and propose personalized treatment. Results We performed tests on a large scale in 3 AML cohorts, 1112 RNAseq samples. The accuracy to distinguish NPM1 mutant and non-mutant patients using machine learning models achieved more than 95% in three different scenarios. Using our approach, we found already described genes associated with NPM1 mutations and new genes to be investigated. Furthermore, we provide a new view to search for signatures/biomarkers and explore diagnosis/prognosis, at the k-mer level. Availability Code available at and . The cohorts used in this article were authorized for use. Contact* therese.commes{at}inserm.fr ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work has been supported by La Ligue Contre le Cancer and the Agence Nationale de la Recherche (TranSipedia and FullRNA projects). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study used only openly available human data that were originally located at GSE49642 and phs001657.v2.p1 accessions. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines and uploaded the relevant EQUATOR Network research reporting checklist(s) and other pertinent material as supplementary files, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors.
Additional file 15 Details and metadata of FANTOM6 dataset used for functional prediction.
Anaplastic large cell lymphomas associated with ALK translocation have a good outcome after CHOP treatment; however, the 2-year relapse rate remains at 30%. Microarray gene-expression profiling of 48 samples obtained at diagnosis was used to identify 47 genes that were differentially expressed between patients with early relapse/progression and no relapse. In the relapsing group, the most significant overrepresented genes were related to the regulation of the immune response and T-cell activation while those in the non-relapsing group were involved in the extracellular matrix. Fluidigm technology gave concordant results for 29 genes, of which FN1, FAM179A, and SLC40A1 had the strongest predictive power after logistic regression and two classification algorithms. In parallel with 39 samples, we used a Kallisto/Sleuth pipeline to analyze RNA sequencing data and identified 20 genes common to the 28 genes validated by Fluidigm technology—notably, the FAM179A and FN1 genes. Interestingly, FN1 also belongs to the gene signature predicting longer survival in diffuse large B-cell lymphomas treated with CHOP. Thus, our molecular signatures indicate that the FN1 gene, a matrix key regulator, might also be involved in the prognosis and the therapeutic response in anaplastic lymphomas.
Background The development of RNA sequencing (RNAseq) and the corresponding emergence of public datasets have created new avenues of transcriptional marker search. The long non-coding RNAs (lncRNAs) constitute an emerging class of transcripts with a potential for high tissue specificity and function. Therefore, we tested the biomarker potential of lncRNAs on Mesenchymal Stem Cells (MSCs), a complex type of adult multipotent stem cells of diverse tissue origins, that is frequently used in clinics but which is lacking extensive characterization. Results We developed a dedicated bioinformatics pipeline for the purpose of building a cell-specific catalogue of unannotated lncRNAs. The pipeline performs ab initio transcript identification, pseudoalignment and uses new methodologies such as a specific k-mer approach for naive quantification of expression in numerous RNAseq data. We next applied it on MSCs, and our pipeline was able to highlight novel lncRNAs with high cell specificity. Furthermore, with original and efficient approaches for functional prediction, we demonstrated that each candidate represents one specific state of MSCs biology. Conclusions We showed that our approach can be employed to harness lncRNAs as cell markers. More specifically, our results suggest different candidates as potential actors in MSCs biology and propose promising directions for future experimental investigations.
The huge body of publicly available RNA-sequencing (RNA-seq) libraries is a treasure of functional information allowing to quantify the expression of known or novel transcripts in tissues. However, transcript quantification commonly relies on alignment methods requiring a lot of computational resources and processing time, which does not scale easily to large datasets. K-mer decomposition constitutes a new way to process RNA-seq data for the identification of transcriptional signatures, as k-mers can be used to quantify accurately gene expression in a less resource-consuming way. We present the Kmerator Suite, a set of three tools designed to extract specific k-mer signatures, quantify these k-mers into RNA-seq datasets and quickly visualize large dataset characteristics. The core tool, Kmerator, produces specific k-mers for 97% of human genes, enabling the measure of gene expression with high accuracy in simulated datasets. KmerExploR, a direct application of Kmerator, uses a set of predictor gene-specific k-mers to infer metadata including library protocol, sample features or contaminations from RNA-seq datasets. KmerExploR results are visualized through a user-friendly interface. Moreover, we demonstrate that the Kmerator Suite can be used for advanced queries targeting known or new biomarkers such as mutations, gene fusions or long non-coding RNAs for human health applications.
Genomic integrity of human pluripotent stem cells (hPSCs) is essential for research and clinical applications. However, genetic abnormalities can accumulate during hPSC generation and routine culture and following gene editing. Their occurrence should be regularly monitored, but the current assays to assess hPSC genomic integrity are not fully suitable for such regular screening. To address this issue, we first carried out a large meta-analysis of all hPSC genetic abnormalities reported in more than 100 publications and identified 738 recurrent genetic abnormalities (i.e., overlapping abnormalities found in at least five distinct scientific publications). We then developed a test based on the droplet digital PCR technology that can potentially detect more than 90% of these hPSC recurrent genetic abnormalities in DNA extracted from culture supernatant samples. This test can be used to routinely screen genomic integrity in hPSCs.
IDH1-mutated gliomas are slow-growing brain tumours which progress into high-grade gliomas. The early molecular events causing this progression are ill-defined. Previous studies revealed that 20% of these tumours already have transformation foci. These foci offer opportunities to better understand malignant progression. We used immunohistochemistry and high throughput RNA profiling to characterize foci cells. These have higher pSTAT3 staining revealing activation of JAK/STAT signaling. They downregulate RNAs involved in Wnt signaling ( DAAM2, SFRP2 ), EGFR signaling ( MLC1 ), cytoskeleton and cell-cell communication ( EZR, GJA1 ). In addition, foci cells show reduced levels of RNA coding for Ethanolamine-Phosphate Phospho-Lyase ( ETNPPL/AGXT2L1 ), a lipid metabolism enzyme. ETNPPL is involved in the catabolism of phosphoethanolamine implicated in membrane synthesis. We detected ETNPPL protein in glioma cells as well as in astrocytes in the human brain. Its nuclear localization suggests additional roles for this enzyme. ETNPPL expression is inversely correlated to glioma grade and we found no ETNPPL protein in glioblastomas. Overexpression of ETNPPL reduces the growth of glioma stem cells indicating that this enzyme opposes gliomagenesis. Collectively, these results suggest that a combined alteration in membrane lipid metabolism and STAT3 pathway promotes IDH1-mutated glioma malignant progression.
The development of RNA sequencing (RNAseq) and corresponding emergence of public datasets have created new avenues of transcriptional marker search. The long non-coding RNAs (lncRNAs) constitute an emerging class of transcripts with a potential for high tissue specificity and function. Using a dedicated bioinformatics pipeline, we propose to construct a cell-specific catalogue of unannotated lncRNAs and to identify the strongest cell markers. This pipeline uses ab initio transcript identification, pseudoalignment and new methodologies such as a specific k-mer approach for naive quantification of expression in numerous RNAseq data. For an application model, we focused on Mesenchymal Stem Cells (MSCs), a type of adult multipotent stem-cells of diverse tissue origins. Frequently used in clinics, these cells lack extensive characterisation. Our pipeline was able to highlight different lncRNAs with high specificity for MSCs. In silico methodologies for functional prediction demonstrated that each candidate represents one specific state of MSCs biology. Together, these results suggest an approach that can be employed to harness lncRNA as cell marker, showing different candidates as potential actors in MSCs biology, while suggesting promising directions for future experimental investigations.
Progress in assisted reproductive technologies strongly relies on understanding the regulation of the dialogue between oocyte and cumulus cells (CCs). Little is known about the role of long non-coding RNAs (lncRNAs) in the human cumulus-oocyte complex (COC). To this aim, publicly available RNA-sequencing data were analyzed to identify lncRNAs that were abundant in metaphase II (MII) oocytes (BCAR4, C3orf56, TUNAR, OOEP-AS1, CASC18, and LINC01118) and CCs (NEAT1, MALAT1, ANXA2P2, MEG3, IL6STP1, and VIM-AS1). These data were validated by RT-qPCR analysis using independent oocytes and CC samples. The functions of the identified lncRNAs were then predicted by constructing lncRNA-mRNA co-expression networks. This analysis suggested that MII oocyte lncRNAs could be involved in chromatin remodeling, cell pluripotency and in driving early embryonic development. CC lncRNAs were co-expressed with genes involved in apoptosis and extracellular matrix-related functions. A bioinformatic analysis of RNA-sequencing data to identify CC lncRNAs that are affected by maternal age showed that lncRNAs with age-related altered expression in CCs are essential for oocyte growth. This comprehensive analysis of lncRNAs expressed in human MII oocytes and CCs could provide biomarkers of oocyte quality for the development of non-invasive tests to identify embryos with high developmental potential.
RNA-Seq approach enables the detection and characterization of fusion or chimeric transcript associated to complex genome rearrangement. Until now, these events are classically identified at DNA level.Here we describe a complete procedure including a novel way of analyzing reads that combines genomic locations and local coverage to directly infer chimeric junctions with a high sensitivity and specificity, allowing identification of different classes of chimeric RNA events. We also recommend the best practices for the bioinformatics analysis and describe the experimental process for RNA validation using real-time PCR and sequencing.
Background: High-throughput next generation sequencing (NGS) technologies enable the detection of biomarkers used for tumor classification, disease monitoring and cancer therapy. Whole-transcriptome analysis using RNA-seq is important, not only as a means of understanding the mechanisms responsible for complex diseases but also to efficiently identify novel genes/exons, splice isoforms, RNA editing, allele-specific mutations, differential gene expression and fusion-transcripts or chimeric RNA (chRNA). Methods: We used Crac, a tool that uses genomic locations and local coverage to classify biological events and directly infer splice and chimeric junctions within a single read. Crac’s algorithm extracts transcriptional chimeric events irrespective of annotation with a high sensitivity, and CracTools was used to aggregate, annotate and filter the chRNA reads. The selected chRNA candidates were validated by real time PCR and sequencing. In order to check the tumor specific expression of chRNA, we analyzed a publicly available dataset using a new tag search approach. Results: We present data related to acute myeloid leukemia (AML) RNA-seq analysis. We highlight novel biological cases of chRNA, in addition to previously well characterized leukemia chRNA. We have identified and validated 17 chRNAs among 3 AML patients: 10 from an AML patient with a translocation between chromosomes 15 and 17 (AML-t(15;17), 4 from patient with normal karyotype (AML-NK) 3 from a patient with chromosomal 16 inversion (AML-inv16). The new fusion transcripts can be classified into four groups according to the exon organization. Conclusions: All groups suggest complex but distinct synthesis mechanisms involving either collinear exons of different genes, non-collinear exons, or exons of different chromosomes. Finally, we check tumor-specific expression in a larger RNA-seq AML cohort and identify new AML biomarkers that could improve diagnosis and prognosis of AML.
Ronnie Alves合作论文数Department of Informatics, University of Minho, Braga, Portugal5
Thierry Lecroq合作论文数LITIS EA 4108
Universite de Rouen4