Next Generation sequencing (NGS) has greatly advanced precision oncology. The growing amount and complexity of NGS data, along with other clinical data such as drug responses and measurable residual disease (MRD), present a big challenge for data integration and interpretation. Machine learning, especially deep neural networks, offer a promising approach for efficiently analyzing intricate relationships within large datasets, with the potential of improving clinical outcomes. We performed targeted RNA sequencing in routine diagnostics and sequenced 1849 cases with hematological neoplasms, primarily pediatric B-cell precursor acute lymphoblastic leukemia (BCP-ALL). We used featureCounts in megSAP pipeline and Uniform Manifold Approximation and Projection (UMAP) to analyze the gene expression data, and Bokeh to create the web interface for data integration and visualization. Scikit-learn and PyTorch were used to develop shallow machine learning models and deep learning neural networks (NN), respectively. Our platform integrates various genetic and clinical data based on UMAP analysis of gene expression patterns. It improves the point-of-care decision making and facilitates the discovery of new patterns such as subpopulations with different genetic and clinical features. The platform is supported by machine learning algorithms for cancer subtyping. Six basic machine learning algorithms were coupled with feature selection methods and the best F1 score achieves 98%. We also built a biologically informed deep NN that can accurately predict BCP-ALL subtypes (F1=97%) with a good interpretability, which helps to identify crucial genetic aberrations associated with disease subtypes. Our machine learning based platform can not only provide support for clinical decision-making but also bring novel translational insights for hematological malignancies.
Background: B-cell precursor acute lymphoblastic leukemia (BCP-ALL) is the most common childhood malignancy. Improvements in the genetic-based risk stratification and treatment adaptations have resulted in an increased overall survival, reaching 90% in the contemporary treatment protocols. However, in 25-30% of patients, known as B-other, no recurrent genetic aberrations relevant for the risk stratification and treatment personalization can be detected using conventional cytogenetic methods. Therefore, identification and implementation of new risk-stratifying markers in the current treatment protocols may aid further improvement in the management of children with BCP-ALL. Aims: Our aim was to investigate the applicability of the integrated use of the whole transcriptome RNA sequencing and conventional cytogenetic methods in the routine diagnostic of the B-other BCP-ALL cases. Methods: We performed RNA sequencing and analyzed gene fusions, expression profiles, and mutations in diagnostic samples of 174 children with B-other ALL. In order to further refine genetic classification and validate findings obtained with RNA sequencing, we integrated our results with the findings obtained using immunophenotyping, FISH, karyotyping, arrayCGH and Sanger sequencing. Results: Our analysis unraveled the presence of risk-stratifying fusion transcripts, in the cases in which these alterations were not detectable using conventional cytogenetic methods. These included six cases with cryptic KMT2A rearrangements and one of each with TCF3-PBX1 and TCF3-HLF fusion genes. In addition, we were able to detect 10 cases with recently described fusions involving ZNF384 gene (ZNF384r), four cases with NUTM1 rearrangements and three cases with fusions involving MEF2D gene. Gene expression-based clustering unraveled a subset of B-other cases which cluster together with the known subtypes, indicating the presence of previously described ZNF384r-like, ETV6-RUNX1-like and BCR-ABL1-like subtypes. Furthermore, we identified 27 previously unassigned B-other cases co-clustering together with 13 DUX4-positive ALL cases. This finding suggests that gene expression-based clustering can identify the cases with fusions involving DUX4 gene, known to be cryptic to most of cytogenetic and NGS approaches. We further assessed the ability of the analysis pipeline to detect fusion transcripts in the samples with <50% of tumor blasts. In five tested cases, with BCR-ABL1 or ETV6-RUNX1 fusions, our analysis pipeline was able to detect the presence of fusion transcripts, with high reliability, even in the samples with down to 7% of tumor blasts. Finally, we assessed the applicability of using commonly available EDTA tubes and RNA stabilizing PAXgene tubes for bone marrow sampling, RNA isolation and whole transcriptome sequencing. We assessed the quality of isolated RNA, library complexity and ability of our analysis pipeline to detect relevant fusion transcripts. PCA analysis of matched samples stored in either EDTA or PAXgene tubes showed high similarity in gene expression, while the correlation between TPM expression values and decay constants indicates that genes over-abundant in EDTA samples are not dominated by slow-degrading transcripts. These findings suggest that for short-term storage EDTA tubes are a viable alternative to the RNA stabilizing PAXgene tubes. Summary/Conclusion: Taken together, our findings demonstrate the applicability of whole transcriptome sequencing for personalized diagnostics in pediatric ALL, including the tentative classification of the B-other cases that are difficult to diagnose using conventional methods.