Myelodysplastic syndromes (MDS) are heterogeneous myeloid neoplasms with an increased risk of progression to secondary acute myeloid leukemia (sAML). This study investigates the genomic correlates of disease progression in MDS by profiling active genomic regulatory regions and their transcriptional impact through H3K27ac ChIP-seq and RNA-seq analysis on CD34+ bone marrow progenitors cells isolated from a prospective cohort of 86 and 357 patients, respectively. Our analysis revealed distinct patterns of genomic region activation and transcriptional regulation across different disease stages (low-risk MDS, high-risk MDS and sAML). Unexpectedly, unsupervised clustering revealed a subset of low-risk MDS patients displaying regulatory and transcriptional profiles similar to those of high-risk MDS and sAML, highlighting early molecular events that may predispose patients to disease progression. This subset is characterized by PU.1 genomic occupancy in regions linked to immune and inflammatory responses, increased T-cell and NK activation, and a higher frequency of SRSF2 mutations. Clinically, patients in this group exhibit greater susceptibility to infections and cardiovascular events, along with an elevated risk of disease progression, resulting in a significantly reduced overall survival. Functional studies demonstrate that PU.1 inhibition suppresses MDS cell proliferation and clonogenicity, as impaired PU.1 binding inhibits the activation of key transcriptional programs involved in disease advancement. Collectively, these findings identify epigenetic factors that predispose low-risk MDS patients to progression into high-risk MDS and, ultimately, sAML.
Abstract Myeloproliferative neoplasms (MPN) impose substantial clinical burden through thrombotic complications and leukemic progression, representing the main determinants of patient morbidity and mortality. These clonal malignancies arise from somatic mutations in hematopoietic stem cells (HSC), most commonly JAK2-V617F. Despite therapeutic advances, molecular mechanisms underlying aberrant megakaryopoiesis and pathologic thrombocytosis remain unmet. This study investigates Mediator complex subunit 12-like (MED12L), a functionally uncharacterized protein, in normal and malignant hematopoiesis, identifying a novel regulatory function in platelet (PLT) turnover dynamics. In JAK2V617F mice genetic ablation of the transcription factor Pbx1 resulted in amelioration of thrombocythemia and erythrocytosis. Transcriptomics analysis of their stem/progenitor cells uncovered downregulation of Med12l. Conversely, in JAK2V617F mice, Med12l was upregulated in common Megakaryocyte (Mk)-Erythroid progenitors and in committed Mk (but not erythroid) precursors, rising the hypothesis that MED12L dysregulation contributes to aberrant PLT homeostasis in myeloid malignancies. To assess the functional significance of MED12L in the hematopoietic system, we studied Med12l knockout (KO) mice and performed systematic hematopoietic characterization under steady-state and stress conditions. We also generated Med12l KO/Jak2V617F compound mutants (JM mice) to assess the impact of MED12L absence in vivo in an MPN setting. Clinical validation utilized MPN and myelodysplastic syndrome cohorts. Our investigations reveal a previously unrecognized role for MED12L in regulating PLT lifespan and turnover. Blood analysis of Med12l-KO mice demonstrated an altered PLT compartment, with increased mean PLT volume, distribution width, area, internal complexity, and elevated surface integrin expression with respect to WT. These findings indicate impaired PLT maturation, supported by a higher frequency of reticulated PLT and increased mitochondrial and endoplasmic reticulum content. In vivo PLT depletion with anti-GPIbα antibody revealed impaired recovery kinetics in KO mice, with delayed bone marrow Mk-primed HSC and Mk-committed progenitor expansion. Critically, in vivo biotinylation studies definitively established accelerated platelet clearance in MED12L absence. Importantly, in JM mice the absence of MED12L counteracted thrombocytosis, a hallmark of MPN, without affecting erythrocytosis. Consistently, scRNA-seq from MPN patients showed a higher proportion of MED12L-expressing platelets compared to healthy donors; similarly, in patients affected by myelodysplastic syndrome we observed a strong correlation between the expression of MED12L in CD34+ cells and platelet features.This study uncovers MED12L as a novel player in platelet biology, highlighting its potential as a therapeutic target in myeloid neoplasms. Citation Format: Matteo Brindisi, Laura Crisafulli, Matteo Zampini, Mirko G. Liturri, Alessia Campagna, Luca B. Lanino, Marta Ubezio, Gabriele Todisco, Giulia Maggioni, Antonio Russo, Monica Bacci, Nicla Manes, Elena Riva, Denise Ventura, Nicole Pinocchio, Elisa Calvetti, Dario Strina, Tata Nageswara Rao, Matteo Della Porta, Francesca Ficara. MED12L: A novel player in platelet dysfunction in myeloid neoplasms [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 3985.
BACKGROUND:Myelodysplastic syndromes are clonal hematopoietic stem cell disorders characterized by multistep molecular evolution and a variable risk of leukemic transformation. Given this prognostic heterogeneity, accurate risk stratification is essential for clinical decision-making. We developed ProgEvo, a proprietary framework that infers molecular evolutionary trajectories and integrates them with clinical data to improve prognostic accuracy. METHODS:ProgEvo was trained on 2519 patients in cBioPortal (https://www.cbioportal.org) and validated using two external cohorts: Genomed4All (2043 patients) and a Moffitt Cancer Center (MCC) cohort (2157 patients). Directional evolutionary routes were inferred and selected for prognostic modeling if they were consistently associated with leukemia-free survival. A multivariable feature selection strategy was applied to integrate evolution-consistent variables into the existing IPSS-M model. RESULTS:ProgEvo identified 1765 gene co-occurrences aggregated into 45 directional evolutionary routes. Of these, 18 were validated in the Genomed4All cohort. Five evolution-informed variables, two directional routes (Additional Sex Combs-Like 1 [ASXL1]→KRAS Proto-Oncogene [KRAS] and Serine and Arginine-Rich Splicing Factor 2 [SRSF2]→NRAS Proto-Oncogene [NRAS]), one co-occurrence (NRAS/RUNX Family Transcription Factor 1 [RUNX1]), and two early mutations (ATRX [ATRX Chromatin Remodeler] and Janus Kinase 2 [JAK2]) were integrated into IPSS-M to generate IPSS-M-Evo. The model with "-Evo" improved discrimination for both leukemia-free survival and overall survival, with over 40% of patients restratified in the Genomed4All data. The performance of the model was further confirmed in the MCC cohort. CONCLUSIONS:ProgEvo enabled inference of a molecular evolution model and integration of evolution-informed covariates into clinical prognostic frameworks, supporting the development of the IPSS-M-Evo model. A free web-based tool allows clinicians to calculate the IPSS-M-Evo score and match individual mutational profiles to cohort-derived evolutionary trajectories (https://evoclin.unimib.it/tools/evolution-graphs.html and https://evoclin.unimib.it/tools/ipssmevo.html). (Funded by the European Union and others.).
Abstract Myelodysplastic syndromes (MDS) are clinically and biologically diverse disorders, emphasizing the need for personalized treatment approaches. The International Working Group for Prognostication of MDS (IWG_PM) recently introduced a molecular classification, referred to as the MDS taxonomy, that categorizes patients into 16 subgroups based on 21 gene mutations, 6 cytogenetic abnormalities, and loss of heterozygosity (LOH) at TP53 and TET2 loci. This study sought to validate and enhance the clinical relevance of the MDS taxonomy by analyzing a large retrospective cohort (n = 5136) and transcriptomic data from a prospective cohort (n = 477). The taxonomy successfully identified subgroups with distinct clinical characteristics and disease progression patterns. However, incorporating gene interactions from taxonomy subgroups did not improve the prognostic performance of the Molecular International Prognostic Scoring System (IPSS‐M). We further assessed whether the taxonomy could guide management in patients receiving disease‐modifying therapies. Except for the “TP53‐complex” subgroup, taxonomy classifications were not predictive of hypomethylating agent response or transplant outcomes. Nonetheless, they correlated with overall survival, suggesting that while both IPSS‐M and the taxonomy capture disease biology, other non‐genetic factors may influence treatment response. RNA sequencing confirmed the biological distinctiveness of the taxonomy groups. Transcriptomic profiling of CD34+ bone marrow cells revealed unique, homogeneous gene expression patterns, particularly within the AML‐like, biTET2, SF3B1, and TP53‐complex subgroups. Further integration of multi‐omics data may refine MDS classification, improving clinical decision‐making and guiding the development of targeted therapies.
Contemporary risk models in chronic myelomonocytic leukemia (CMML) focus on the prognostic relevance of individual rather than concurrent mutations. In the current study of 605 Mayo Clinic patients with CMML, we applied machine-learning algorithms in order to examine the influence of cooperative mutational interactions on blast transformation (BT). A hierarchical clustering algorithm was developed and tailored for patient stratification using survival outcomes and co-occurrence of genomic alterations. Five molecular clusters were identified with 3-year blast BT rates ranging from 0% to 100% (AUC at 3 years 0.78). A subsequent Cox regression analysis confirmed independent detrimental impact of specific mutations or their combinations including NPM1 (HR 26.7; p < 0.01), "NRAS + SETBP1" (HR 12.6; p < 0.01), "ASXL1 + BCOR" (HR 8.4; p < 0.01), "ASXL1 + RUNX1" (HR 2.2, p < 0.01), JAK2 (HR 2.1; p < 0.01), and "ASXL1 + TET2" (HR 1.7; p = 0.02) while "PHF6+wild-type ASXL1" (HR 5.61e-10; p < 0.01) had a favorable impact. Furthermore, compared to NPM1 wild-type cases, NPM1-mutated patients were less likely to have co-occurring mutations involving ASXL1 (0% vs. 43%, p < 0.01), RUNX1 (0% vs. 17%, p = 0.02), and SRSF2 (7% vs. 39%, p < 0.01) and were more likely DNMT3A (71% vs. 7%, p < 0.01). The prognostic relevance of "NRAS + SETBP1", "ASXL1 + RUNX1", NPM1 and BCOR was validated in an external cohort from Italy (N = 501). Taken together, these observations highlight i) the possibility of prognostic interaction of mutations in CMML that should be considered in the development of future risk models and ii) the distinct genotypic and prognostic characteristics of NPM1-mutated CMML.
Abstract Secondary acute myeloid leukemia (AML) comprises heterogeneous entities, unified by poor prognosis. We evaluated the associations of genetic profiles with blast counts and patients’ history in a cohort of 924 patients with myelodysplastic syndrome (MDS)/AML or AML, classified according to the International Consensus Classification (ICC). The cohort included 109 patients with “mutated TP53,” 497 with “myelodysplasia-related (MDR) gene mutation,” 93 with “MDR-cytogenetic abnormality,” 77 were therapy-related, and 136 controls, categorized as “not otherwise specified” (NOS) AML. Exploring the ICC hierarchy, AML and MDS/AML categories with “mutated TP53” and “MDR-cytogenetic abnormality” presented similar biology and prognosis, irrespective of blast counts. Conversely, in MDS/AML with “MDR gene mutation” and NOS, profiles significantly differed from AML and were characterized by a higher number of mutations in STAG2, SRSF2, ASXL1 and TET2. This corresponded to improved survival in MDS/AML vs AML (MDR-gene mutation: median overall survival 24.8 vs 13.6 months, P< .0001; and NOS: 49.9 vs 19.2 months, P = .028). Within each ICC-defined AML category, a prior MDS history vs de novo onset did not impact on patients’ prognosis. We then analyzed secondary AML, defined by “prior MDS or MDS/MPN” or “therapy-related” (t-AML), as diagnostic qualifiers. According to European LeukemiaNet (ELN) 2022, AML progressing from MDS or MDS/myeloproliferative neoplasm (MPN) (AML post-MDS) mostly clustered in the adverse-risk group (84.1%), whereas t-AML showed more heterogeneous ELN profiles (12.9% favorable, 33.8% intermediate, and 53.3% adverse risk) reflecting diverse overall survival. Our findings underscore that genetic features and the ICC classification reliably capture disease biology, refine risk stratification, and ultimately guide treatment decisions in most secondary AML and MDS/AML.
Abstract We explored the impact of luspatercept therapy on overall survival (OS) and possible predictors of response in low‐risk (LR) myelodysplastic syndrome (MDS) patients. We evaluated 331 anemic patients treated with luspatercept. Hematological response (HI) was defined as (i) hemoglobin (Hb) increase of ≥1.5 g/dL in nontransfusion‐dependent (NTD) patients, and (ii) red blood cell (RBC) transfusion independence (TI) with a concomitant Hb increase of ≥1.5 g/dL, or RBC‐TI without an Hb increase of 1.5 g/dL, or >50% reduction in RBC transfusion burden (TB) for TD patients. Response was observed in 166 patients (50.2%), with significantly higher response in NTD and low TB versus high TB patients (p < 0.001). A significant correlation between lower Molecular International Prognostic Scoring System (IPSS‐M) risk scores and response was observed. No statistically significant difference in HI was found in SF3B1‐mutated versus wild‐type MDS patients (53.8% vs. 40.1%, respectively). SF3B1mut hotspots (K700E vs. others) and variant allele frequencies (VAFs; <38% VAF vs. ≥38% VAF) did not impact on HI. SF3B1‐mutated MDS with del5q showed inferior HI compared to other LR‐MDS (p = 0.046). The median treatment duration overall was 35 weeks (20.86–90.29), the median time to response was 11 weeks (8.71–21.86), and the median duration of response was 65 weeks (26.5–114). After a median follow‐up of 13 months, median OS was not reached (NR) for responders and 24 months for nonresponders (hazard ratio [HR] 0.25, 95% confidence interval 0.14–0.44, p < 0.001). This analysis of 331 luspatercept real‐life‐treated LR‐MDS patients demonstrated a significant OS benefit upon luspatercept response. Low baseline RBC‐TB and lower risk IPSS‐M scores correlated with higher HI and could constitute predictive markers of response.
Background Effective communication between clinicians and patients is vital in hematology for personalized treatment. However, traditional practice is time-consuming and prioritizes clinical data over patient-reported outcomes (PROs), limiting direct patient interaction. Existing speech-to-text AI-based solutions promise to solve these issues but still present significant barriers: these solutions are often costly due to reliance on third-party models, lack of specialization on specific medical domains, are not easy to install on-premise and integrated with hospital systems, raising privacy and compliance issues. To overcome these limitations, we have developed a hematology-tailored platform called Ambient AI. This work presents results of an ongoing prospective study, currently under validation with real patients. The aim is to assess the platform's ability to support medical documentation and optimize patient enrollment and management in clinical trials within a real-world hematology setting. Methods Ambient AI records, transcribes in real-time clinician-patient conversations and extracts clinical information while filtering out non-essential details. It leverages four key modules: real-time conversation transcription, an AI-powered system for generating medical reports, a clinical data extraction module for research dataset and a dedicated component for optimizing patient management in clinical trials. The platform generates structured medical reports for physician review, modification, and approval. Once validated, reports are saved, and intermediate data is deleted to ensure privacy and compliance. The platform integrates local AI models using speech-to-text technology and open-source large language models (LLM) fine-tuned for hematology. The validation framework assesses: 1) technological performance, using metrics as Jaccard Similarity (JS) for text accuracy and Word Mover Distance (WMD) for contextual understanding, and an LLM as a judge for expert assessment; 2) clinical fidelity and physician satisfaction; and 3) patient experience and engagement measured through surveys. Results A preliminary analysis of 100 simulated reports, using an open-source Gemma3 12B quantized model, showed that AI-assisted transcription improves documentation efficiency by reducing manual data entry. AI-generated reports had high clinical relevance, requiring minimal physician edits. Strong transcription accuracy (JS 0.85) and effective contextual interpretation (WMD 0.81) were observed. However, some medical terminology or sentences were not properly captured. To address this, the model was fine-tuned with domain knowledge, and a sample of 30 reports from the initial 100 was used for performance assessment. Results showed improved metrics (JS 0.88, WMD 0.73). For a final assessment, a team of five physicians was presented with transcribed text and two reports, and tasked with selecting the more appropriate. The fine-tuned model's report was preferred on average in 28 out of 30 instances (93.3%). The Ambient AI platform validation is currently undergoing at Humanitas Research Hospital, Italy, involving 1,000 patients. Early results indicate increased physician satisfaction, improved workflow efficiency, and enhanced patient engagement. Moreover, we developed an AI tool leveraging an LLM for automatic data extraction from medical records, enabling structured dataset generation for research. To optimize clinical trial management we implemented a module that: 1) automatically scans medical records to identify eligible patients for ongoing clinical trials; 2) reports and grades adverse events during trial visits according to international guidelines; 3) suggests potential drug modification schedules based on study protocols; and 4) generates automated patient visit reports for Clinical Research Organizations (CROs), reducing the need for in-person monitoring visits. Conclusion The Ambient AI platform aims to improve hematology workflows by enhancing communication, streamlining documentation, and integrating PROs into clinical practice. This fosters better doctor-patient interaction, a key aspect of patient-centered care, providing support and empowering individuals. Preliminary findings show its effectiveness in accurate data collection, with ongoing validation. The technology may also facilitate patient selection and management in clinical trials, supporting precision medicine.
Background. The emergence of Generative AI is expanding the use of Synthetic Data (SD) for ground-breaking applications, such as digital twins for evidence generation and synthetic control arms. However, their adoption is limited by technical barriers and unclear regulatory validation process, particularly due to the absence of robust tools and standardized approaches to assess its clinical applicability. This is particularly evident in hematology, where leveraging on large-scale, multimodal data is essential to develop personalized treatments and address unmet needs in rare diseases like Myeloid Neoplasms (MN). This study presents SAFE (Synthetic vAlidation FramEwork), a comprehensive framework for evaluating multimodal SD based on statistical fidelity, clinical utility and privacy preservability, validated in the MN clinical setting. Methods. We applied SAFE on SD generated from the extensive TITAN cohort (n=20,054), a retrospective multimodal dataset comprising 7104 AML, 8410 MDS, 2986 MDS/MPN and 1554 MF cases, including clinical, genomic, transcriptomic (bulk RNA-seq) and histopathological images data. SAFE was developed within the SYNTHEMA and SYNTHIA consortia as a modular, extensible, Python-based solution comprising three main analysis modules tailored for distinct data modalities: safe.tabular, safe.series (for longitudinal data) and safe.images. Each module evaluates statistical fidelity through different metrics specific for each clinical modality and data type, ensuring privacy preservability by preventing any link with real patients and their replication. SAFE introduces an innovative synthetic RNA-seq validation pipeline, specifically designed to tackle the biological complexity of transcriptomic data. By integrating a clinically-driven layer across multiple data modalities, SAFE provides disease-specific and interpretable insights on SD usability. In the MN setting, this was demonstrated by using SD for disease classification and personalized prognostic evaluation through the MOSAIC framework (PMID: 38875514). Results. We generated a multimodal synthetic cohort (n=20,054) using a TRAIN SD platform (www.train-ai.eu), accurately mirroring real dataset's disease stratification. We applied safe.tabular framework, summarizing validation performance on each modality through key metrics: Clinical Synthetic Fidelity (CSF), Genomic Synthetic Fidelity (GSF), Clinical Synthetic Utility (CSU), Transcriptomics Synthetic Fidelity (TSF) and Privacy Synthetic Score (PSS). These were combined into an overall SAFE score, using optimal thresholds between 85–95% to balance data accuracy and privacy. The analysis revealed high concordance for clinical feature distributions and correlations (CSF: 91%), as well as for genomic alterations and pairwise gene associations (GSF: 88%). Clinical utility was evaluated using the MOSAIC framework, demonstrating that synthetic patients had comparable outcomes to real patients in unsupervised patient stratification, prognostic scoring and survival analysis (log-rank p-value=0.8) and when applying conventional scoring systems (CSU: 90.2%). Synthetic RNA-seq quality and its biological fidelity were validated by transcriptomic profiles distribution, differential expression and enrichment analyses, with Jaccard, Dice and Spearman correlation metrics, confirming that synthetic expression accurately reflects functional alterations associated with clinical conditions (TSF: 88%). Privacy was assessed across modalities via Distance to Closest Record and Nearest Neighbor Distance Ratio, confirming a low re-identification risk (PSS: 86%). For digital pathology slides, safe.images achieved a Fréchet Inception Distance of 8.3 and Multi-Scale Structural Similarity Index values ranging from 0.028 to 0.216, indicating good realism and intra-class diversity. Extracted morphological, color and Haralick features also showed comparable distributions. The overall SAFE score of 89% reflected a high-quality generation, with all results compiled into an automated, interpretable report. Conclusions. SAFE advances standard validation by embedding clinical expertise into its design, proving its robustness on the TITAN dataset. By offering a comprehensive, disease-specific evaluation of fidelity and clinical utility, SAFE stands out from existing tools, supporting reliable clinical research and potentially informing regulatory adoption of AI-generated evidence in hematology.
Background Bulk RNA sequencing is a powerful and cost-effective high-throughput technology that provides comprehensive insights into gene expression profiles. However, one key limitation is the lack of cellular resolution as bulk RNA-seq measures an average expression across all cell types in a sample. In complex biological systems, particularly in diseases where immunity plays a crucial role, a more granular understanding of the immune microenvironment is key to identifying novel therapeutic targets, improving patient stratification and guiding personalized treatment strategies. Methods To estimate cellular compositions from bulk transcriptomics, we can use a conventional deconvolution approach, but challenges remain for less characterized tissues: lack of disease-specific immune signatures, background predictions and rigid references. To mitigate some of those limitations, we developed a novel strategy that integrates bulk RNA-seq and immune signatures derived from our newly generated single-cell CITE-seq dataset, using the single-sample Gene Set Enrichment Analysis (ssGSEA) scoring method. By leveraging single-cell and bulk RNA-seq, we identify robust, disease-specific immune signatures and incorporate them into risk stratification models. Using myelodysplastic syndromes (MDS) as a case study, we show that this approach improves patient prognostication and identifies potential immunological targets. Results We utilized bulk RNA-seq and clinical data from three independent MDS cohorts encompassing 723 (cohort 1), 324 (cohort 2) and 432 (cohort 3) diagnostic bone marrow (BM) samples. Patients were treatment-naïve at the time of sequencing. Cohorts 1 and 2 were used for primary analyses and cohort 3 for validation. To estimate the relative immune cell content, we first computed immune signature scores using ssGSEA. These scores were highly correlated to cell-type proportions estimated by flow cytometry in 79 matched samples, improving upon other deconvolution methods. We then used LASSO to select the most informative signatures to incorporate in a Cox proportional hazard model. By multiplying regression coefficients by their respective signature scores, we generated a single immune-based risk score for each patient. We assigned risk categories to patients based on their risk scores and evaluated survival outcomes across the different categories. To pinpoint biological differences between these categories, we compared various signature scores, immune checkpoints and immune-related hallmark pathways. Because our risk score is calculated independently of any other clinical data, it can be combined with other annotations to stratify patients on additional levels. For example, while IPSS-M is a widely used prognostic scoring method for MDS, it does not consider the immune context. To address this knowledge gap, we used our immune risk score to further stratify patients within the three lowest IPSS-M categories into 3 immune risk groups, namely low, medium and high risk (median OS low=63mo, medium=80mo, high=107mo, p<2.9e-6). We observed that a subset of lowest IPSS-M high-risk patients had a survival probability similar to that of higher IPSS-M patients, suggesting a refinement of prognostic classification. Notably, the newly identified “immune” high-risk group exhibited an increase in exhausted CD8 memory (p<1.0e-6) and effector T cells (p<2.3e-6), highlighting the relevance of immune dysfunction to disease progression. We validated these findings in cohort 3, where we defined similar risk groups and recapitulated enrichment in immune cell populations. Additionally, the immune characteristics of the identified low and high immune risk groups are in line with an independent study based on flow data (Riva et al., ASH2024, Presentation 665). Lastly, PBMC samples matched with cohort 2 are being prepared as a second validation to determine whether similar markers and risk categories can also be identified in blood. Conclusion We have developed a computational workflow that integrates high dimensional immune-related data extracted from bulk and single-cell RNA-seq into risk models, introducing a refined immune risk scoring method. This approach is highly flexible and can be easily combined with clinically established prognostic and classification systems. We have shown here how signatures derived from BM samples can be combined with IPSS-M, enhancing patient stratification and identifying patients more likely to respond to immune therapies.
Myelodysplastic Syndromes (MDS) present an increased risk of progression to Acute Myeloid Leukemia (AML). The complex interactions between neoplastic clone, bone marrow (BM) microenvironment and immune cells during disease evolution remain poorly understood. We used multi-omics single-cell approach to define patterns of clonal expansion and microenvironment shifts associated with MDS disease progression. We analyzed paired BM samples at diagnosis and at time of AML transformation from 20 MDS patients who had not received disease-modifying treatments before progression. Single-cell analysis was performed by CITE-seq, integrating transcriptomic and protein expression data from hematopoietic stem and progenitor cells (HSPC), myeloid, T and NK cells, in combination with single-cell genotyping (TAPESTRI). To study longitudinal dynamics of cell states, we projected each cell into gene expression space and quantified the fold-enrichment of transcriptionally similar cells between diagnosis and AML by k-nearest neighbor analysis. Differential gene/protein expression analyses were performed by linear mixed-effects models accounting for inter-patient variability. We identified two evolution patterns in HSPC compartment. In 9 patients (pts), progression was marked by emergence of novel HSPC clusters with leukemic stem cell (LSC)-like phenotype (absent/minimally detectable at diagnosis), showing upregulation of LSC markers (CD99, CD44) and immune evasion proteins (CD47, CD276) and downregulation of TGF-β and interferon (INF) response programs. In the remaining pts, progression was associated with expansion of a multipotent progenitor (MPP)-like population (already present at diagnosis). MPP-like cells exhibited increased activity of proliferative and INF-related inflammatory pathways, as well as downregulation of HLA molecules, suggesting the involvement of distinct immune escape mechanisms. Patients with NPM1, RUNX1, or TP53 mutations were more likely to show emergence of LSC-like clusters, whereas MDS with spliceosome gene mutations had heterogeneous patterns of HSPC evolution. Notably, pts showing LSC-like cluster emergence progressed more rapidly to AML (p=0.01). Considering BM microenvironment, across all pts, disease progression was associated with increased inflammatory monocytes (CD14⁺ CD86⁺ and high expression of INF-related genes) and neutrophils, suggesting that mature myeloid cells contribute to shape a pro-inflammatory marrow niche. Longitudinal analysis of immune cell states in all pts revealed widespread remodeling from diagnosis to evolution: 1) NK cells reduced their cytotoxic activity (GZMK/B-, PRF1-) and upregulated pro-inflammatory programs (NF-kb, IFN-γ); 2) T-regs acquired a highly immunosuppressive phenotype, with increased ICOS expression, downregulation of BACH2, and a switch to CD45RO⁺; 3) CD4⁺ effector memory T cells showed lower cytotoxic potential reducing GZMB/GZMK expression. Notably, in a subset of pts, small populations of these dysfunctional immune subsets—particularly highly suppressive T-regs—were already detectable at diagnosis and were associated with a shorter time to progression (p = 0.02). When comparing immunological changes based on the type of HSPC expansion, pts with LSC-like cluster displayed a more exhausted immune microenvironment, characterized by reduced frequencies of naïve T cells and increased terminally differentiated effector memory T cells, potentially supporting the selective advantage of LSC-like clones. Conversely, pts with MPP-like expansion showed increased IFN signaling across multiple immune cell populations. TP53-mutated MDS exhibited a distinct inflammatory signature, independent of IFN signaling, in both mature myeloid cells and T-regs. These myeloid cells showed HLA downregulation, while T-regs were enriched for a CD161⁺ subset with enhanced suppressive function—indicating a specific pattern of immune dysregulation driven by myeloid inflammation and impaired antigen presentation. MDS follow distinct evolutionary trajectories within the HSPC compartment. Consistent alterations in the BM microenvironment emerged as a potential common driver of disease progression. Early detection of rare, aberrant myeloid and immune cell populations at diagnosis may help identify pts at higher risk of rapid transformation to AML. TP53-mutated MDS exhibited a unique immunosuppressive profile, which may be a driver of their poor prognosis.
Blast transformation (BT) occurs in approximately 15-30% of patients with chronic myelomonocytic leukemia (CMML) and remains a leading cause of death. Allogeneic stem cell transplantation (ASCT) is currently the only treatment modality with the potential to cure the disease or prolong survival. Optimal timing of ASCT is critical for maximizing benefit while minimizing risks. To that end, contemporary risk models have focused on the prognostic relevance of individual, as opposed to concurrent mutations. In the current study, we looked into the possibility of prognostic prominence from concurrent mutations in predicting BT in CMML. CMML diagnostic criteria were according to the International Consensus Classification (Arber et al. Blood 2022; 140:1200). A machine-learning hierarchical clustering algorithm was developed and tailored for patient stratification using survival outcomes and co-expression of genomic alterations. To reduce complexity and improve interpretability, we generalized each cluster using existing rule induction algorithms, including the in Trees framework and JRip. The final output was a set of mutation-based cluster definitions, each representing a distinct patient subgroup with similar survival trajectories. Competing risk analysis and cumulative incidence functions were used in downstream evaluation to validate the clinical distinctiveness of the clusters. For all survival analysis, patients were censored at the time of ASCT. Time-to-BT was calculated from the date of diagnosis to the date of BT or ASCT or last contact. BT-free-survival was calculated from the date of diagnosis to the date of BT, ASCT, death, or last contact. The core study cohort included 605 patients from the Mayo Clinic, USA, and the external validation cohort 501 patients from Humanitas Cancer Center, Milan, Italy. Using the patient cohort from the Mayo Clinic, machine-learning algorithms identified five molecular clusters with 3-year blast transformation (BT) rates ranging from 0% to 100% (AUC at 3 years 0.78): the order of molecular signature assignment (probability of BT/death from another cause) was i) PHF6MUT/ASXL1WT (0%/17% at 3 years; N=32), ii) NPM1MUT OR BCORMUT/ASXL1MUT OR SETBP1MUT/NRASMUT (48%/30% at 1 year; N=24), iii) RUNX1MUT/ASXL1MUT OR SRSF2MUT/NRASMUT OR EZH2MUT/ASXL1MUT OR SETBP1MUT OR BCORMUT (31%/55% at 3 years; N=132), iv) ASXL1MUT/TET2MUT OR DNMT3AMUT OR JAK2MUT (24%/28% at 3 years; N=123), and v) all other permutations (10%/40% at 3 years; N=294) [Figure 1]. Additional analysis confirmed significant differences in survival between RUNX1MUT/ASXL1MUT vs. RUNX1MUT/ASXL1WT (p<0.01) OR RUNX1WT/ASXL1WT (p<0.01) OR RUNX1WT/ASXL1MUT (p=0.038) [Figure 2]. Similar patterns of differences in survival were also documented for NRAS/SETBP1, ASXL1/EZH2, and NRAS/SRSF2 mutation combinations (Figure 2). A subsequent Cox regression analysis confirmed independent prognostic contributions from “PHF6MUT/ASXL1WT” (HR 5.43e-10; p<0.01), NPM1MUT (HR 26.7; p<0.01), “SETBP1MUT/NRASMUT” (HR 12.7; p<0.01), BCORMUT(HR 5.8; p<0.01), “RUNX1MUT/ASXL1MUT” (HR 2.3, p<0.01), JAK2MUT (HR 2.1; p<0.01), and “ASXL1MUT/TET2MUT” (HR 1.7; p=0.02). The prognostic relevance of “SETBP1MUT/NRASMUT”, “RUNX1MUT/ASXL1MUT”, NPM1MUT, and BCORMUT was validated in the external cohort from Italy (N=501). In the Mayo Clinic cohort, presence of any of the latter mutations was associated with 1-, 3-, and 5-year BT (death from other cause) rates of 27% (21%), 44% (49%), and 44% (51%), respectively (Figure 3). The corresponding values in the absence of high risk mutations were 7% (15%), 15% (37%), and 18% (52%) [Gray's p value <0.01 for both BT and death from another cause; Figure 3]. Similarly, the 1-, 3-, and 5-year BT (death from another cause) rates in the Italian cohort were 21% (19%), 37% (47%), and 42% (54%) in the presence and 8% (14%), 20% (35%), and 24% (45%) in the absence of high-risk mutations (Gray's p value <0.01 for BT and 0.18 for death from another cause; Figure 4). In the current study, machine-learning algorithms enabled the discovery of concurrent mutations in CMML that were shown to be prognostically more significant than their individual constituents. Such prognostic interaction might have contributed to some of the discrepancies noted in current literature regarding the prognostic relevance of certain mutations in CMML and should be accounted for in the development of future risk models.
ABSTRACT:Acquired somatic mutations are incorporated in the classification and prognosis of myelodysplastic syndromes/neoplasms (MDSs). However, the predictive role of molecular features in MDS needs to be elucidated, especially in the lower-risk subtypes (LR-MDS), where treatment has become heterogeneous and predictive biomarkers are lacking. In this study, we investigated genetic markers associated with erythropoiesis-stimulating agents (ESAs) response in LR-MDS. A European cohort of 535 patients with LR-MDS was analyzed using targeted next-generation sequencing (t-NGS) to calculate molecular prognostic scores (International Prognostic Scoring System, molecular [IPSS-M]). The integration of IPSS-M score among the 2 known variables, serum erythropoietin (sEPO) and transfusion dependence (TD), refined the capability to predict response (area under the curve [AUC], 0.71 vs 0.63, P = .0004). Based on these 3 variables, a molecular predictive score, which we named ESA-PSS-M (-0.05 × [sEPO U/L] -4.5 × [IPSS-M score] -5 × [TD (yes = 1; no = 0)]; specificity 76%; sensitivity 57%), was generated and validated in an external cohort (n = 223 patients with LR-MDS). Despite the impact of IPSS-M score, no single mutated gene was linked to ESA response; however, when we stratified cases by sex at birth, the X-linked STAG2 gene mutations were significantly associated with ESA resistance in males with LR-MDS (odds ratio, 0.13; P = .003). To our knowledge, this is the first study based on a large multicenter cohort of patients suggesting that the integration of IPSS-M score and sex-specific mutations can characterize ESA resistance and guide first-line (1L) therapeutic choices for anemic LR-MDS (ie, ESAs vs luspatercept).
BACKGROUND AND OBJECTIVES:Several computational pipelines for biomedical data have been proposed to stratify patients and to predict their prognosis through survival analysis. However, these analyses are usually performed independently, without integrating the information derived from each of them. Clustering of survival data is an underexplored problem, and current approaches are limited for biomedical applications, whose data are usually heterogeneous and multimodal, with poor scalability for high-dimensionality. METHODS:We introduce VAE-Surv, a multimodal computational framework for patients' stratification and prognosis prediction. VAE-Surv integrates a Variational Autoencoder (VAE), which reduces the high-dimensional space characterizing the molecular data, with a deep survival model, which combines the embedded information with the clinical features. The VAE embedding step prioritizes local coherence within the feature space to detect potential nonlinear relationships among the molecular markers. The latent representation is then exploited to perform K-means clustering. To test the clinical robustness of the algorithm, VAE-Surv was applied to the Genomed4all cohort of Myelodysplastic Syndromes (MDS), comparing the identified subtypes with the World Health Organization (WHO) classification. The survival outcome was compared with the state-of-the-art Cox model and its penalized versions. Finally, to assess the generalizability of the results, the method was also validated on an external MDS cohort. RESULTS:Tested on 2,043 patients in the GenomMed4All cohort, VAE-Surv achieved a median C-index of 0.78, outperforming classical approaches. In addition, the latent space enhanced the clustering performance compared to a traditional approach that applies the clustering directly to the input data. Compared to the WHO 2016 MDS subtypes, the analysis of the identified clusters showed that the proposed framework can capture existing clinical categorizations while also suggesting novel, data-driven patient groups. Even tested in an external MDS cohort of 2,384 patients, VAE-Surv achieved a good prediction performance (median C-index=0.74), preserving the interpretability of the main clinical and genetic features. CONCLUSIONS:VAE-Surv enables automatic identification of patients' clusters, while outperforming the traditional CoxPH model in survival prediction tasks at the same time. Applied to MDS use case, the obtained genetic-based clusters exhibit a clear survival stratification, and the application of the clinical information allowed high performance in prognosis prediction.