Five patients received a matched targeted therapy based on the molecular profiling results.
Assessment of ITH using the Shannon diversity index score. Box plots show the distribution of ITH scores (y-axis), in which each dot corresponds to an individual tumor sample. The ITH scores were assessed across sarcoma subtypes, disease status, and tumor location, with meaningful differences observed in osteosarcoma (OS) and Ewing sarcoma (EW) but not in RMS or SS. ITH distribution stratified by (A) tumor location in OSs and (B) disease status in EWs.
Abstract Sarcomas in pediatric and adolescent–young adult (AYA) populations represent rare and biologically heterogeneous tumors with complex genetic underpinnings. Genomic profiling reveals subtype-specific alterations and therapeutic targets. Such tumors still represent an unmet clinical need due to limited treatment options and poorer outcomes, especially in advanced stages. Here, we present the SAR-GEN2016 and SAR-GEN_ITA clinical trials, conducted across 12 Italian centers, which enrolled 201 patients, including 158 bone and soft-tissue sarcoma samples collected at diagnosis or relapse. Whole-exome sequencing was successfully performed on 120 tumor samples. The most representative histotypes were osteosarcoma (n = 53), Ewing sarcoma (n = 39), rhabdomyosarcoma (n = 13), and synovial sarcoma (n = 5), and the genomic analyses were mainly focused on these subtypes. Overall, our cohort showed genomic differences between subtypes, highlighting how genomic complex sarcomas and fusion-driven sarcomas are distinct entities. The genomic complex histotypes, such as osteosarcoma, were characterized by a lower tumor mutational burden (TMB) and higher copy-number variation burden with enrichment of the CN2 signature. Recurrent and metastatic Ewing sarcomas have a higher TMB compared with treatment-naïve primary tumors, along with increased intratumoral heterogeneity. Oncogenic pathway analyses revealed dysregulation of the RTK–RAS and NOTCH pathways across subtypes, particularly in metastatic and recurrent tumors. In 71 of 120 analyzed samples (59%), at least one potentially actionable genomic alteration was identified, and 16% of those patients with relapsed disease received a matched targeted therapy based on the molecular profiling results. All findings were classified as ESCAT tier II or III. Our findings support the value of integrating genomic and clinical data to accelerate translational research in rare tumors. Significance: Pediatric and AYA sarcomas are rare with poor outcomes in advanced stages and limited treatment options. Through the SAR-GEN2016 and SAR-GEN_ITA multicenter trials, we performed whole-exome sequencing on 120 tumor samples with matched normal tissue from 158 patients with bone and soft-tissue sarcoma. Our integrative genomic analysis supports the genomic stratification and precision oncology in rare pediatric sarcomas.
Oncogenic signaling pathways. A–D, Mirrored circular bar plots depict the overall enrichment of oncogenic pathways in osteosarcoma (OS; A), Ewing sarcoma (EW; B), RMS (C), and SS (D). Each bar represents a single pathway, with the height indicating the proportion of samples harboring corresponding somatic alterations. Somatic variants are shown in yellow and CNVs in light blue, with the corresponding percentage of affected samples. E and F, The heatmaps provide a detailed view of (E) variants and (F) CNVs in genes of each pathway stratified by disease status and, grouped by disease subtype. Top color bar indicates disease status. Bar plots on the right represent the proportion of each alteration type across pathways, cells, and corresponding color code.
CNV across sarcoma subtypes. A, CNV burden is shown with a box plot highlighting significant differences across all tumor samples from osteosarcoma (OS), Ewing sarcoma (EW), RMS, and SS, without stratification by disease status. A statistically significant difference was observed between EW and OS samples (two-tailed Wilcoxon test). B–D, Recurrent and significant CNVs were identified in OS and EW samples using GISTIC2 for primary (B and C) and metastatic (D) tumors. The cytobands are visualized using the R package maftools. Blue bars indicate deletions, red bars indicate amplifications, and gray bars represent nonsignificant CNVs. The G-score represents the amplitude and frequency of the CNVs across tumors of interest. E, Distribution of CN signatures across the cohorts, stratified by disease status for each tumor subtype. Left, signature names. Each dot represents a mutational signature. The color of the dot indicates the number (N) of samples harboring the signature, whereas the size of the dot reflects the percentage of samples carrying that signature. Signatures are grouped horizontally based on their classification and vertically by disease. The color bar at the bottom indicates the disease status.
Significant cytobands identified in osteosarcoma (primary and recurrent) and Ewing’s sarcoma (primary, recurrent, metastasis) with the respective genes
e13002 Background: Selecting the most effective anticancer drugs remains challenging. Comprehensive molecular profiling identifies alterations actionable with targeted agents, but genomic and clinical data may also predict benefit across different drug classes. We developed machine learning–based models (Support Vector Machines - SVM and Random Forest - RF) to predict treatment outcomes using real-world clinical, pathological, genomic, and prior-therapy data at the treatment line level. Genomic alterations were summarized at the pathway level. Methods: We retrospectively analyzed 46 homogeneous patients with breast cancer who received PARP inhibitors (PARPi) discussed at the IEO Molecular Tumor Board between 2019 and 2023. Outcomes were progression-free survival (PFS) and objective response rate (ORR). As a quality control measure, models were trained to predict overall survival in first-line setting to assess data reliability. The most important predictors were expected (hormone receptor (HR), Ki-67 and age), thus supporting data reliability. We then trained outcome models for the four most represented treatment classes: PARPi, taxanes (Tx), endocrine therapies (ET), and pyrimidine analogues (Pyr). Results: RF models outperformed linear SVMs, particularly for PFS prediction, suggesting non linear interactions among predictors. Overall discrimination was modest (Table 1), reflecting limited sample size and the scarcity of true negative controls, especially for targeted treatments. Prior treatments emerged as the strongest predictor: longer PFS of previous platinum therapy (pPFS) was among the most important features to predict longer PFS and better ORR for PARPi, but similar results were obtained for all the treatment classes. Features such as Ki-67, performance status (PS), age, HR status and treatment history added predictive value, while genomic pathway alterations provided weaker but detectable signals. Conclusions: Our study demonstrates the feasibility of integrating data to predict treatment benefit and prioritize therapeutic options. Although current limited performance, the approach highlights the dominant yet still underestimated importance of prior treatment response on subsequent therapies’ outcome prediction, establishing a foundation for a methodological refinement, expansion across tumors, and external validation. It could evolve into a clinical decision support tool optimizing both personalized and standard treatments, which is especially relevant for low-income countries without access to expensive drugs. Drug PFS(RMSE) ORR(ROC) Top featuresORR Direction RF SVM RF SVM PARPi46 9 11.5 0.63 0.59 RAS wt + pPFS platinum + BRCA1, BRCA2 and PALB2 wt - Lobular - ET35 14.4 15.9 0.48 0.45 PGR% + pPFS ET + RAS altered - NOTCH wt - Tx35 8.2 10.6 0.75 0.58 pPFS anthracyclines + pPFS TROP2 ADC + PS - pPFS PARPi - Pyr34 15.5 22.8 0.6 0.58 pPFS alkylating + Lobular + Ki-67 - pPFS HER2 ADC -
Abstract Motif discovery and binding-site prediction in DNA and RNA sequences are central tasks in regulatory genomics, yet the methodological landscape is split between interpretable but rigid position weight matrices (PWMs) and high-performing but opaque machine-learning models. We present KDM, a unifying framework in which both motifs and sequences are represented as probability distributions over a shared k-mer dictionary, embedded via the Hellinger transformation. This common geometry enables motif-sequence scoring, motif-motif comparison, de novo discovery, and binding prediction with a single primitive, the Bhattacharyya coefficient. We instantiate four tools on this representation: KDMMap for positional enrichment analysis, KDMMatch for information-content-aware motif matching, KDMFind for unsupervised motif discovery via projective non-negative matrix factorization, and KDM-LRLM for binding prediction with Lasso-regularized logistic regression. Across 1,324 transcription-factor ChIP-seq and 161 RBP eCLIP experiments, KDMMap matches CentriMo’s motif rankings in 84% of TF and 79% of RBP experiments, and KDMMatch agrees with Tomtom on motif annotation in 74.5% of TFs. On binding prediction across four datasets covering 2,475 experiments, KDM-LRLM matches or exceeds eight deep-learning and three k-mer-based competitors. Notably, AI methods overtake k-mer methods only in the top quartile of training-set size, indicating that data scale, not architecture, drives the recent dominance of deep models. KDM provides a single interpretable representation across the full motif-analysis workflow.
Alternative splicing expands proteomic diversity and is shaped by interactions between RNA‑binding proteins (RBPs) and multivalent RNA motifs. Linking sequence elements to regulatory proteins remains difficult from sequence information alone. Here we present RNAMaRs, a interpretable statistical framework that combines motif discovery with in vivo binding and splicing responses to infer motif-RBP relationships. RNAMaRs learns RBP binding principles, weights signal quality, and optimizes motif discovery in an RBP-specific manner. Across ENCODE datasets RNAMaRs consistently prioritizes the perturbed regulator, especially for large splicing effects. Independent validation in prostate cancer cells recapitulates HNRNPK binding signatures, supporting transferability across an unseen cellular context. ### Competing Interest Statement The authors have declared no competing interest. AIRC, BRIDGE 2023 ID 28739
Overview of SAR-GEN2016 and SAR-GEN_ITA prospective multicentric trials. A, Workflow of the clinical trials, illustrating the process from patient enrollment through sample collection, genomic profiling, and analysis, to the multidisciplinary discussion at the MTB for therapeutic decision-making, with excluded samples indicated at each stage. B, Combination of percent stacked barcharts and box plot, from left to right: number of samples for each sarcoma subtype, tissue type of the biopsy (fresh or FFPE), disease status at the moment of the surgical procedure (primary, recurrent, and metastasis), purity of the tumor calculated after sequencing, and sex and age of the samples. EW, Ewing sarcoma; NA, not available; OS, osteosarcoma. [A, Created in BioRender. Grieco, M. (2025) https://BioRender.com/c7w0fbz.]
Genomic landscape of the four sarcoma subtypes. A and B, Box plots show the distribution of TMB on the y-axis, in which each dot represents an individual tumor sample. The horizontal dotted line at 1 mutation per megabase (mut/Mb) denotes the threshold used to separate lowly and highly mutated tumors. A, TMB distribution across all tumor samples of osteosarcoma (OS), Ewing sarcoma (EW), RMS, and SS. Statistically significant difference was observed between OS and EW samples (two-tailed Wilcoxon test). B, TMB distribution stratified by disease status (primary, recurrent, and metastasis) within the EW cohort. Significant differences were observed between primary and recurrent tumors (two-tailed Wilcoxon test) and between primary and metastatic tumors (two-tailed Wilcoxon test). C–F, Bar plots summarizing the top 25 mutated genes for each sarcoma subtypes, based on all tumor samples and without stratification by disease status. Genes are listed on the left side of each plot, with the corresponding percentage of mutated samples shown on the right. The percentages represent the number of unique samples harboring a mutation divided by the total number of samples in each sarcoma subtype (OS = 53, EW = 39, RMS = 13, and SS = 5). The x-axis shows the total number of mutations identified per gene. Color code represents the type of variants. G, Summary of the SBS and small indel signatures, stratified by disease status for each tumor subtype. On the left, the signature names. Each dot represents a mutational signature. The color of the dot (N) indicates the number of samples in which the signature is present, whereas the dot size represents the percentage of samples carrying that signature within that cohort. Signatures are grouped horizontally based on their classification and vertically by disease. The color bar at the bottom indicates the disease status.
Double positive OCT3/4-PD1 subsets in the bulk NSCLC cells after cisplatin treatment. OCT3/4 expression preferentially co-localized on PD-1+ NSCLC cells as compared with the PD-1- counterpart (14% vs 0.1%). Representative flow cytometry histograms of OCT3/4 expression are reported in figure.
PD-L1/2 and phospho-p38 MAPK expression in NSCLC cells. Tumor membrane expression of PD-L1 and PD-L2 in NSCLC cell lines was intensely enhanced by treatment with Interferon γ (n=9, PD-L1 p=0.02, PD-L2 p<0.0001) (A). Western blot analysis of phosphorylated (p) and total p38 MAPK in NSCLC cells (H1975) with or without treatment by anti-PD-1 and/or PD-L1s (B).
Treatment with PD-L1s is associated with the up-regulation of expression of key protein coding genes involved in processes of tumorigenesis, EMT, tumor aggressiveness and chemoresistance. Heatmap showing z-scaled DESeq2-normalized expression values of the 14 differentially expressed genes between PD-L1s and untreated cells. Right annotation indicates the mean DESeq2-normalized gene expression values. Top annotations depict the pneumosphere of origin and the treatment. Hierarchical clustering of genes and samples is shown on the left and top, respectively (A). Volcano plot showing the log2 fold change (log2FC) and adjusted p-value values of all genes between PD-L1s and untreated samples. Differentially expressed genes are depicted in red (B).
Somatic CAG instability in the mutant Huntingtin (HTT) gene is increasingly recognized as a key hallmark of Huntington’s disease (HD). Using our novel human CAGinSTEM platform, we manipulated cis genetic elements influencing instability in human HD neurons, monitoring repeat length. Quality-controlled CRISPR-engineered stem cells with increasing CAG lengths and clinical haplotypes were analyzed using third-generation sequencing. Our findings link interruptions in the CAG repeat, especially the loss or duplication of the penultimate CAA of canonical alleles, to significant instability modulation. Notably, four internal CAA interruptions completely abolish CAG instability, reversing HD phenotypes such as altered striatal fate acquisition and nuclear disorganization. This platform highlights the role of cis modifiers, emphasizing the direct influence of HTT DNA repeat composition on CAG instability and providing a robust framework for modeling HTT repeat instability in vitro.