Short tandem repeats (STRs) are a major source of genetic variation, yet their potential for genome-wide population structure inference remains underexplored. Here we present a multi-modal framework for STR-based population inference, integrating unsupervised clustering, supervised population assignment, and a novel admixture inference model, Directional Non-negative Matrix Factorization (dNMF). Applying this framework to thousands of genomes from multiple global cohorts, we first demonstrate that genome-wide STR variations provide substantially finer resolution of human population structure than single-nucleotide polymorphisms (SNPs), particularly at regional levels. The dNMF model estimates ancestry coefficients under the hypothesis that true ancestral populations are consistently encoded in the bidirectional mutation dynamics of STRs. Population structures inferred by dNMF are fine-grained, reproducible across datasets, and robust to technical artifacts. Motif-specific analyses further reveal directionbiased mutational tendencies and show that distinct STR motif classes encode complementary layers of population structure at different evolutionary scales. These results establish STRs as powerful and biologically interpretable markers for population structure inference, offering a mutation-aware perspective that complements traditional SNP-based frameworks and refines understanding of human demographic history.
Many cancers and most childhood cancers are rare. To conduct research on infrequent tumor types, a central infrastructure for secure, cross-border discovery of cancer cases with genomic variant data is essential. Currently, cancer centers lack an accessible, federated and standardized framework for exchanging sensitive clinical and genotypic data. This is driven by fragmented privacy legislation, non-standard terminologies, and significant financial barriers. A federated network of cancer cohort data catalogs based on globally standardized specifications with low- cost implementations is needed to allow clinicians and researchers to query all cancer centers simultaneously without the source data having to cross national borders. The BeaconV2 API is a lightweight data discovery protocol, endorsed by the Global Alliance for Genomics and Health, that enables medical institutes to make their genomic and biomedical data globally discoverable. A collection of BeaconV2 API nodes builds a Beacon network that helps researchers compose research cohorts from multiple centers for genotype first research. Last year, we introduced the "OncoBeacon" specification that complements BeaconV2 specification and addresses the challenges of describing the rich meta-data associated with cancer genomics. The OncoBeacon specification provides researchers with essential context to select samples within a patient's treatment history and aids in genetic variation interpretation. Beacon implementations supporting the OncoBeacon specification have emerged in the Hartwig Medical Foundation and Princess Máxima Center, allowing the medical centers to establish their local data catalogue nodes in their centers and form an OncoBeacon network. The OncoBeacon network will provide pre-clinical researchers with unprecedented opportunities to assemble rare (pediatric) cancer research cohorts, across centers, based on both detailed clinical and molecular selection criteria. Additionally clinicians will be able to rapidly find suitable patients for clinical trials with increased accuracy. The open specification, low-threshold implementation, and privacy-friendly design make the OncoBeacon infrastructure broadly applicable including centers in middle-income countries, making it a pillar in the international mission to offer a treatment perspective for rare cancers.
Somatic copy number aberrations (CNAs) represent a distinct class of genomic mutations associated with oncogenetic effects. Over the past three decades, significant volumes of CNA data have been generated through molecular-cytogenetic and genome sequencing-based techniques. These data have been pivotal in identifying cancer-related genes and advancing research on the relationship between CNAs and histopathologically defined cancer types. However, comprehensive studies of CNA landscapes and disease parameters are challenging due to the vast diagnostic and genomic heterogeneity encountered in "pan-cancer" approaches. In this study, we introduce CNAttention, an attention-based deep multiple instance learning method designed to comprehensively analyze CNAs across different cancers and uncover specific CNA patterns within integrated gene-level CNA profiles of 30 cancer types. CNAttention effectively learns CNA features unique to each cancer type and generates CNA signatures for 30 cancer types using attention mechanisms, highlighting the distinctiveness of their CNA landscapes. CNAttention demonstrates high accuracy and exhibits stable performance even with the incorporation of external datasets or parameter adjustments, underscoring its effectiveness in tumor identification. Expanding these signatures to cancer classification trees reveals common patterns not only among physiologically related cancer types but also among clinico-pathologically distant types, such as different cancers originating from neural crest derived cells. Additionally, detected signatures also uncover genomic heterogeneity in individual cancer types, for instance in brain lower grade glioma. Additional experiments with classification models underscore the efficacy of these signatures in representing various cancer types and their potential utility in clinical diagnosis.
Short tandem repeats (STRs) have been reported to influence gene expression across various human tissues. While STR variations are enriched in colorectal, stomach, and endometrial cancers, particularly in microsatellite instable tumors, their functional effects and regulatory mechanisms on gene expression remain poorly understood across these cancer types. Here, we leverage whole-exome sequencing and gene expression data to identify STRs for which repeat lengths are associated with the expression of nearby genes (eSTRs) in colorectal, stomach, and endometrial tumors. While most eSTRs are cancer-specific, shared eSTRs across multiple cancers exhibit consistent effects on gene expression. Notably, coding-region eSTRs identified in all three cancer types show positive correlations with nearby gene expression. We further validate the functional effects of eSTRs by demonstrating associations between somatic eSTR mutations and gene expression changes during the transition from normal to tumor tissues, suggesting their potential roles in tumorigenesis. Combined with DNA methylation data, we perform the first quantitative analysis of the interplay between STR variations and DNA methylation in tumors. We identify eSTRs where repeat lengths are associated with methylation levels of nearby CpG sites (meSTRs) and show that >70% of eSTRs are significantly linked to local DNA methylation. Importantly, the effects of meSTRs on DNA methylation remain consistent across cancer types. Overall, our findings enhance the understanding of how functional STR variations influence gene expression and DNA methylation. Our study highlights shared regulatory mechanisms of STRs across multiple cancers, offering a foundation for future research into their broader implications in tumor biology.
Accurate determination of the genomic copy number baseline is crucial for identifying copy number alterations (CNAs) in cancer, yet it remains a significant challenge in tumors with complex karyotypes. To address this, we present CNAdjust, an integrated method to systematically detect and correct baseline inaccuracies in CNA data. CNAdjust employs a Bayesian framework that integrates cohort-specific CNA frequency priors with a data-driven plausibility score, ensuring that adjusted calls align with both biological cohort patterns and study-specific data. Performance validation using the TCGA pan-cancer dataset demonstrated improved alignment with absolute copy number estimates and enhanced CNA pattern interpretation. Furthermore, we revealed a strong correlation between chromosomal aneuploidy and baseline abnormalities, underscoring the prevalence of this issue in cancer genomics. By systematically improving the precision of CNA calls, CNAdjust serves as a critical tool for constructing harmonized reference datasets and advancing the progress of precision oncology. Its implementation as a standard, portable workflow enables the reproducible and scalable analysis of large, heterogeneous datasets, supporting large-scale genomic research. Source codes are available at: https://github.com/baudisgroup/CNAdjust.
Somatic copy number aberrations (CNAs) represent a distinct class of genomic mutations associated with oncogenetic effects. Over the past three decades, significant volumes of CNA data have been generated through molecular-cytogenetic and genome sequencing-based techniques. This data has been pivotal in identifying cancer-related genes and advancing research on the relationship between CNAs and histopathologically defined cancer types. However, comprehensive studies of CNA landscapes and disease parameters are challenging due to the vast diagnostic and genomic heterogeneity encountered in “pan-cancer” approaches. In this study, we introduce CNAttention , an attention-based deep multiple instance learning method designed to comprehensively analyze CNAs across different cancers and uncover specific CNA patterns within integrated gene-level CNA profiles of 30 cancer types. CNAttention effectively learns CNA features unique to each cancer type and generates CNA signatures for 30 cancer types using attention mechanisms, highlighting the distinctiveness of their CNA landscapes. CNAttention demonstrates high accuracy and exhibits stable performance even with the incorporation of external datasets or parameter adjustments, underscoring its effectiveness in tumor identification. Expanding these signatures to cancer classification trees reveals common patterns not only among physiologically related cancer types but also among clinico-pathologically distant types, such as different cancers originating from neural crest derived cells. Additionally, detected signatures also uncover genomic heterogeneity in individual cancer types, for instance in brain lower-grade glioma. Additional experiments with classification models underscore the efficacy of these signatures in representing various cancer types and their potential utility in clinical diagnosis. ### Competing Interest Statement The authors have declared no competing interest.
Motivation:The Beacon v2 specification, established by the Global Alliance for Genomics and Health (GA4GH), consists of a standardized framework and data models for genomic and phenotypic data discovery. By enabling secure, federated data sharing, it fosters interoperability across genomic resources. Progenetix, a Beacon v2 reference implementation, exemplifies its potential for large-scale genomic data integration, offering open access to genomic mutation data across diverse cancer types. Results:We present pgxRpi, an open-source R/Bioconductor package that provides a streamlined interface to the Progenetix Beacon v2 REST API, facilitating efficient and flexible genomic data retrieval. Beyond data access, pgxRpi offers integrated visualization and analysis functions, enabling users to explore, interpret, and process queried data effectively. Leveraging the flexibility of the Beacon v2 standard, pgxRpi extends beyond Progenetix, supporting interoperable data access across multiple Beacon-enabled resources, thereby enhancing data-driven discovery in genomics. Availability and Implementation:pgxRpi is freely available under the Artistic-2.0 license from Bioconductor (https://doi.org/10.18129/B9.bioc.pgxRpi), with actively maintained source code on GitHub (https://github.com/progenetix/pgxRpi). Comprehensive usage instructions and example workflows are provided in the package vignettes, available at https://github.com/progenetix/pgxRpi/tree/devel/vignettes.
Cancer cell lines are an important component in biological and medical research, enabling studies of cellular mechanisms as well as the development and testing of pharmaceuticals. Genomic alterations in cancer cell lines are widely studied as models for oncogenetic events and are represented in a wide range of primary resources. We have created a comprehensive, curated knowledge resource-cancercelllines.org-with the aim to enable easy access to genomic profiling data in cancer cell lines, curated from a variety of resources and integrating both copy number and single nucleotide variants data. We have gathered over 5600 copy number profiles as well as single nucleotide variant annotations for 16 000 cell lines and provide these data with mappings to the GRCh38 reference genome. Both genomic variations and associated curated metadata can be queried through the GA4GH Beacon v2 Application Programming Interface (API) and a graphical user interface with extensive data retrieval enabled using GA4GH data schemas under a permissive licensing scheme.Database URL: https://cancercelllines.org
In the age of data-driven biomedical research and clinical practice, the sharing of genomic and clinical data for health research and personalized medicine has become an important contributor to improved diagnosis and treatment.From the data owner's perspective, potential benefits include improved treatments, personalization of healthcare practice, and more effective control of disease proliferation.However, the requirement for high levels of data security to protect sensitive information presents a barrier to data discovery and sharing [1].Beacon is designed to enable the benefits of data discovery while minimizing the associated risks.It is a Global Alliance for Genomics and Health (GA4GH) API specification (see Box 1 for a definition of key Beacon terminology), allowing easy discovery of sensitive data that require controlled (authorized) access [2].It uses simple concepts and can be adapted to different use cases.The protocol is designed to respond to queries, such as the following:"Can you provide data about males, diagnosed with Type 2 diabetes, whose age of onset is below 30 years, and who carry mutations in the APOE gene?"Depending on the data controller's preferences over response granularity, the response options range from, "Yes, our data includes one or more" (boolean response), "Yes, we have 125" (count response), to "Yes, and here are some details about the 125 individuals that match your request" (detailed "record level" response).Discovery is the necessary first step in the sharing and reuse of data and other assets, and Beacon facilitates this by enabling federated discovery in any number of networks, as a complement to unwieldy central catalogs.Many beacons have already been successfully "lit" (deployed) across the globe (https://public.tableau.com/app/profile/elixir/viz/ELIXIRBeaconNetwork/Sheet1).This article is written to support data owners that might be interested in deploying a beacon to make their data discoverable while keeping them secure.Whatever your background is, this article will provide you with some tips to get you started.Specifically, we will review important steps to complete-and pitfalls to avoid-when deploying a beacon. Tip #1: Evaluate the value of your data to your research or clinical domainDeploying a Beacon instance will help you keep the data both as open as possible and as restricted as necessary (Tip #9).That said, you might want to take into account the perceived value of your dataset, which depends on the target communities.For example, a dataset on the
Cancer cell lines are frequently used in biological and translational research to study cellular mechanisms and explore treatment options. However, cancer cell lines may display mutational profiles divergent from native cancers or may be misidentified or contaminated. We explored how similar cancer cell lines are to native cancers to find the most suitable representations for the corresponding diseases by utilising large collections of copy number variation (CNV) profiles and applied machine learning (ML) algorithms to predict cell line classifications. Our results confirm that cancer cell lines indeed accumulate more mutations compared to native cancers but retain similar CNV profiles. We demonstrate that many relevant oncogenes and tumor suppressor genes are altered by CNV events in both cancers and their corresponding cell lines. Based on the similarities between the two groups and the predictions of the ML model, we provide some recommendations about cell lines with good potential to represent selected cancer types in in vitro studies.
Motivation:With the proliferation of research means and computational methodologies, published biomedical literature is growing exponentially in numbers and volume. Cancer cell lines are frequently used models in biological and medical research that are currently applied for a wide range of purposes, from studies of cellular mechanisms to drug development, which has led to a wealth of related data and publications. Sifting through large quantities of text to gather relevant information on cell lines of interest is tedious and extremely slow when performed by humans. Hence, novel computational information extraction and correlation mechanisms are required to boost meaningful knowledge extraction.Results:In this work, we present the design, implementation, and application of a novel data extraction and exploration system. This system extracts deep semantic relations between textual entities from scientific literature to enrich existing structured clinical data concerning cancer cell lines. We introduce a new public data exploration portal, which enables automatic linking of genomic copy number variants plots with ranked, related entities such as affected genes. Each relation is accompanied by literature-derived evidences, allowing for deep, yet rapid, literature search, using existing structured data as a springboard.Availability and implementation:Our system is publicly available on the web at https://cancercelllines.org.
Short tandem repeat (STR) mutations are prevalent in colorectal cancer (CRC), especially in tumours with the microsatellite instability (MSI) phenotype. While STR length variations are known to regulate gene expression under physiological conditions, the functional impact of STR mutations in CRC remains unclear. Here, we integrate STR mutation data with clinical information and gene expression data to study the gene regulatory effects of STR mutations in CRC. We confirm that STR mutability in CRC highly depends on the MSI status, repeat unit size, and repeat length. Furthermore, we present a set of 1244 putative expression STRs (eSTRs) for which the STR length is associated with gene expression levels in CRC tumours. The length of 73 eSTRs is associated with expression levels of cancer-related genes, nine of which are CRC-specific genes. We show that linear models describing eSTR-gene expression relationships allow for predictions of gene expression changes in response to eSTR mutations. Moreover, we found an increased mutability of eSTRs in MSI tumours. Our evidence of gene regulatory roles for eSTRs in CRC highlights a mostly overlooked way through which tumours may modulate their phenotypes. Future extensions of these findings could uncover new STR-based targets in the treatment of cancer.
Somatic copy number alterations (SCNAs) are a widespread type of genomic alteration in cancer samples that affect a large proportion of the genome. High-throughput technologies enable the measurement of SCNAs, generating a large number of SCNA profiles. However, annotating and integrating these profiles is challenging due to the heterogeneity of signal scales, measurement noise, and segmentation noise. In this poster, we present an algorithm called LabelSeg that improves the interpretation of segment profiles by assigning relative copy number state labels to each SCNA segment. We have validated the performance of LabelSeg using simulated data and real biological data. Our algorithm makes the profiling of large numbers of tumor CNA profiles faster and more comprehensive, promoting the discovery of interesting SCNA regions. This goal is also shared by the hCNV community of ELIXIR, which aims to develop standards and best practices for the identification and interpretation of copy number variation in human disease. By providing a tool for more accurate and efficient profiling of SCNA profiles, LabelSeg contributes to the advancement of the hCNV community's efforts.
With the large amount of genomic and associated data generated for biomedical research and personalized health applications, a move towards shared standards for genomic data discovery and federated analyses has become a matter of eminent importance. After earlier conceptual versions, the recent "v2" editions of the GA4GH Beacon and Phenopackets standards as well as the VRS variant representation format - with contributions from the ELXIR hCNV community - and various ontology improvements have laid a solid foundation for distributed and federated genomics resources and discovery tools. Since the start of the ELIXIR Beacon project the Progenetix data resource with its large collection of curated, publicly accessible oncogenomic screening data has been a driver for the Beacon protocol expansion, e.g. for the support of structural variant data or the definition of the handover concept for Beacon- initiated data delivery. Current Beacon development support here focusses on the refinement of structural variant definitions and representation of variant annotations; here, our new cancercelllines resource combines sample derived CNV profiles with the majority of annotated genomic variants for a deep representation of genomic variation landscapes in cancer cell lines. Beyond the Beacon v2 standard model, the Progenetix resources additionally implement the Phenopackets v2 standard for data delivery, thereby demonstrating the integration of major GA4GH standards as prototypes for an "Internet of Genomics". Underlying these developments is the open source bycon pure Python project, open for adoption and inviting contributions. Links: progenetix.org cancercelllines.org docs.genomebeacons.org cnvar.org
Genome variation is the direct cause of cancer and driver of its clonal evolution. While the impact of many point mutations can be evaluated through their modification of individual genomic elements, even a single copy number aberration (CNA) may encompass hundreds of genes and therefore pose challenges to untangle potentially complex functional effects. However, consistent, recurring and disease-specific patterns in the genome-wide CNA landscape imply that particular CNA may promote cancer-type-specific characteristics. Discerning essential cancer-promoting alterations from the inherent co-dependency in CNA would improve the understanding of mechanisms of CNA and provide new insights into cancer biology and potential therapeutic targets. Here we implement a model using segmental breakpoints to discover non-random gene coverage by copy number deletion (CND). With a diverse set of cancer types from multiple resources, this model identified common and cancer-type-specific oncogenes and tumor suppressor genes as well as cancer-promoting functional pathways. Confirmed by differential expression analysis of data from corresponding cancer types, the results show that for most cancer types, despite dissimilarity of their CND landscapes, similar canonical pathways are affected. In 25 analyses of 17 cancer types, we have identified 19 to 169 significant genes by copy deletion, including RB1, PTEN and CDKN2A as the most significantly deleted genes among all cancer types. We have also shown a shared dependence on core pathways for cancer progression in different cancers as well as cancer type separation by genome-wide significance scores. While this work provides a reference for gene specific significance in many cancers, it chiefly contributes a general framework to derive genomewide significance and molecular insights in CND profiles with a potential for the analysis of rare cancer types as well as non-coding regions.
Methods: In total, 76 longitudinal cfDNA samples from 21 patients with advanced NSCLC receiving first-line treatment were sequenced at 10x coverage.In parallel, we performed deep-targeted sequencing (approx.1250x coverage) for 47 matched time points.By applying genome-wide fragmentation-based statistics, a trained model was used to determine tumor fractions at baseline and their dynamic changes throughout treatment.Results: Histological subtypes were: adenocarcinoma (n = 11), squamous cell carcinoma (n = 5), and others (n = 5).15/21 (71%) patients were treated with immune checkpoint inhibitors alone or in combination with chemotherapy.Genome-wide fragmentation profiles in nonoverlapping regions showed marked heterogeneity at baseline when compared to time points associated with stable disease or partial response.We observed a strong correlation between tumor fractions determined by DELFI-TF and MAF (Pearson, r = 0.92, p < 0.001).Longitudinal DELFI-TF dynamics overlapped MAF changes and were associated with treatment response as assessed by conventional CT scans.Conclusion: DELFI-TF accurately tracked tumor fractions in patients with advanced NSCLC receiving immune checkpoint blockade.Our results show the potential of using an inexpensive and reliable method for monitoring treatment responses independent of a mutation-based approach.
After approval of the GA4GH Beacon v2 standard, a first Beacon v2 Network prototype implementation was used to demonstrate a good level of interoperability. Based on the learnings from several new Beacon v2, an updated Beacon Network Aggregator has been implemented, connected with dedicated adjustments to the Beacon specification itself. To support the emerging GA4GH Beacon v2 Networks design, the Centre for Genomic Regulation (CRG) developed a new Beacon Network v2 User Interface which, in conjunction with the Barcelona Supercomputing Center (BSC) Beacon Network Aggregator, provides a complete federated Beacon v2 querying solution with improved user experience. Attention was put on the interoperability, as new implementations were tested. Although the reference implementation provides the Beacon v2 compatibility verification tool, substantial divergences in the implementation were detected. As a result, the Beacon Network v2 working group has prepared a set of guidelines for the Beacon v2 developers that should improve the interoperability within the Beacon Network. In summary, the current Beacon Network v2 project provides an important milestone to facilitate cooperation in terms of sharing genomic data and associated annotations. We expect that the project will have a big impact on how researchers, especially practitioners in medical genetics and cancer genomics, will approach the sharing and discovery of such data and utilize the internet to empower future discoveries. This current implementation should also serve as the stepping stone for further developments coupled with new data types becoming available in Beacon, e.g. clinical and medical imaging data.
Cancer cell lines are important models for studying the disease as well as testing for potential therapeutics. However, many of these cell lines used in research today have been found to be misidentified or contaminated. To enable better identification of misidentified or contaminated cell lines as well as cancer cell lines that are the best representations of their origins, we have created a novel resource for cancer cell line variants. In this database we have combined knowledge of both copy number variants (CNVs) and single nucleotide variants (SNVs). cancercelllines.org is a daughter of progenetix.org - a knowledge resource on cancer CNVs. All cancer cell line CNVs from Progenetix have been incorporated to cancercelllines.org. Additionally, SNVs have been included as well, from ClinVar and CCLE resources. Currently, we have over 16 000 human cancer cell lines, 5600 CNV profiles from over 2000 samples representing 256 tumor types based on NCIt classification and 1157 unique ClinVar variants. Associated metadata on cell lines, such as NCIt diagnostic codes, and cell line ancestries are displayed. Cancer cell line variants mapped from ClinVar include information about known pathogenicity and linked disease codes. Here, we present this novel resource built on Beacon framework.