Africa is entering a new era of cancer research, driven by renewed commitments to genomic innovation, data equity, and strengthened cancer registry systems. The “Harnessing Functional Genomics in Cancer Research: Opportunities for Diagnosis and Treatment” conference, held in Windhoek, Namibia, from 23 to 26 September 2025, convened leading experts to evaluate current progress and identify priorities for advancing cancer genomics and precision oncology across the continent. This conference synthesis report and narrative review drew on four days of expert-led presentations, interactive panel discussions, and structured delegate engagement sessions. Conference proceedings were documented through session minutes, presenter slide decks, recorded discussions, and post-conference feedback. Perspectives from more than 160 participants representing 21 countries were synthesized into priority thematic areas through an iterative drafting and multi-author review process. Conference findings were integrated with supporting peer-reviewed and grey literature to identify opportunities, challenges, and strategic priorities for advancing precision oncology in Africa. Discussions highlighted Africa’s unparalleled human genomic diversity and its continued underrepresentation in global genomic datasets, emphasizing the need for African-led genomic discovery and precision oncology strategies. Key priorities included strengthening cancer registries and surveillance systems, expanding genomic infrastructure and workforce capacity, establishing ethical and sovereign data-governance frameworks, and improving the translation of genomic discoveries into clinical practice. Next-generation sequencing, multi-omics approaches, artificial intelligence, and collaborative research networks were identified as important enablers of equitable and sustainable precision oncology implementation across Africa. Africa’s genetic diversity represents a critical resource for advancing precision oncology. Conference insights, supported by the literature, underscore the need for coordinated investments in genomic infrastructure, cancer surveillance systems, workforce development, ethical data governance, and translational implementation. Advancing these priorities will strengthen genomics-informed cancer prevention, diagnosis, and treatment while positioning Africa as an increasingly important contributor to global oncology innovation.
Compound birth-death processes are widely used to model the age-incidence curves of many cancers. There are efficient schemes for directly computing the relevant probability distributions in the context of linear multi-stage clonal expansion (MSCE) models. However, these schemes have not been generalised to models on arbitrary graphs, forcing the use of either full stochastic simulations or mean-field approximations, which can become inaccurate at late times or old ages. Here, we present a numerical integration scheme for directly computing survival probabilities of a first-order birth-death process on an arbitrary directed graph, without the use of stochastic simulations. As a concrete application, we show that this new numerical method can be used to infer the parameters of an example graphical model from simulated data.
Abstract Background Kataegis, the focal hypermutation of single base positions in tumour genomes, has received little attention with regards to prostate cancer (PCa) molecular features, tumour evolution and associated clinical presentation. Most notably, the impact of this phenomenon is yet to be explored across ancestral lineages representing the extremities of PCa presentation and outcomes, with men of African ancestry disproportionately disadvantaged. The purpose of this study is to address the knowledge gap through African inclusive multi-ancestral interrogation. Methods We assessed for ancestrally shared and unique molecular, evolutionary and clinical features of kataegis in 669 multi-ancestral whole PCa genomes. Access to raw whole-genome sequenced data allowed for direct single-pipeline comparative analysis between 109 southern African and 57 European derived treatment naïve high-risk-biased primary tumours (74% and 88%) with paired blood samples, further assessed against publicly available 207 Asian high-risk-leaning comparative (65%) and 296 European low-risk-biased alternative (79%) resources. Comparisons between ancestries and risk groups were through Wilcoxon’s rank sum test and Fisher’s exact tests, with P values adjusted by false discovery rate. Results Confirming relatively low burdens, we found kataegis to be significantly associated with genomic instability, cancer drivers, and clinical adversity across ancestries (false discovery rate = $$\:7.44\times\:{10}^{-6}-0.04$$ ). Notably, kataegis-postive tumours were associated with elevated prostate-specific antigen levels at presentation in African (false discovery rate = $$\:1.66\times\:{10}^{-3}$$ ) and higher risk for metastatic progression in European patients (Kaplan-Meier estimator, $$\:P=0.03$$ ). Enrichment of APOBEC’s context preferences showed more attribution from APOBEC3B than APOBEC3A. Further through analyses of evolution and structural variant (SV) co-occurrence, commonly the ancestry agnostic SV-associated kataegis predominated in the clonal evolutionary state, while the less common the SV-independent kataegis ( $$\:P=0.03$$ ) and subclonal kataegis ( $$\:P=1.67\times\:{10}^{-3}$$ ) showed African specificity. Conclusions We found kataegis-positivity to be associated with poor PCa presentation and prognosis, irrespective of patient ancestry. Kataegis-related genomic instability occurring early and late during African derived tumourigenesis, may partly explain the heightened tumour and clinical heterogeneity observed for patients of African ancestry.
Productivity disparity exists in cancer research systems between High Income Countries (HICs) versus Lower-Middle-Income-Countries (LMICs). Capacity and productivity for cancer genomics research is an exemplar metric for overall cancer research capacity. Kenyan cancer genomics research capacity was used as a case study for an LMIC benchmarked against the HIC countries of the UK and USA. A conceptual framework was used to analyse drivers of productivity of cancer genomics research within institutions based on research environment resources. Preliminary recommendations which can help close the research productivity gap were then made. This research study was carried out by conducting sixteen semi-structured interviews with researchers based across the United Kingdom, United States, Australia, South Africa, and Kenya, and analysis of this qualitative data was then conducted via NVivo. Further primary research data was sought via surveying of genomics researchers and specific institutions based in case study countries. Secondary research was conducted through online and database searches, and in-depth literature reviews on cancer genomics research and global disparities in cancer research. We applied and considered ‘The Matthew Effect’, that relates to cumulative advantage, and theories of productivity. This study found that drivers of the productivity gap between national cancer research systems across HICs and LMICs can be attributed to interrelated factors within three key areas: Data and Technology, Funding and Partnerships, and Knowledge and Capacity Build. These are critical to the productivity and eventual impact of cancer research in a particular environment. This study identified that external context factors are as important as the drivers of research productivity, and that research needs to be context relevant and designed to be relative to the local environment. This study concludes that The Matthew Effect is clearly observable within cancer research systems. Whilst this has amplified global disparity to date, the correct application of stated recommendations can change this trajectory and help close the gap to ultimately improve global progress and outcomes for comparative cancer research. Rachel Chown, Georgina Curtis, Colin Quigley, Nicklas Jansler, Roberto Jacome, Robert G. Bristow, George F. Njoroge, Esther Njoki Mwangi Maina, David C. Wedge. Drivers of cancer genomics research productivity in high income countries (HICs) versus lower middle income countries (LMICs): Which interventions can help close the gap. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 3571.
Most solid tumors harbor somatic mutations attributed to off-target activities of APOBEC3A (A3A) and/or APOBEC3B (A3B). However, how APOBEC3A/B enzymes affect tumor evolution in the presence of exogenous mutagenic processes is largely unknown. Here, multi-omics profiling of 309 lung cancers from smokers identifies two subtypes defined by low (LAS) and high (HAS) APOBEC mutagenesis. LAS are enriched for A3B-like mutagenesis and KRAS mutations; HAS for A3A-like mutagenesis and TP53 mutations. Compared to LAS, HAS have older age at onset and high proportions of newly generated progenitor-like cells likely due to the combined tobacco smoking- and APOBEC3A-associated DNA damage and apoptosis. Consistently, HAS exhibit high expression of pulmonary healing signaling pathway, stemness markers, distal cell-of-origin, more neoantigens, slower clonal expansion, but no smoking-associated genomic/epigenomic changes. With validation in 184 lung tumor samples, these findings show how heterogeneity in mutational burden across co-occurring mutational processes and cell types contributes to tumor development.
Understanding lung cancer evolution can identify tools for intercepting its growth. In a landscape analysis of 1024 lung adenocarcinomas (LUAD) with deep whole-genome sequencing integrated with multiomic data, we identified 542 LUAD that displayed diverse clonal architecture. In this group, we observed an interplay between mobile elements, endogenous and exogenous mutational processes, distinct driver genes, and epidemiological features. Our results revealed divergent evolutionary trajectories based on tobacco smoking exposure, ancestry, and sex. LUAD from smokers showed an abundance of tobacco-related C:G>A:T driver mutations in KRAS plus short subclonal diversification. LUAD in never smokers showed early occurrence of copy number alterations and EGFR mutations associated with SBS5 and SBS40a mutational signatures. Tumors harboring EGFR mutations exhibited long latency, particularly in females of European-ancestry (EU_N). In EU_N, EGFR mutations preceded the occurrence of other driver genes, including TP53 and RBM10. Tumors from Asian never smokers showed a short clonal evolution and presented with heterogeneous repetitive patterns for the inferred mutational order. Importantly, we found that the mutational signature ID2 is a marker of a previously unrecognized mechanism for LUAD evolution. Tumors with ID2 showed short latency and high L1 retrotransposon activity linked to L1 promoter demethylation. These tumors exhibited an aggressive phenotype, characterized by increased genomic instability, elevated hypoxia scores, low burden of neoantigens, propensity to develop metastasis, and poor overall survival. Reactivated L1 retrotransposition-induced mutagenesis can contribute to the origin of the mutational signature ID2, including through the regulation of the transcriptional factor ZNF695, a member of the KZFP family. The complex nature of LUAD evolution creates both challenges and opportunities for screening and treatment plans.
Approximately 30% of patients with chronic myelomonocytic leukemia (CMML) undergo transformation to a chemo-refractory blastic phase (BP-CMML). Seeking novel therapeutic approaches, we profiled blast transcriptomes from 42 BP-CMMLs, observing extensive transcriptional heterogeneity and poor alignment to current acute myeloid leukemia (AML) classifications. BP-CMMLs display distinctive transcriptomic profiles, including enrichment for quiescence and variability in drug response signatures. Integrating clinical, immunophenotype, and transcriptome parameters, Random Forest unsupervised clustering distinguishes immature and mature subtypes characterized by differential expression of transcriptional modules, oncogenes, apoptotic regulators, and patterns of surface marker expression. Subtypes differ in predicted response to AML drugs, validated ex vivo in primary samples. Iteratively refined stratification resolves a classification structure comprising five subtypes along a maturation spectrum, predictive of response to novel agents including consistent patterns for receptor tyrosine kinase (RTK), cyclin-dependent kinase (CDK), mechanistic target of rapamycin (MTOR), and mitogen-activated protein kinase (MAPK) inhibitors. Finally, we generate a prototype decision tree to stratify BP-CMML with high specificity and sensitivity, requiring validation but with potential clinical applicability to guide personalized drug selection for improved outcomes.
Background: Breast cancer (BC) stands as a prominent contributor to cancer-related fatalities on a global scale. Traditionally, the primary focus in assessing recurrence risk and guiding treatment decisions has revolved around scrutinizing the primary tumor’s histopathological characteristics and molecular compositions. However, recent findings suggest that exploring normal tissue adjacent to the tumor (NAT) could unveil valuable insights into the intricate characteristics of BC advancement and patient prognoses. Specifically, in our recent work, we discovered three phylogenetic tree groups of NAT and tumor tissues that were associated with distinct tumor and tumor microenvironment (TME) features using the whole genome sequencing data collected from 43 Hong Kong BC patients (Zhu et al. under review). Here, we aimed to further characterize the molecular heterogeneity of NAT by integrating methylation and gene expression profiling data in an expanded sample set, which may deepen our understanding of tumor evolution. Methods: Genome-wide DNA was extracted from fresh frozen paired tumor/NAT tissues in 188 patients diagnosed with breast cancer and treated in Hong Kong, China. DNA was then profiled using the Infinium Methylation 850K BeadChip array. Horvath clock was utilized to estimate epigenetic age, and MethylCIBERSORT was used to estimate cellular composition using the methylation data. For a subset of the patients (N= 76), we also analyzed RNA sequencing (RNA-Seq) data and used CIBERSORTx to estimate the proportions of immune cell subpopulations in NAT samples. Risk factor and clinical information was available for most patients. Results: The uniform manifold approximation and projection (UMAP) identified two distinct clusters of NAT samples (cluster 1, N= 139; cluster 2, N= 49) based on the top 10K most variable methylation probes. The two clusters were recapitulated by unsupervised clustering. The two sets of patients did not differ by chronological age at diagnosis, but when investigating epigenetic aging - as captured by Horvath clock – we found that cluster 2 displayed a lower age acceleration in NAT compared to cluster 1. Among the subset of patients with phylogenetic trees constructed in our previous work, we observed cluster 1 had a higher frequency of patients (64.3%) that showed distinct evolutionary trajectories in matched NAT and tumor samples (multiple-tree group) compared to cluster 2 (33.3%). Consistent with previous findings, we found that cluster 1, enriched with multiple-tree NAT samples, was associated with lower proportions of CD14 cell population (P=1.6x10-6), monocytes (P=7.1x10-6), and M2 macrophages (P=2.3x10-7) compared to cluster 2. Interestingly, genomic features in the matched tumors such as PAM50 subtype, TP53 mutation status, or tumor mutational burden, did not vary significantly between clusters 1 and 2. Lastly, established BC risk factors examined in this study did not show significant differences by this methylation-based cluster assignment. Conclusion: Using methylation data in an expanded dataset, we extended our previous analysis of paired tumor and NAT samples to characterize breast cancer field effect. Our main findings illuminate the diverse nature of the microenvironment surrounding the tumor and its potential influence on the genomic evolutionary path of the tumor. Citation Format: Hela Koka, Bin Zhu, Priscilla Lee, Kevin Wang, Avraam Tapinos, Difei Wang, Gary M. Tse, Koon-ho Tsang, Cherry Wu, Chad A. Highfill, Kristine Jones, Belynda Hicks, Amy Hutchinson, Montserrat Garcia-Closas, Stephen Chanock, David C. Wedge, Lap Ah Tse, Xiaohong R. Yang. Molecular heterogeneity in adjacent normal tissue among Chinese breast cancer patients [abstract]. In: Proceedings of the San Antonio Breast Cancer Symposium 2024; 2024 Dec 10-13; San Antonio, TX. Philadelphia (PA): AACR; Clin Cancer Res 2025;31(12 Suppl):Abstract nr P2-02-06.
Breast cancer (BC) is a heterogenous disease, and the rising global epidemic of premenopausal BC remains poorly understood. Mutational landscape and evolutionary dynamics of BC from diverse geography and populations are largely unknown, preventing the global acceleration of progress towards precision prevention. To examine etiology and heterogeneity of BC in indigenous Africans, 43 multi-region tumor samples and blood samples from 18 Nigerian women with BC (mean age 53 +/-12.7) were analyzed. Whole Genome Sequencing (WGS) was performed to identify somatic single nucleotide variants, insertions and deletions (ID), copy number alterations (CN), and structural variants (SV). Five mutational signature types (single base substitutions (SBS), double base substitutions (DBS), ID, CN, and SV), were analyzed and compared to COSMIC human cancer signatures. Multi-DPclust was used to classify mutations to clonal and subclonal cluster and create mutational phylogenetic trees. Driver gene analysis highlighted key driver genes, such as TP53, GATA3, and PIK3CA, corroborating prior findings. The most common signature was clock-like (CL) SBS5, followed by APOBEC-related SBS2 and SBS13. Other notable signatures included hypoxia-linked SBS18 and HRD-related DBS13. Signature profiles showed moderate heterogeneity across samples, with distinct patterns observed across mutational clusters. Intra-cluster correlation coefficients for signatures within samples from same patient range from 0.37 to 0.99 (median 0.84). TP53 and GATA3 mutations are common in clonal (early) clusters, with GATA3 often appearing in consecutive subclones. The CL SBS5 signature was prevalent early, while SBS18, linked to hypoxia, became prominent late, highlighting evolving mutational processes across disease progression. Distinct BC subtypes also displayed unique mutational profiles. HR+/HER2+ tumors exhibited higher levels of signature SBS91, while HR-/HER2- tumors exhibited higher levels of the HRD signature SBS3 compared to the rest of the tumors. HR-/HER2- tumors demonstrated lower levels of the mismatch repair-related signature ID1 but higher levels of the HRD signature ID6 and TOP2A signature ID8. Significant differences were observed in SV signatures, with higher SV9 signature activities in HR+/HER2+ tumors and elevated BRCA-related SV3 signature activities in HR-/HER2- tumors. Heterogeneity was also evident in the DBS and CN mutational signatures, further highlighting tumor complexity. This analysis sheds light on the diverse mutational dynamics within and across molecular subtypes, providing insights into the mutational evolution of BC in a non-screen detected young onset population. Ongoing work integrating WGS and transcriptome data will be presented at the conference. Avraam Tapinos, Toshio Yoshimatsu, Ilona Siljander, Mustapha A. Ajan, Ayodele Sanni, Atara Ntekim, Abayomi Odetunde, Elisabeth Sveen, Jeffrey Mueller, Galina Khramtsova, Sulin Wu, Dorothy Nyamai, Mihai Giurcanu, Dezheng Huo, Yonglan Zheng, David C. Wedge, Olufunmilayo I. Olopade. Multi-sample whole genome sequencing unveils complex mutational dynamics and clonal evolutionary patterns in young onset breast cancer from Nigeria [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 3892.
Chromothripsis, the chaotic shattering and repair of chromosomes, is common in cancer. Whether chromothripsis generates actionable therapeutic targets remains an open question. In a cohort of 64 patients in blast phase of a myeloproliferative neoplasm (BP-MPN), we describe recurrent amplification of a region of chromosome 21q ('chr. 21amp') in 25%, driven by chromothripsis in a third of these cases. We report that chr. 21amp BP-MPN has a particularly aggressive and treatment-resistant phenotype. DYRK1A, a serine threonine kinase, is the only gene in the 2.7-megabase minimally amplified region that showed both increased expression and chromatin accessibility compared with non-chr. 21amp BP-MPN controls. DYRK1A is a central node at the nexus of multiple cellular functions critical for BP-MPN development and is essential for BP-MPN cell proliferation in vitro and in vivo, and represents a druggable axis. Collectively, these findings define chr. 21amp as a prognostic biomarker in BP-MPN, and link chromothripsis to a therapeutic target.
Diffuse gliomas are the commonest malignant primary brain tumour in adults. Herein, we present analysis of the genomic landscape of adult glioma, by whole genome sequencing of 403 tumours (256 glioblastoma, 89 astrocytoma, 58 oligodendroglioma; 338 primary, 65 recurrence). We identify an extended catalogue of recurrent coding and non-coding genetic mutations that represents a source for future studies and provides a high-resolution map of structural variants, copy number changes and global genome features including telomere length, mutational signatures and extrachromosomal DNA. Finally, we relate these to clinical outcome. As well as identifying drug targets for treatment of glioma our findings offer the prospect of improving treatment allocation with established targeted therapies.
Prostate cancer commonly presents as multifocal disease, with distinct tumours that are often genetically independent. Metastases, however, typically derive from a single dominant clone, highlighting the clinical challenge of identifying which lesion drives progression. Focal therapy, an emerging treatment strategy, aims to target only the most aggressive tumour focus or foci, sparing normal prostate structures and reducing morbidity. Yet, the lack of consistent genomic drivers in primary tumours complicates efforts to identify the focus most likely to progress. While prostate cancers are known to harbour widespread methylation alterations, the extent to which these epigenetic changes track with clonal origin remains poorly understood. To address this, we performed multiregional epigenomic sequencing of 189 prostate tissue samples, including 109 tumour foci and 80 non-tumour regions from 44 patients (2-8 tumour samples per case). Differential methylation analysis revealed widespread alterations distinguishing benign from tumour tissue, including a core ‘epigenetic trunk’ of >20,000 recurrently altered CpGs shared across cases. DNA variants called from non-methylation sites of the epigenomic sequencing closely matched variants from WGS in cases with WGS available. Phylogenetic trees constructed from DNA variants and phyloepigenetic trees from methylation sequencing revealed striking patterns. Some cases lacked shared DNA variants across all tumour foci, indicating a polyclonal tumour origin, yet these genetically unrelated foci shared extensive methylation alterations. This suggests widespread methylation convergence occurs in clonally distinct tumour foci. To dissect clonal relationships further, we applied two orthogonal lineage tracing approaches using the methylation sequencing data: fluctuating CpGs (sites that change stochastically over time) and highly entropic 8-mer ‘methylation barcodes’ within the protocadherin gene cluster. Fluctuating CpGs showed concordant states across tumours that shared DNA sequence variants, while the methylation barcodes likewise revealed high similarity among such tumours. Together, these methods support the independent evolutionary origins of polyclonal prostate tumour foci. Pathway analysis of recurrently altered methylation loci highlighted consistent enrichment in epithelial–mesenchymal transition, MYC signalling, DNA damage response, and hormonal signalling across cases, suggesting functional convergence at the pathway level despite divergent genetic origins, consistent with selective pressures acting on shared programmes. Together, these findings demonstrate that DNA methylation alterations represent fundamental and recurrent events in prostate cancer, more consistent than DNA sequence variants across multifocal disease. Lineage-tracing analyses further highlight that methylation processes record tumour evolutionary history, encoding both clonal relationships and convergent biology. These observations may inform methylation-based biomarkers and the rational selection of lesions for focal therapy. Tamsin J. Robb, Melissa Cheung, Rajbir Batra, Henson Lee Yu, John C. Thomas, Anne Warren, Andrew Lynch, Daniel Brewer, David Wedge, CRUK-ICGC Prostate Cancer Group, Charles Massie, Harveer Dev. Widespread methylation convergence in clonally distinct foci of multifocal prostate cancer [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Cancer Evolution: The Dynamics of Progression and Persistence; 2025 Dec 4-6; Albuquerque, NM. Philadelphia (PA): AACR; Cancer Res 2025;85(23_Suppl):Abstract nr B003.
The role of extrachromosomal DNA (ecDNA) in lung cancer, particularly in subjects who never smoked (LCINS), remains unclear. Examination of 1,216 whole-genome-sequenced lung cancers identified ecDNA in 18.9% of patients. Recurrent amplification of MDM2 and other oncogenes via ecDNA possibly drives a LCINS subset. Tumors harboring ecDNA showed worse overall survival than tumors harboring other focal amplifications. A strong association with whole-genome doubling suggests most ecDNA reflects genomic instability in treatment-naïve lung cancer.
Lung cancer in never smokers (LCINS) accounts for around 25% of all lung cancers1,2 and has been associated with exposure to second-hand tobacco smoke and air pollution in observational studies3-5. Here we use data from the Sherlock-Lung study to evaluate mutagenic exposures in LCINS by examining the cancer genomes of 871 treatment-naive individuals with lung cancer who had never smoked, from 28 geographical locations. KRAS mutations were 3.8 times more common in adenocarcinomas of never smokers from North America and Europe than in those from East Asia, whereas a higher prevalence of EGFR and TP53 mutations was observed in adenocarcinomas of never smokers from East Asia. Signature SBS40a, with unknown cause6, contributed the largest proportion of single base substitutions in adenocarcinomas, and was enriched in cases with EGFR mutations. Signature SBS22a, which is associated with exposure to aristolochic acid7,8, was observed almost exclusively in patients from Taiwan. Exposure to secondhand smoke was not associated with individual driver mutations or mutational signatures. By contrast, patients from regions with high levels of air pollution were more likely to have TP53 mutations and shorter telomeres. They also exhibited an increase in most types of mutations, including a 3.9-fold increase in signature SBS4, which has previously been linked with tobacco smoking9, and a 76% increase in the clock-like10 signature SBS5. A positive dose-response effect was observed with air-pollution levels, correlating with both a decrease in telomere length and an increase in somatic mutations, mainly attributed to signatures SBS4 and SBS5. Our results elucidate the diversity of mutational processes shaping the genomic landscape of lung cancer in never smokers.
Cancer progression involves the sequential accumulation of genetic alterations that cumulatively shape the tumour phenotype. In prostate cancer, tumours can follow divergent evolutionary trajectories that lead to distinct subtypes, but the causes of this divergence remain unclear. While causal inference could elucidate the factors involved, conventional methods are unsuitable due to the possibility of unobserved confounders and ambiguity in the direction of causality. Here, we propose a method that circumvents these issues and apply it to genomic data from 829 prostate cancer patients. We identify several genetic alterations that drive divergence as well as others that prevent this transition, locking tumours into one trajectory. Further analysis reveals that these genetic alterations may cause each other, implying a positive-feedback loop that accelerates divergence. Our findings provide insights into how cancer subtypes emerge and offer a foundation for genomic surveillance strategies aimed at monitoring the progression of prostate cancer.
Prostate cancer (PCa) germline testing, while gaining momentum, is ancestry restrictive and African exclusive. Through whole genome sequencing for 217 African ancestral cases (186 southern African, 31 Pan representative), we identify 172 potentially pathogenic variants in 78 DNA damage repair or PCa related genes. Prevalence for reported (13/217, 5.99%) and cumulative predicted (24/217, 11.06%) variants of significance (11 genes) falls below that reported for non-Africans. Conversely, BRCA1, HOXB13, CDK12, MLH1, MSH2, and BRIP1 remain unimpacted. Through pathogenic ranking based on variant frequency and functionality, clinical presentation and tumour-matched biallelic inactivation, top-ranked candidates include PREX2, POLE, FAT1, BRCA2, POLQ, LRP1B and ATM. Besides notable impact of DNA polymerases, including POLG, Fanconi anaemia genes include FANCD2, FANCA, FANCG, ERCC4, FANCE and FANCI, while DNA mismatch repair genes MSH3 and PMS1 outranked known namesakes MSH6 and PMS2. This study provides insights into the spectrum of African-relevant potentially pathogenic PCa variants, highlighting much-needed gene candidates for ancestry-inclusive germline testing.
Oncomicrobes are estimated to cause 15% of cancers worldwide. When cancer whole-genome sequencing (WGS) data are collected, the microbes present are also sequenced, allowing the investigation of potential etiological and clinical associations. Interrogating the microbial community for 8908 patients encompassing 22 cancer types from the Genomics England WGS dataset revealed that only colorectal tumors exhibited unmistakably distinct microbial communities that can reliably be used to distinguish anatomical site [positive predictive value (PPV) = 0.95]. This pattern was validated in two independent datasets. Potential clinical relevance uncovered by our analyses included accurate detection of alphapapillomaviruses [human papillomavirus (HPV)] in oral cancers, when compared with current clinical standards, and the detection of rare, highly pathogenic viruses such as human T-lymphotropic virus-1. Biomarker investigations demonstrated statistically significant associations (P < 0.05) between a subset of anaerobic bacteria and survival in certain subtypes of sarcoma. Our results contradict previous claims that each cancer type has a distinct microbiological signature but highlight the potential value of microbial analysis for certain cancers as WGS of tumor samples becomes common in the clinic.
Mutational signature analysis has greatly enhanced our understanding of the mutagenic processes found in cancer and normal tissues. As part of a recent study, we analyzed 802 treatment-naïve, microsatellite-stable colorectal cancers (CRC) and identified a de novo signature, SBS_D, which was conservatively decomposed into SBS18, a signature associated with reactive oxygen species. Here, we re-evaluate this decomposition and provide evidence that SBS_D represents a distinct mutational process from that of SBS18. Through an independent analysis of 2,616 whole-genome sequenced microsatellite-stable CRCs across three distinct cohorts, we demonstrate that SBS_D is consistently present at a similar prevalence, suggesting that this signature may have been previously overlooked. Using a naïve decomposition approach, we demonstrate that the pattern of SBS_D better aligns with signatures previously associated with deficiencies in DNA polymerase delta (POLD1) proofreading and mismatch repair. However, multiple lines of evidence, including the absence of pathogenic mutations in the exonuclease domain of POLD1 or in mismatch repair-associated genes, indicate that SBS_D is not driven by canonical defects in these DNA repair pathways. Overall, this study identifies a previously unrecognized mutational signature in microsatellite-stable CRC and proposes that its etiology may be linked to DNA repair infidelity emerging late in tumor development in samples without canonical defects in DNA repair pathways.