Multiple sclerosis (MS) is a disorder of the CNS in which autoreactive immune cells migrate through a damaged blood-brain barrier, resulting in focal demyelinating lesions. Beyond focal lesions, there are also diffuse 'surface-in' gradients of pathology in MS, wherein damage is most severe directly adjacent to CSF-contacting surfaces, such as the subpial and periventricular areas. This observation suggests that toxic factors within MS CSF contribute to the emergence and/or evolution of surface-in gradients. Directly separating the CSF from the periventricular parenchyma are ependymal cells-a glial epithelium-that are equipped with tufts of motile cilia, which are critical for circulating CSF solutes and regulating local fluid flow. While damage to ependymal cilia has the potential to drastically modify CSF homeostasis and thus contribute to the damage of CSF exposed regions, these motile cellular structures have yet to be investigated in the context of MS. We first conducted single-cell RNA sequencing of fresh human periventricular brain tissue containing ependymal cells from patients with MS and non-MS disease controls. We subsequently collected CSF from patients with MS and exposed cultured rodent ependymal cells to this CSF to evaluate the impact on ependymal ciliary function. To complement our direct evaluation of cilia in the context of MS, we also confirmed whether cilia were altered in an animal model of MS, experimental autoimmune encephalomyelitis (EAE), and designed a novel transgenic animal model to evaluate the cellular and behavioural effect(s) of adult ependymal ciliary disruption. Single-cell RNA sequencing analysis of human ependymal cells in MS demonstrated large-scale dysregulation of ciliary genes, and in situ stains of MS brain tissue confirmed a loss of ependymal cilia. Exposure of ependymal cells to MS CSF led to transcriptional modification of ciliary gene and protein expression and reduced ciliary beating frequency. Likewise, analysis of ependymal cells in EAE demonstrated altered cilia gene and protein expression. We showed that IFNγ, which is elevated in MS CSF, could alter cilia protein expression and motility. Lastly, conditional knockout of Ccdc39 in ependymal cells of adult mice led to transient ventricular enlargement, increased periventricular microglial density and alterations in nesting behaviour. These data suggest that motile cilia in ependymal cells are dysregulated in CNS autoimmunity. More importantly, they suggest that ependymal cilia disruption could play a role in periventricular pathology formation in MS and be associated with behavioural deficits underlying non-motor symptomatology.
Anastrepha fraterculus is a cryptic species complex with at least eight morphotypes distributed across the Americas. Among them, A. fraterculus sp.1, present in Argentina, is a major pest impacting fresh fruit production. Integrated pest management strategies, including chemical control and trapping, are currently employed to mitigate its effects. Genetic sexing strains of A. fraterculus sp.1 are being evaluated for use in sterile insect technique programs. To support traditional and emerging control methods, this study aimed to enhance the genomic understanding of this morphotype. Individual female and male samples were sequenced using long- and short-read technologies. The female genome (760 Mb) was de novo assembled into 58 scaffolds and the male genome (750 Mb) into 68 scaffolds, with BUSCO completeness scores of 98.8% and 98.7%, respectively. Synteny analysis revealed complete scaffolds of the five autosomes and enabled near-complete reconstruction of the X and Y chromosomes. Gene prediction identified 17 751 and 16 535 protein-coding genes (for female and male genomes, respectively), with repetitive regions representing 46% of both genomes. Additionally, the mitochondrial genome was fully assembled and annotated. This comprehensive genomic resource reveals candidate genes for functional studies, including gene editing and RNA interference, as successfully applied in related tephritid species. These findings lay the foundation for innovative, complementary biocontrol tools against A. fraterculus.
Advances in DNA sequencing have transformed genomics, enabling comprehensive insights into human genetic variation. While short-read sequencing (SRS) remains dominant due to its high accuracy and affordability, its limitations in complex genomic regions have spurred the adoption of long-read sequencing (LRS) platforms, such as those from Pacific Biosciences (PacBio) and Oxford Nanopore Technologies (ONT). Despite these advances, there is still a lack of systematic, large-scale benchmarking of variant calling performance across diverse platforms, variant types, sequencing depths, and genomic contexts. Here, we present a comprehensive benchmark of sequencing technologies and variant calling algorithms, evaluating their performance in detecting single nucleotide polymorphisms (SNPs), insertions/deletions (indels), and structural variants (SVs). We show that while SRS combined with DeepVariant or DRAGEN offers excellent small variant detection in well-mapped regions, LRS technologies significantly outperform SRS in complex regions and SV detection. PacBio achieves high SNP and smaller SV accuracy even at moderate coverage, while ONT excels in detecting large SVs. The SV callers Dysgu and SVIM emerged as top performers across LRS datasets. Our results highlight that no single platform is optimal for all variant types or regions: SRS remains optimal for high-throughput small variant detection in accessible regions, whereas LRS is critical for capturing SVs and resolving difficult-to-map loci. These findings offer practical guidance for selecting sequencing technologies, coverage and variant calling strategies tailored to specific research or clinical goals, contributing to more accurate and cost-effective genomic analyses. Highlights ● State-of-the-art short variant calling algorithms are highly comparable but focus on precision and sensitivity differently ● Short-read technologies outperform long-read technologies for small SNP and InDels, with the exception of difficult-to-map variants. ● To capture the majority of variants, a minimum coverage of 15x for PacBio, 20x for SRS, or 30x for ONT is required. However, optimal coverage depends on zygosity, variant type, and the region of interest. ● Long-read technologies outperform short-read technologies for all validation sets tested for deletions and insertions in all size categories. ### Competing Interest Statement The authors have declared no competing interest. Canada Institute of Health Research (CIHR) project grant, PJT-191707 Genome Canada Genome Technology Platform grant
Esophageal adenocarcinoma arises from Barrett's esophagus, a metaplastic condition. Multi-omics profiling, integrating single-cell transcriptomics, extracellular matrix proteomics, tissue mechanics and spatial proteomics of the paths of progression from squamous epithelium through metaplasia, dysplasia to adenocarcinoma, in 107 samples from 26 patients in two independent cohorts, defined shared and patient-specific progression characteristics. Metaplastic replacement of epithelial cell composition and architecture was paralleled by changes in stromal cells, extracellular matrix (ECM) and tissue stiffness. This change in pre-cancerous metaplasia was already accompanied by appearance of fibroblasts with the molecular characteristics of carcinoma-associated fibroblasts. These fibroblasts produced the immunosuppressive protein POSTN, whose expression shifted from vascular to stromal cells, consistent with the emergence of an immunosuppressive microenvironment evident in cell neighborhoods enriched for immunoregulatory NK and Treg cells. Thus, Barrett's esophagus progresses as a coordinated multi-component system, supporting treatment paradigms that go beyond targeting cancerous cells to incorporate stromal reprogramming.
The tumor microenvironment (TME) of chronic inflammation-associated cancers (CIACs) is shaped by cycles of injury and maladaptive repair, yet the principles organizing fibrotic stroma in these tumors remain unclear. Here, we applied the concept of hot versus cold fibrosis, originally credentialed in non-cancerous fibrosis of heart and kidney, to lung squamous cell carcinoma (LUSC), a prototypical CIAC. Single-cell transcriptomics of matched tumor and adjacent-normal tissue from 16 treatment-naive LUSC patients identified a cold fibrotic architecture in the LUSC TME: cancer-associated fibroblasts (CAFs) expanded and adopted myofibroblast and stress-response states, while macrophages were depleted. This macrophage-poor, CAF-rich stroma was maintained by CAF autocrine growth factor loops, including TIMP1, INHBA, TGFB1, and GMFB. In parallel, the immune compartment exhibited a hot tumor phenotype with abundant T and B cells, forming spatially distinct but molecularly engaged networks with CAFs. CAF gene programs typifying cold fibrosis in LUSC were conserved in other CIACs, including esophageal and gastric adenocarcinomas. These results redefine desmoplastic regions of tumors through the lens of a non-cancer fibrosis model, demonstrating that conserved stromal circuits constitute therapeutic vulnerabilities in CIACs.
Brook trout (Salvelinus fontinalis) is a socioeconomically important fish species for fisheries, aquaculture, and aquatic conservation. We produced a 2.5-Gb reference assembly by combining Hi-C chromosome conformation capture with high-coverage short- and long-read sequencing of a fully homozygous mitotic gynogenic doubled-haploid fish, which facilitates the assembly of highly complex salmonid genomes. The assembly has an N50 of 50.98 Mb and 88.9% of the total assembled sequence length is anchored into 42 main chromosomes, of which 63.44% represents repeated contents, including 1,461,010 DNA transposons. 56,058 genes were found, with 98.6% of the 3,640 expected conserved orthologs BUSCO genes (actinopterygii_odb10 lineage database). Additionally, we found significant homology within the 42 chromosomes, as expected for this pseudo-tetraploid species, as well as with the sister species lake trout (Salvelinus namaycush) and Atlantic salmon (Salmo salar). This assembly will serve as a reliable genomic resource for brook trout, thus enabling a wider range of reference-based applications to support ongoing research and management decision-making for the species.
The genome of the hemlock woolly adelgid (Adelges tsugae) is presented in 10 chromosomal pseudomolecules along with 132 unplaced scaffolds for a combined length of 216.56 Mb. The genome is highly contiguous with an N50 of 20.54 Mb and its longest scaffold 37.41 Mb in length. The assembly recovered 97.7% of benchmarked universal single-copy Hemiptera orthologs (BUSCO) with an overall completion score of 98.6%. On the basis of chromosomal synteny with other aphid genomes, the hemlock woolly adelgid's genome is composed of eight autosomes and two sex (X) chromosomes. Annotation of the assembly identified 11,800 coding genes, 1,930 noncoding genes, and 20,403 mRNA transcripts. The mitogenome is also presented in a single, annotated, circular contig 25,980 bases in length.
Metastatic breast cancer with complex molecular mechanisms of progression accounts for most cancer related deaths in women. To improve diagnosis and drug development, it is important to identify novel biomarkers and critical molecular pathways involved in tumor initiation and progression. Here, we profiled and analyzed the expression of lncRNAs from three distinct stages of tumor initiation and progression (hyperplasia, adenoma, and carcinoma) in mammary tumors. We performed RNAseq on tumor and mammary epithelial cells derived from ROSAmT/mG; MMTV-PyMT mice and ROSAmT/mG; non-cancer mice derived from the same strain, respectively. We identified 1913 differentially expressed protein coding genes and 324 lncRNAs in breast cancer cells of all stages compared with normal mammary epithelial cells. Pearson correlation analysis correlated 93 differentially expressed lncRNAs with protein coding genes, providing a comprehensive lncRNA-coding gene co-expression network. Among them, we focused on Gm19303 which was paired with the differentially expressed protein coding gene, Trps1, and identified its human counterpart as LINC00536. Both LINC00536 and human TRPS1 are exclusively overexpressed in breast cancer and correlate with poor patient prognosis from the TCGA and GTEx databases. Single cell RNAseq data from the Atlas of Human breast cancers further confirmed that TRPS1 is upregulated in human breast cancer compared to normal human mammary tissue with highest expression in ER+ subgroup. In summary, our study explored the potential role of lncRNAs in breast cancer initiation and progression, presented the lncRNA-protein coding genes co-expression profiles, and identified human LINC00536/TRPS1 as a potential diagnostic and therapeutic target for breast cancer. Rui Zhang, Jiarong Li, Dunarel Badescu, Jiannis Ragoussis, Richard Kremer. New insights from transgenic mouse models of PyMT-induced breast cancer: identifying novel long non-coding RNA biomarkers [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 3913.
Friend of GATA1 (FOG-1) is an essential transcriptional co-factor of the master erythroid transcription factor GATA1. The knockout of the Zfpm1 gene, coding for FOG-1, results in early embryonic lethality due to anemia in mice, similar to the embryonic lethal phenotype of the Gata1 gene knockout. However, a detailed molecular analysis of the Zfpm1 knockout phenotype in erythropoiesis is presently incomplete. To this end, we used CRISPR/Cas9 to knockout Zfpm1 in mouse erythroleukemic (MEL) cells. Phenotypic characterization of DMSO-induced terminal erythroid differentiation showed that the Zfpm1 knockout MEL cells did not progress past the proerythroblast stage of differentiation. Expression profiling of the Zfpm1 knockout MEL cells by RNAseq showed a lack of upregulation of erythroid-related gene expression profiles. Bioinformatic analysis highlighted cholesterol transport as a pathway affected in the Zfpm1 knockout cells. Moreover, we show that the cholesterol transporters Abca1 and Ldlr fail to be repressed during erythroid differentiation in Zfpm1 knockout cells, resulting in higher intracellular lipid levels and higher membrane fluidity. We also show that in FOG-1 knockout cells, the nuclear levels of SREBP2, a key transcriptional regulator of cholesterol biosynthesis and transport, are markedly increased. On the basis of these findings we propose that FOG-1 (and, potentially, GATA1) regulate cholesterol homeostasis during erythroid differentiation directly through the downregulation of cholesterol transport genes and indirectly, through the repression of the SREBP2 transcriptional activator of cholesterol homeostasis. Taken together, our work provides a molecular basis for understanding FOG-1 functions in eryhropoiesis and reveals a novel role for FOG-1 in cholesterol transport.
The persistence of SARS-CoV-2 despite the development of vaccines and a degree of herd immunity is partly due to viral evolution reducing vaccine and treatment efficacy. Serial infections of wild-type (WT) SARS-CoV-2 in Balb/c mice yield mouse-adapted strains with greater infectivity and mortality. We investigate if passaging unmodified B.1.351 (Beta) and B.1.617.2 (Delta) 20 times in K18-ACE2 mice, expressing the human ACE2 receptor, in a BSL-3 laboratory without selective pressures, drives human health-relevant evolution and if evolution is lineage-dependent. Late-passage virus causes more severe disease, at organism and lung tissue scales, with late-passage Delta demonstrating antibody resistance and interferon suppression. This resistance co-occurs with a de novo spike S371F mutation, linked with both traits. S371F, an Omicron-characteristic mutation, is co-inherited at times with spike E1182G per Nanopore sequencing, existing in different within-sample viral variants at others. Both S371F and E1182G are linked to mammalian GOLGA7 and ZDHHC5 interactions, which mediate viral-cell entry and antiviral response. This study demonstrates SARS-CoV-2’s tendency to evolve with phenotypic consequences, its evolution varying by lineage, and suggests non-dominant quasi-species contribution.
BACKGROUND:Cardiomyopathy is a clinically and genetically heterogeneous heart condition that can lead to heart failure and sudden cardiac death in childhood. While it has a strong genetic basis, the genetic aetiology for over 50% of cardiomyopathy cases remains unknown. METHODS:In this study, we analyse the characteristics of tandem repeats from genome sequence data of unrelated individuals diagnosed with cardiomyopathy from Canada and the United Kingdom (n = 1216) and compare them to those found in the general population. We perform burden analysis to identify genomic and epigenomic features that are impacted by rare tandem repeat expansions (TREs), and enrichment analysis to identify functional pathways that are involved in the TRE-associated genes in cardiomyopathy. We use Oxford Nanopore targeted long-read sequencing to validate repeat size and methylation status of one of the most recurrent TREs. We also compare the TRE-associated genes to those that are dysregulated in the heart tissues of individuals with cardiomyopathy. FINDINGS:We demonstrate that tandem repeats that are rarely expanded in the general population are predominantly expanded in cardiomyopathy. We find that rare TREs are disproportionately present in constrained genes near transcriptional start sites, have high GC content, and frequently overlap active enhancer H3K27ac marks, where expansion-related DNA methylation may reduce gene expression. We demonstrate the gene silencing effect of expanded CGG tandem repeats in DIP2B through promoter hypermethylation. We show that the enhancer-associated loci are found in genes that are highly expressed in human cardiomyocytes and are differentially expressed in the left ventricle of the heart in individuals with cardiomyopathy. INTERPRETATION:Our findings highlight the underrecognized contribution of rare tandem repeat expansions to the risk of cardiomyopathy and suggest that rare TREs contribute to ∼4% of cardiomyopathy risk. FUNDING:Government of Ontario (RKCY), The Canadian Institutes of Health Research PJT 175329 (RKCY), The Azrieli Foundation (RKCY), SickKids Catalyst Scholar in Genetics (RKCY), The University of Toronto McLaughlin Centre (RKCY, SM), Ted Rogers Centre for Heart Research (SM), Data Sciences Institute at the University of Toronto (SM), The Canadian Institutes of Health Research PJT 175034 (SM), The Canadian Institutes of Health Research ENP 161429 under the frame of ERA PerMed (SM, RL), Heart and Stroke Foundation of Ontario & Robert M Freedom Chair in Cardiovascular Science (SM), Bitove Family Professorship of Adult Congenital Heart Disease (EO), Canada Foundation for Innovation (SWS, JR), Canada Research Chair (PS), Genome Canada (PS, JR), The Canadian Institutes of Health Research (PS).
Four main medulloblastoma (MB) molecular subtypes have been identified based on transcriptional, DNA methylation, and genetic profiles. However, it is currently not known whether 3D genome architecture differs between MB subtypes. To address this question, we performed in situ Hi-C to reconstruct the 3D genome architecture of MB subtypes. In total, we generated Hi-C and matching transcriptome data for 28 surgical specimens and Hi-C data for one patient-derived xenograft. The average resolution of the Hi-C maps was 6,833 bp. Using these data, we found that insulation scores of topologically associating domains (TADs) were effective at distinguishing MB molecular subgroups. TAD insulation score differences between subtypes were globally not associated with differential gene expression, although we identified few exceptions near genes expressed in the lineages of origin of specific MB subtypes. Our study therefore supports the notion that TAD insulation scores can distinguish MB subtypes independently of their transcriptional differences.
Posterior fossa group A (PFA) ependymoma is a lethal brain cancer diagnosed in infants and young children. The lack of driver events in the PFA linear genome led us to search its 3D genome for characteristic features. Here, we reconstructed 3D genomes from diverse childhood tumor types and uncovered a global topology in PFA that is highly reminiscent of stem and progenitor cells in a variety of human tissues. A remarkable feature exclusively present in PFA are type B ultra long-range interactions in PFAs (TULIPs), regions separated by great distances along the linear genome that interact with each other in the 3D nuclear space with surprising strength. TULIPs occur in all PFA samples and recur at predictable genomic coordinates, and their formation is induced by expression of EZHIP. The universality of TULIPs across PFA samples suggests a conservation of molecular principles that could be exploited therapeutically.
During the COVID-19 pandemic, the monitoring of SARS-CoV-2 RNA in wastewater was used to track the evolution and emergence of variant lineages and gauge infection levels in the community, informing appropriate public health responses without relying solely on clinical testing. As more sublineages were discovered, it increased the difficulty in identifying distinct variants in a mixed population sample, particularly those without a known lineage. Here, we compare the sequencing technology from Illumina and from Oxford Nanopore Technologies, in order to determine their efficacy at detecting variants of differing abundance, using 248 wastewater samples from various Quebec and Ontario cities. Our study used two analytical approaches to identify the main variants in the samples: the presence of signature and marker mutations and the co-occurrence of signature mutations within the same amplicon. We observed that each sequencing method detected certain variants at different frequencies as each method preferentially detects mutations of distinct variants. Illumina sequencing detected more mutations with a predominant lineage that is in low abundance across the population or unknown for that time period, while Nanopore sequencing had a higher detection rate of mutations that are predominantly found in the high abundance B.1.1.7 (Alpha) lineage as well as a higher sequencing rate of co-occurring mutations in the same amplicon. We present a workflow that integrates short-read and long-read sequencing to improve the detection of SARS-CoV-2 variant lineages in mixed population samples, such as wastewater.
Central tolerance of thymocytes to self-antigen depends on the medullary thymic epithelial cell (mTEC) transcription factor autoimmune regulator (Aire), which drives tissue-restricted antigen (TRA) gene expression. Vitamin D signaling regulates Aire and TRA expression in mTECs, providing a basis for links between vitamin D deficiency and autoimmunity. We find that mice lacking Cyp27b1, which cannot produce hormonally active vitamin D, display profoundly reduced thymic cellularity, with a reduced proportion of Aire + mTECs, attenuated TRA expression, and poorly defined cortical-medullary boundaries. Markers of T cell negative selection are diminished, and organ-specific autoantibodies are present in knockout (KO) mice. Single-cell RNA sequencing revealed that loss of Cyp27b1 skews mTEC differentiation toward Ccl21 + intertypical TECs and generates a gene expression profile consistent with premature aging. KO thymi display accelerated involution and reduced expression of thymic longevity factors. Thus, loss of thymic vitamin D signaling disrupts normal mTEC differentiation and function and accelerates thymic aging.
Ovarian and endometrial cancers come within the top-4 for incident cancers as well as deaths in North American women. Cure rates have not improved in 30 years as high-grade subtypes continue to be diagnosed in Stage III/IV. Attempts at early diagnosis have failed because high-grade cancer cells exfoliate and metastasize while the primary cancer is small and undetectable by existing tests based on imaging and blood-based tumour markers. DOvEEgene (Detecting Ovarian and Endometrial cancers Early using genomics) is a genomic uterine pap test developed by a McGill team to screen and detect these cancers while they are confined to the gynecologic organs and curable by surgery. The test identifies pathogenic somatic mutations in uterine brush samples A high sensitivity error-reducing capture technology (DOvEEgene-SureSelectHS) utilizing duplex sequencing interrogates the exons of 23 genes involved in the development of sporadic and hereditary ovarian and endometrial cancers. We apply a combination of germline gene panel testing on saliva samples with deep duplex sequencing to detect somatic mutations at <0.1% VAF, interrogation of microsatellite loci for instability and low coverage WGS for copy number analysis of uterine brush samples. Currently, DOvEEgene is the only test that can discriminate ovarian and endometrial cancers in peri- and postmenopausal women from benign gynecologic diseases common in that age group. This is important because pathogenic somatic driver mutations are also associated with increasing age and benign disease. DOvEEgene incorporates a deep machine-learning derived classifier that can discriminate the mutational signature of these cancers from benign disease aiming for a sensitivity of 70% and a specificity of 100% in a population with high background mutational burden. Here we tested the Onso system, a highly accurate sequencing technology from PacBio in order to potentially increase sensitivity while driving down sequencing costs by reducing required sequencing depth vs the current NGS standard. We sequenced 15 duplex Illumina sequencing libraries produced using the DovEE assay at PE 100bp mode and compared Onso data in non- duplex sequencing mode as well as duplex sequencing mode to the original duplex sequencing method. Here, we present this comparison and highlight the benefits of high accuracy sequencing for the detection of very low frequency (<0.1%) somatic mutations. Citation Format: Jiannis Ragoussis, Nairi Pezeshkian, Lucy Gilbert. Improved detection of low frequency mutations in ovarian and endometrial cancers by utilizing a highly accurate sequencing platform [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 6522.
Whole genome sequencing (WGS) at high-depth (30X) allows the accurate discovery of variants in the coding and non-coding DNA regions and helps elucidate the genetic underpinnings of human health and diseases. Yet, due to the prohibitive cost of high-depth WGS, most large-scale genetic association studies use genotyping arrays or high-depth whole exome sequencing (WES). Here we propose a cost-effective method which we call “Whole Exome Genome Sequencing” (WEGS), that combines low-depth WGS and high-depth WES with up to 8 samples pooled and sequenced simultaneously (multiplexed). We experimentally assess the performance of WEGS with four different depth of coverage and sample multiplexing configurations. We show that the optimal WEGS configurations are 1.7–2.0 times cheaper than standard WES (no-plexing), 1.8–2.1 times cheaper than high-depth WGS, reach similar recall and precision rates in detecting coding variants as WES, and capture more population-specific variants in the rest of the genome that are difficult to recover when using genotype imputation methods. We apply WEGS to 862 patients with peripheral artery disease and show that it directly assesses more known disease-associated variants than a typical genotyping array and thousands of non-imputable variants per disease-associated locus.
The COVID-19 pandemic led to a large global effort to sequence SARS-CoV-2 genomes from patient samples to track viral evolution and inform the public health response. Millions of SARS-CoV-2 genome sequences have been deposited in global public repositories. The Canadian COVID-19 Genomics Network (CanCOGeN - VirusSeq), a consortium tasked with coordinating expanded sequencing of SARS-CoV-2 genomes across Canada early in the pandemic, created the Canadian VirusSeq Data Portal, with associated data pipelines and procedures, to support these efforts. The goal of VirusSeq was to allow open access to Canadian SARS-CoV-2 genomic sequences and enhanced, standardized contextual data that were unavailable in other repositories and that meet FAIR standards (Findable, Accessible, Interoperable and Reusable). In addition, the portal data submission pipeline contains data quality checking procedures and appropriate acknowledgement of data generators that encourages collaboration. From inception to execution, the portal was developed with a conscientious focus on strong data governance principles and practices. Extensive efforts ensured a commitment to Canadian privacy laws, data security standards, and organizational processes. This portal has been coupled with other resources, such as Viral AI, and was further leveraged by the Coronavirus Variants Rapid Response Network (CoVaRR-Net) to produce a suite of continually updated analytical tools and notebooks. Here we highlight this portal (https://virusseq-dataportal.ca/), including its contextual data not available elsewhere, and the Duotang (https://covarr-net.github.io/duotang/duotang.html), a web platform that presents key genomic epidemiology and modelling analyses on circulating and emerging SARS-CoV-2 variants in Canada. Duotang presents dynamic changes in variant composition of SARS-CoV-2 in Canada and by province, estimates variant growth, and displays complementary interactive visualizations, with a text overview of the current situation. The VirusSeq Data Portal and Duotang resources, alongside additional analyses and resources computed from the portal (COVID-MVP, CoVizu), are all open source and freely available. Together, they provide an updated picture of SARS-CoV-2 evolution to spur scientific discussions, inform public discourse, and support communication with and within public health authorities. They also serve as a framework for other jurisdictions interested in open, collaborative sequence data sharing and analyses.
The basal breast cancer subtype is enriched for triple-negative breast cancer (TNBC) and displays consistent large chromosomal deletions. Here, we characterize evolution and maintenance of chromosome 4p (chr4p) loss in basal breast cancer. Analysis of The Cancer Genome Atlas data shows recurrent deletion of chr4p in basal breast cancer. Phylogenetic analysis of a panel of 23 primary tumor/patient-derived xenograft basal breast cancers reveals early evolution of chr4p deletion. Mechanistically we show that chr4p loss is associated with enhanced proliferation. Gene function studies identify an unknown gene, C4orf19, within chr4p, which suppresses proliferation when overexpressed—a member of the PDCD10-GCKIII kinase module we name PGCKA1. Genome-wide pooled overexpression screens using a barcoded library of human open reading frames identify chromosomal regions, including chr4p, that suppress proliferation when overexpressed in a context-dependent manner, implicating network interactions. Together, these results shed light on the early emergence of complex aneuploid karyotypes involving chr4p and adaptive landscapes shaping breast cancer genomes.