Single-cell RNA sequencing (scRNA-seq) has transformed our understanding of phenotypic heterogeneity. Although the predominant focus of scRNA-seq analyses has been assessing gene expression changes, several approaches have been proposed in recent years to identify changes at the DNA level from scRNA-seq data. In this study, we evaluated the relative performance of six strategies for calling single-nucleotide variants from scRNA-seq data using 381 single-cell transcriptomes from five cancer patients. Specifically, we focused on the quality of the inferred genotypes and the resulting single-cell phylogenies. We found that scAllele, Monopogen, and Monovar consistently returned phylogenetically informative genotype calls, providing more precise signals of discrimination between tumor and normal cells within heterogeneous samples and among distinct subclonal lineages in longitudinal samples. In addition, we evaluated the evolution of gene expression along the cell phylogenies. While most transcriptomic variation was very plastic and did not correlate with the cell phylogeny, a group of genes associated with cell cycle processes showed a strong phylogenetic signal in one of the patients, underscoring a potential link between gene expression patterns and lineage-specific traits in the context of cancer progression. In summary, our study highlights the potential of scRNA-seq data for inferring cell phylogenies to decipher the evolutionary dynamics of cell populations.
With rapid advancements in single-cell DNA sequencing (scDNA-seq), various computational methods have been developed to study evolution and call variants on single-cell level. However, modeling deletions remains challenging because they affect total coverage in ways that are difficult to distinguish from technical artifacts. We present DelSIEVE, a statistical method that infers cell phylogeny and single-nucleotide variants, accounting for deletions, from scDNA-seq data. DelSIEVE distinguishes deletions from mutations and artifacts, detecting more evolutionary events than previous methods. Simulations show high performance, and application to cancer samples reveals varying amounts of deletions and double mutants in different tumors.
The genomic diversity of circulating tumor cells (CTCs) and its clinical implications remain poorly understood. In this study, we characterized the mutational landscape of CTC pools stemming from 29 metastatic colorectal cancer (mCRC) patients and examined its relationship with disease progression. Our analysis revealed substantial variation in mutational burden among patients, with all CTC pools harboring non-silent mutations in key CRC driver genes. Importantly, higher genomic diversity in CTC pools was significantly associated with reduced overall survival. Furthermore, the presence of non-silent mutations in BCL9L emerged as a strong predictor of patient survival. Taken together, these findings underscore the potential of CTC genomic profiling as a promising prognostic tool in mCRC and highlight the need for further research into its clinical applications. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported by an AXA Research Fund postdoctoral grant (awarded to J.M.A), and by the Spanish Ministry of Science and Innovation - MICINN (PID2019-106247GB-I00 awarded to D.P.). J.M.A. is currently supported by the AECC (INVES20007FERN). D.P. receives further support from Xunta de Galicia. J.C. received grants from Spain's Carlos III Health Care Institute (Co-funded by European Regional Development Fund/European Social Fund: A way to make Europe/Investing in your future), No. PI17/00837 and PI21/01771. J.C. is additionally funded by the Agencia Gallega de Innovacion (N607B-2020/02). R.P. received support from Roche-Chus Joint Unit (IN853B 2018/03) funded by Axencia Galega de Innovacion (GAIN), Conselleria de Economia, Emprego e Industria. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: All samples were obtained and collected after written informed consent from all subjects using a protocol approved by the Clinical Ethics Committee of Pontevedra-Vigo-Ourense (2018/301 approved 19/06/2018). This study was approved by the Clinical Ethics Committee of Pontevedra-Vigo-Ourense I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study will be made publicly available upon publication
Metastatic colorectal cancer (mCRC) remains a major cause of cancer-related mortality, but few noninvasive biomarkers exist to track disease progression or inform treatment strategies. Circulating tumor cells (CTCs) offer a minimally invasive source of tumor material, yet the prognostic significance of their genomic diversity remains unclear. We conducted whole-exome sequencing of CTC pools from 29 mCRC patients to characterize their mutational landscape and assess associations with overall survival. Our analysis revealed substantial variation in mutational burden among patients, with all CTC pools harboring non-silent mutations in key CRC driver genes. Higher genomic diversity in CTC pools was significantly associated with reduced overall survival. Additionally, non-silent mutations in BCL9L emerged as a strong predictor of patient survival. Genomic diversity and BCL9L mutational status in CTC pools emerged as strong predictors of survival in mCRC, underscoring the potential of CTC genomic profiling as a minimally invasive and clinically relevant prognostic tool in mCRC.
Cancer cell lines are valuable models for studying tumor biology, yet their genomic evolution during culture can compromise experimental reproducibility. We conducted a detailed genomic analysis of the triple-negative breast cancer cell line MDA-MB-231-luc-GFP, examining sublines obtained from different sources, at various time points, and across distinct passages. We introduce the concept of intraline heterogeneity (ILH) to highlight the genomic variability observed among these sublines. Our analyses revealed extensive genomic diversity, including differences in single-nucleotide variants (SNVs) and copy number alterations (CNAs). In particular, CNAs exhibited remarkable heterogeneity, with pronounced chromosomal gains and losses between sublines, underscoring the impact of genomic instability on ILH. These findings suggest that ILH may influence experimental outcomes, emphasizing the importance of considering passage-specific genomic characterization to ensure consistency and reliability in cancer research.
Single-cell genomics enables studying tissues and organisms at the highest resolution. However, since a cell contains a small amount of DNA, single-cell DNA sequencing (scDNA-seq) typically requires single-cell whole-genome amplification (scWGA). Unfortunately, scWGA methods introduce technical biases that complicate the interpretation of scDNA-seq data. We compared six scWGA methods, three MDA (multiple displacement amplification; GenomiPhi, REPLI-g, and TruePrime) and three non-MDA (Ampli1, MALBAC, and PicoPLEX), on 206 tumoral and 24 healthy human cells. scWGA methods performed differently depending on the parameter of interest. REPLI-g minimized regional amplification bias, while non-MDA methods showed a more uniform and reproducible amplification. Ampli1 exhibited the lowest allelic imbalance and dropout, the most accurate insertion or deletion (indel) and copy-number detection, and a low polymerase error rate. However, REPLI-g yielded higher DNA quantities, longer amplicons, and greater genome coverage. We offer a comprehensive guide for selecting a scWGA approach, outlining trade-offs that influence the interpretation of scDNA-seq data.
BACKGROUND:The spread of SARS-CoV-2 has been influenced by multiple factors, from the inherent transmission capabilities of the different variants to the control measurements put in place. Understanding how new variants enter a country is essential for managing future outbreaks. This study investigates how three major variants-Alpha, Delta, and Omicron (BA.1)-entered Spain and how different restrictions potentially affected their introduction. METHODS:We collected Spanish and international SARS-CoV-2 genomes from the GISAID database. Leveraging connectivity data from different countries with Spain, we performed a phylodynamic Bayesian analysis of the SARS-CoV-2 introductions into Spain. RESULTS:Most introductions of the Alpha variant originated from France. As travel restrictions eased, the number of introductions from different countries increased. During the Delta and Omicron waves, the United Kingdom and Germany became important sources of the virus. The highest number of introductions occurred during the Delta wave, coinciding with fewer travel restrictions and the summer season, when Spain receives a considerable number of tourists. CONCLUSIONS:Our findings highlight the role of international travel in the spread of new variants. They underscore the importance of monitoring travel patterns and implementing targeted public health measures to manage the spread of SARS-CoV-2.
The dynamics of SARS-CoV-2 transmission are influenced by a variety of factors, including social restrictions and the emergence of distinct variants. In this study, we delve into the origins and dissemination of the Alpha, Delta, and Omicron variants of concern in Galicia, northwest Spain. For this, we leveraged genomic data collected by the EPICOVIGAL Consortium and from the GISAID database, along with mobility information from other Spanish regions and foreign countries. Our analysis indicates that initial introductions during the Alpha phase were predominantly from other Spanish regions and France. However, as the pandemic progressed, introductions from Portugal and the USA became increasingly significant. Notably, Galicia's major coastal cities emerged as critical hubs for viral transmission, highlighting their role in sustaining and spreading the virus. This research emphasizes the critical role of regional connectivity in the spread of SARS-CoV-2 and offers essential insights for enhancing public health strategies and surveillance measures.
Wastewater surveillance for SARS-CoV-2 provides early warnings of emerging variants of concerns and can be used to screen for novel cryptic linked-read mutations, which are co-occurring single nucleotide mutations that are rare, or entirely missing, in existing SARS-CoV-2 databases. While previous approaches have focused on specific regions of the SARS-CoV-2 genome, there is a need for computational tools capable of efficiently tracking cryptic mutations across the entire genome and investigating their potential origin. We present Crykey, a tool for rapidly identifying rare linked-read mutations across the genome of SARS-CoV-2. We evaluated the utility of Crykey on over 3,000 wastewater and over 22,000 clinical samples; our findings are three-fold: i) we identify hundreds of cryptic mutations that cover the entire SARS-CoV-2 genome, ii) we track the presence of these cryptic mutations across multiple wastewater treatment plants and over three years of sampling in Houston, and iii) we find a handful of cryptic mutations in wastewater mirror cryptic mutations in clinical samples and investigate their potential to represent real cryptic lineages. In summary, Crykey enables large-scale detection of cryptic mutations in wastewater that represent potential circulating cryptic lineages, serving as a new computational tool for wastewater surveillance of SARS-CoV-2. Wastewater surveillance has the potential to be used for early detection of new SARS-CoV-2 lineages. Here, the authors present Crykey, a computational method for detecting cryptic SARS-CoV-2 mutations in wastewater that co-occur on the same sequencing read, potentially representing new lineages.
Different factors influence the spread of SARS-CoV-2, from the inherent transmission capabilities of the different variants to the control measurements put in place. Here we studied the introduction of the Alpha, Delta, and Omicron-BA.1 variants of concern (VOCs) into Spain. For this, we collected genomic data from the GISAID database and combined it with connectivity data from different countries with Spain to perform a phylodynamic Bayesian analysis of the introductions. Our findings reveal that the introductions of these VOCs predominantly originated from France, especially in the case of Alpha. As travel restrictions were eased during the Delta and Omicron-BA.1 waves, the number of introductions from distinct countries increased, with the United Kingdom and Germany becoming significant sources of the virus. The largest number of introductions detected corresponded to the Delta wave, which was associated with fewer restrictions and the summer period, when Spain receives a considerable number of tourists. This research underscores the importance of monitoring international travel patterns and implementing targeted public health measures to manage the spread of SARS-CoV-2.
Transmissible cancers are malignant cell lineages that spread clonally between individuals. Several such cancers, termed bivalve transmissible neoplasia (BTN), induce leukemia-like disease in marine bivalves. This is the case of BTN lineages affecting the common cockle, Cerastoderma edule, which inhabits the Atlantic coasts of Europe and northwest Africa. To investigate the evolution of cockle BTN, we collected 6,854 cockles, diagnosed 390 BTN tumors, generated a reference genome and assessed genomic variation across 61 tumors. Our analyses confirmed the existence of two BTN lineages with hemocytic origins. Mitochondrial variation revealed mitochondrial capture and host co-infection events. Mutational analyses identified lineage-specific signatures, one of which likely reflects DNA alkylation. Cytogenetic and copy number analyses uncovered pervasive genomic instability, with whole-genome duplication, oncogene amplification and alkylation-repair suppression as likely drivers. Satellite DNA distributions suggested ancient clonal origins. Our study illuminates long-term cancer evolution under the sea and reveals tolerance of extreme instability in neoplastic genomes.
Wastewater-based epidemiology has been widely used as a cost-effective method for tracking the COVID-19 pandemic at the community level. Here we describe COVIDBENS, a wastewater surveillance program running from June 2020 to March 2022 in the wastewater treatment plant of Bens in A Coruña (Spain). The main goal of this work was to provide an effective early warning tool based in wastewater epidemiology to help in decision-making at both the social and public health levels. RT-qPCR procedures and Illumina sequencing were used to weekly monitor the viral load and to detect SARS-CoV-2 mutations in wastewater, respectively. In addition, own statistical models were applied to estimate the real number of infected people and the frequency of each emerging variant circulating in the community, which considerable improved the surveillance strategy. Our analysis detected 6 viral load waves in A Coruña with concentrations between 103 and 106 SARS-CoV-2 RNA copies/L. Our system was able to anticipate community outbreaks during the pandemic with 8–36 days in advance with respect to clinical reports and, to detect the emergence of new SARS-CoV-2 variants in A Coruña such as Alpha (B.1.1.7), Delta (B.1.617.2), and Omicron (B.1.1.529 and BA.2) in wastewater with 42, 30, and 27 days, respectively, before the health system did. Data generated here helped local authorities and health managers to give a faster and more efficient response to the pandemic situation, and also allowed important industrial companies to adapt their production to each situation. The wastewater-based epidemiology program developed in our metropolitan area of A Coruña (Spain) during the SARS-CoV-2 pandemic served as a powerful early warning system combining statistical models with mutations and viral load monitoring in wastewater over time.
ABSTRACT Genomic surveillance and epidemiology have shed light on the viral diversity driving coronavirus disease 2019 (COVID-19) outbreaks and are important during waves of highly transmissible and immune-escaping variants of interest or of concern (VOCs). We analyzed the epidemiological data of the understudied country of Malta and related the patterns observed with viral genetic sequences obtained through the surveillance system headed by the Mater Dei Hospital and the University of Malta. We reconstructed the evolutionary history and spatiotemporal dynamics of Maltese severe acute respiratory syndrome coronavirus 2 viruses using a phylodynamics framework. Our findings suggest that the number of cases associated with B.1.1.7/Alpha, B.1.617.2.X/Delta, and B.1.1.529.X/Omicron VOCs was nine times higher than those associated with wild-type variants. The positivity rates in Malta remained low to moderate (<10%). A combination of public health interventions appeared to have allowed Malta to mitigate the impact of COVID-19. Our phylodynamic reconstruction traced most of the 173 viral introductions inferred to countries in Northern Europe, which is consistent with flight connectivity patterns. We also observed prolonged periods of cryptic transmission (median = 102 days) until expansion into larger outbreaks. These larger outbreaks were more easily detected by the intermittent genomic surveillance in Malta, characterized by periods of sequencing hiatus. Our study demonstrates that integrating epidemiological and genomic data are crucial for uncovering the COVID-19 dynamics of understudied locations, particularly when genomic surveillance is suboptimal. Accordingly, strengthening the genomic surveillance system in Malta should help in the earlier detection of introductions and minimize viral expansion in the country while informing public health interventions. IMPORTANCE Our study provides insights into the evolution of the coronavirus disease 2019 (COVID-19) pandemic in Malta, a highly connected and understudied country. We combined epidemiological and phylodynamic analyses to analyze trends in the number of new cases, deaths, tests, positivity rates, and evolutionary and dispersal patterns from August 2020 to January 2022. Our reconstructions inferred 173 independent severe acute respiratory syndrome coronavirus 2 introductions into Malta from various global regions. Our study demonstrates that characterizing epidemiological trends coupled with phylodynamic modeling can inform the implementation of public health interventions to help control COVID-19 transmission in the community.
In recent years, many algorithmic strategies have been developed to exploit single-cell mutational profiles generated via sequencing experiments of cancer samples and return reliable models of cancer evolution. Here, we introduce the COB-tree algorithm, which summarizes the solutions explored by state-of-the-art methods for clonal tree inference, to return a unique consensus optimum branching tree. The method proves to be highly effective in detecting pairwise temporal relations between genomic events, as demonstrated by extensive tests on simulated datasets. We also provide a new method to visualize and quantitatively inspect the solution space of the inference methods, via Principal Coordinate Analysis. Finally, the application of our method to a single-cell dataset of patient-derived melanoma xenografts shows significant differences between the COB-tree solution and the maximum likelihood ones.
Certain sacoglossan sea slugs can sequester and maintain photosynthetically active chloroplasts through algae feeding, a phenomenon called kleptoplasty. The period while these plastids remain active inside the slug’s body is species- and environment-dependent and can span from a few days to more than three weeks. Here we report for the first time the transcriptome of sea slug Elysia viridis (Montagu, 1804), which can maintain kleptoplasts for more than two weeks and is distributed along all the Atlantic European coastline. The obtained transcriptome of E. viridis comprised 12,884 protein-coding sequences (CDS). The shortest one was 261bp, and the longest 8,766bp; the whole transcriptome has a total length of 9.3Mb (Table S4 and Fig. S2). Analysing these CDS, we identified 9,422 different proteins, with best hits mainly from two genera: Elysia (87.2%), and Plakobranchus (11.0%) (Fig. S2); the other 2.3% corresponded to multiple genera of sea slugs and snails (Tectipleura) (Kano et al., 2016). We got the functional annotation (Gene Ontologies, GO) corresponding to 9,333 CDS: 4,755 CDS associated with 2,583 Biological Process (BP); 5,466 CDS linked to 683 Cellular Components (CC); and 6,693 related to 1,606 Molecular Functions (MF). We identified 201 CDS related to response to stress (GO:0006950) and 10 CDS associated with the regulation of response to stress (GO:0080134). Focussing on the ROS-quenching toolkit, we found 24 CDS related to oxidoreductase complex (GO:1990204) and 560 annotated with oxidoreductase activity (GO:0016491) acting in a large number of donors, e.g., CH-OH, CH=O, C=O, CH and CH2. In addition, we found 39 CDS with antioxidant activity (GO:00162099) and other CDS with ROS-quenching function: superoxide dismutase (GO:0004784), peroxidase (GO:0004601), glutathione oxidoreductase (GO:0097573) and peroxidase (GO:0004602); and thioredoxin peroxidase (GO:0008379) activity. Furthermore, we found 8 CDS related to the symbiont response (GO:0140546) and nine related to the pattern recognition receptor signalling pathway (GO:0002221).
Cell lineages accumulate somatic mutations during organismal development, potentially leading to pathological states. The rate of somatic evolution within a cell population can vary due to multiple factors, including selection, a change in the mutation rate, or differences in the microenvironment. Here, we developed a statistical test called the Poisson Tree (PT) test to detect varying evolutionary rates among cell lineages, leveraging the phylogenetic signal of single-cell DNA sequencing (scDNA-seq) data. We applied the PT test to 24 healthy and cancer samples, rejecting a constant evolutionary rate in 11 out of 15 cancer and five out of nine healthy scDNA-seq datasets. In six cancer datasets, we identified subclonal mutations in known driver genes that could explain the rate accelerations of particular cancer lineages. Our findings demonstrate the efficacy of scDNA-seq for studying somatic evolution and suggest that cell lineages often evolve at different rates within cancer and healthy tissues.
A detailed understanding of how and when severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) transmission occurs is crucial for designing effective prevention measures. Other than contact tracing, genome sequencing provides information to help infer who infected whom. However, the effectiveness of the genomic approach in this context depends on both (high enough) mutation and (low enough) transmission rates. Today, the level of resolution that we can obtain when describing SARS-CoV-2 outbreaks using just genomic information alone remains unclear. In order to answer this question, we sequenced forty-nine SARS-CoV-2 patient samples from ten local clusters in NW Spain for which partial epidemiological information was available and inferred transmission history using genomic variants. Importantly, we obtained high-quality genomic data, sequencing each sample twice and using unique barcodes to exclude cross-sample contamination. Phylogenetic and cluster analyses showed that consensus genomes were generally sufficient to discriminate among independent transmission clusters. However, levels of intrahost variation were low, which prevented in most cases the unambiguous identification of direct transmission events. After filtering out recurrent variants across clusters, the genomic data were generally compatible with the epidemiological information but did not support specific transmission events over possible alternatives. We estimated the effective transmission bottleneck size to be one to two viral particles for sample pairs whose donor-recipient relationship was likely. Our analyses suggest that intrahost genomic variation in SARS-CoV-2 might be generally limited and that homoplasy and recurrent errors complicate identifying shared intrahost variants. Reliable reconstruction of direct SARS-CoV-2 transmission based solely on genomic data seems hindered by a slow mutation rate, potential convergent events, and technical artifacts. Detailed contact tracing seems essential in most cases to study SARS-CoV-2 transmission at high resolution.
Transmissible cancers are malignant cell clones that spread among individuals through transfer of living cancer cells. Several such cancers, collectively known as bivalve transmissible neoplasia (BTN), are known to infect and cause leukaemia in marine bivalve molluscs. This is the case of BTN clones affecting the common cockle, Cerastoderma edule , which inhabits the Atlantic coasts of Europe and north-west Africa. To investigate the origin and evolution of contagious cancers in common cockles, we collected 6,854 C. edule specimens and diagnosed 390 cases of BTN. We then generated a reference genome for the species and assessed genomic variation in the genomes of 61 BTN tumours. Analysis of tumour-specific variants confirmed the existence of two cockle BTN lineages with independent clonal origins, and gene expression patterns supported their status as haemocyte-derived marine leukaemias. Examination of mitochondrial DNA sequences revealed several mitochondrial capture events in BTN, as well as co-infection of cockles by different tumour lineages. Mutational analyses identified two lineage-specific mutational signatures, one of which resembles a signature associated with DNA alkylation. Karyotypic and copy number analyses uncovered genomes marked by pervasive instability and polyploidy. Whole-genome duplication, amplification of oncogenes CCND3 and MDM2 , and deletion of the DNA alkylation repair gene MGMT , are likely drivers of BTN evolution. Characterization of satellite DNA identified elements with vast expansions in the cockle germ line, yet absent from BTN tumours, suggesting ancient clonal origins. Our study illuminates the evolution of contagious cancers under the sea, and reveals long-term tolerance of extreme instability in neoplastic genomes.