BACKGROUND:The risk of developing advanced neoplasia (AN; colorectal cancer and/or high-grade dysplasia) in ulcerative colitis (UC) patients with a low-grade dysplasia (LGD) lesion is variable and difficult to predict. This is a major challenge for effective clinical management. OBJECTIVE:We aimed to provide accurate AN risk stratification in UC patients with LGD. We hypothesised that the pattern and burden of somatic genomic copy number alterations (CNAs) in LGD lesions could predict future AN risk. DESIGN:We performed a retrospective multicentre validated case-control study using n=270 LGD samples from n=122 patients with UC. Patients were designated progressors (n=40) if they had a diagnosis of AN in the ~5 years following LGD diagnosis or non-progressors (n=82) if they remained AN-free during follow-up. DNA was extracted from the baseline LGD lesion, low-coverage whole genome sequencing performed and data processed to detect CNAs. Survival analysis was used to evaluate CNAs as predictors of future AN risk. RESULTS:CNA burden was significantly higher in progressors than non-progressors (p=2×10-6 in discovery cohort) and was a very significant predictor of AN risk in univariate analysis (OR=36; p=9×10-7), outperforming existing clinical risk factors such as lesion size, shape and focality. Optimal risk prediction was achieved with a multivariate model combining CNA burden with the known clinical risk factor of incomplete LGD resection. Within-LGD lesion genetic heterogeneity did not confound risk prediction. CONCLUSION:Measurement of CNAs in LGD is an accurate predictor of AN risk in inflammatory bowel disease and is likely to support clinical management.
Deficiency in the mismatch repair system (MMRd) causes microsatellite instability (MSI) in cancers and determines eligibility for immunotherapy. Here, we show that MMRd tumours harbour long-deletion signatures (≥2-5+ base pairs deleted in repetitive regions), which provide new insights into MSI evolution and enable sensitive MSI detection particularly in challenging clinical samples. Long deletions, accumulated through stepwise DNA slippage errors, are significantly more prevalent in metastatic MMRd tumours compared to primary tumours. Importantly, we show that long-deletion signatures harbour features that are distinct from background noise, making them robustly detectable even in shallow whole genome sequencing (sWGS, ~0.1X coverage) of formalin-fixed samples. We constructed a machine learning classifier that uses these distinct features to detect Microsatellite Instability in LOw-quality (MILO) samples. MILO achieved 100% accuracy in detecting MSI in sWGS data with only 2%-15% tumour purity and demonstrated promise in identifying MMRd clones in precancerous intestinal lesions. We propose that MILO could be clinically used for the sensitive monitoring of MMRd cancer evolution from early to late stages, using minimal sequencing data from both archival and fresh-frozen samples with low tumour content. ### Competing Interest Statement TG is named as a co-inventor on patent applications that describe a method for TCR sequencing (GB2305655.9), and a method to measure evolutionary dynamics in cancers using DNA methylation (GB2317139.0). TG has received honorarium from Genentech and DAiNA therapeutics. AB is also co-inventor on the TCR patent. All other authors declare no conflicts of interest.
Abstract PISCA, originally introduced by Martinez et al. in 2018, represents a pivotal Bayesian phylogenetics tool for the modeling of tumor evolution using multi-region somatic chromosomal alteration (SCA) data. PISCA takes allele-specific copy number data, typically obtained from deep genome sequencing or SNP arrays, or absolute copy number data from low-pass genome sequencing methodologies. It extends the classic BEAST1 framework and inherits a rich repertoire of evolutionary models. Importantly, PISCA leverages longitudinal sampling to estimate SCA mutational clock rates, either employing strict clock models, where mutation rates remain constant throughout the evolutionary tree, or relaxed clocks, which allow each branch or subtree to possess its distinct mutation rate. This nuanced approach empowers PISCA to account for the heterogeneous rates of mutations, a pivotal consideration in understanding tumor evolution dynamics. However, PISCA has historically posed a formidable entry barrier for wider community adoption due to platform and Java dependencies for installation, as well as the need for manual or custom-scripted XML file generation. Here we showcase PISCA-box, a user-friendly interface designed to streamline the generation and testing of XML files. Our PISCA-box Docker image works on any desktop machine to create a locally hosted webpage, where users can input SCA data, sampling dates, select clock and demographic models, and set priors for key parameters - similar to the BEAST XML generator, BEAUTi. The Docker or Singularity installation can then be used for longer analyses using high-performance computing resources. To demonstrate its practical utility, we present novel colorectal cancer data. First, we examine data from a patient with a long-standing history of inflammatory bowel disorder (IBD), a known high-risk factor for colorectal cancer. This patient had undergone surveillance colonoscopies for many years, which provided 38 samples, and subsequently an additional 118 samples were collected from a total colectomy. All were analyzed with low-coverage whole genome sequencing. We find some samples from surveillance and colectomy form lineages with overlapping copy number events, that we used to construct a phylogeny, finding a cancer-adjacent clade that appears to evolve rapidly. A second dataset of multi-region samples from a cohort of exceptional survivors of oligometastatic colorectal cancer who lived >60 months from metastatic diagnosis with biopsies/resections across 3-10 time points is examined and temporal models are used to estimate the ages of metastatic clades. PISCA therefore harnesses the power of longitudinal SCA data to enable a comprehensive study of SCA dynamics. We are currently expanding this framework to include fluctuating methylation clocks which PISCA-box will soon include. By providing this accessible tool, we enable researchers to readily apply Bayesian phylogenetics to real-world clinical datasets to better understand tumor evolution. Citation Format: Heather E. Grant, Rachel Alcraft, Pablo Bousquets-Muñoz, Calum Gabbutt, Alison Berner, Mehmet Yalchin, Carlo C. Maley, Trevor A. Graham, Diego Mallo. PISCA-box: A user-friendly interface for Phylogenetic Inference using Somatic Chromosomal Alterations (PISCA) [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Translating Cancer Evolution and Data Science: The Next Frontier; 2023 Dec 3-6; Boston, Massachusetts. Philadelphia (PA): AACR; Cancer Res 2024;84(3 Suppl_2):Abstract nr B004.
Patients with inflammatory bowel disease (IBD) are at increased risk of colorectal cancer (CRC), and this risk increases dramatically in those who develop low-grade dysplasia (LGD). However, there is currently no accurate way to risk-stratify patients with LGD, leading to both over- and under-treatment of cancer risk. Here we show that the burden of somatic copy number alterations (CNAs) within resected LGD lesions strongly predicts future cancer development. We performed a retrospective multi-centre validated case-control study of n=122 patients (40 progressors, 82 non-progressors, 270 LGD regions). Low coverage whole genome sequencing revealed CNA burden was significantly higher in progressors than non-progressors (p=2×10-6 in discovery cohort) and was a very significant predictor of CRC risk in univariate analysis (odds ratio = 36; p=9×10-7), outperforming existing clinical risk factors such as lesion size, shape and focality. Optimal risk prediction was achieved with a multivariate model combining CNA burden with the known clinical risk factor of incomplete LGD resection. The measurement of CNAs in LGD lesions is a robust, low-cost and rapidly translatable predictor of CRC risk in IBD that can be used to direct management and so prevent CRC in high-risk individuals whilst sparing those at low-risk from unnecessary intervention.
Abstract Introduction: Cancer is a dynamic evolutionary process; thus, the evolutionary history of a cancer should strongly influence its future trajectory and so the clinical outcome of patients. However, accurately assessing an individual cancer’s evolutionary history in a scalable and clinically relevant manner is currently an open problem. To address this, we introduce a novel methodology (EVOFLUx) to quantitively infer the history of a cancer from single time-point bulk samples using inexpensive, standard methylation arrays. Description of Methodology: We recently identified fluctuating CpGs (fCpGs), which are CpGs that neutrally and stochastically change methylation state over time in individual cells, uniquely barcoding cells and thus enabling high temporal-resolution lineage tracing approaches. In a bulk population of cancer cells, the distribution of fCpG methylation is determined by the evolution of the population and thus encode a cancer’s past clonal dynamics. Here, we have developed a general computational approach to model how the patterns of fCpGs vary according to the evolutionary history of a growing cancer, and a Bayesian inference method to infer these parameters from data. Results We applied EVOFLUx to quantify the evolutionary dynamics of 1,976 lymphoid cancers, finding that the tumor growth rates, malignancy age and epimutation rates varied by orders of magnitude across disease types. In addition, EVOFLUx allowed us to test for the presence of ongoing subclonal selection and validated our inference with matched whole exome sequencing data. Alongside characterizing the evolutionary dynamics within single bulk samples, fCpGs also allow for the inference of the phylogenetic relationships between longitudinal samples, which we validated with whole genome sequencing. Within cancer types, more aggressive subtypes typically had elevated growth rates. In our clinically annotated cohort of chronic lymphocytic leukemia (CLL) samples, we found that patient-specific evolutionary dynamics are strongly associated with outcome. Conclusions: Hence, we present a powerful new method that uses widely available, low-cost bulk DNA methylation data to precisely measures cancer evolutionary dynamics in patient samples with clinical implications. Citation Format: Calum Gabbutt, Martí Duran-Ferrer, Heather Grant, Diego Mallo, Ferran Nadeu, Jacob Househam, Neus Villamor, Olga Krali, Jessica Nordlund, Thorsten Zenz, Elias Campo, Armando Lopez-Guillermo, Jude Fitzgibbon, Chris P. Barnes, Darryl Shibata, José I. Martin-Subero, Trevor A. Graham. Large-scale, low-cost, and accurate measurement of cancer evolutionary dynamics from clinical patient samples [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 4300.
Abstract The development of cancer from healthy somatic cells is fundamentally an evolutionary process. Understanding this process is vital to improving its clinical management and developing new treatment strategies. Tracing individual human cells’ fates would simultaneously advance somatic evolutionary theory and our understanding of normal tissue function and its change through development, aging, and cancer progression. We cannot use human cell-labeling techniques in vivo. Still, population genetic and phylogenetic models can leverage the information recorded in heritable alterations to reconstruct the unobserved somatic cells’ history. Most such methods are limited by the few parameters they can estimate and restricted to neoplastic samples due to the lack of adequate evolutionary signal in healthy somatic cells. Gabbutt et al. recently ameliorated this limitation by introducing fluctuating methylation clocks (FMCs)—tissue-specific CpG sites that randomly change their methylation state—and a model that uses them to reconstruct the evolution of a clonal stem-cell population. In this application, all FMCs in a bulk sample are interrogated using a single standard methylation array to yield estimates of rates of methylation, demethylation and cell replacement, and the size of the stem-cell niche. This technique is especially suited to studying stem-cell niche dynamics in healthy or pre-malignant tissues organized in clonal cell populations where the whole population can be sampled together, like crypts (e.g., intestine, Barrett’s esophagus) and glands (e.g., endometrium). We have developed a multi-sample extension of this model that simultaneously reconstructs the evolutionary dynamics of stem cells within a niche and the niches within the tissue. This phylodynamic model produces additional estimates like ancestral relationships between sampled clonal populations, their calendar divergence times, the number of effective evolutionary niches, and how some of these parameters change over time. Here, we will present the method, demonstrate its validity, characterize its accuracy under a broad simulation parameter space, and showcase its application in analyzing human data from the following tissues: colon, small intestine, endometrium, and Barrett’s esophagus. We will also discuss the current model's limitations and future improvements that will discern between mechanisms of stem-cell-niche reproduction and incorporate geographical information. The stochastic nature of neoplastic progression makes it unlikely for any particular somatic alteration to be highly predictive of its outcome. This can explain why developing reliable traditional biomarkers has been so difficult. Alternatively, evolutionary biomarkers measure the characteristics of the process itself and thus should apply to all neoplasms. Parametric phylogenetic reconstruction models like the one introduced here will enable us to develop such universal biomarkers. Our ultimate goal is for evolutionary biomarkers to join evolutionary therapies in the revolution of the clinical management of cancer. Citation Format: Diego Mallo, Pablo Bousquets-Muñoz, Heather E. Grant, Darryl Shibata, Trevor A. Graham, Carlo C. Maley, Calum Gabbutt. Tick tock trees: Reconstructing the evolutionary dynamics of human tissues using fluctuating methylation clocks [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Translating Cancer Evolution and Data Science: The Next Frontier; 2023 Dec 3-6; Boston, Massachusetts. Philadelphia (PA): AACR; Cancer Res 2024;84(3 Suppl_2):Abstract nr PR009.
Parental care is considered crucial for the enhanced survival of offspring and evolutionary success of many metazoan groups. Most bryozoans incubate their young in brood chambers or intracoelomically. Based on the drastic morphological differences in incubation chambers across members of the order Cheilostomatida (class Gymnolaemata), multiple origins of incubation were predicted in this group. This hypothesis was tested by constructing a molecular phylogeny based on mitogenome data and nuclear rRNA genes 18S and 28S with the most complete sampling of taxa with various incubation devices to date. Ancestral character estimation suggested that distinct types of brood chambers evolved at least 10 times in Cheilostomatida. In Eucratea loricata and Aetea spp. brooding evolved unambiguously from a zygote-spawning ancestral state, as it probably did in Tendra zostericola, Neocheilostomata, and 'Carbasea' indivisa. In two further instances, brooders with different incubation chamber types, skeletal and non-skeletal, formed clades (Scruparia spp., Leiosalpinx australis) and (Catenicula corbulifera (Steginoporella spp. (Labioporella spp., Thalamoporella californica))), each also probably evolved from a zygote-spawning ancestral state. The modular nature of bryozoans probably contributed to the evolution of such a diverse array of embryonic incubation chambers, which included complex constructions made of polymorphic heterozooids, and maternal zooidal invaginations and outgrowths.
PHYLOGENIES All_genes_alignment.nex The concatenated mixed alignment consisting of, 13 mitochondrial protein-coding genes as amino acids, mitochondrial ribosomal RNA genes 12S+16S, and nuclear 18S+28S rRNA genes. Gene boundaries and excludes sites are indicated. Fig_2.nex Topology of the Bayesian phylogenetic analysis of the mixed concatenated alignment consisting of three partitions: (i) 13 mitochondrial protein-coding genes as amino acids, (ii) mitochondrial ribosomal RNA genes 12S+16S, (iii) nuclear 18S+28S rRNA genes. The analysis was performed in MrBayes5D v. 3.2.6 under the GTR+G model of nucleotide evolution (nucleotides) and the MTZOA+G model (amino acids). The analysis was run for 2.4 million generations; 1.5 million generations were discarded as burn-in. Fig_S3 Topology of the Bayesian phylogenetic analysis of the mixed concatenated alignment consisting of three partitions: (i) 13 mitochondrial protein-coding genes (PCGs) as amino acids, (ii) mitochondrial ribosomal RNA genes 12S+16S, (iii) nuclear 18S+28S rRNA genes. The analysis was performed in p4 under the GTR+G model of nucleotide evolution (nucleotides) and the MTZOA+G+F model (amino acids). The +F model component accommodates empirical composition in the amino acid model. The analysis used three separate runs for 300,000 generations; 200,000 generations were discarded as burn-in. Fig_S4 Topology of the maximum likelihood phylogenetic analysis of the mixed concatenated alignment consisting of three partitions: (i) 13 mitochondrial protein-coding genes as amino acids, (ii) mitochondrial ribosomal RNA genes 12S+16S, (iii) nuclear 18S+28S rRNA genes. The analysis was performed in RAxML HPC-PTHREADS-SSE3 v. 8.2.12 under the GTR+G (nucleotides) and the MTZOA+G+F models (amino acids). Fig_S5 Topology of the Bayesian phylogenetic analysis of the 12S+16S rRNA gene partition constructed using MrBayes v. 3.2.6 under the GTR + G model. The analysis was run for 20 million generations; 10 million generations were discarded as burn-in. Fig_S6 Topology of the maximum likelihood phylogenetic analysis of the 12S+16S rRNA gene partition constructed using RAxML HPC-PTHREADS-SSE3 v. 8.2.12 under the GTRCAT model. Fig_S7 Topology of the Bayesian phylogenetic analysis of the 18S+28S rRNA gene partition constructed using MrBayes v. 3.2.6 under the GTR + G model. The analysis was run for 20 million generations; 10 million generations were discarded as burn-in. Fig_S8 Topology of the maximum likelihood phylogenetic analysis of the 18S+28S rRNA gene partition constructed using RAxML HPC-PTHREADS-SSE3 v. 8.2.12 under the GTRCAT model. Fig_S9 Topology of the Bayesian phylogenetic analysis of 13 mitochondrial protein-coding genes as amino acids constructed using MrBayes5D v. 3.2.6 under the MTZOA+G model. The analysis was run for 3.7 million generations; 2.5 million generations were discarded as burn-in. Fig_S10 Topology of the maximum likelihood phylogenetic analysis of 13 mitochondrial protein-coding genes as amino acids constructed using RAxML HPC-PTHREADS-SSE3 v. 8.2.12 under the PROTGAMMAMTZOA model. Fig_S11 Topology of the Bayesian phylogenetic analysis of the mixed concatenated alignment consisting of three partitions: (i) 13 mitochondrial protein-coding genes (PCGs) as amino acids, (ii) mitochondrial ribosomal RNA genes 12S+16S, (iii) nuclear 18S+28S rRNA genes. The analysis was performed in p4 under the NDCH-C2 model. The analysis used four separate runs for 300,000 generations; 200,000 generations were discarded as burn-in. The NDCH model accommodates compositional tree-heterogeneity and was used because there was a large amount of compositional heterogeneity over the sequences, especially in the PCGs and 12S+16S rRNA data partitions. This is an NDCH model with two composition vectors on each of the three data partitions. Fig_S12 Topology of the Bayesian phylogenetic analysis of the mixed concatenated alignment consisting of three partitions: (i) 13 mitochondrial protein-coding genes as amino acids, (ii) mitochondrial ribosomal RNA genes 12S+16S, (iii) nuclear 18S+28S rRNA genes. This analysis excluded all terminals for which less than half of mitogenome genes were available, or which only had one of the two nuclear rRNA genes. The analysis was performed in MrBayes5D v. 3.2.6 under the GTR+G model of nucleotide evolution (nucleotides) and the MTZOA+G model (amino acids). The analysis was run for 350,000 generations; 125,000 generations were discarded as burn-in. Fig_S13 Topology of the maximum likelihood phylogenetic analysis of the mixed concatenated alignment consisting of three partitions: (i) 13 mitochondrial protein-coding genes as amino acids, (ii) mitochondrial ribosomal RNA genes 12S+16S, (iii) nuclear 18S+28S rRNA genes. This analysis excluded all terminals for which less than half of mitogenome genes were available, or which only had one of the two nuclear rRNA genes. The analysis was performed in RAxML HPC-PTHREADS-SSE3 v. 8.2.12 under the GTR+G (nucleotides) and the MTZOA+G+F models (amino acids). ANCESTRAL CHARACTER ESTIMATION: ACE.R R script of the ancestral character estimation carried out in phytools. Reproductive_strategy_numbers.csv Data input file for ACE analysis (reproductive strategies coded as numbers) Reproductive_strategies.xlsx List of reproductive strategies per taxon with the corresponding numerical codes used in the file 'Reproductive_stategies_numbers.csv'. Tree.tre Input tree for ACE analysis.
Cancer development, progression, and response to treatment are evolutionary processes, but characterising the evolutionary dynamics at sufficient scale to be clinically-meaningful has remained challenging. Here, we develop a new methodology called EVOFLUx, based upon natural DNA methylation barcodes fluctuating over time, that quantitatively infers evolutionary dynamics using only a bulk tumour methylation profile as input. We apply EVOFLUx to 1,976 well-characterised lymphoid cancer samples spanning a broad spectrum of diseases and show that tumour growth rates, malignancy age and epimutation rates vary by orders of magnitude across disease types. We measure that subclonal selection occurs only infrequently within bulk samples and detect occasional examples of multiple independent primary tumours. Clinically, we observe that tumour growth rates are higher in more aggressive disease subtypes, and in two series of chronic lymphocytic leukaemia patients, evolutionary histories are independent prognostic factors. Phylogenetic analyses of longitudinal CLL samples using EVOFLUx detect the seeds of future Richter transformation many decades prior to presentation. We provide orthogonal verification of EVOFLUx inferences using additional genetic and clinical data. Collectively, we show how widely- available, low-cost bulk DNA methylation data precisely measures cancer evolutionary dynamics, and provides new insights into cancer biology and clinical behaviour.
The Sustainable East Africa Research in Community Health (SEARCH) trial was a universal test-and-treat (UTT) trial in rural Uganda and Kenya, aiming to lower regional HIV-1 incidence. Here, we quantify breakthrough HIV-1 transmissions occurring during the trial from population-based, dried blood spot samples. Between 2013 and 2017, we obtained 549 gag and 488 pol HIV-1 consensus sequences from 745 participants: 469 participants infected prior to trial commencement and 276 SEARCH-incident infections. Putative transmission clusters, with a 1.5% pairwise genetic distance threshold, were inferred from maximum likelihood phylogenies; clusters arising after the start of SEARCH were identified with Bayesian time-calibrated phylogenies. Our phylodynamic approach identified nine clusters arising after the SEARCH start date: eight pairs and one triplet, representing mostly opposite-gender linked (6/9), within-community transmissions (7/9). Two clusters contained individuals with non-nucleoside reverse transcriptase inhibitor (NNRTI) resistance, both linked to intervention communities. The identification of SEARCH-incident, within-community transmissions reveals the role of unsuppressed individuals in sustaining the epidemic in both arms of a UTT trial setting. The presence of transmitted NNRTI resistance, implying treatment failure to the efavirenz-based antiretroviral therapy (ART) used during SEARCH, highlights the need to improve delivery and adherence to up-to-date ART recommendations, to halt HIV-1 transmission.
We present 109 near full-length HIV genomes amplified from blood serum samples obtained during early 1986 from across Uganda, which to our knowledge is the earliest and largest population sample from the initial phase of the HIV epidemic in Africa. Consensus sequences were made from paired-end Illumina reads with a target-capture approach to amplify HIV material following poor success with standard approaches. In comparisons with a smaller 'intermediate' genome dataset from 1998 to 1999 and a 'modern' genome dataset from 2007 to 2016, the proportion of subtype D was significantly higher initially, dropping from 67% (73/109), to 57% (26/46) to 17% (82/465) respectively (p < 0.0001). Subtype D has previously been shown to have a faster rate of disease progression than other subtypes in East African population studies, and to have a higher propensity to use the CXCR4 co-receptor ( "X4 tropism "); associated with a decrease in time to AIDS. Here we find significant differences in predicted tropism between A1 and D subtypes in all three sample periods considered, which is particularly striking the 1986 sample: 66% (53/80) of subtype D env sequences were predicted to be X4 tropic compared with none of the 24 subtype A1. We also analysed the frequency of subtype in the envelope region of inter-subtype recombinants, and found that subtype A1 is over-represented in env, suggesting recombination and selection have acted to remove subtype D env from circulation. The reduction of subtype D frequency over three decades therefore appears to be a result of selective pressure against X4 tropism and its higher virulence. Lastly, we find a subtype D specific codon deletion at position 24 of the V3 loop, which may explain the higher propensity for subtype D to utilise X4 tropism.
A new abundance of full-length HIV-1 genome sequences provides an opportunity to revisit the standard model of HIV-1/M diversity that clusters genomes into largely non-recombinant subtypes, which is not consistent with recent evidence of deep recombinant histories for SIV and other HIV-1 groups. Here we develop an unsupervised non-parametric clustering approach, which does not rely on predefined non-recombinant genomes, by adapting a community detection method developed for dynamic social network analysis. We show that this method (DSBM) attains a significantly lower mean error rate in detecting recombinant breakpoints in simulated data (quasibinomial GLM, P < 8 × 10 −8 ), compared to other reference-free recombination detection programs (GARD, RDP4 and RDP5). Applied to a representative sample of n = 525 actual HIV-1 genomes, we determined k = 25 as the optimal number of DSBM clusters, and used change point detection to estimate that at least 95% of these genomes are recombinant. Further, we identified both known and novel recombination hotspots in the HIV-1 genome, and evidence of inter-subtype recombination in HIV-1 subtype reference genomes. We propose that clusters generated by DSBM can provide an informative new framework for HIV-1 classification.
Recombination is an important feature of HIV evolution, occurring both within and between the major branches of diversity (subtypes). The Ugandan epidemic is primarily composed of two subtypes, A1 and D, that have been co-circulating for 50 years, frequently recombining in dually infected patients. Here, we investigate the frequency of recombinants in this population and the location of breakpoints along the genome. As part of the PANGEA-HIV consortium, 1,472 consensus genome sequences over 5 kb have been obtained from 1,857 samples collected by the MRC/UVRI & LSHTM Research unit in Uganda, 465 (31.6 per cent) of which were near full-length sequences (>8 kb). Using the subtyping tool SCUEAL, we find that of the near full-length dataset, 233 (50.1 per cent) genomes contained only one subtype, 30.8 per cent A1 (n = 143), 17.6 per cent D (n = 82), and 1.7 per cent C (n = 8), while 49.9 per cent (n = 232) contained more than one subtype (including A1/D (n = 164), A1/C (n = 13), C/D (n = 9); A1/C/D (n = 13), and 33 complex types). K-means clustering of the recombinant A1/D genomes revealed a section of envelope (C2gp120-TMgp41) is often inherited intact, whilst a generalized linear model was used to demonstrate significantly fewer breakpoints in the gag-pol and envelope C2-TM regions compared with accessory gene regions. Despite similar recombination patterns in many recombinants, no clearly supported circulating recombinant form (CRF) was found, there was limited evidence of the transmission of breakpoints, and the vast majority (153/164; 93 per cent) of the A1/D recombinants appear to be unique recombinant forms. Thus, recombination is pervasive with clear biases in breakpoint location, but CRFs are not a significant feature, characteristic of a complex, and diverse epidemic.
In order to learn more about the evolution of shell colour in molluscs, 20 characters related to shell colour were recorded for 81 bivalve clades using the entire collection of dry shells at the Natural History Museum in London (> 44000 lots) and plotted onto a published phylogeny. Statistical tests for phylogenetic signal show that coloured shells are not distributed evenly across the class. The phylogenetic distribution of colour is statistically significant, as are the distributions of individual shell and periostracal colours. Blue and green shells are rare, with non-iridescent blue and green coloration occurring on the outer side of the valve more commonly in the periostracum than in the shell matrix and on the inside of valves in species that produce proteinaceous sheets overlying calcareous shell. These findings suggest that these colours in bivalves are attributable to pigments that are more easily incorporated into organic material than into calcareous shell. Ancestral state reconstructions show that the ancestral bivalves likely had a coloured shell. Although similar colours can arise from different pigments or from structural elements, the broad-scale phylogenetic distribution of colour in Bivalvia, along with the co-occurrence of 'sets' of colours, probably reflects, in part, the distribution of classes of pigments. This confirms the idea that the major classes of pigments found in molluscan shells are evolutionarily ancient and continue to contribute to shell coloration, despite recent genomic evidence suggesting that protein moieties associated with pigments might be evolutionarily diverse.