An important feature of the evolution of the SARS-CoV-2 virus has been the emergence of highly mutated novel variants, which are characterised by the gain of multiple mutations relative to viruses circulating in the general global population. Cases of chronic viral infection have been suggested as an explanation for this phenomenon, whereby an extended period of infection, with an increased rate of evolution, creates viruses with substantial genetic novelty. However, measuring a rate of evolution during chronic infection is made more difficult by the potential existence of compartmentalisation in the viral population, whereby the viruses in a host form distinct subpopulations. We here describe and apply a novel statistical method to study within-host virus evolution, identifying the minimum number of subpopulations required to explain sequence data observed from cases of chronic infection, and inferring rates for within-host viral evolution. Across nine cases of chronic SARS-CoV-2 infection in hospitalised patients we find that non-trivial population structure is relatively common, with five cases showing evidence of more than one viral population evolving independently within the host. The detection of non-trivial population structure was more common in severely immunocompromised individuals (p = 0.04, Fisher’s Exact Test). We find cases of within-host evolution proceeding significantly faster, and significantly slower, than that of the global SARS-CoV-2 population, and of cases in which viral subpopulations in the same host have statistically distinguishable rates of evolution. Non-trivial population structure was associated with high rates of within-host evolution that were systematically underestimated by a more standard inference method.
The population structure of the malaria parasite Plasmodium falciparum can reveal underlying demographic and adaptive evolutionary processes. Here, we analyse population structure in 4,376 P. falciparum genomes from 21 countries across Africa. We identified a strongly differentiated cluster of parasites, comprising ∼1.2% of samples analysed, geographically distributed over 13 countries across the continent. Members of this cluster, named AF1, carry a genetic background consisting of a large number of highly differentiated variants, rarely observed outside this cluster, at a multitude of genomic loci distributed across most chromosomes. At these loci, the AF1 haplotypes appear to have common ancestry, irrespective of the sampling location; outside the shared loci, however, AF1 members are genetically similar to their sympatric parasites. AF1 parasites sharing up to 23 genomic co-inherited regions were found in all major regions of Africa, at locations over 7,000 km apart. We coined the term cryptotype to describe a complex common background which is geographically widespread, but concealed by genomic regions of local origin. Most AF1 differentiated variants are functionally related, comprising structural variations and single nucleotide polymorphisms in components of the MSP1 complex and several other genes involved in interactions with red blood cells, including invasion and erythrocyte antigen export. We propose that AF1 parasites have adapted to some as yet unidentified evolutionary niche, by acquiring a complex compendium of interacting variants that rarely circulate separately in Africa. As the cryptotype spread across the continent, it appears to have been maintained mostly intact in spite of recombination events, suggesting a selective advantage. It is possible that other cryptotypes circulate in Africa, and new analysis methods may be needed to identify them.### Competing Interest StatementThe authors have declared no competing interest.
Despite substantial progress in machine learning for scientific discovery in recent years, truly de novo design of small molecules which exhibit a property of interest remains a significant challenge. We introduce LambdaZero, a generative active learning approach to search for synthesizable molecules. Powered by deep reinforcement learning, LambdaZero learns to search over the vast space of molecules to discover candidates with a desired property. We apply LambdaZero with molecular docking to design novel small molecules that inhibit the enzyme soluble Epoxide Hydrolase 2 (sEH), while enforcing constraints on synthesizability and drug-likeliness. LambdaZero provides an exponential speedup in terms of the number of calls to the expensive molecular docking oracle, and LambdaZero de novo designed molecules reach docking scores that would otherwise require the virtual screening of a hundred billion molecules. Importantly, LambdaZero discovers novel scaffolds of synthesizable, drug-like inhibitors for sEH. In in vitro experimental validation, a series of ligands from a generated quinazoline-based scaffold were synthesized, and the lead inhibitor N-(4,6-di(pyrrolidin-1-yl)quinazolin-2-yl)-N-methylbenzamide (UM0152893) displayed sub-micromolar enzyme inhibition of sEH.
ABSTRACT The population structure of the malaria parasite Plasmodium falciparum can reveal underlying demographic and adaptive evolutionary processes. Here, we analyse population structure in 4,376 P. falciparum genomes from 21 countries across Africa. We identified a strongly differentiated cluster of parasites, comprising ∼1.2% of samples analysed, geographically distributed over 13 countries across the continent. Members of this cluster, named AF1, carry a genetic background consisting of a large number of highly differentiated variants, rarely observed outside this cluster, at a multitude of genomic loci distributed across most chromosomes. At these loci, the AF1 haplotypes appear to have common ancestry, irrespective of the sampling location; outside the shared loci, however, AF1 members are genetically similar to their sympatric parasites. AF1 parasites sharing up to 23 genomic co-inherited regions were found in all major regions of Africa, at locations over 7,000 km apart. We coined the term cryptotype to describe a complex common background which is geographically widespread, but concealed by genomic regions of local origin. Most AF1 differentiated variants are functionally related, comprising structural variations and single nucleotide polymorphisms in components of the MSP1 complex and several other genes involved in interactions with red blood cells, including invasion and erythrocyte antigen export. We propose that AF1 parasites have adapted to some as yet unidentified evolutionary niche, by acquiring a complex compendium of interacting variants that rarely circulate separately in Africa. As the cryptotype spread across the continent, it appears to have been maintained mostly intact in spite of recombination events, suggesting a selective advantage. It is possible that other cryptotypes circulate in Africa, and new analysis methods may be needed to identify them.
Intermittent demand patterns are commonly present in business aircraft spare parts supply chains. Because of the infrequent arrivals and large variations in demand, aircraft aftermarket demand is difficult to forecast, which often leads to shortages or overstocking of spare parts. In this paper, we present the development and implementation of an advanced analytics framework at Bombardier Aerospace, which is carried out by the Bombardier inventory planning team and IVADO Labs to improve the aftermarket demand forecasting process. This integrated predictive analytics pipeline leverages machine-learning (ML) models and traditional time series models in a single framework in a systematic fashion. We also make use of a tree-based machine-learning method with a large set of input features to estimate two components of intermittent demand, namely demand sizes and interdemand intervals. Through the ML models, we incorporate different features, including those derived from flight data. Outputs of different forecasting models are combined using an ensemble technique that enhances the robustness and accuracy of the forecasts for different groups of aftermarket spare parts categorized by demand patterns. The validation results show an improvement in forecast accuracy of approximately 7% and in unbiased forecast of 5%. The ML-based Bombardier Aftermarket forecasting system has been successfully deployed and used to forecast the aftermarket demand at Bombardier of more than 1 billion Canadian dollars on a regular basis. History: This paper was refereed.
Escherichia coli is a ubiquitous component of the human gut microbiome, but is also a common pathogen, causing around 40, 000 bloodstream infections (BSI) in the United Kingdom (UK) annually. The number of E. coli BSI has increased over the last decade in the UK, and emerging antimicrobial resistance (AMR) profiles threaten treatment options. Here, we combined clinical, epidemiological, and whole genome sequencing data with high content imaging to characterise over 300 E. coli isolates associated with BSI in a large teaching hospital in the East of England. Overall, only a limited number of sequence types (ST) were responsible for the majority of organisms causing invasive disease. The most abundant (20 % of all isolates) was ST131, of which around 90 % comprised the pandemic O25b:H4 group. ST131-O25b:H4 isolates were frequently multi-drug resistant (MDR), with a high prevalence of extended spectrum β-lactamases (ESBL) and fluoroquinolone resistance. There was no association between AMR phenotypes and the source of E. coli bacteraemia or whether the infection was healthcare-associated. Several clusters of ST131 were genetically similar, potentially suggesting a shared transmission network. However, there was no clear epidemiological associations between these cases, and they included organisms from both healthcare-associated and non-healthcare-associated origins. The majority of ST131 isolates exhibited strong binding with an anti-O25b antibody, raising the possibility of developing rapid diagnostics targeting this pathogen. In summary, our data suggest that a restricted set of MDR E. coli populations can be maintained and spread across both community and healthcare settings in this location, contributing disproportionately to invasive disease and AMR.
Knowledge Graphs (KG) and associated Knowledge Graph Embedding (KGE) models have recently begun to be explored in the context of drug discovery and have the potential to assist in key challenges such as target identification. In the drug discovery domain, KGs can be employed as part of a process which can result in lab-based experiments being performed, or impact on other decisions, incurring significant time and financial costs and most importantly, ultimately influencing patient healthcare. For KGE models to have impact in this domain, a better understanding of not only of performance, but also the various factors which determine it, is required. In this study we investigate, over the course of many thousands of experiments, the predictive performance of five KGE models on two public drug discovery-oriented KGs. Our goal is not to focus on the best overall model or configuration, instead we take a deeper look at how performance can be affected by changes in the training setup, choice of hyperparameters, model parameter initialisation seed and different splits of the datasets. Our results highlight that these factors have significant impact on performance and can even affect the ranking of models. Indeed these factors should be reported along with model architectures to ensure complete reproducibility and fair comparisons of future work, and we argue this is critical for the acceptance of use, and impact of KGEs in a biomedical setting.
Conventional representation learning algorithms for knowledge graphs (KG) map each entity to a unique embedding vector. Such a shallow lookup results in a linear growth of memory consumption for storing the embedding matrix and incurs high computational costs when working with real-world KGs. Drawing parallels with subword tokenization commonly used in NLP, we explore the landscape of more parameter-efficient node embedding strategies with possibly sublinear memory requirements. To this end, we propose NodePiece, an anchor-based approach to learn a fixed-size entity vocabulary. In NodePiece, a vocabulary of subword/sub-entity units is constructed from anchor nodes in a graph with known relation types. Given such a fixed-size vocabulary, it is possible to bootstrap an encoding and embedding for any entity, including those unseen during training. Experiments show that NodePiece performs competitively in node classification, link prediction, and relation prediction tasks while retaining less than 10% of explicit nodes in a graph as anchors and often having 10x fewer parameters. To this end, we show that a NodePiece-enabled model outperforms existing shallow models on a large OGB WikiKG 2 graph having 70x fewer parameters.
Breakthrough infections with SARS-CoV-2 Delta variant have been reported in doubly-vaccinated recipients and as re-infections. Studies of viral spread within hospital settings have highlighted the potential for transmission between doubly-vaccinated patients and health care workers and have highlighted the benefits of high-grade respiratory protection for health care workers. However the extent to which vaccination is preventative of viral spread in health care settings is less well studied. Here, we analysed data from 118 vaccinated health care workers (HCW) across two hospitals in India, constructing two probable transmission networks involving six HCWs in Hospital A and eight HCWs in Hospital B from epidemiological and virus genome sequence data, using a suite of computational approaches. A maximum likelihood reconstruction of transmission involving known cases of infection suggests a high probability that doubly vaccinated HCWs transmitted SARS-CoV-2 between each other and highlights potential cases of virus transmission between individuals who had received two doses of vaccine. Our findings show firstly that vaccination may reduce rates of transmission, supporting the need for ongoing infection control measures even in highly vaccinated populations, and secondly we have described a novel approach to identifying transmissions that is scalable and rapid, without the need for an infection control infrastructure.
A bstract RNA secondary structure prediction is a fundamental task in computational and molecular biology. While machine learning approaches in this area have been shown to improve upon traditional RNA folding algorithms, performance remains limited for several reasons such as the small number of experimentally determined RNA structures and suboptimal use of evolutionary information. To address these challenges, we introduce a practical and effective pretraining strategy that enables learning from a larger set of RNA sequences with computationally predicted structures and in the meantime, tapping into the rich evolutionary information available in databases such as Rfam. Coupled with a flexible and scalable neural architecture that can navigate different learning scenarios while providing ease of integrating evolutionary information, our approach significantly improves upon state-of-the-art across a range of benchmarks, including both single sequence and alignment based structure prediction tasks, with particularly notable benefits on new, less well-studied RNA families. Our source code, data and packaged RNA secondary structure prediction software RSSMFold can be accessed at https://github.com/HarveyYan/RSSMFold .
A review was undertaken of all genomic epidemiology studies on COVID-19 in long term care facilities (LTCF) that have been published to date. It was found that staff and residents were usually infected with identical, or near identical, SARS-CoV-2 genomes. Outbreaks usually involved one predominant lineage, and the same lineages persisted in LTCFs despite infection control measures. Outbreaks were most commonly due to single or few introductions followed by spread rather than a series of seeding events from the community into LTCFs. Sequencing of samples taken consecutively from the same cases showed persistence of the same genome sequence indicating that the sequencing technique was robust over time. When combined with local epidemiology, genomics facilitated likely transmission sources to be better characterised. Transmission between LTCFs was detected in multiple studies. The mortality rate amongst residents was high in all cases, regardless of the lineage. Bioinformatics methods were inadequate in one third of the studies reviewed, and reproducing the analyses was difficult as sequencing data were not available in many cases.
Malaria is a global public health priority causing over 600,000 deaths annually, mostly young children living in Sub-Saharan Africa. Molecular surveillance can provide key information for malaria control, such as the prevalence and distribution of antimalarial drug resistance. However, genome sequencing capacity in endemic countries can be limited. Here, we have implemented an end-to-end workflow for P. falciparum genomic surveillance in Ghana using Oxford Nanopore Technologies, targeting antimalarial resistance markers and the leading vaccine antigen circumsporozoite protein ( csp ). The workflow was rapid, robust, accurate, affordable and straightforward to implement, and could be deployed using readily collected dried blood spot samples. We found that P. falciparum parasites in Ghana had become largely susceptible to chloroquine, with persistent sulfadoxine-pyrimethamine (SP) resistance, and no evidence of artemisinin resistance. Multiple Single Nucleotide Polymorphism (SNP) differences from the vaccine csp sequence were identified, though their significance is uncertain. This study demonstrates the potential utility and feasibility of malaria genomic surveillance in endemic settings using Nanopore sequencing.
Adversarial attacks expose important vulnerabilities of deep learning models, yet little attention has been paid to settings where data arrives as a stream. In this paper, we formalize the online adversarial attack problem, emphasizing two key elements found in real-world use-cases: attackers must operate under partial knowledge of the target model, and the decisions made by the attacker are irrevocable since they operate on a transient data stream. We first rigorously analyze a deterministic variant of the online threat model by drawing parallels to the well-studied $k$-secretary problem in theoretical computer science and propose Virtual+, a simple yet practical online algorithm. Our main theoretical result shows Virtual+ yields provably the best competitive ratio over all single-threshold algorithms for $k<5$ -- extending the previous analysis of the $k$-secretary problem. We also introduce the \textit{stochastic $k$-secretary} -- effectively reducing online blackbox transfer attacks to a $k$-secretary problem under noise -- and prove theoretical bounds on the performance of Virtual+ adapted to this setting. Finally, we complement our theoretical results by conducting experiments on MNIST, CIFAR-10, and Imagenet classifiers, revealing the necessity of online algorithms in achieving near-optimal performance and also the rich interplay between attack strategies and online attack selection, enabling simple strategies like FGSM to outperform stronger adversaries.
Drug discovery and development is a complex and costly process. Machine learning approaches are being investigated to help improve the effectiveness and speed of multiple stages of the drug discovery pipeline. Of these, those that use Knowledge Graphs (KG) have promise in many tasks, including drug repurposing, drug toxicity prediction and target gene-disease prioritisation. In a drug discovery KG, crucial elements including genes, diseases and drugs are represented as entities, whilst relationships between them indicate an interaction. However, to construct high-quality KGs, suitable data is required. In this review, we detail publicly available sources suitable for use in constructing drug discovery focused KGs. We aim to help guide machine learning and KG practitioners who are interested in applying new techniques to the drug discovery field, but who may be unfamiliar with the relevant data sources. The datasets are selected via strict criteria, categorised according to the primary type of information contained within and are considered based upon what information could be extracted to build a KG. We then present a comparative analysis of existing public drug discovery KGs and a evaluation of selected motivating case studies from the literature. Additionally, we raise numerous and unique challenges and issues associated with the domain and its datasets, whilst also highlighting key future research directions. We hope this review will motivate KGs use in solving key and emerging questions in the drug discovery domain.
RNA 3D architectures are stabilized by sophisticated networks of (non-canonical) base pair interactions, which can be conveniently encoded as multi-relational graphs and efficiently exploited by graph theoretical approaches and recent progresses in machine learning techniques. RNAglib is a library that eases the use of this representation, by providing clean data, methods to load it in machine learning pipelines and graph-based deep learning models suited for this representation. RNAglib also offers other utilities to model RNA with 2.5D graphs, such as drawing tools, comparison functions or baseline performances on RNA applications. The method and data is distributed as a fully documented pip package. Availability: this https URL
Understanding SARS-CoV-2 transmission in higher education settings is important to limit spread between students, and into at-risk populations. In this study, we sequenced 482 SARS-CoV-2 isolates from the University of Cambridge from 5 October to 6 December 2020. We perform a detailed phylogenetic comparison with 972 isolates from the surrounding community, complemented with epidemiological and contact tracing data, to determine transmission dynamics. We observe limited viral introductions into the university; the majority of student cases were linked to a single genetic cluster, likely following social gatherings at a venue outside the university. We identify considerable onward transmission associated with student accommodation and courses; this was effectively contained using local infection control measures and following a national lockdown. Transmission clusters were largely segregated within the university or the community. Our study highlights key determinants of SARS-CoV-2 transmission and effective interventions in a higher education setting that will inform public health policy during pandemics.
Identifying linked cases of infection is a critical component of the public health response to viral infectious diseases. In a clinical context, there is a need to make rapid assessments of whether cases of infection have arrived independently onto a ward, or are potentially linked via direct transmission. Viral genome sequence data are of great value in making these assessments, but are often not the only form of data available. Here, we describe A2B-COVID, a method for the rapid identification of potentially linked cases of COVID-19 infection designed for clinical settings. Our method combines knowledge about infection dynamics, data describing the movements of individuals, and evolutionary analysis of genome sequences to assess whether data collected from cases of infection are consistent or inconsistent with linkage via direct transmission. A retrospective analysis of data from two wards at Cambridge University Hospitals NHS Foundation Trust during the first wave of the pandemic showed qualitatively different patterns of linkage between cases on designated COVID-19 and non-COVID-19 wards. The subsequent real-time application of our method to data from the second epidemic wave highlights its value for monitoring cases of infection in a clinical context.
Recent work on training neural retrievers for open-domain question answering (OpenQA) has employed both supervised and unsupervised approaches. However, it remains unclear how unsupervised and supervised methods can be used most effectively for neural retrievers. In this work, we systematically study retriever pre-training. We first propose an approach of unsupervised pre-training with the Inverse Cloze Task and masked salient spans, followed by supervised finetuning using question-context pairs. This approach leads to absolute gains of 2+ points over the previous best result in the top-20 retrieval accuracy on Natural Questions and TriviaQA datasets. We also explore two approaches for end-to-end supervised training of the reader and retriever components in OpenQA models. In the first approach, the reader considers each retrieved document separately while in the second approach, the reader considers all the retrieved documents together. Our experiments demonstrate the effectiveness of these approaches as we obtain new state-of-the-art results. On the Natural Questions dataset, we obtain a top-20 retrieval accuracy of 84, an improvement of 5 points over the recent DPR model. In addition, we achieve good results on answer extraction, outperforming recent models like REALM and RAG by 3+ points. We further scale up end-to-end training to large models and show consistent gains in performance over smaller models.
The lack of anisotropic kernels in graph neural networks (GNNs) strongly limits their expressiveness, contributing to well-known issues such as over-smoothing. To overcome this limitation, we propose the first globally consistent anisotropic kernels for GNNs, allowing for graph convolutions that are defined according to topologicaly-derived directional flows. First, by defining a vector field in the graph, we develop a method of applying directional derivatives and smoothing by projecting node-specific messages into the field. Then, we propose the use of the Laplacian eigenvectors as such vector field. We show that the method generalizes CNNs on an n-dimensional grid and is provably more discriminative than standard GNNs regarding the Weisfeiler-Lehman 1-WL test. We evaluate our method on different standard benchmarks and see a relative error reduction of 8% on the CIFAR10 graph dataset and 11% to 32% on the molecular ZINC dataset, and a relative increase in precision of 1.6% on the MolPCBA dataset. An important outcome of this work is that it enables graph networks to embed directions in an unsupervised way, thus allowing a better representation of the anisotropic features in different physical or biological problems.
MalariaGEN is a data-sharing network that enables groups around the world to work together on the genomic epidemiology of malaria. Here we describe a new release of curated genome variation data on 7,000 Plasmodium falciparum samples from MalariaGEN partner studies in 28 malaria-endemic countries. High-quality genotype calls on 3 million single nucleotide polymorphisms (SNPs) and short indels were produced using a standardised analysis pipeline. Copy number variants associated with drug resistance and structural variants that cause failure of rapid diagnostic tests were also analysed. Almost all samples showed genetic evidence of resistance to at least one antimalarial drug, and some samples from Southeast Asia carried markers of resistance to six commonly-used drugs. Genes expressed during the mosquito stage of the parasite life-cycle are prominent among loci that show strong geographic differentiation. By continuing to enlarge this open data resource we aim to facilitate research into the evolutionary processes affecting malaria control and to accelerate development of the surveillance toolkit required for malaria elimination.