BACKGROUND:CD8+ T-cells are crucial for controlling and resolving SARS-CoV-2 infection, yet their epitope specificity and relationship to COVID-19 disease severity remain incompletely understood. METHODS:We performed comprehensive longitudinal profiling of antigen-specific CD8+ T-cell populations using DNA-barcoded peptide-HLA multimers, analysing 553 SARS-CoV-2 epitopes across globally prevalent HLA alleles in patients with mild and severe COVID-19. Functional and phenotypic characterisation was performed using multidimensional single-cell analysis and detailed cytokine profiling. The impact of post-infection COVID-19 vaccination on T-cell memory was also assessed. FINDINGS:Severe and mild COVID-19 were associated with robust yet distinct patterns of CD8+ T-cell activation. In the acute phase, severe disease was characterised by a broader T-cell repertoire (139 unique epitopes) with a median frequency of 1.4% (IQR 0.2-5.0) and a high-frequency of immunodominant epitope-specific T-cells that exhibited reduced cytotoxic profile. In contrast, patients with mild COVID-19 mounted responses against a more limited set of epitopes (98 unique epitopes), partially overlapping with those observed in severe disease, with a median T-cell frequency of 0.7% (IQR 0-1.9) and displayed a stronger cytotoxic phenotype and functional state. Over time, the memory T-cell compartment contracted to a restricted subset of immunodominant epitopes in the two patient groups and COVID-19 vaccination further enhanced frequencies of spike-specific T-cells independent of prior disease severity. INTERPRETATION:These findings delineate the epitope-specific frequency, function, and persistence of antigen-specific T-cell populations during SARS-CoV-2 infection, highlighting how differential activation, rather than magnitude alone, shapes immune outcomes across disease severities and other viral infections. FUNDING:This work was supported by the Independent Research Fund Denmark (DFF-Sapere Aude, 2066-00044B), the EU Horizon Europe REACT project (101057129), the European Research Council (ERC) Starting Grant MIMIC (101045517), and the Danish National Research Foundation (DNRF170).
Transmission of influenza A viruses (IAVs) between pigs and humans can trigger pandemics but more often cease as isolated infections without further spread in the new host species population. In Denmark, a major pig-producing country, the first two detections of human infections with swine-like IAVs were reported in 2021. These zoonotic IAVs were reassortants of the H1N1 pandemic 2009 lineage ("H1N1pdm09," H1 lineage 1A, clade 1A.3.3.2) introduced to swine farms in Denmark through humans approximately 11 years prior. However, predicting the likelihood and outcome of such IAV spillovers is challenging without a better understanding of the viral determinants. This study traced the evolution of H1N1pdm09 from 207 sequenced genomes as the virus propagated across Danish swine farms over a decade. H1N1pdm09 diverged into several genetically distinct viral populations, largely prompted by reassortments with neuraminidase (NA) segments from other enzootic IAV lineages. The genomic segments encoding the viral envelope glycoproteins, hemagglutinin (HA) and NA, evolved at the fastest rates, while the M and NS genomic segments were among the lowest evolutionary rates. The two zoonotic IAVs emerged from separate viral populations and shared the highest number of amino acid mutations in the PB2 and HA proteins. Acquisition of additional predicted glycosylation sites on the HA proteins of the zoonotic IAVs may have facilitated infection of the human patients. Ultimately, the analysis provides a foundation from which to further explore viral genetic indicators of host adaptation and zoonotic risk.
The impact of hepatitis B virus (HBV) diversity and evolution on disease progression is not well-understood. This study aims to compare intra-individual viral evolution in two groups of chronic hepatitis B (CHB) patients, using antiviral treatment initiation as a measure of lack of immunological control. From the Danish Database for Hepatitis B and C (DANHEP), 25 CHB patients were included; 14 with antiviral treatment initiation (TI group), and 11 without (NTI group). For each patient, three serial plasma samples taken before potential treatment initiation were selected. HBV DNA was amplified by PCR and analyzed by next-generation sequencing. HBV DNA and alanine transaminase were elevated in the TI group throughout the study period. Significantly higher substitution rates in the NTI group versus the TI group were found both within the viral population and at consensus level. Putative predicted CD8+ T cell epitopes contained significantly more substitutions in the NTI group. Genome-wide association analysis revealed several amino acid residues in the HBV genome associated with treatment initiation. This study shows that HBV has a higher rate of substitutions in CHB patients not requiring treatment. This could be linked to host immune pressure leading to disease control.
Accurate identification of translation initiation sites is essential for the proper translation of mRNA into functional proteins. In eukaryotes, the choice of the translation initiation site is influenced by multiple factors, including its proximity to the 5 ^' end and the local start codon context. Translation initiation sites mark the transition from non-coding to coding regions. This fact motivates the expectation that the upstream sequence, if translated, would assemble a nonsensical order of amino acids, while the downstream sequence would correspond to the structured beginning of a protein. This distinction suggests potential for predicting translation initiation sites using a protein language model. We present NetStart 2.0, a deep learning-based model that integrates the ESM-2 protein language model with the local sequence context to predict translation initiation sites across a broad range of eukaryotic species. NetStart 2.0 was trained as a single model across multiple species, and despite the broad phylogenetic diversity represented in the training data, it consistently relied on features marking the transition from non-coding to coding regions. By leveraging “protein-ness”, NetStart 2.0 achieves state-of-the-art performance in predicting translation initiation sites across a diverse range of eukaryotic species. This success underscores the potential of protein language models to bridge transcript- and peptide-level information in complex biological prediction tasks. The NetStart 2.0 webserver is available at: https://services.healthtech.dtu.dk/services/NetStart-2.0/ .
BACKGROUND:Preschool wheeze is a heterogenous and poorly understood clinical syndrome. As a result, current treatments are insufficient, and prevention is not possible. OBJECTIVE:We sought to increase understanding of the genetic susceptibility and underlying disease mechanisms of wheeze phenotypes in early childhood through large-scale genome-wide association study analyses. METHODS:We performed meta-analyses of genome-wide association study on early-onset wheeze, defined as recurrent wheeze or asthma in the first 3 years of life, and its subtypes, including early transient and persistent wheeze, defined by asthma/wheeze at age 3 and subsequent remission or persistence at age 6, respectively. The discovery analyses included data on more than 13,000 children from 15 cohorts; replication was sought through meta-analyses of data from 7 additional cohorts including up to 5000 children. Genetic variants associated with asthma-related traits in adulthood (adult asthma, atopy, eosinophils, and lung function) were used to quantify the degree to which genetic risk influencing asthma-related adult traits also influences genetic risk of preschool wheeze. RESULTS:Variants near the GSDMB gene in the 17q region showed genome-wide significant association with early-onset wheeze (rs2305480; odds ratio [95% confidence interval] = 1.26 [1.17-1.33], P = 2.30E-16) and persistent wheeze (rs11078926; 1.43 [1.30-1.578], P = 2.14E-11), but not with early transient wheeze (rs1054609; 1.08 [0.98-1.18], P = .094). Other known asthma loci were associated with early-onset wheeze, particularly CDHR3. Additionally, increased genetic risk to early-onset wheeze was associated with genetic risk for asthma at older ages, atopy, eosinophil count, and lower adult lung function. This was driven by persistent wheeze, whereas transient early wheeze was only associated with low lung function. CONCLUSIONS:Preschool wheeze phenotypes displayed distinct patterns of single nucleotide polymorphism associations and genetic enrichment with asthma-related traits. These results indicate distinct etiologies of wheeze phenotypes, which could inform studies in optimization of prevention and treatment strategies.
Influenza A viruses (IAVs) in swine have zoonotic potential and pose a continuous threat of causing human pandemics, as demonstrated by the H1N1 pandemic in 2009. Despite increased genomic surveillance, we have limited knowledge of the IAV evolutionary dynamics leading to such zoonotic events and no clear understanding of genetic markers associated with interspecies transmission of IAV between humans and swine. To explore this, we analysed a comprehensive publicly available whole-genome dataset of human and swine IAV sequences. We conducted phylogenetic analyses and inference of ancestral host and sequence states for each IAV segment to map inferred mutations associated with hypothetical representative transmissions within and between swine and human hosts. We developed a custom python library to combine information from host and ancestral sequence annotated trees and applied statistical models to identify genetic markers associated with intra- or interspecies transmissions between swine and humans. This included analysing mutation rates and the selective pressures acting on the viral proteins following intra- and interspecies transmissions and using a scalable, gradient-boosted decision tree machine learning approach to predict key amino acid positions critical for different transmission types. Our analyses not only indicated complex mutational patterns within and across viral proteins, but also suggested that specific protein regions and amino acid positions of especially several of the internal gene segments were more important for interspecies transmission. Our findings identify potential genetic signatures across the IAV proteins associated with host adaptation and zoonotic potential, offering valuable markers for early-warning genomic surveillance systems to enhance animal health and minimize the potential for zoonotic transmission of IAV.
Polyphenol oxidases (PPOs) are coupled binuclear copper proteins that catalyze the oxidation of phenols. New functions of PPOs are continuously being discovered, latest with several fungal o-methoxy phenolases, which are active on lignin-derived compounds. Here, we perform a comprehensive phylogenetic analysis of PPOs from a wide taxonomic origin and define 12 PPO groups. We find that a deep gene duplication has led to two distinct PPO types. Type 1 includes PPOs from chordates and molluscs, as well as the fungal o-methoxy phenolases. Type 2 includes plant PPOs, molluscan hemocyanins, and fungal tyrosinases. Most of the type 2 proteins have a C-terminal shielding domain and a thioether bond in the copper-binding site. We also find that most ascomycetes contain high numbers of the PPO type 1 that includes the o-methoxy phenolases, which may indicate a role in the lignin conversion strategy of these fungi.
Most heritable diseases are polygenic. To comprehend the underlying genetic architecture, it is crucial to discover the clinically relevant epistatic interactions (EIs) between genomic single nucleotide polymorphisms (SNPs) (1-3). Existing statistical computational methods for EI detection are mostly limited to pairs of SNPs due to the combinatorial explosion of higher-order EIs. With NeEDL (network-based epistasis detection via local search), we leverage network medicine to inform the selection of EIs that are an order of magnitude more statistically significant compared to existing tools and consist, on average, of five SNPs. We further show that this computationally demanding task can be substantially accelerated once quantum computing hardware becomes available. We apply NeEDL to eight different diseases and discover genes (affected by EIs of SNPs) that are partly known to affect the disease, additionally, these results are reproducible across independent cohorts. EIs for these eight diseases can be interactively explored in the Epistasis Disease Atlas (https://epistasis-disease-atlas.com). In summary, NeEDL demonstrates the potential of seamlessly integrated quantum computing techniques to accelerate biomedical research. Our network medicine approach detects higher-order EIs with unprecedented statistical and biological evidence, yielding unique insights into polygenic diseases and providing a basis for the development of improved risk scores and combination therapies.
The severe acute respiratory syndrome coronavirus-2 (SARS-CoV-2) not only caused the COVID-19 pandemic but also had a major impact on farmed mink production in several European countries. In Denmark, the entire population of farmed mink (over 15 million animals) was culled in late 2020. During the period of June to November 2020, mink on 290 farms (out of about 1100 in the country) were shown to be infected with SARS-CoV-2. Genome sequencing identified changes in the virus within the mink and it is estimated that about 4000 people in Denmark became infected with these mink virus variants. However, the routes of transmission of the virus to, and from, the mink have been unclear. Phylogenetic analysis revealed the generation of multiple clusters of the virus within the mink. Detailed analysis of changes in the virus during replication in mink and, in parallel, in the human population in Denmark, during the same time period, has been performed here. The majority of cases in mink involved variants with the Y453F substitution and the H69/V70 deletion within the Spike (S) protein; these changes emerged early in the outbreak. However, further introductions of the virus, by variants lacking these changes, from the human population into mink also occurred. Based on phylogenetic analysis of viral genome data, we estimate, using a conservative approach, that about 17 separate examples of mink to human transmission occurred in Denmark but up to 59 such events (90% credible interval: (39-77)) were identified using parsimony to count cross-species jumps on transmission trees inferred using Bayesian methods. Using the latter approach, 136 jumps (90% credible interval: (117-164)) from humans to mink were found, which may underlie the farm-to-farm spread. Thus, transmission of SARS-CoV-2 from humans to mink, mink to mink, from mink to humans and between humans were all observed.
Ancient environmental DNA (aeDNA) is becoming a powerful tool to gain insights about past ecosystems, overcoming the limitations of conventional fossil records. However, several methodological challenges remain, particularly for classifying the DNA to species level and conducting phylogenetic analysis. Current methods, primarily tailored for modern datasets, fail to capture several idiosyncrasies of aeDNA, including species mixtures from closely related species and ancestral divergence. We introduce soibean, a novel tool that utilizes mitochondrial pangenomic graphs for identifying species from aeDNA reads. It outperforms existing methods in accurately identifying species from multiple closely related sources within a sample, enhancing phylogenetic analysis for aeDNA. soibean employs a damage-aware likelihood model for precise identification at low coverage with a high damage rate. Additionally, we reconstructed ancestral sequences for soibean's database to handle aeDNA that is highly diverged from modern references. soibean demonstrates effectiveness through simulated data tests and empirical validation. Notably, our method uncovered new empirical results in published datasets, including using porpoise whales as food in a Mesolithic community in Sweden, demonstrating its potential to reveal previously unrecognized findings in aeDNA studies.
The recent expansion of mpox in Africa is characterized by a dramatic increase in zoonotic transmission (clade Ia) and the emergence of a new clade Ib that is transmitted from human to human by close contact. Clade Ia does not pose a threat in areas without zoonotic reservoirs. But clade Ib may spread widely, as did clade IIb which has spread globally since 2022 among men who have sex with men. It is not clear whether controlling clade Ib will be more difficult than clade IIb. The population at risk potentially counts 100 million but only a million vaccine doses are expected in the next year. Surveillance is needed with exhaustive case detection, polymerase chain reaction confirmation, clade determination, and about severe illness. Such data is needed to identify routes of transmission and core transmitters, such as sex workers. Health care workers are vaccinated to ensure their protection, but this will not curb mpox transmission. With the recent inequitable distribution of COVID-19 vaccines in mind, it is a global responsibility to ensure that low-income nations in the mpox epicenter have meaningful access to vaccines. Vaccination serves not only to reduce mortality in children but limit the risk of future mpox variants emerging that may spread in human populations globally.
The gut microbiome is shaped through infancy and impacts the maturation of the immune system, thus protecting against chronic disease later in life. Phages, or viruses that infect bacteria, modulate bacterial growth by lysis and lysogeny, with the latter being especially prominent in the infant gut. Viral metagenomes (viromes) are difficult to analyse because they span uncharted viral diversity, lacking marker genes and standardized detection methods. Here we systematically resolved the viral diversity in faecal viromes from 647 1-year-olds belonging to Copenhagen Prospective Studies on Asthma in Childhood 2010, an unselected Danish cohort of healthy mother-child pairs. By assembly and curation we uncovered 10,000 viral species from 248 virus family-level clades (VFCs). Most (232 VFCs) were previously unknown, belonging to the Caudoviricetes viral class. Hosts were determined for 79% of phage using clustered regularly interspaced short palindromic repeat spacers within bacterial metagenomes from the same children. Typical Bacteroides-infecting crAssphages were outnumbered by undescribed phage families infecting Clostridiales and Bifidobacterium. Phage lifestyles were conserved at the viral family level, with 33 virulent and 118 temperate phage families. Virulent phages were more abundant, while temperate ones were more prevalent and diverse. Together, the viral families found in this study expand existing phage taxonomy and provide a resource aiding future infant gut virome research.
The application of multiple omics technologies in biomedical cohorts has the potential to reveal patient-level disease characteristics and individualized response to treatment. However, the scale and heterogeneous nature of multi-modal data makes integration and inference a non-trivial task. We developed a deep-learning-based framework, multi-omics variational autoencoders (MOVE), to integrate such data and applied it to a cohort of 789 people with newly diagnosed type 2 diabetes with deep multi-omics phenotyping from the DIRECT consortium. Using in silico perturbations, we identified drug-omics associations across the multi-modal datasets for the 20 most prevalent drugs given to people with type 2 diabetes with substantially higher sensitivity than univariate statistical tests. From these, we among others, identified novel associations between metformin and the gut microbiota as well as opposite molecular responses for the two statins, simvastatin and atorvastatin. We used the associations to quantify drug-drug similarities, assess the degree of polypharmacy and conclude that drug effects are distributed across the multi-omics modalities.
Abstract Ancient environmental DNA (aeDNA) is a crucial source of information for past environmental reconstruction. However, the computational analysis of aeDNA involves the inherited challenges of ancient DNA (aDNA) and the typical difficulties of eDNA samples, such as taxonomic identification and abundance estimation of identified taxonomic groups. Current methods for aeDNA fall into those that only perform mapping followed by taxonomic identification and those that purport to do abundance estimation. The former leaves abundance estimates to users, while methods for the latter are not designed for large metagenomic datasets and are often imprecise and challenging to use. Here, we introduce euka, a tool designed for rapid and accurate characterisation of aeDNA samples. We use a taxonomy‐based pangenome graph of reference genomes for robustly assigning DNA sequences and use a maximum‐likelihood framework for abundance estimation. At the present time, our database is restricted to mitochondrial genomes of tetrapods and arthropods but can be expanded in future versions. We find euka to outperform current taxonomic profiling tools and their abundance estimates. Crucially, we show that regardless of the filtering threshold set by existing methods, euka demonstrates higher accuracy. Furthermore, our approach is robust to sparse data, which is idiosyncratic of aeDNA, detecting a taxon with an average of 50 reads aligning. We also show that euka is consistent with competing tools on empirical samples. euka's features are fine‐tuned to deal with the challenges of aeDNA, making it a simple‐to‐use, all‐in‐one tool. It is available on GitHub: https://github.com/grenaud/vgan. euka enables researchers to quickly assess and characterise their sample, thus allowing it to be used as a routine screening tool for aeDNA.
BACKGROUND:Asthma with severe exacerbation is one of the most common causes of hospitalization among young children. Exacerbations are typically triggered by respiratory infections, but the host factors causing recurrent infections and exacerbations in some children are poorly understood. As a result, current treatment options and preventive measures are inadequate.OBJECTIVE:We sought to identify genetic interaction associated with the development of childhood asthma.METHODS:We performed an exhaustive search for pairwise interaction between genetic single nucleotide polymorphisms using 1204 cases of a specific phenotype of early childhood asthma with severe exacerbations in patients aged 2 to 6 years combined with 5328 nonasthmatic controls. Replication was attempted in 3 independent populations, and potential underlying immune mechanisms were investigated in the COPSAC2010 and COPSAC2000 birth cohorts.RESULTS:We found evidence of interaction, including replication in independent populations, between the known childhood asthma loci CDHR3 and GSDMB. The effect of CDHR3 was dependent on the GSDMB genotype, and this interaction was more pronounced for severe and early onset of disease. Blood immune analyses suggested a mechanism related to increased IL-17A production after viral stimulation.CONCLUSIONS:We found evidence of interaction between CDHR3 and GSDMB in development of early childhood asthma, possibly related to increased IL-17A response to viral infections. This study demonstrates the importance of focusing on specific disease subtypes for understanding the genetic mechanisms of asthma.
The discovery of microorganisms capable of complete ammonia oxidation to nitrate (comammox) has prompted a paradigm shift in our understanding of nitrification, an essential process in N cycling, hitherto considered to require both ammonia oxidizing and nitrite oxidizing microorganisms. This intriguing metabolism is unique to the genus Nitrospira, a diverse taxon previously known to only contain canonical nitrite oxidizers. Comammox Nitrospira have been detected in diverse environments; however, a global view of the distribution, abundance, and diversity of Nitrospira species is still incomplete. In this study, we retrieved 55 metagenome-assembled Nitrospira genomes (MAGs) from newly obtained and publicly available metagenomes. Combined with publicly available MAGs, this constitutes the largest Nitrospira genome database to date with 205 MAGs, representing 132 putative species, most without cultivated representatives. Mapping of metagenomic sequencing reads from various environments against this database enabled an analysis of the distribution and habitat preferences of Nitrospira species. Comammox Nitrospira’s ecological success is evident as they outnumber and present higher species-level richness than canonical Nitrospira in all environments examined, except for marine and wastewaters samples. The type of environment governs Nitrospira species distribution, without large-scale biogeographical signal. We found that closely related Nitrospira species tend to occupy the same habitats, and that this phylogenetic signal in habitat preference is stronger for canonical Nitrospira species. Comammox Nitrospira eco-evolutionary history is more complex, with subclades achieving rapid niche divergence via horizontal transfer of genes, including the gene encoding hydroxylamine oxidoreductase, a key enzyme in nitrification. Our study expands the genomic inventory of the Nitrospira genus, exposes the ecological success of complete ammonia oxidizers within a wide range of habitats, identifies the habitat preferences of (sub)lineages of canonical and comammox Nitrospira species, and proposes that horizontal transfer of genes involved in nitrification is linked to niche separation within a sublineage of comammox Nitrospira.
Since the influenza pandemic in 2009, there has been an increased focus on swine influenza A virus (swIAV) surveillance. This paper describes the results of the surveillance of swIAV in Danish swine from 2011 to 2018. In total, 3800 submissions were received with a steady increase in swIAV-positive submissions, reaching 56% in 2018. Full-genome sequences were obtained from 129 swIAV-positive samples. Altogether, 17 different circulating genotypes were identified including six novel reassortants harboring human seasonal IAV gene segments. The phylogenetic analysis revealed substantial genetic drift and also evidence of positive selection occurring mainly in antigenic sites of the hemagglutinin protein and confirmed the presence of a swine divergent cluster among the H1pdm09Nx (clade 1A.3.3.2) viruses. The results provide essential data for the control of swIAV in pigs and emphasize the importance of contemporary surveillance for discovering novel swIAV strains posing a potential threat to the human population.
The gut microbiome (GM) is shaped through infancy and plays a major role in determining susceptibility to chronic diseases later in life. Bacteriophages (phage) are known to modulate bacterial populations in numerous ecosystems, including the gut. However, virome data is difficult to analyse because it mostly consists of unknown viruses, i.e. viral dark matter. Here, we manually resolved the viral dark matter in the largest human virome study published to date. Fecal viromes from a cohort of 647 infants at 1 year of age were deeply sequenced and analysed through successive rounds of clustering and curation. This uncovered more than ten thousand viral species distributed over 248 viral families falling within 17 viral order-level clusters. Most of the defined viral families and orders were novel and belonged to the Caudoviricetes viral class. Bacterial hosts were predicted for 79% of the viral species using CRISPR spacers in metagenomes from the same fecal samples. While Bacteroides-infecting Crassphages were present, novel viral families were more predominant, including phages infecting Clostridiales and Bifidobacterium. Phage lifestyles were determined for more than three thousand caudoviral species. Lifestyles were homogeneous at the family level for 149 caudiviral families. 32 families were found to be virulent, while 117 families were temperate. Virulent phage families were more abundant but temperate phage families were more diverse and widespread. Together, the viral families found in this study represent a major expansion of current bacteriophage taxonomy, and the sequences have been put online for use and validation by the community.
Søren Brunak合作论文数Rigshospitalet;Novo Nordisk Foundation Center for Protein Research, University of Copenhagen;Department of Systems Biology, Technical University of Denmark18
Pierre Baldi合作论文数Department of Information and Computer Science, School of Information and Computer Sciences, University of California, Irvine;Center for Machine Learning and Intelligent Systems, Bren School of Information and Computer Science, University of California, Irvine;Mohamed bin Zayed University of Artificial Intelligence8