Antiretroviral therapy is the standard treatment for HIV, but it requires daily use and can cause side effects. Despite being available for decades, there are still 1.5 million new infections and 700,000 deaths each year, highlighting the need for better therapies. Broadly neutralizing antibodies (bNAbs), which are highly active against HIV-1, represent a promising new approach and clinical trials have demonstrated the potential of bNAbs in the treatment and prevention of HIV-1 infection. However, HIV-1 antibody resistance (HIVAR) due to variants in the HIV-1 envelope glycoproteins (HIV-1 Env) is not well understood yet and poses a critical problem for the clinical use of bNAbs in treatment. HIVAR also plays an important role in the future development of an HIV-1 vaccine, which will require elicitation of bNAbs to which the circulating strains are sensitive. In recent years, a variety of methods have been developed to detect, characterize and predict HIVAR. Structural analysis of antibody-HIV-1 Env complexes has provided insight into viral residues critical for neutralization, while testing of viruses for antibody susceptibility has verified the impact of some of these residues. In addition, in vitro viral neutralization and adaption assays have shaped our understanding of bNAb susceptibility based on the envelope sequence. Furthermore, in vivo studies in animal models have revealed the rapid emergence of escape variants to mono- or combined bNAb treatments. Finally, similar variants were found in the first clinical trials testing bNAbs for the treatment of HIV-1-infected patients. These structural, in vitro, in vivo and clinical studies have led to the identification and validation of HIVAR for almost all available bNAbs. However, defined assays for the detection of HIVAR in patients are still lacking and for some novel, highly potent and broad-spectrum bNAbs, HIVAR have not been clearly defined. Here, we review currently available approaches for the detection, characterization and prediction of HIVAR.
Background Antiretroviral therapy (ART) is a life saving option for people living with HIV-1 (PLWH) and is effective against many viral strains. The most common ARTs involve combinations of drugs targeting viral or cellular proteins. Most of these drugs have to be taken daily. An alternative to ARTs with established inhibitors comprises broadly neutralizing antibodies (bNAbs). However, bNAbs share the problem of viral resistance with protein inhibitors. We developed a web service geno2pheno[bNAbs] that allows users to upload viral genotypes and estimates the respective resistance to many common bNAbs. The service uses trained statistical models to classify the virus into sensitive and resistant, respectively or to regress the IC50. Methods We used two linear models as well as two neural nets for each task and multi-task (MT) learning to train both models for IC50 prediction and classification simultaneously. During multi-task learning we penalize divergence of class and IC50 score in addition to the loss individual to each of the models. Findings We compared the linear models of geno2pheno[bNAbs] to other state-of-the-art methods like recurrent neural nets and self-attention, and found them to be competitive in regard to accuracy and have the benefit of fast computation and being easily interpretable in regard to features, i.e., positions on the envelope. Interpretation We developed a web service for the prediction of antibody resistance (geno2pheno[bNAbs]) to HIV-1, which is free to use and can be extended to other viruses, like Sars-Cov2, in the future. ### Competing Interest Statement The authors have declared no competing interest.
ABSTRACTHepatitis C virus infection is a significant global health concern, affecting millions worldwide. Although direct‐acting antivirals achieve over 90% success rate, treatment failures still occur, particularly when pan‐genotypic DAAs are unavailable, and drugs need to be chosen based on the present HCV genotype. Genotyping tests can be misleading, especially in cases involving the 2k/1b recombinant variant. The 2k/1b variant was first discovered in Saint Petersburg in 2002 and is most commonly observed in Eastern European countries, including Russia, Georgia, and Ukraine. Due to migration, the 2k/1b variant has spread to Western Europe and other regions, potentially increasing HCV transmission and changing the virus's epidemiological landscape. The situation highlights the importance of molecular epidemiology in monitoring the spread of the 2k/1b variant. Accurate detection and characterization of the 2k/1b variant are crucial for an effective treatment if no pan‐genotypic DAAs are available. To address this need, machine learning models were developed to predict the 2k/1b variant based on 1b and 2k/1b sequence data from nonstructural proteins. They were integrated into the geno2pheno[HCV] tool, providing physicians and researchers with an open‐access resource for determining HCV genotypes, including the 2k/1b variant.
Summary:Modern biological research critically depends on public databases. The introduction and propagation of errors within and across databases can lead to wasted resources as scientists are led astray by bad data or have to conduct expensive validation experiments. The emergence of generative artificial intelligence systems threatens to compound this problem owing to the ease with which massive volumes of synthetic data can be generated. We provide an overview of several key issues that occur within the biological data ecosystem and make several recommendations aimed at reducing data errors and their propagation. We specifically highlight the critical importance of improved educational programs aimed at biologists and life scientists that emphasize best practices in data engineering. We also argue for increased theoretical and empirical research on data provenance, error propagation, and on understanding the impact of errors on analytic pipelines. Furthermore, we recommend enhanced funding for the stewardship and maintenance of public biological databases. Availability and implementation:Not applicable.
MOTIVATION:In predicting HIV therapy outcomes, a critical clinical question is whether using historical information can enhance predictive capabilities compared with current or latest available data analysis. This study analyses whether historical knowledge, which includes viral mutations detected in all genotypic tests before therapy, their temporal occurrence, and concomitant viral load measurements, can bring improvements. We introduce a method to weigh mutations, considering the previously enumerated factors and the reference mutation-drug Stanford resistance tables. We compare a model encompassing history (H) with one not using this information (NH). RESULTS:The H-model demonstrates superior discriminative ability, with a higher ROC-AUC score (76.34%) than the NH-model (74.98%). Wilcoxon test results confirm significant improvement of predictive accuracy for treatment outcomes through incorporating historical information. The increased performance of the H-model might be attributed to its consideration of latent HIV reservoirs, probably obtained when leveraging historical information. The findings emphasize the importance of temporal dynamics in acquiring mutations. However, our result also shows that prediction accuracy remains relatively high even when no historical information is available. AVAILABILITY AND IMPLEMENTATION:This analysis was conducted using the Euresist Integrated DataBase (EIDB). For further validation, we encourage reproducing this study with the latest release of the EIDB, which can be accessed upon request through the Euresist Network.
Trained as an engineer, he turned to computational biology at the beginning of the new millennium.Together with his team, he is developing data fusion methods for the identification of candidate disease genes and variants [1-3], especially in rare genetic diseases.He also investigates data fusion methods for drug design and drug discovery [4,5].A third focus topic of his work is federated and privacy-preserving analysis of chemical or clinical genetic data [6][7][8].Towards these purposes, he employs methods of machine learning, such as Bayesian matrix factorization and deep learning.Going beyond mere methodical work, he strives for clinical or industrial applicability of his methods.Thus, he cofounded the company Cartagenia, which focused on bioinformatical solutions for clinical genetic diagnosis and has since become part of Agilent Technologies.
Even after three decades of antiretroviral therapy for HIV-1 (human immunodeficiency virus 1), therapy failure is a continual challenge. This is especially so if the viral variant is a recombinant of subtypes. Thus, improved diagnosis of recombined subtypes can help with the selection of therapy. We are using a new implementation of the previously published computational method recco to detect de novo recombination of known subtypes, independent of and in addition to known circulating recombinant forms (CRFs). We detect an optimal path in a multiple alignment of viral reference sequences based on mutation calls and probable breakpoints for recombination. A tuning parameter is used to favor either mutation calls or breakpoints. Besides novel recombinants, our tool g2p-recco integrated in the geno2pheno web service (https://geno2pheno.org) can successfully detect known recombinant events given only the full consensus references (without CRFs) of the involved subtypes with breakpoints. In addition, the tool can be applied to other viruses, i.e. hepatitis E virus (HEV). In this fashion, we could also detect several previously unknown recombinations in HEV.
BACKGROUND:Torque teno virus (TTV) is part of the human virome. TTV load was related to the immune status in patients after organ transplantation. We hypothesize that TTV load could be an additional marker for immune function in people living with HIV (PLWH). METHODS:In this analysis, serum samples of PLWH from the RESINA multicenter cohort were reanalyzed for TTV. Investigated clinical and epidemiological parameters included human pegivirus load, patient age and sex, HIV load, CD4+ T-cell count (Centers for Disease Control and Prevention [CDC] stage 1, 2, or 3), and CDC clinical stage (1993 CDC classification system; stage A, B, or C) before initiation of antiretroviral therapy. Regression analysis was used to detect possible associations among parameters. RESULTS:Our analysis confirmed TTV as a strong predictor of CD4+ T-cell count and CDC class 3. This relationship was used to propose a first classification of TTV load with regard to clinical stage. We found no association with clinical CDC stages A-C. The human pegivirus load was inversely correlated with HIV load but not TTV load. CONCLUSIONS:TTV load was associated with immunodeficiency in PLWH. Neither TTV nor HIV load were predictive for the clinical categories of HIV infection.
Human immunodeficiency virus (HIV) can develop resistance to all antiretroviral drugs. Multidrug resistance, however, is a rare event in modern HIV treatment, but can be life‐threatening, particular in patients with very long therapy histories and in areas with limited access to novel drugs. To understand the evolution of multidrug resistance, we analyzed the EuResist database to uncover the accumulation of mutations over time. We hypothesize that the accumulation of resistance mutations is not acquired simultaneously and randomly across viral genotypes but rather tends to follow a predetermined order. The knowledge of this order might help to elucidate potential mechanisms of multidrug resistance. Our evolutionary model shows an almost monotonic increase of resistance with each acquired mutation, including less well‐known nucleoside reverse transcriptase (RT) inhibitor‐related mutations like K223Q, L228H, and Q242H. Mutations within the integrase (IN) (T97A, E138A/K G140S, Q148H, N155H) indicate high probability of multidrug resistance. Hence, these IN mutations also tend to be observed together with mutations in the protease (PR) and RT. We followed up with an analysis of the mutation‐specific error rates of our model given the data. We identified several mutations with unusual rates (PR: M41L, L33F, IN: G140S). This could imply the existence of previously unknown virus variants in the viral quasispecies. In conclusion, our bioinformatics model supports the analysis and understanding of multidrug resistance.
The EuResist cohort was established in 2006 with the purpose of developing a clinical decision-support tool predicting the most effective antiretroviral therapy (ART) for persons living with HIV (PLWH), based on their clinical and virological data. Further to continuous extensive data collection from several European countries, the EuResist cohort later widened its activity to the more general area of antiretroviral treatment resistance with a focus on virus evolution. The EuResist cohort has retrospectively enrolled PLWH, both treatment-naïve and treatment-experienced, under clinical follow-up from 1998, in nine national cohorts across Europe and beyond, and this article is an overview of its achievement. A clinically oriented treatment-response prediction system was released and made available online in 2008. Clinical and virological data have been collected from more than one hundred thousand PLWH, allowing for a number of studies on the response to treatment, selection and spread of resistance-associated mutations and the circulation of viral subtypes. Drawing from its interdisciplinary vocation, EuResist will continue to investigate clinical response to antiretroviral treatment against HIV and monitor the development and circulation of HIV drug resistance in clinical settings, along with the development of novel drugs and the introduction of new treatment strategies. The support of artificial intelligence in these activities is essential.
Background Lower respiratory tract infections are among the main causes of death. Although there are many respiratory viruses, diagnostic efforts are focused mainly on influenza. The Respiratory Viruses Network (RespVir) collects infection data, primarily from German university hospitals, for a high diversity of infections by respiratory pathogens. In this study, we computationally analysed a subset of the RespVir database, covering 217,150 samples tested for 17 different viral pathogens in the time span from 2010 to 2019. Methods We calculated the prevalence of 17 respiratory viruses, analysed their seasonality patterns using information-theoretic measures and agglomerative clustering, and analysed their propensity for dual infection using a new metric dubbed average coinfection exclusion score (ACES). Results After initial data pre-processing, we retained 206,814 samples, corresponding to 1,408,657 performed tests. We found that Influenza viruses were reported for almost the half of all infections and that they exhibited the highest degree of seasonality. Coinfections of viruses are frequent; the most prevalent coinfection was rhinovirus/bocavirus and most of the virus pairs had a positive ACES indicating a tendency to exclude each other regarding infection. Conclusions The analysis of respiratory viruses dynamics in monoinfection and coinfection contributes to the prevention, diagnostic, treatment, and development of new therapeutics. Data obtained from multiplex testing is fundamental for this analysis and should be prioritized over single pathogen testing.
AbstractMotivationA common practice in the analysis of pathogens and their strains is using single-nucleotide polymorphisms (SNPs) to reconstruct their evolutionary history. However, genome-wide SNP-based phylogenetic trees are rarely analyzed without any further information. Including the underlying SNP data together with further metadata on the respective samples in the exploration process can facilitate linking the genomic and phenotypic properties of the samples.ResultsWe introduce Efficient VIsual analytics tool for Data ENrichment in phylogenetic TreEs (Evidente), a web-application that provides an interactive visual analysis interface for the simultaneous interrogation of phylogenetic relationships, genome-wide SNP data and metadata for samples of an organism. Besides visualizing the phylogenetic tree, Evidente classifies SNPs as supporting or non-supporting of the tree structures and shows the distribution of both types of SNPs among samples and clades of interest. Furthermore, additional metadata can be included in the visualization. Lastly, Evidente includes an enrichment analysis to identify over-represented genomic features encoded by GO-terms within the clades of the tree. We demonstrate the usability of Evidente with the data of the pathogens Treponema pallidum and Mycobacterium leprae.Availability and implementationEvidente is available at the TueVis visualization web server at https://evidente-tuevis.cs.uni-tuebingen.de/, it can also be run locally.Supplementary informationSupplementary data are available at Bioinformatics Advances online.
Background Non-pharmaceutical measures to control the spread of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) should be carefully tuned as they can impose a heavy social and economic burden. To quantify and possibly tune the efficacy of these anti-SARS-CoV-2 measures, we have devised indicators based on the abundant historic and current prevalence data from other respiratory viruses. Methods We obtained incidence data of 17 respiratory viruses from hospitalized patients and outpatients collected by 37 clinics and laboratories between 2010-2020 in Germany. With a probabilistic model for Bayes inference we quantified prevalence changes of the different viruses between months in the pre-pandemic period 2010-2019 and the corresponding months in 2020, the year of the pandemic with noninvasive measures of various degrees of stringency. Results We discovered remarkable reductions δ in rhinovirus (RV) prevalence by about 25% (95% highest density interval (HDI) [−0.35,−0.15]) in the months after the measures against SARS-CoV-2 were introduced in Germany. In the months after the measures began to ease, RV prevalence increased to low pre-pandemic levels, e.g. in August 2020 δ =−0.14 (95% HDI [−0.28,0.12]). Conclusions RV prevalence is negatively correlated with the stringency of anti-SARS-CoV-2 measures with only a short time delay. This result suggests that RV prevalence could possibly be an indicator for the efficiency for these measures. As RV is ubiquitous at higher prevalence than SARS-CoV-2 or other emerging respiratory viruses, it could reflect the efficacy of noninvasive measures better than such emerging viruses themselves with their unevenly spreading clusters.
Mario Albrecht合作论文数Research Group Computational Biology, Max Planck Institute for Informatics59