Nonsense variations, characterized by premature termination codons, play a major role in human genetic diseases as well as in cancer susceptibility. Despite their high prevalence, effective therapeutic strategies targeting premature termination codons remain a challenge. To understand and explore the intricate mechanisms involved, we developed StopKB, a comprehensive knowledgebase aggregating data from multiple sources on nonsense variations, associated genes, diseases, and phenotypes. StopKB identifies 637 317 unique nonsense variations, distributed across 18 022 human genes and linked to 3206 diseases and 7765 phenotypes. Notably, similar to 32% of these variations are classified as nonsense-mediated mRNA decay-insensitive, potentially representing suitable targets for nonsense suppression therapies. We also provide an interactive web interface to facilitate efficient and intuitive data exploration, enabling researchers and clinicians to navigate the complex landscape of nonsense variations. StopKB represents a valuable resource for advancing research in precision medicine and more specifically, the development of targeted therapeutic interventions for genetic diseases associated with nonsense variations.Database URL: https://lbgi.fr/stopkb/
Proteins interact with each other in complex ways to perform significant biological functions. These interactions, known as protein-protein interactions (PPIs), can be depicted as a graph where proteins are nodes and their interactions are edges. The development of high-throughput experimental technologies allows for the generation of numerous data which permits increasing the sophistication of PPI models. However, despite significant progress, current PPI networks remain incomplete. Discovering missing interactions through experimental techniques can be costly, time-consuming, and challenging. Therefore, computational approaches have emerged as valuable tools for predicting missing interactions. In PPI networks, a graph is usually used to model the interactions between proteins. An edge between two proteins indicates a known interaction, while the absence of an edge means the interaction is not known or missed. However, this binary representation overlooks the reliability of known interactions when predicting new ones. To address this challenge, we propose a novel approach for link prediction in weighted protein-protein networks, where interaction weights denote confidence scores. By leveraging data from the yeast Saccharomyces cerevisiae obtained from the STRING database, we introduce a new model that combines similarity-based algorithms and aggregated confidence score weights for accurate link prediction purposes. Our model significantly improves prediction accuracy, surpassing traditional approaches in terms of Mean Absolute Error, Mean Relative Absolute Error, and Root Mean Square Error. Our proposed approach holds the potential for improved accuracy in predicting PPIs, which is crucial for better understanding the underlying biological processes.
In fungi, the most abundant transcription factor (TF) class contains a fungal-specific ‘GAL4-like’ Zn2C6 DNA binding domain (DBD), while the second class contains another fungal-specific domain, known as ‘fungal_trans’ or middle homology domain (MHD), whose function remains largely uncharacterized. Remarkably, almost a third of MHD-containing TFs in public sequence databases apparently lack DNA binding activity, since they are not predicted to contain a DBD. Here, we reassess the domain organization of these ‘MHD-only’ proteins using an in silico error-tracking approach. In a large-scale analysis of ~17,000 MHD-only TF sequences present in all fungal phyla except Microsporidia and Cryptomycota, we show that the vast majority (>90%) result from genome annotation errors and we are able to predict a new DBD sequence for 14,261 of them. Most of these sequences correspond to a Zn2C6 domain (82%), with a small proportion of C2H2 domains (4%) found only in Dikarya. Our results contradict previous findings that the MHD-only TF are widespread in fungi. In contrast, we show that they are exceptional cases, and that the fungal-specific Zn2C6–MHD domain pair represents the canonical domain signature defining the most predominant fungal TF family. We call this family CeGAL, after the highly characterized members: Cep3, whose 3D structure is determined, and GAL4, a eukaryotic TF archetype. We believe that this will not only improve the annotation and classification of the Zn2C6 TF but will also provide critical guidance for future fungal gene regulatory network analyses.
The comparison of protein domain architectures provides insight into the evolution and function of proteins. By comparing the domains of different proteins, scientists can identify common domains, classify proteins based on their domain architecture, and highlight proteins that have evolved differently in one or more species or clades. Such proteins are often thought to represent genetic novelty underlying unique adaptations. However, genome-wide identification of different protein domain architectures involves a complex error-prone pipeline that includes genome sequencing, prediction of gene exon/intron structures, and inference of protein sequences and domain annotations. Here we developed an automated fact-checking approach to distinguish true domain loss/gain events from false events caused by errors that occur during the annotation process. Using genome-wide ortholog sets and taking advantage of the high-quality human and Saccharomyces cerevisiae genome annotations, we analyzed the domain gain and loss events in the predicted proteomes of 9 non-human primates (NHP) and 20 non- S. cerevisiae fungi (NSF) as annotated in the Uniprot and Interpro databases. Our approach allowed us to quantify the impact of errors on estimates of protein domain gains and losses, and we show that domain losses are over-estimated ten-fold and three-fold in the NHP and NSF proteins respectively. This is in line with previous studies of gene-level losses, where sequencing issues or incorrect gene prediction led to genes being falsely inferred as absent. For the first time, to our knowledge, we show that domain gains are also over-estimated by three-fold and two-fold respectively in NHP and NSF proteins. Based on our more accurate estimates, we infer that true domain losses and gains in NHP with respect to humans are observed at similar rates, while domains gains in the more divergent NSF are observed twice as frequently as domain losses with respect to S. cerevisiae . This study highlights the need to critically examine the scientific validity of protein annotations, and represents a significant step toward scalable computational fact-checking methods that may one day mitigate the propagation of wrong information in protein databases.
The widespread use of high throughput genome sequencing technologies has resulted in a significant increase in the number of available sequences, creating new challenges for genome annotation and prediction of protein-coding genes in terms of error detection and quality control. Multiple Sequence Alignments (MSAs) of the predicted protein sequences provide important contextual information that can be used to distinguish errors (caused by artifacts in the raw genome data, badly predicted gene sequences, or the alignment methods themselves) from true biological events. This can be achieved either by human expertise or by statistical analysis of the sequence data. Here, we propose a new approach that uses visual representations of MSAs as inputs for Convolutional Neural Networks (CNN) to classify MSAs into erroneous and non-erroneous categories. The MSAs are extracted from a unique in-house dataset, in which errors are carefully identified. Our model, called De-MISTED (Deep learning for MultIple Sequence alignmenTs Error Detection) identifies MSAs containing erroneous sequences with high accuracy (87%) and sensitivity (92%). Visual explanation techniques show that our model correctly identifies the position of multiple errors of different types (insertions, deletions and mismatches). Close examination of the data showed that our model can also identify errors that were not previously annotated in the data. The De-MISTED method thus contributes to a more robust exploitation of the genome data.
Protein annotation errors can have significant consequences in a wide range of fields, ranging from protein structure and function prediction to biomedical research, drug discovery, and biotechnology. By comparing the domains of different proteins, scientists can identify common domains, classify proteins based on their domain architecture, and highlight proteins that have evolved differently in one or more species or clades. However, genome-wide identification of different protein domain architectures involves a complex error-prone pipeline that includes genome sequencing, prediction of gene exon/intron structures, and inference of protein sequences and domain annotations. Here we developed an automated fact-checking approach to distinguish true domain loss/gain events from false events caused by errors that occur during the annotation process. Using genome-wide ortholog sets and taking advantage of the high-quality human and Saccharomyces cerevisiae genome annotations, we analyzed the domain gain and loss events in the predicted proteomes of 9 non-human primates (NHP) and 20 non-S. cerevisiae fungi (NSF) as annotated in the Uniprot and Interpro databases. Our approach allowed us to quantify the impact of errors on estimates of protein domain gains and losses, and we show that domain losses are over-estimated ten-fold and three-fold in the NHP and NSF proteins respectively. This is in line with previous studies of gene-level losses, where issues with genome sequencing or gene annotation led to genes being falsely inferred as absent. In addition, we show that insistent protein domain annotations are a major factor contributing to the false events. For the first time, to our knowledge, we show that domain gains are also over-estimated by three-fold and two-fold respectively in NHP and NSF proteins. Based on our more accurate estimates, we infer that true domain losses and gains in NHP with respect to humans are observed at similar rates, while domain gains in the more divergent NSF are observed twice as frequently as domain losses with respect to S. cerevisiae. This study highlights the need to critically examine the scientific validity of protein annotations, and represents a significant step toward scalable computational fact-checking methods that may 1 day mitigate the propagation of wrong information in protein databases.
BACKGROUND:The powerful 'graft versus leukemia' effect thought partly responsible for the therapeutic effect of allogeneic hematopoietic cell transplantation in acute myeloid leukemia (AML) provides rationale for investigation of immune-based therapies in this high-risk blood cancer. There is considerable preclinical evidence for potential synergy between PD-1 immune checkpoint blockade and the hypomethylating agents already commonly used for this disease.METHODS:We report here the results of 17 H-0026 (PD-AML, NCT02996474), an investigator sponsored, single-institution, single-arm open-label 10-subject pilot study to test the feasibility of the first-in-human combination of pembrolizumab and decitabine in adult patients with refractory or relapsed AML (R-AML).RESULTS:In this cohort of previously treated patients, this novel combination of anti-PD-1 and hypomethylating therapy was feasible and associated with a best response of stable disease or better in 6 of 10 patients. Considerable immunological changes were identified using T cell receptor β sequencing as well as single-cell immunophenotypic and RNA expression analyses on sorted CD3+ T cells in patients who developed immune-related adverse events (irAEs) during treatment. Clonal T cell expansions occurred at irAE onset; single-cell sequencing demonstrated that these expanded clones were predominately CD8+ effector memory T cells with high cell surface PD-1 expression and transcriptional profiles indicative of activation and cytotoxicity. In contrast, no such distinctive immune changes were detectable in those experiencing a measurable antileukemic response during treatment.CONCLUSION:Addition of pembrolizumab to 10-day decitabine therapy was clinically feasible in patients with R-AML, with immunological changes from PD-1 blockade observed in patients experiencing irAEs.
Multiple Sequence Alignments set the basis for many biological sequence analysis methods. However, they are susceptible to irregularities that result either from the predicted sequences or from natural biological events. In this paper, we propose MERLIN (Msa ERror Localization and IdentificatioN), an object detector that consists in identifying such irregularities using visual representations of MSAs. Our model is developed using a state-of-the-art deep learning object detector, YOLOv4, and trained on a set of MSA images from an in-house built dataset with automatically annotated errors. Our object detector exhibits a mean Average Precision of 71.18% in predicting different types of errors within MSAs. We conducted a thorough examination of the obtained results which showed that our method correctly identifies certain inconsistencies that were missed by the automatic annotation algorithm.
Transcription factors (TF) regulate gene activity in eukaryotic cells by binding specific regions of genomic DNA. In fungi, the most abundant TF class contains a fungal-specific ‘GAL4-like’ Zn2C6 DNA binding domain (DBD), while the second class contains another fungal-specific domain, known as ‘fungal_trans’ or Middle Homology Domain (MHD), whose function remains largely uncharacterized. Remarkably, almost a third of MHD-containing TF in public sequence databases apparently lack DNA binding activity, since they are not predicted to contain a DBD. Here, we reassess the domain organization of these ‘MHD-only’ proteins using an in silico error-aware approach. Our large-scale analysis of ~17000 MHD-only TF sequences showed that the vast majority (>90%) result from gene annotation errors, thus contradicting previous findings that the MHD-only TF are widespread in fungi. We show that they are in fact exceptional cases, and that the Zn2C6-MHD domain pair represents the canonical domain signature defining a new TF family composed of two fungal-specific domains. We call this family CeGAL, after the most characterized members: Ce p3, whose 3D structure has been determined and GAL 4, an archetypal eukaryotic TF. This definition should improve the classification of the Zn2C6 TF and provide critical insights into fungal gene regulatory networks.IMPORTANCE In fungi, extensive efforts focus on genome-wide characterization of potential Transcription Factors (TFs) and their targets genes to provide a better understanding of fungal processes and a rational for transcriptional manipulation. The second most abundant families of fungal-specific TFs, characterized by a Middle Homology Domain, are major regulators of primary and secondary metabolisms, multidrug resistance and virulence. Remarkably, one third of these TFs do not have a DNA Binding Domain (DBD-orphan) and thus are excluded from genome-wide studies. This particularity has been the subject of debate for many years. By computationally inspecting the close genomic environment of about 20,000 DBD-orphan TFs from a wide range of fungal species, we reveal that more than 90% contained sequences encoding a zinc-finger DBD. This analysis implies that the arrays of DBD containing TFs and their control DNA-sequences in target genes need to be reconsidered and expands the combinatorial regulation degree of the crucial fungal processes controlled by this TF family.
The X circular code is a set of 20 trinucleotides (codons) that has been identified in the protein-coding genes of most organisms (bacteria, archaea, eukaryotes, plasmids, viruses). It has been shown previously that the X circular code has the important mathematical property of being an error-correcting code. Thus, motifs of the X circular code, i.e. a series of codons belonging to X and called X motifs, allow identification and maintenance of the reading frame in genes. X motifs are significantly enriched in protein-coding genes, but have also been identified in many transfer RNA (tRNA) genes and in important functional regions of the ribosomal RNA (rRNA), notably in the peptidyl transferase center and the decoding center. Here, we investigate the potential role of X motifs as functional elements of protein-coding genes. First, we identify the codons of the X circular code which are frequent or rare in each domain of life (archaea, bacteria, eukaryota) and show that, for the amino acids with the highest codon bias, the preferred codon is often an X codon. We also observe a correlation between the 20 X codons and the optimal codons/dicodons that have been shown to influence translation efficiency. Then, we examined recently published experimental results concerning gene expression levels in diverse organisms. The approach used is the analysis of X motifs according to their density ds(X), i.e. the number of X motifs per kilobase in a gene sequence s. Surprisingly, this simple parameter identifies several unexpected relations between the X circular code and gene expression. For example, the X motifs are significantly enriched in the minimal gene set belonging to the three domains of life, and in codon-optimized genes. Furthermore, the density of X motifs generally correlates with experimental measures of translation efficiency and mRNA stability. Taken together, these results lead us to propose that the X motifs may represent a genetic signal contributing to the maintenance of the correct reading frame and the optimization and regulation of gene expression.
Omics analyses are powerful methods to obtain an integrated view of complex biological processes, disease progression, or therapy efficiency. However, few studies have compared different disease forms and different therapy strategies to define the common molecular signatures representing the most significant implicated pathways. In this study, we used RNA sequencing and mass spectrometry to profile the transcriptomes and proteomes of mouse models for three forms of centronuclear myopathies (CNMs), untreated or treated with either a drug (tamoxifen), antisense oligonucleotides reducing the level of dynamin 2 (DNM2), or following modulation of DNM2 or amphiphysin 2 (BIN1) through genetic crosses. Unsupervised analysis and differential gene and protein expression were performed to retrieve CNM molecular signatures. Longitudinal studies before, at, and after disease onset highlighted potential disease causes and consequences. Main pathways in the common CNM disease signature include muscle contraction, regeneration and inflammation. The common therapy signature revealed novel potential therapeutic targets, including the calcium regulator sarcolipin. We identified several novel biomarkers validated in muscle and/or plasma through RNA quantification, western blotting, and enzyme-linked immunosorbent assay (ELISA) assays, including ANXA2 and IGFBP2. This study validates the concept of using multi-omics approaches to identify molecular signatures common to different disease forms and therapeutic strategies.
Background Ab initio prediction of splice sites is an essential step in eukaryotic genome annotation. Recent predictors have exploited Deep Learning algorithms and reliable gene structures from model organisms. However, Deep Learning methods for non-model organisms are lacking. Results We developed Spliceator to predict splice sites in a wide range of species, including model and non-model organisms. Spliceator uses a convolutional neural network and is trained on carefully validated data from over 100 organisms. We show that Spliceator achieves consistently high accuracy (89-92%) compared to existing methods on independent benchmarks from human, fish, fly, worm, plant and protist organisms. Conclusions Spliceator is a new Deep Learning method trained on high-quality data, which can be used to predict splice sites in diverse organisms, ranging from human to protists, with consistently high accuracy.
Purpose DYRK1A syndrome is among the most frequent monogenic forms of intellectual disability (ID). We refined the molecular and clinical description of this disorder and developed tools to improve interpretation of missense variants, which remains a major challenge in human genetics. Methods We reported clinical and molecular data for 50 individuals with ID harboring DYRK1A variants and developed (1) a specific DYRK1A clinical score; (2) amino acid conservation data generated from 100 DYRK1A sequences across different taxa; (3) in vitro overexpression assays to study level, cellular localization, and kinase activity of DYRK1A mutant proteins; and (4) a specific blood DNA methylation signature. Results This integrative approach was successful to reclassify several variants as pathogenic. However, we questioned the involvement of some others, such as p.Thr588Asn, still reported as likely pathogenic, and showed it does not cause an obvious phenotype in mice. Conclusion Our study demonstrated the need for caution when interpreting variants in DYRK1A , even those occurring de novo. The tools developed will be useful to interpret accurately the variants identified in the future in this gene. Graphic abstract
Abstract Genetic mutations associated with acute myeloid leukemia (AML) also occur in age-related clonal hematopoiesis, often in the same individual. This makes confident assignment of detected variants to malignancy challenging. The issue is particularly crucial for AML posttreatment measurable residual disease monitoring, where results can be discordant between genetic sequencing and flow cytometry. We show here that it is possible to distinguish AML from clonal hematopoiesis and to resolve the immunophenotypic identity of clonal architecture. To achieve this, we first design patient-specific DNA probes based on patient's whole-genome sequencing and then use them for patient-personalized single-cell DNA sequencing with simultaneous single-cell antibody–oligonucleotide sequencing. Examples illustrate AML arising from DNMT3A- and TET2-mutated clones as well as independently. The ability to personalize single-cell proteogenomic assessment for individual patients based on leukemia-specific genomic features has implications for ongoing AML precision medicine efforts. Significance: This study offers a proof of principle of patient-personalized customized single-cell proteogenomics in AML including whole-genome sequencing–defined structural variants, currently unmeasurable by commercial “off-the-shelf” panels. This approach allows for the definition of genetic and immunophenotype features for an individual patient that would be best suited for measurable residual disease tracking.
Three-base periodicity (TBP), where nucleotides and higher order n-tuples are preferentially spaced by 3, 6, 9, etc. bases, is a well-known intrinsic property of protein-coding DNA sequences. However, its origins are still not fully understood. One hypothesis is that the periodicity reflects a primordial coding system that was used before the emergence of the modern standard genetic code (SGC). Recent evidence suggests that the X circular code, a set of 20 trinucleotides allowing the reading frames in genes to be retrieved locally, represents a possible ancestor of the SGC. Motifs from the X circular code have been found in the reading frame of protein-coding regions in extant organisms from bacteria to eukaryotes, in many transfer RNA (tRNA) genes and in important functional regions of the ribosomal RNA (rRNA), notably in the peptidyl transferase centre and the decoding centre. Here, we have used a powerful correlation function to search for periodicity patterns involving the 20 trinucleotides of the X circular code in a large set of bacterial protein-coding genes, as well as in the translation machinery, including rRNA and tRNA sequences. As might be expected, we found a strong circular code periodicity 0 modulo 3 in the protein-coding genes. More surprisingly, we also identified a similar circular code periodicity in a large region of the 16S rRNA. This region includes the 3' major domain corresponding to the primordial proto-ribosome decoding centre and containing numerous sites that interact with the tRNA and messenger RNA (mRNA) during translation. Furthermore, 3D structural analysis shows that the periodicity region surrounds the mRNA channel that lies between the head and the body of the SSU. Our results support the hypothesis that the X circular code may constitute an ancestral translation code involved in reading frame retrieval and maintenance, traces of which persist in modern mRNA, tRNA and rRNA despite their long evolution and adaptation to the SGC.
Premature infants are poor regulators of body temperature and are subjected to environmental factors that can lead to rapid heat loss, leaving them vulnerable to an increased risk of morbidity and mortality from hypothermia. Thermoregulation protocols have proven to increase survival in preterm infants.To evaluate a Plan-Do-Study-Act (PDSA) cycle on a previously implemented Golden Hour protocol at a military medical care facility for infants born at less than 32 weeks of gestation and weighing less than1500 g. Specific aims included the use of increased delivery/operating room temperatures and proper use of thermoregulatory devices (polyethylene bags and thermal mattress).Outcomes were analyzed and compared using a pre/postdesign. The data was collected using the neonatal intensive care unit admission worksheet.Although statistical analysis was not significant, clinical significance was illustrated by a decrease in hypothermia rates on admission and at 1 hour of life. There was a 100% compliance rate with increasing delivery room/operating room temperatures and thermal mattress use. Polyethylene bag use compliance was 50%.Golden Hour protocols have proven to be an effective tool. Thermoregulation is a significant component of these protocols, and it is imperative that every step is taken to manage the environmental temperature during the birth and admission process.There is a need for continued research on the impacts of thermoregulatory devices and protocols, with resulting practice and device recommendations.
The standard genetic code (SGC) describes how 64 trinucleotides (codons) encode 20 amino acids and the stop translation signal. Biochemical and statistical studies have shown that the standard genetic code is optimized to reduce the impact of errors caused by incorporation of wrong amino acids during translation. This is achieved by mapping codons that differ by only one nucleotide to the same amino acid or one with similar biochemical properties, so that if misincorporation occurs, the structure and function of the translated protein remain relatively unaltered. Some previous studies have extended the analysis of SGC optimality to the effect of frameshift errors on the conservation of amino acids. Here, we compare the optimality of the SGC with a set of circular codes, and in particular the X circular code identified in genes, on the basis of various biochemical properties over all possible frameshift errors. We show that the X circular code is more optimized to minimize the impact of frameshift errors than the SGC for the chosen amino acid properties. Furthermore, in the context of a problem that has been unresolved since 1996, we also demonstrate that the X circular code has a frameshift optimality in its combinatorial class of 216 maximal self-complementary C3 circular codes. To our knowledge, this is the first demonstration of the role of the X circular code in mitigation of translation errors. These results lead us to discuss the potential role of the X circular code in the evolution of the standard genetic code.