Loss-of-function variants in PALB2 give rise to defects in DNA damage repair by homologous recombination (HR), increasing the risk of breast cancer in female carriers. However, genetic testing frequently reveals missense variants of uncertain significance (VUS) for which the impact on protein function and cancer risk are unclear. Here we assay 84% of all possible missense variants in 11 out of 13 PALB2 exons using site-saturation functional screens with PARP inhibitor sensitivity as a readout for HR. These exons encode the coiled-coil and WD40 domains, which we identify as the minimal regions required for HR. Furthermore, we reveal the functional impact of 6718 missense variants, classifying 3904 variants as functional (58%), 2422 as intermediate (36%), and 392 as damaging (6%). A burden-type analysis shows that damaging missense variants in PALB2 are associated with a significantly increased risk of breast cancer, similar to that observed for truncating variants. These results will be valuable for the classification of PALB2 missense VUS and clinical management of carriers.
Transaminases are essential biocatalysts for asymmetric synthesis in the pharmaceutical and fine chemical industries. Here, we report the application of 3DM Engineering an AI-driven protein engineering platform-to optimize transaminase function by systematically exploring sequence-activity landscapes beyond those represented in the training data set. Our approach integrated the identification of hotspots from substrate tunnel analysis, enabling the construction of a focused, high-quality variant library targeting 53 residues for mutagenesis, which were subsequently used to train a protein language model. Further exploration of the sequence space identified mutations with previously unknown functional utility as salient targets for combination. The resulting higher-order variants displayed up to 21-fold improvement in catalytic efficiency and superior performance in the stereoselective synthesis of (S)-1-(2-chlorophenyl)ethanamine, achieving complete conversion and high enantiomeric excess (>99% ee). These results highlight the power of combining systematic hotspot identification with AI-driven exploration to discover unseen enzyme variants.
In digenic inheritance, pathogenic variants in two genes must be inherited together to cause disease. Only very few examples of digenic inheritance have been described in the neuromuscular disease field. Here we show that predicted deleterious variants in SRPK3 , encoding the X-linked serine/argenine protein kinase 3, lead to a progressive early onset skeletal muscle myopathy only when in combination with heterozygous variants in the TTN gene. The co-occurrence of predicted deleterious SRPK3 / TTN variants was not seen among 76,702 healthy male individuals, and statistical modeling strongly supported digenic inheritance as the best-fitting model. Furthermore, double-mutant zebrafish ( srpk3 −/− ; ttn.1 +/− ) replicated the myopathic phenotype and showed myofibrillar disorganization. Transcriptome data suggest that the interaction of srpk3 and ttn.1 in zebrafish occurs at a post-transcriptional level. We propose that digenic inheritance of deleterious changes impacting both the protein kinase SRPK3 and the giant muscle protein titin causes a skeletal myopathy and might serve as a model for other genetic diseases.
Predicting pathogenicity of missense variants in molecular diagnostics remains a challenge despite the available wealth of data, such as evolutionary information, and the wealth of tools to integrate that data. We describe DeepRank-Mut, a configurable framework designed to extract and learn from physicochemically relevant features of amino acids surrounding missense variants in 3D space. For each variant, various atomic and residue-level features are extracted from its structural environment, including sequence conservation scores of the surrounding amino acids, and stored in multi-channel 3D voxel grids which are then used to train a 3D convolutional neural network (3D-CNN). The resultant model gives a probabilistic estimate of whether a given input variant is disease-causing or benign. We find that the performance of our 3D-CNN model, on independent test datasets, is comparable to other widely used resources which also combine sequence and structural features. Based on the 10-fold cross-validation experiments, we achieve an average accuracy of 0.77 on the independent test datasets. We discuss the contribution of the variant neighborhood in the model’s predictive power, in addition to the impact of individual features on the model’s performance. Two key features: evolutionary information of residues in the variant neighborhood and their solvent accessibilities were observed to influence the predictions. We also highlight how predictions are impacted by the underlying disease mechanisms of missense mutations and offer insights into understanding these to improve pathogenicity predictions. Our study presents aspects to take into consideration when adopting deep learning approaches for protein structure-guided pathogenicity predictions.
Bacteriophages encode a wide variety of cell wall disrupting enzymes that aid the viral escape in the final stages of infection. These lytic enzymes have accumulated notable interest due to their potential as novel antibacterials for infection treatment caused by multiple-drug resistant bacteria. Here, the detailed functional and structural characterization of Thermus parvatiensis prophage peptidoglycan lytic amidase AmiP, a globular Amidase_3 type lytic enzyme adapted to high temperatures is presented. The sequence and structure comparison with homologous lytic amidases reveals the key adaptation traits that ensure the activity and stability of AmiP at high temperatures. The crystal structure determined at a resolution of 1.8 Å displays a compact α/β-fold with multiple secondary structure elements omitted or shortened compared with protein structures of similar proteins. The functional characterization of AmiP demonstrates high efficiency of catalytic activity and broad substrate specificity toward thermophilic and mesophilic bacteria strains containing Orn-type or DAP-type peptidoglycan. The here presented AmiP constitutes the most thermoactive and ultrathermostable Amidase_3 type lytic enzyme biochemically characterized with a temperature optimum at 85°C. The extraordinary high melting temperature Tm 102.6°C confirms fold stability up to approximately 100°C. Furthermore, AmiP is shown to be more active over the alkaline pH range with pH optimum at pH 8.5 and tolerates NaCl up to 300 mM with the activity optimum at 25 mM NaCl. This set of beneficial characteristics suggests that AmiP can be further exploited in biotechnology.
Supplementary Table from Functional Analysis Identifies Damaging CHEK2 Missense Variants Associated with Increased Cancer Risk
AbstractHeterozygous carriers of germline loss-of-function variants in the tumor suppressor gene checkpoint kinase 2 (CHEK2) are at an increased risk for developing breast and other cancers. While truncating variants in CHEK2 are known to be pathogenic, the interpretation of missense variants of uncertain significance (VUS) is challenging. Consequently, many VUS remain unclassified both functionally and clinically. Here we describe a mouse embryonic stem (mES) cell–based system to quantitatively determine the functional impact of 50 missense VUS in human CHEK2. By assessing the activity of human CHK2 to phosphorylate one of its main targets, Kap1, in Chek2 knockout mES cells, 31 missense VUS in CHEK2 were found to impair protein function to a similar extent as truncating variants, while 9 CHEK2 missense VUS resulted in intermediate functional defects. Mechanistically, most VUS impaired CHK2 kinase function by causing protein instability or by impairing activation through (auto)phosphorylation. Quantitative results showed that the degree of CHK2 kinase dysfunction correlates with an increased risk for breast cancer. Both damaging CHEK2 variants as a group [OR 2.23; 95% confidence interval (CI), 1.62–3.07; P < 0.0001] and intermediate variants (OR 1.63; 95% CI, 1.21–2.20; P = 0.0014) were associated with an increased breast cancer risk, while functional variants did not show this association (OR 1.13; 95% CI, 0.87–1.46; P = 0.378). Finally, a damaging VUS in CHEK2, c.486A>G/p.D162G, was also identified, which cosegregated with familial prostate cancer. Altogether, these functional assays efficiently and reliably identified VUS in CHEK2 that associate with cancer.Significance:Quantitative assessment of the functional consequences of CHEK2 variants of uncertain significance identifies damaging variants associated with increased cancer risk, which may aid in the clinical management of patients and carriers.
Background Protein truncating variants in ATM , BRCA1 , BRCA2 , CHEK2 , and PALB2 are associated with increased breast cancer risk, but risks associated with missense variants in these genes are uncertain. Methods We analyzed data on 59,639 breast cancer cases and 53,165 controls from studies participating in the Breast Cancer Association Consortium BRIDGES project. We sampled training (80%) and validation (20%) sets to analyze rare missense variants in ATM (1146 training variants), BRCA1 (644), BRCA2 (1425), CHEK2 (325), and PALB2 (472). We evaluated breast cancer risks according to five in silico prediction-of-deleteriousness algorithms, functional protein domain, and frequency, using logistic regression models and also mixture models in which a subset of variants was assumed to be risk-associated. Results The most predictive in silico algorithms were Helix ( BRCA1 , BRCA2 and CHEK2 ) and CADD ( ATM ). Increased risks appeared restricted to functional protein domains for ATM (FAT and PIK domains) and BRCA1 (RING and BRCT domains). For ATM , BRCA1 , and BRCA2 , data were compatible with small subsets (approximately 7%, 2%, and 0.6%, respectively) of rare missense variants giving similar risk to those of protein truncating variants in the same gene. For CHEK2 , data were more consistent with a large fraction (approximately 60%) of rare missense variants giving a lower risk (OR 1.75, 95% CI (1.47–2.08)) than CHEK2 protein truncating variants. There was little evidence for an association with risk for missense variants in PALB2 . The best fitting models were well calibrated in the validation set. Conclusions These results will inform risk prediction models and the selection of candidate variants for functional assays and could contribute to the clinical reporting of gene panel testing for breast cancer susceptibility.
In this white paper we introduce Helix, an AI based solution for missense pathogenicity prediction. With recent advances in the sequencing of human genomes, massive amounts of genetic data have become available. This has shifted the burden of labor for genetic diagnostics and research from the gathering of data to its interpretation. Helix presents a state of the art platform for the prediction of pathogenicity in human missense variants. In addition to offering best-in-class predictive performance, Helix offers a platform that allows researchers to analyze and interpret variants in depth that can be accessed at helixlabs.ai.
The Virus-X-Viral Metagenomics for Innovation Value-project was a scientific expedition to explore and exploit uncharted territory of genetic diversity in extreme natural environments such as geothermal hot springs and deep-sea ocean ecosystems. Specifically, the project was set to analyse and exploit viral metagenomes with the ultimate goal of developing new gene products with high innovation value for applications in biotechnology, pharmaceutical, medical, and the life science sectors. Viral gene pool analysis is also essential to obtain fundamental insight into ecosystem dynamics and to investigate how viruses influence the evolution of microbes and multicellular organisms. The Virus-X Consortium, established in 2016, included experts from eight European countries. The unique approach based on high throughput bioinformatics technologies combined with structural and functional studies resulted in the development of a biodiscovery pipeline of significant capacity and scale. The activities within the Virus-X consortium cover the entire range from bioprospecting and methods development in bioinformatics to protein production and characterisation, with the final goal of translating our results into new products for the bioeconomy. The significant impact the consortium made in all of these areas was possible due to the successful cooperation between expert teams that worked together to solve a complex scientific problem using state-of-the-art technologies as well as developing novel tools to explore the virosphere, widely considered as the last great frontier of life.
Despite advances in the field of missense variant effect prediction, the real clinical utility of current computational approaches remains rather limited. There is a large difference in performance metrics reported by developers and those observed in the real world. Most currently available predictors suffer from one or more types of circularity in their training and evaluation strategies that lead to overestimation of predictive performance. We present a generic strategy that is independent of dataset properties and algorithms used, to deal with circularity in the training phase. This results in more robust predictors and evaluation scores that accurately reflect the real-world performance of predictive models. Additionally, we show that commonly used training methods can have an adverse impact on model performance and lead to gross overestimation of true predictive performance.
Heterozygous carriers of germ-line loss-of-function variants in the DNA repair gene PALB2 are at a highly increased lifetime risk for developing breast cancer. While truncating variants in PALB2 are known to increase cancer risk, the interpretation of missense variants of uncertain significance (VUS) is in its infancy. Here we describe the development of a relatively fast and easy cDNA-based system for the semi high-throughput functional analysis of 48 VUS in human PALB2. By assessing the ability of PALB2 VUS to rescue the DNA repair and checkpoint defects in Palb2 knockout mouse embryonic stem (mES) cells, we identify various VUS in PALB2 that impair its function. Three VUS in the coiled-coil domain of PALB2 abrogate the interaction with BRCA1, whereas several VUS in the WD40 domain dramatically reduce protein stability. Thus, our functional assays identify damaging VUS in PALB2 that may increase cancer risk.
Congenital myasthenic syndrome (CMS) is a heterogeneous disorder that causes fatigable muscle weakness. CMS has been associated with variants in the MuSK gene and, to date, 16 patients have been reported. MuSK-CMS patients present a different phenotypic pattern of limb girdle weakness. Here, we describe four additional patients and discuss the phenotypic and clinical relationship with those previously reported. Two novel damaging missense variants are described: c.1742T > A; p.I581N found in homozygosis, and c.1634T > C; p.L545P found in compound heterozygosis with p.R166*. The reported patients had predominant limb girdle weakness with symptom onset at 12, 17, 18, and 30 years of age, and the majority exhibited a good clinical response to Salbutamol therapy, but not to esterase inhibitors. Meta-analysis including previously reported variants revealed an increased likelihood of a severe, respiratory phenotype with null alleles. Missense variants exclusively affecting the kinase domain, but not the catalytic site, are associated with late onset. These data refine the phenotype associated with MuSK-related CMS.
Dominant mutations in STIM1 are a cause of three allelic conditions: tubular aggregate myopathy, Stormorken syndrome (a complex phenotype including myopathy, hyposplenism, hypocalcaemia and bleeding diathesis), and a platelet dysfunction disorder, York platelet syndrome. Previous reports have suggested a genotype phenotype correlation with mutations in the N -terminal EF-hand domain associated with tubular aggregate myopathy, and a common mutation at p.R304W in a coiled coil domain associated with Stormorken syndrome. In this study individuals with STIMI variants were identified by exome sequencing or STIMI direct sequencing, and assessed for neuromuscular, haematological and biochemical evidence of the allelic disorders of STIMI. STIMI mutations were investigated by fibroblast calcium imaging and 3D modelling. Six individuals with STIM1 mutations, including two novel mutations (c.262A>G (p.S88G) and c.911G>A (p.R304Q)), were identified. Extra neuromuscular symptoms including thrombocytopenia, platelet dysfunction, hypocalcaemia or hyposplenism were present in 5/6 patients with mutations in both the EF-hand and CC domains. 3/6 patients had psychiatric disorders, not previously reported in STIMI disease. Review of published STIM1 patients (n = 49) confirmed that neuromuscular symptoms are present in most patients. We conclude that the phenotype associated with activating STIM1 mutations frequently includes extra -neuromuscular features such as hypocalcaemia, hypo-/asplenia and platelet dysfunction regardless of mutation domain. (C) 2017 Elsevier B.V. All rights reserved.
CorNet is a web-based tool for the analysis of co-evolving residue positions in protein super-family sequence alignments. CorNet projects external information such as mutation data extracted from literature on interactively displayed groups of co-evolving residue positions to shed light on the functions associated with these groups and the residues in them. We used CorNet to analyse six enzyme super-families and found that groups of strongly co-evolving residues tend to consist of residues involved in a same function such as activity, specificity, co-factor binding, or enantioselectivity. This finding allows to assign a function to residues for which no data is available yet in the literature. A mutant library was designed to mutate residues observed in a group of co-evolving residues predicted to be involved in enantioselectivity, but for which no literature data is available yet. The resulting set of mutations indeed showed many instances of increased enantioselectivity.
The prediction of missense variant pathogenicity is normally performed using analyses of multiple sequence alignments optionally augmented with analyses of the (predicted) protein structure. The most straightforward way, though, is to search the literature to see whether this variant has already been described. Variant data from homologous proteins are also valuable because mutations in a homologous protein often have similar effects as mutations at the equivalent residues of the protein of interest. Transferring variant data seems trivial but is seriously hampered by the fact that homologous residue positions have different numbers in different species. This problem is even bigger when to proteins have such low sequence identities that they can no longer be aligned based on their sequences only and their structures need to be compared to align them accurately. The protein superfamily analysis software suite 3DM solves these problems, because 3DM is a system that combines high quality structure based multiple sequence alignments in which aligned residues have the same number, with all published mutant and variant data for human and all other species. We have used 3DM to analyze nine human proteins for which many disease-related variants are known. This study reveals that mutation data can be transferred even between very distant homologous proteins. Thus, protein superfamily information systems, such as 3DM, offer a wealth of unused information that can be used in the analysis of human variants.
ABSTRACT The GPCRDB is a Molecular Class-Specific Information System (MCSIS) that collects, combines, validates, and disseminates large amounts of heterogeneous data on G protein-coupled receptors (GPCRs). The GPCRDB contains experimental data on sequences, ligand binding constants, mutations, and oligomers, as well as many different types of computationally derived data such as multiple sequence alignments and homology models. The GPCRDB provides access to the data via a number of different access methods. It offers visualization and analysis tools, and a number of query systems. The data is updated automatically on a monthly basis. The GPCRDB can be found online at http://www.gpcr.org/7tm/ INTRODUCTION G protein-coupled receptors constitute a large family of cell surface receptors. They regulate a wide range of cellular processes, including the senses of taste, smell, and vision, and control a myriad of intracellular signalling systems in response to external stimuli. GPCRs are a major target for the pharmaceutical industry as is reflected by the fact that more than a quarter of all FDA approved drugs act on a GPCR (1). GPCRs are arguably one of the most-researched classes of proteins, but despite intensive academic and industrial research efforts over the past three decades, little is known about the structural basis of GPCR function. From about 350 genes that code for the non-olfactorial receptors in the human species (2), only about 30 are truly validated therapeutic targets (3), indicating this family’s immense potential for future drug development. The fact that GPCRs can form homo-oligomeric and hetero-oligomeric complexes (4) has created a lot of new challenges and opportunities in the rational drug design process. In addition, a number of high-resolution crystal structures recently became available, providing new insights in receptor structure and function and giving the GPCR field a big stimulus.Researchers who focus on one particular protein or a class of proteins are confronted with the fact that both the number and the size of databases are expanding at an ever-increasing pace. Although many databases like PDB (5), UniProtKB (6), KEGG (7), EMBL (8), GenBank (9), etcetera are invaluable for their research, for the average wet-lab scientist these databases are less suitable for gathering, integrating, and updating different types of data in an easy and efficient manner. Studies that involve carrying over information from one protein to the other seem simple at a first glance, however, the amount of data that needs to be collected from heterogeneous sources, converted to syntactic and semantic homogeneity, validated, curated, stored, and indexed, is enormous.The GPCRDB is a data source that holds a large amount of heterogeneous data in a well-organized and easily accessible form. This data is validated, internally consistent, and updated regularly. In addition to being a one-stop GPCR resource, the data in the GPCRDB facilitates inferring new information using a wide spectrum of bioinformatics techniques.
The NucleaRDB is a Molecular Class-Specific Information System that collects, combines, validates and disseminates large amounts of heterogeneous data on nuclear hormone receptors. It contains both experimental and computationally derived data. The data and knowledge present in the NucleaRDB can be accessed using a number of different interactive and programmatic methods and query systems. A nuclear hormone receptor-specific PDF reader interface is available that can integrate the contents of the NucleaRDB with full-text scientific articles. The NucleaRDB is freely available at http://www.receptors.org/nucleardb.