Template-based modelling, also known as homology modelling, is a powerful approach to predict the structure of a protein from its amino acid sequence. The approach requires one to identify a sequence similarity between the query sequence and that of a known structure as they will adopt a similar conformation, and the known structure can be used as the template for modelling the query sequence. Recently several approaches, most notably AlphaFold, have employed enhanced machine learning and have yielded accurate models irrespective of whether there is an identifiable template. Here we report Phyre2.2 which incorporates several enhancements to our widely-used template modelling portal Phyre2. The main development is facilitating a user to submit their sequence and then Phyre2.2 identifies the most suitable AlphaFold model to be used as a template. In Phyre2.2 the user searches a template library of known structures. We have now included in our library a representative structure for every protein sequence in the protein databank (PDB). In addition, there are representatives for an apo and a holo structure if they are in the PDB. The ranking of hits has been modified to highlight to the user if there are different domains spanning the sequence. Phyre2.2 continues to support batch processing where a user can submit up to 100 sequences facilitating processing of proteomes. Phyre2.2 is freely available to all users, including commercial users, at https://www.sbg.bio.ic.ac.uk/phyre2/. (c) 2025 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY license (http://creativecom-mons.org/licenses/by/4.0/).
The AlphaFold database provides 200M protein structures predicted by AlphaFold2 in 2022 from sequences in UniProt. However, of the 20,504 full-length human structures in the AlphaFold database, 575 entries conflict with UniProt (version 2024_05); and there is a similar discrepancy for other species. This highlights how bioinformatics resources as exemplified by the AlphaFold database can rapidly age. ### Competing Interest Statement The authors have declared no competing interest. Wellcome Trust, 218242/Z/19/z Biotechnology and Biological Sciences Research Council, BB/T010487/, BB/V018558/1 Medical Research Council, MR/Y031091/1
Variant effect predictors assess if a substitution is pathogenic or benign. Most predictors, including those that are structure-based, are designed for globular proteins in aqueous environments and do not consider that the variant residue is located within the membrane. We report Missense3D-TM that provides a structure-based assessment of the impact of a missense variant located within a membrane. On a data -set of 2,078 pathogenic and 1,060 benign variants, spanning 711 proteins from 706 structures, Missense3D-TM achieved an accuracy of 66%, Mathews correlation coefficient of 0.37, sensitivity of 58% and specificity of 81%. Missense3D-TM performed similarly to mCSM-membrane: accuracy 66% vs 61% (p = 0.02) on an unbalanced test set and 70% vs 67% (p = 0.20) on a balanced test set. The Missense3D-TM website provides an analysis of the structural effects of the variant along with its pre-dicted position within the membrane. The web server is available at http://missense3d.bc.ic.ac.uk/. (c) 2023 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY license (http://creativecom-mons.org/licenses/by/4.0/).
We provide an overview of the methods that can be used for protein structure-based evaluation of missense variants. The algorithms can be broadly divided into those that calculate the difference in free energy (ΔΔG) between the wild type and variant structures and those that use structural features to predict the damaging effect of a variant without providing a ΔΔG. A wide range of machine learning approaches have been employed to develop those algorithms. We also discuss challenges and opportunities for variant interpretation in view of the recent breakthrough in three-dimensional structural modelling using deep learning.
ABSTRACT In 2019, we released Missense3D which identifies stereochemical features that are disrupted by a missense variant, such as introducing a buried charge. Missense3D analyses the effect of a missense variant on a single structure and thus may fail to identify as damaging surface variants disrupting a protein interface i.e., a protein-protein interaction (PPI) site. Here we present Missense3D-PPI designed to predict missense variants at PPI interfaces. Our development dataset comprised of 1,279 missense variants (pathogenic n=733, benign n=546) in 434 proteins and 545 experimental structures of PPI complexes. Benchmarking of Missense3D-PPI was performed after dividing the dataset in training (320 benign and 320 pathogenic variants) and testing (226 benign and 413 pathogenic). Structural features affecting PPI, such as disruption of interchain bonds and introduction of unbalanced charged interface residues, were analysed to assess the impact of the variant at PPI. Missense3D-PPI’s performance was superior to that of Missense3D: sensitivity 42% versus 8% and accuracy 58% versus 40%, p=4.23×10 −16 However, the specificity of Missense3D-PPI was slightly lower compared to Missense3D (84% versus 98%). On our dataset, Missense3D-PPI’s accuracy was superior to BeAtMuSiC (p=2.3×10 −5 ), mCSM-PPI2 (p=3.2×10 −12 ) and MutaBind2 (p=0.003). Missense3D-PPI represents a valuable tool for predicting the structural effect of missense variants on biological protein networks and is available at the Missense3D web portal ( http://missense3d.bc.ic.ac.uk/missense3d/indexppi.html ).
Endeavors in the field of dye-sensitized solar cells (DSCs) have shown great promise when adopting a data-driven approach to materials discovery, such as successful molecular-scale predictions of light-harvesting chromophores. However, predictions of DSC dyes would become much more sophisticated if a molecular-to-macroscopic DSC device prediction methodology existed. Thereby, a fully computational pipeline is presented that predicts device-performance parameters of DSCs which contain varying dye combinations. Optimal pairing of complementary dyes is identified via a data-driven workflow that affords cosensitized DSCs with maximum power-conversion efficiencies. Six high-performing DSC dyes are paired with partner dyes that are screened from a database of 8488 compounds using sequential heuristic filters. Existing models that predict short-circuit-current density (J(SC)) and open-circuit voltage (V-OC) parameters are adapted to predict singly sensitized and cosensitized DSC performance. The predictions for J(sc) values of singly sensitized devices match experimental literature values with comparable accuracy to more computationally costly methods. Five out of six dye pairings are predicted to have greater J(SC) values when cosensitized compared to their corresponding singly sensitized devices, including two pairs that show strong J(sc) boosts of +13% and +12% when cosensitized. Thus, the prospect of an entirely in-silico prediction pipeline for DSC performance that can be used to realize the fully automated design of optimized cosensitized DSCs is demonstrated.
Background: The human protein transmembrane protease serine type 2 (TMPRSS2) plays a key role in SARS-CoV-2 infection, as it is required to activate the virus' spike protein, facilitating entry into target cells. We hypothesized that naturally-occurring TMPRSS2 human genetic variants affecting the structure and function of the TMPRSS2 protein may modulate the severity of SARS-CoV-2 infection. Methods: We focused on the only common TMPRSS2 non-synonymous variant predicted to be damaging (rs12329760 C>T, p.V160M), which has a minor allele frequency ranging from 0.14 in Ashkenazi Jewish to 0.38 in East Asians. We analysed the association between the rs12329760 and COVID-19 severity in 2,244 critically ill patients with COVID-19 from 208 UK intensive care units recruited as part of the GenOMICC (Genetics Of Mortality In Critical Care) study. Logistic regression analyses were adjusted for sex, age and deprivation index. For in vitro studies, HEK293 cells were co-transfected with ACE2 and either TMPRSS2 wild type or mutant (TMPRSS2(V160M)). A SARS-CoV-2 pseudovirus entry assay was used to investigate the ability of TMPRSS2(V160M) to promote viral entry. Results: We show that the T allele of rs12329760 is associated with a reduced likelihood of developing severe COVID-19 (OR 0.87, 95%CI:0.79-0.97, p = 0.01). This association was stronger in homozygous individuals when compared to the general population (OR 0.65, 95%CI:0.50-0.84, p = 1.3 x 10(-3)). We demonstrate in vitro that this variant, which causes the amino acid substitution valine to methionine, affects the catalytic activity of TMPRSS2 and is less able to support SARS-CoV-2 spike-mediated entry into cells. Conclusion: TMPRSS2 rs12329760 is a common variant associated with a significantly decreased risk of severe COVID-19. Further studies are needed to assess the expression of TMPRSS2 across different age groups. Moreover, our results identify TMPRSS2 as a promising drug target, with a potential role for camostat mesilate, a drug approved for the treatment of chronic pancreatitis and postoperative reflux esophagitis, in the treatment of COVID-19. Clinical trials are needed to confirm this. (C) 2022 The Authors. Published by Elsevier Masson SAS.
3DLigandSite is a web tool for the prediction of ligand-binding sites in proteins. Here, we report a significant update since the first release of 3DLigandSite in 2010. The overall methodology remains the same, with candidate binding sites in proteins inferred using known binding sites in related protein structures as templates. However, the initial structural modelling step now uses the newly available structures from the AlphaFold database or alternatively Phyre2 when AlphaFold structures are not available. Further, a sequence-based search using HHSearch has been introduced to identify template structures with bound ligands that are used to infer the ligand-binding residues in the query protein. Finally, we introduced a machine learning element as the final prediction step, which improves the accuracy of predictions and provides a confidence score for each residue predicted to be part of a binding site. Validation of 3DLigandSite on a set of 6416 binding sites obtained 92% recall at 75% precision for non-metal binding sites and 52% recall at 75% precision for metal binding sites. 3DLigandSite is available at https://www.wass-michaelislab.org/3dligandsite. Users submit either a protein sequence or structure. Results are displayed in multiple formats including an interactive Mol* molecular visualization of the protein and the predicted binding sites.
Thailand was the first country outside China to officially report COVID-19 cases. Despite the strict regulations for international arrivals, up until February 2021, Thailand had been hit by two major outbreaks. With a large number of SARS-CoV-2 sequences collected from patients, the effects of many genetic variations, especially those unique to Thai strains, are yet to be elucidated. In this study, we analysed 439,197 sequences of the SARS-CoV-2 spike protein collected from NCBI and GISAID databases. 595 sequences were from Thailand and contained 52 variants, of which 6 had not been observed outside Thailand (p.T51N, p.P57T, p.I68R, p.S205T, p.K278T, p.G832C). These variants were not predicted to be of concern. We demonstrate that the p.D614G, although already present during the first Thai outbreak, became the prevalent strain during the second outbreak, similarly to what was described in other countries. Moreover, we show that the most common variants detected in Thailand (p.A829T, p.S459F and p.S939F) do not appear to cause any major structural change to the spike trimer or the spike-ACE2 interaction. Among the variants identified in Thailand was p.N501T. This variant, which involves an asparagine critical for spike-ACE2 binding, was not predicted to increase SARS-CoV-2 binding, thus in contrast to the variant of global concern p.N501Y. In conclusion, novel variants identified in Thailand are unlikely to increase the fitness of SARS-CoV-2. The insights obtained from this study could aid SARS-CoV-2 variants prioritisations and help molecular biologists and virologists working on strain surveillance.
The main goal of molecular simulation is to accurately predict experimental observables of molecular systems. Another long-standing goal is to devise models for arbitrary neutral organic molecules with little or no reliance on experimental data. While separately these goals have been met to various degrees, for an arbitrary system of molecules they have not been achieved simultaneously. For biophysical ensembles that exist at room temperature and pressure, and where the entropic contributions are on par with interaction strengths, it is the free energies that are both most important and most difficult to predict. We compute the free energies of solvation for a diverse set of neutral organic compounds using a polarizable force field fitted entirely to ab initio calculations. The mean absolute errors (MAE) of hydration, cyclohexane solvation, and corresponding partition coefficients are 0.2 kcal/mol, 0.3 kcal/mol and 0.22 log units, i.e . within chemical accuracy. The model (ARROW FF) is multipolar, polarizable, and its accompanying simulation stack includes nuclear quantum effects (NQE). The simulation tools’ computational efficiency is on a par with current state-of-the-art packages. The construction of a wide-coverage molecular modelling toolset from first principles, together with its excellent predictive ability in the liquid phase is a major advance in biomolecular simulation.
We introduce a multi-reward reinforcement learning (RL) approach to train a flexible bond-order potential (BOP) for 2D phosphorene based on ab initio training data sets. Our approach is based on a continuous action space Monte Carlo tree search algorithm that is general and scalable and presents an efficient multiobjective optimization scheme for high-dimensional materials design problems. As a proof-of-concept, we deploy this scheme to parametrize multiple structural and dynamical properties of 2D phosphorene polymorphs. Our RL-trained BOP model adequately captures the structure, energetics, transformation barriers, equation of state, elastic constants, and phonon dispersions of various 2D P polymorphs. We use this model to probe the impact of temperature and strain rate on the phase transition from black (α-P) to blue phosphorene (β-P) through molecular dynamics simulations. A decrease in critical strain for this phase transition with increase in temperature is observed, and the underlying atomistic mechanisms are discussed.
AlphaFold, the deep learning algorithm developed by DeepMind, recently released the three-dimensional models of the whole human proteome to the scientific community. Here we discuss the advantages, limitations and the still unsolved challenges of the AlphaFold models from the perspective of a biologist, who may not be an expert in structural biology.
We present the development process of Bioblox2-5D, an educational biology game aimed at teenagers. The game content refers to protein docking and aims to improve learning about molecular shape complexity, the roles of charges in molecular docking and the scoring function to calculate binding affinity. We developed the game as part of a collaboration between the Computing Department at Goldsmiths, University of London, and the Structural Bioinformatics group at Imperial College London. The team at Imperial provided the content requirements and validated the technical solution adopted in the game. The team at Goldsmiths designed and implemented the content requirements into a fun and stimulating educational puzzle game that supports teaching and motivates students to engage with biology. We illustrate the game design choices, the compromises and solutions that we applied to accomplish the desired learning outcomes. This paper aims to illustrate useful insights and inspirations in the context of educational game development for biology students.
Reinforcement learning (RL) approaches that combine a tree search with deep learning have found remarkable success in searching exorbitantly large, albeit discrete action spaces, as in chess, Shogi and Go. Many real-world materials discovery and design applications, however, involve multi-dimensional search problems and learning domains that have continuous action spaces. Exploring high-dimensional potential energy models of materials is an example. Traditionally, these searches are time consuming (often several years for a single bulk system) and driven by human intuition and/or expertise and more recently by global/local optimization searches that have issues with convergence and/or do not scale well with the search dimensionality. Here, in a departure from discrete action and other gradient-based approaches, we introduce a RL strategy based on decision trees that incorporates modified rewards for improved exploration, efficient sampling during playouts and a “window scaling scheme" for enhanced exploitation, to enable efficient and scalable search for continuous action space problems. Using high-dimensional artificial landscapes and control RL problems, we successfully benchmark our approach against popular global optimization schemes and state of the art policy gradient methods, respectively. We demonstrate its efficacy to parameterize potential models (physics based and high-dimensional neural networks) for 54 different elemental systems across the periodic table as well as alloys. We analyze error trends across different elements in the latent space and trace their origin to elemental structural diversity and the smoothness of the element energy surface. Broadly, our RL strategy will be applicable to many other physical science problems involving search over continuous action spaces.
Rapid progress in structural modeling of proteins and their interactions is powered by advances in knowledge-based methodologies along with better understanding of physical principles of protein struc-ture and function. The pool of structural data for modeling of proteins and protein-protein complexes is constantly increasing due to the rapid growth of protein interaction databases and Protein Data Bank. The GWYRE (Genome Wide PhYRE) project capitalizes on these developments by advancing and apply-ing new powerful modeling methodologies to structural modeling of protein-protein interactions and genetic variation. The methods integrate knowledge-based tertiary structure prediction using Phyre2 and quaternary structure prediction using template-based docking by a full-structure alignment protocol to generate models for binary complexes. The predictions are incorporated in a comprehensive public resource for structural characterization of the human interactome and the location of human genetic vari-ants. The GWYRE resource facilitates better understanding of principles of protein interaction and struc-ture/function relationships. The resource is available at http://www.gwyre.org.(c) 2022 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY license (http://creativecom-mons.org/licenses/by/4.0/).
Rita Casadio合作论文数Bologna Biocomputing Unit6