Understanding the structures of protein complexes is pivotal for breakthroughs in health, agriculture, bioengineering, and beyond. MultiFOLD2 and ModFOLDdock2 are leading servers for protein quaternary structure prediction and model quality assessment, respectively. MultiFOLD2 includes integrated stoichiometry prediction for quaternary structures and improved sampling and scoring, leading to high performance in continuous independent benchmarks such as CAMEO. ModFOLDdock2 uses a hybrid consensus approach to generate global and local quality scores for predicted quaternary structures. ModFOLDdock2 is integrated with MultiFOLD2 while also being available as a stand-alone server, enabling the independent evaluation of quaternary structure models from any source. Both servers have been independently rigorously evaluated, demonstrating high performance and ranking among the top servers in their respective categories in the recent CASP16 experiment. The MultiFOLD2 and ModFOLDdock2 servers are freely accessible through user-friendly web interfaces at https://www.reading.ac.uk/bioinf/.
Connexin (Cx) gap junction proteins are expressed by a multitude of cells and function as plasma membrane hemichannels or dock to form intercellular communication tunnels. Whilst Cx43 has garnered considerable attention, less is known about the structure and function of Cx62 channels. Platelets and megakaryocytes express Cx37, Cx40 and Cx62, which contribute to hemostatic and thrombotic responses. Our study explores an unexpected finding that following platelet activation, an extracellular region of Cx62 undergoes proteolytic cleavage by calpain-1. We adopted an interdisciplinary approach to evaluate structural and functional consequences of calpain-mediated cleavage of Cx62. Cellular signaling was assayed by immunoblotting, aggregation and calcium flux assays. Gap junction function and thrombus formation were assessed under arteriolar flow. In silico modeling was used to predict calpain-mediated changes to the pore diameter and design a decoy peptide (62Pept-NT). Mechanistically, Cx62 cleavage is Ca2+-dependent and requires calpain-1 externalization. Modeling a predicted calpain-1 cleavage site on the first extracellular loop, shows that calpain can dock to Cx62 monomers, promoting stepwise channel cleavage. Consequently, we predict a significant pore dilation enhancing diffusion of signaling molecules between cells and into the extracellular milieu. We designed a decoy peptide that abrogated calpain-1-mediated cleavage, reduced intercellular communication and restricted thrombus growth. Cx62 cleavage was dependent upon sequential action of protein kinase A, protein phosphatase 2A and Ca2+ release from intracellular stores. Extracellular calpain cleavage represents a fundamentally new regulatory mechanism for Cx62, culminating in an irreversible open state.
Model quality assessment (MQA) remains a critical component of structural bioinformatics for both structure predictors and experimentalists seeking to use predictions for downstream applications. In CASP16, the Evaluation of Model Accuracy (EMA) category featured both global and local quality estimation for multimeric assemblies (QMODE1 and QMODE2), as well as a novel QMODE3 challenge-requiring predictors to identify the best five models from thousands generated by MassiveFold. This paper presents detailed results from several leading CASP16 EMA methods, highlighting the strengths and limitations of the approaches.
Protein structure prediction is fundamental to molecular biology and has numerous applications in areas such as drug discovery and protein engineering. Machine learning techniques have greatly advanced protein 3D modeling in recent years, particularly with the development of AlphaFold2 (AF2), which can analyze sequences of amino acids and predict 3D structures with near experimental accuracy. Since the release of AF2, numerous studies have been conducted, either using AF2 directly for large-scale modeling or building upon the software for other use cases. Many reviews have been published discussing the impact of AF2 in the field of protein bioinformatics, particularly in relation to neural networks, which have highlighted what AF2 can and cannot do. It is evident that AF2 and similar approaches are open to further development and several new approaches have emerged, in addition to older refinement approaches, for improving the quality of predictions. Here we provide a brief overview, aimed at the general biologist, of how machine learning techniques have been used for improvement of 3D models of proteins following AF2, and we highlight the impacts of these approaches. In the most recent experiment on the Critical Assessment of Techniques for Protein Structure Prediction (CASP15), the most successful groups all developed their own tools for protein structure modeling that were based at least in some part on AF2. This improvement involved employing techniques such as generative modeling, changing parameters such as dropout to generate more AF2 structures, and data-driven approaches including using alternative templates and MSAs.
Accurate models of protein tertiary structures are now available from numerous advanced prediction methods, although the accuracy of each method often varies depending on the specific protein target. Additionally, many models may still contain significant local errors. Therefore, reliable, independent model quality estimates are essential both for identifying errors and selecting the very best models for further biological investigations. ModFOLD9 is a leading independent server for detecting the local errors in models produced by any method, and it can accurately discriminate between high-quality models from multiple alternative approaches. ModFOLD9 incorporates several new scores from deep learning-based approaches, leading to greatly improved prediction accuracy compared with earlier versions of the server. ModFOLD9 is continuously independently benchmarked, and it is shown to be highly competitive with other public servers. ModFOLD9 is freely available at https://www.reading.ac.uk/bioinf/ModFOLD/.
ABSTRACT Motivation Despite an increase in the accuracy of predicted protein structures following the development of AlphaFold2, there remains a gap in the accuracy of predicted model quality assessment scores when compared to those generated with reference to experimental structures. The predictions of model accuracy scores generated by AlphaFold2, plDDT and pTM, have become familiar descriptors of model quality. However, at CASP15 some modelling groups noticed a variation in these scores for models of very similar observed quality, particularly for quaternary structures. There have also been a number of methods describing adaptations of the AlphaFold2 algorithm to purposes such as refinement by custom template recycling and model quality assessment using a similar method of template input. In this study we compare plDDT and pTM to their observed counterparts lDDT (including lDDT-Cα and lDDT-oligo) and TM-score to examine whether they retain their reliability across the whole scoring range for both tertiary and quaternary structures and in situations where the AlphaFold2 algorithm is adapted to customised functionality. In addition, we explore the accuracy with which plDDT and pTM rank AlphaFold2 tertiary and quaternary models and whether these can be improved by the independent model quality assessment programs ModFOLD9 and ModFOLDdock. Results For tertiary structures it was found that plDDT was an accurate descriptor of model quality when compared to observed lDDT-Cα scores (Pearson ρ = 0.97). Additionally, plDDT achieved a tertiary structure ranking agreement with observed scores of 0.34 as measured by true positive rate (TPR) and ModFOLD9 offered similar but not improved performance. However, the accuracy of plDDT (Pearson ρ = 0.67) and pTM (Pearson ρ = 0.70) became more variable for quaternary structures quality assessment where overprediction was seen with both scores for models of lower quality and underprediction was also seen with pTM for models of higher quality. Importantly, ModFOLDdock was able to improve upon AF2-Multimer quaternary structure model ranking as measured by both TM-score (TPR 0.34) and lDDT-oligo (TPR 0.43). Finally, evidence is presented for an increase in variability of both plDDT and pTM when custom template recycling is used, and that this variation is more pronounced for quaternary structures.
Postnatal growth failure is often attributed to dysregulated somatotropin action, however marked genetic and phenotypic heterogeneity exist. We report five patients from three families who present with short stature, immune dysfunction, atopic eczema and gastrointestinal pathology associated with recessive variants in QSOX2. QSOX2 encodes a nuclear membrane protein linked to disulphide isomerase and oxidoreductase activity. Loss of QSOX2 disrupts Growth hormone-mediated STAT5B nuclear translocation despite enhanced Growth hormone-induced STAT5B phosphorylation. Moreover, patient-derived dermal fibroblasts demonstrate Growth hormone-induced mitochondriopathy and reduced mitochondrial membrane potential. Located at the nuclear membrane, QSOX2 acts as a gatekeeper for regulating stabilisation and import of phosphorylated-STAT5B. Altogether, QSOX2 deficiency modulates human growth by impairing Growth hormone-STAT5B downstream activities and mitochondrial dynamics, which contribute to multi-system dysfunction. Furthermore, our work suggests that therapeutic recombinant insulin-like growth factor-1 may circumvent the Growth hormone-STAT5B dysregulation induced by pathological QSOX2 variants and potentially alleviate organ specific disease. Defects in growth hormone (GH) action account for a substantial percentage of endocrine causes of growth failure. Here, the authors report that QSOX2 deficiency modulates human growth by impairing GH-STAT5B downstream activities and mitochondrial dynamics, contributing to multi-system dysfunction.
In CASP15, there was a greater emphasis on multimeric modeling than in previous experiments, with assembly structures nearly doubling in number (41 up from 22) since the previous round. CASP15 also included a new estimation of model accuracy (EMA) category in recognition of the importance of objective quality assessment (QA) for quaternary structure models. ModFOLDdock is a multimeric model QA server developed by the McGuffin group at the University of Reading, which brings together a range of single-model, clustering, and deep learning methods to form a consensus of approaches. For CASP15, three variants of ModFOLDdock were developed to optimize for the different facets of the quality estimation problem. The standard ModFOLDdock variant produced predicted scores optimized for positive linear correlations with the observed scores. The ModFOLDdockR variant produced predicted scores optimized for ranking, that is, the top-ranked models have the highest accuracy. In addition, the ModFOLDdockS variant used a quasi-single model approach to score each model on an individual basis. The scores from all three variants achieved strongly positive Pearson correlation coefficients with the CASP observed scores (oligo-lDDT) in excess of 0.70, which were maintained across both homomeric and heteromeric model populations. In addition, at least one of the ModFOLDdock variants was consistently ranked in the top two methods across all three EMA categories. Specifically, for overall global fold prediction accuracy, ModFOLDdock placed second and ModFOLDdockR placed third; for overall interface quality prediction accuracy, ModFOLDdockR, ModFOLDdock, and ModFOLDdockS were placed above all other predictor methods, and ModFOLDdockR and ModFOLDdockS were placed second and third respectively for individual residue confidence scores. The ModFOLDdock server is available at: https://www.reading.ac.uk/bioinf/ModFOLDdock/. ModFOLDdock is also available as part of the MultiFOLD docker package: https://hub.docker.com/r/mcguffin/multifold.
Abstract The IntFOLD server based at the University of Reading has been a leading method over the past decade in providing free access to accurate prediction of protein structures and functions. In a post-AlphaFold2 world, accurate models of tertiary structures are widely available for even more protein targets, so there has been a refocus in the prediction community towards the accurate modelling of protein-ligand interactions as well as modelling quaternary structure assemblies. In this paper, we describe the latest improvements to IntFOLD, which maintains its competitive structure prediction performance by including the latest deep learning methods while also integrating accurate model quality estimates and 3D models of protein-ligand interactions. Furthermore, we also introduce our two new server methods: MultiFOLD for accurately modelling both tertiary and quaternary structures, with performance which has been independently verified to outperform the standard AlphaFold2 methods, and ModFOLDdock, which provides world-leading quality estimates for quaternary structure models. The IntFOLD7, MultiFOLD and ModFOLDdock servers are available at: https://www.reading.ac.uk/bioinf/.
The refinement of predicted 3D models aims to bring them closer to the native structure by fixing errors including unusual bonds and torsion angles and irregular hydrogen bonding patterns. Refinement approaches based on molecular dynamics (MD) simulations using different types of restraints have performed well since CASP10. ReFOLD, developed by the McGuffin group, was one of the many MD-based refinement approaches, which were tested in CASP 12. When the performance of the ReFOLD method in CASP12 was evaluated, it was observed that ReFOLD suffered from the absence of a reliable guidance mechanism to reach consistent improvement for the quality of predicted 3D models, particularly in the case of template-based modelling (TBM) targets. Therefore, here we propose to utilize the local quality assessment score produced by ModFOLD6 to guide the MD-based refinement approach to further increase the accuracy of the predicted 3D models. The relative performance of the new local quality assessment guided MD-based refinement protocol and the original MD-based protocol ReFOLD are compared utilizing many different official scoring methods. By using the per-residue accuracy (or local quality) score to guide the refinement process, we are able to prevent the refined models from undesired structural deviations, thereby leading to more consistent improvements. This chapter will include a detailed analysis of the performance of the local quality assessment guided MD-based protocol versus that deployed in the original ReFOLD method.
Protein structure modeling is one of the most advanced and complex processes in computational biology. One of the major problems for the protein structure prediction field has been how to estimate the accuracy of the predicted 3D models, on both a local and global level, in the absence of known structures. We must be able to accurately measure the confidence that we have in the quality predicted 3D models of proteins for them to become widely adopted by the general bioscience community. To address this major issue, it was necessary to develop new model quality assessment (MQA) methods and integrate them into our pipelines for building 3D protein models. Our MQA method, called ModFOLD, has been ranked as one of the most accurate MQA tools in independent blind evaluations. This chapter discusses model quality assessment in the protein modeling field, demonstrating both its strengths and limitations. We also present some of the best methods according to independent benchmarking data, which has been gathered in recent years.
Motivation The accuracy gap between predicted and experimental structures has been significantly reduced following the development of AlphaFold2. However, for further studies, such as drug discovery and protein design, AlphaFold2 structures need to be representative of proteins in solution, yet AlphaFold2 was trained to generate only a few structural conformations rather than a conformational landscape. In previous CASP experiments, MD simulation-based methods have been widely used to improve the accuracy of single 3D models. However, these methods are highly computationally intensive and less applicable for practical use in large-scale applications. Despite this, the refinement concept can still provide a better understanding of conformational dynamics and improve the quality of 3D models at a modest computational cost. Here, our ReFOLD4 pipeline was adopted to provide the conformational landscape of AlphaFold2 predictions while maintaining high model accuracy. In addition, the AlphaFold2 recycling process was utilised to improve 3D models by using them as custom template inputs for tertiary and quaternary structure predictions. Results According to the Molprobity score, 94% of the generated 3D models by ReFOLD4 were improved. As measured by average change in lDDT, AlphaFold2 recycling showed an improvement rate of 87.5% (using MSAs) and 81.25% (using single sequences) for monomeric AF2 models and 100% (MSA) and 97.8% (single sequence) for monomeric non-AF2 models. By the same measure, the recycling of multimeric models showed an improvement rate of as much as 80% for AF2 models and 94% for non-AF2 models. The AlphaFold2 recycling processes and ReFOLD4 method can be combined very efficiently to provide conformational landscapes at the AlphaFold2-accuracy level, while also significantly improving the global quality of 3D models for both tertiary and quaternary structures, with much less computational complexity than traditional refinement methods.
The previous studies on the RGD motif (aa403-405) within the SARS CoV-2 spike (S) protein receptor binding domain (RBD) suggest that the RGD motif binding integrin(s) may play an important role in infection of the host cells. We also discussed the possible role of two other integrin binding motifs that are present in S protein: LDI (aa585-587) and ECD (661-663), the motifs used by some other viruses in the course of infection. The MultiFOLD models for protein structure analysis have shown that the ECD motif is clearly accessible in the S protein, whereas the RGD and LDI motifs are partially accessible. Furthermore, the amino acids that are present in Epstein-Barr virus protein (EBV) gp42 playing very important role in binding to the HLA-DRB1 molecule and in the subsequent immune response evasion, are also present in the S protein heptad repeat-2. Our MultiFOLD model analyses have shown that these amino acids are clearly accessible on the surface in each S protein chain as monomers and in the homotrimer complex and bind to HLA-DRB1 β chain. Therefore, they may have the identical role in SARS CoV-2 immune evasion as in EBV infection. The prediction analyses of the MHC class II binding peptides within the S protein have shown that the RGD motif is present in the core 9-mer peptide IRGDEVRQI within the two HLA-DRB1*03:01 and HLA-DRB3*01.01 strong binding 15-mer peptides suggesting that RGD motif may be the potential immune epitope. Accordingly, infected HLA-DRB1*03:01 or HLA-DRB3*01.01 positive individuals may develop high affinity anti-RGD motif antibodies that react with the RGD motif in the host proteins, like fibrinogen, thrombin or von Willebrand factor, affecting haemostasis or participating in autoimmune disorders.
Motivation The accuracy gap between predicted and experimental structures has been significantly reduced following the development of AlphaFold2. However, for further studies, such as drug discovery and protein design, AlphaFold2 structures need to be representative of proteins in solution, yet AlphaFold2 was trained to generate only a few structural conformations rather than a conformational landscape. In previous CASP experiments, MD simulation-based methods have been widely used to improve the accuracy of single 3D models. However, these methods are highly computationally intensive and less applicable for practical use in large-scale applications. Despite this, the refinement concept can still provide a better understanding of conformational dynamics and improve the quality of 3D models at a modest computational cost. Here, our ReFOLD4 pipeline was adopted to provide the conformational landscape of AlphaFold2 predictions while maintaining high model accuracy. In addition, the AlphaFold2 recycling process was utilised to improve 3D models by using them as custom template inputs for tertiary and quaternary structure predictions. Results According to the Molprobity score, 94% of the generated 3D models by ReFOLD4 were improved. As measured by average change in lDDT, AlphaFold2 recycling showed an improvement rate of 87.5% (using MSAs) and 81.25% (using single sequences) for monomeric AF2 models and 100% (MSA) and 97.8% (single sequence) for monomeric non-AF2 models. By the same measure, the recycling of multimeric models showed an improvement rate of as much as 80% for AF2 models and 94% for non-AF2 models. The AlphaFold2 recycling processes and ReFOLD4 method can be combined very efficiently to provide conformational landscapes at the AlphaFold2-accuracy level, while also significantly improving the global quality of 3D models for both tertiary and quaternary structures, with much less computational complexity than traditional refinement methods. ### Competing Interest Statement The authors have declared no competing interest.
Abstract ReFOLD3 is unique in its application of gradual restraints, calculated from local model quality estimates and contact predictions, which are used to guide the refinement of theoretical 3D protein models towards the native structures. ReFOLD3 achieves improved performance by using an iterative refinement protocol to fix incorrect residue contacts and local errors, including unusual bonds and angles, which are identified in the submitted models by our leading ModFOLD8 model quality assessment method. Following refinement, the likely resulting improvements to the submitted models are recognized by ModFOLD8, which produces both global and local quality estimates. During the CASP14 prediction season (May–Aug 2020), we used the ReFOLD3 protocol to refine hundreds of 3D models, for both the refinement and the main tertiary structure prediction categories. Our group improved the global and local quality scores for numerous starting models in the refinement category, where we ranked in the top 10 according to the official assessment. The ReFOLD3 protocol was also used for the refinement of the SARS-CoV-2 targets as a part of the CASP Commons COVID-19 initiative, and we provided a significant number of the top 10 models. The ReFOLD3 web server is freely available at https://www.reading.ac.uk/bioinf/ReFOLD/.
The Ser/Thr kinase MAP4K4, like other GCKIV kinases, has N-terminal kinase and C-terminal citron homology (CNH) domains. MAP4K4 can activate c-Jun N-terminal kinases (JNKs), and studies in the heart suggest it links oxidative stress to JNKs and heart failure. In other systems, MAP4K4 is regulated in striatin-interacting phosphatase and kinase (STRIPAK) complexes, in which one of three striatins tethers PP2A adjacent to a kinase to keep it dephosphorylated and inactive. Our aim was to understand how MAP4K4 is regulated in cardiomyocytes. The rat MAP4K4 gene was not properly defined. We identified the first coding exon of the rat gene using 5′-RACE, we cloned the full-length sequence and confirmed alternative-splicing of MAP4K4 in rat cardiomyocytes. We identified an additional α-helix C-terminal to the kinase domain important for kinase activity. In further studies, FLAG-MAP4K4 was expressed in HEK293 cells or cardiomyocytes. The Ser/Thr protein phosphatase inhibitor calyculin A (CalA) induced MAP4K4 hyperphosphorylation, with phosphorylation of the activation loop and extensive phosphorylation of the linker between the kinase and CNH domains. This required kinase activity. MAP4K4 associated with myosin in untreated cardiomyocytes, and this was lost with CalA-treatment. FLAG-MAP4K4 associated with all three striatins in cardiomyocytes, indicative of regulation within STRIPAK complexes and consistent with activation by CalA. Computational analysis suggested the interaction was direct and mediated via coiled-coil domains. Surprisingly, FLAG-MAP4K4 inhibited JNK activation by H2O2 in cardiomyocytes and increased myofibrillar organisation. Our data identify MAP4K4 as a STRIPAK-regulated kinase in cardiomyocytes, and suggest it regulates the cytoskeleton rather than activates JNKs.
Methods for estimating the quality of 3D models of proteins are vital tools for driving the acceptance and utility of predicted tertiary structures by the wider bioscience community. Here we describe the significant major updates to ModFOLD, which has maintained its position as a leading server for the prediction of global and local quality of 3D protein models, over the past decade (>20 000 unique external users). ModFOLD8 is the latest version of the server, which combines the strengths of multiple pure-single and quasi-single model methods. Improvements have been made to the web server interface and there has been successive increases in prediction accuracy, which were achieved through integration of newly developed scoring methods and advanced deep learning-based residue contact predictions. Each version of the ModFOLD server has been independently blind tested in the biennial CASP experiments, as well as being continuously evaluated via the CAMEO project. In CASP13 and CASP14, the ModFOLD7 and ModFOLD8 variants ranked among the top 10 quality estimation methods according to almost every official analysis. Prior to CASP14, ModFOLD8 was also applied for the evaluation of SARS-CoV-2 protein models as part of CASP Commons 2020 initiative. The ModFOLD8 server is freely available at: https://www.reading.ac.uk/bioinf/ModFOLD/.