Intrinsically disordered regions (IDRs) in proteins play a pivotal role in protein-protein interactions (PPIs). Using AlphaFold2 and enriched multiple sequence alignments, we predicted and investigated PPIs across the human proteome, focusing on those involving disordered regions. Our predictions show that disordered regions predominantly interact with ordered domains, whereas predicted disordered-disordered interactions are relatively rare. Although disordered regions typically lack annotated domains, certain regions-such as the keratin type II head domain and the Krüppel-associated box (KRAB)-mediate specific interactions. In contrast, their predicted binding partners frequently feature diverse Pfam domains, including protein kinase, WD40 repeat, and nuclear hormone receptor domains. These domains are enriched in nuclear localization and α-helical repeat motifs. Disordered regions involved in predicted PPIs exhibit higher sequence conservation than non-interacting disordered regions, suggesting evolutionary constraints at interaction interfaces. Moreover, certain posttranslational modifications (e.g., phosphorylation and acetylation) in disordered regions are enriched within predicted interaction interfaces, likely modulating binding affinities. Notably, we identified a significant enrichment of disease-associated mutations in predicted PPI interfaces involving disordered regions, underscoring their functional and pathological relevance. Together, these findings highlight the intricate interplay between disordered and ordered regions in mediating PPIs and provide insights into their structural and functional contributions to human health and disease.
Large-scale computational predictions of human protein-protein interactions (PPIs) have recently enabled systematic mapping of the human interactome at near-atomic resolution. In our previous work, we reported nearly 18,000 high-confidence predicted PPIs, along with selected examples that demonstrated the ability of these predictions to reveal both pairwise associations across diverse pathways and multi-subunit complex organization. Here, we extend that study by presenting additional examples of biological and biomedical interest. We highlight novel PPIs across DNA repair, mitochondrial function, and biogenesis of cilia and other organelles, where predicted interfaces provide mechanistic hypotheses for protein function and pathway crosstalk. We also illustrate how binary predictions can be assembled into models of higher-order assemblies, uncovering new candidate subunits and organizational principles for multi-protein complexes. Importantly, we integrate structural mapping of disease-associated mutations, many of which lie directly within predicted interaction interfaces, thereby offering explanations for how genetic variation may disrupt protein networks and contribute to pathology. Together, these case studies underscore how large-scale computational predictions can be distilled into detailed mechanistic insights, generating hypotheses for functional studies and expanding the structural and functional annotation of the human interactome.
Proteins are built from modular domains that serve as fundamental units of structure and evolution. While individual domains have been extensively cataloged, their collective distribution across the lineages of life has remained poorly resolved. Here, we use the Evolutionary Classification of Protein Domains (ECOD) to chart the occurrence of domain homology groups (H-groups) across 44 model proteomes representing Eukaryota, Bacteria, and Archaea, in which 1.16 million domains are assigned to 3320 H-groups. H-groups are categorized as universal (occupying all three superkingdoms), shared between superkingdoms, or lineage-specific. The fold architecture distributions were examined: α/β sandwiches and other mixed architectures were abundant in universal H-groups, whereas α-rich architectures are expanded in eukaryotic H-groups and β-rich folds in bacterial H-groups. 126 (3.8%) H-groups occur in all organisms, forming a universal structural core that supports central processes of energy conversion, metabolism, and information flow. These widely distributed folds coincide with canonical superfolds-robust, adaptable architectures repeatedly repurposed for key biochemical roles. Two superkingdom groups trace evolutionary connections between lineages: bacterial metabolic and chaperone systems inherited by eukaryotes, archaeal informational machinery conserved in eukaryotic nuclei, and ancient redox scaffolds linking bacteria and archaea. Lineage-exclusive domains, in turn, highlight distinct adaptive strategies-regulatory and cytoskeletal innovation in eukaryotes, envelope and motility specialization in bacteria, and redox or replication refinements in archaea. Together, these data provide a quantitative, structure-based view of protein domain evolution across the tree of life, showing that the essential architecture of life relies on a conserved set of ancient folds, while lineage-specific diversity has largely arisen through the recombination and functional diversification of pre-existing domains.
Abstract Archaea represent one of the three domains of cellular life and yet account for fewer than 1% of experimentally determined protein structures, leaving the extent of their structural novelty unknown. Here we present a systematic domain-level classification of 124,075 proteins from 65 archaeal classes spanning 21 phyla and all major lineages, using both AFDB and newly predicted AlphaFold3 structures classified against the Evolutionary Classification of protein Domains (ECOD). We assigned 204,758 domains, of which 76.8% received high-confidence classifications, spanning 987 ECOD X-groups; 40% of known structural diversity within a single domain of life. Clustering by Foldseek recovered structural relationships for 63% of domains that are singletons by sequence comparison. To characterize the 21% of proteins lacking high-confidence classification, we applied successive filters for structure prediction confidence, protein length, and structural cluster context, reducing 8,452 domain-free proteins to a small number of well-folded structural orphans (less than 0.1% of the dataset). The unclassified fraction is dominated by sub-threshold matches to known folds (14% of all proteins) and low-confidence structure predictions (5%), not by novel structures. These results demonstrate that the protein fold repertoire at the single-domain level is broadly conserved across the deepest phylogenetic distances in cellular life, and that the gap between archaeal and well-characterized proteomes reflects classification sensitivity for divergent sequences rather than unexplored structural diversity.
Recent advances in protein structure prediction, such as AlphaFold2, have enabled identification of vast numbers of putative novel protein domains across the sequence space, many of which adopt structures dissimilar to known folds. Based on structural segmentation and classification, the Encyclopedia of Domains (TED) project recently cataloged more than 7400 low-symmetry, structure-based domains as candidate novel-fold (CNF) domains. To place these domains in their broader evolutionary and structural context, we applied DPAM (Domain Parser for AlphaFold Models), a complementary method that combines AlphaFold-derived confidence metrics with sensitive sequence and structure similarity searches, to parse domains for the AlphaFold models of proteins containing TED CNF domains. We identified 8044 DPAM domains with significant overlap with TED CNF domains, among which 2490 were confidently assigned to entries in the ECOD (Evolutionary Classification of protein Domains) structural classification hierarchy. Our results suggest that a substantial subset of TED candidate novel-fold domains are distant homologs of existing ECOD domains. Comparison of domain boundaries between TED and DPAM showed varied patterns: more than one-third of cases featured TED CNF domains largely embedded within DPAM domains-often representing insertions or extensions into enzymatic or repeat folds. A smaller fraction (17%) exhibited consistent domain boundaries between TED and DPAM. These consistently defined domains are often characterized by significant structural diversity, including long insertions and duplications. An even smaller subset showed the reverse relationship, with DPAM domains largely embedded within TED CNF domains. In these cases, DPAM effectively separated multiple structural units that TED grouped as single domains. Together, these findings highlight the complementarity of structural and evolutionary approaches for domain annotation and demonstrate the power of integrative methods, such as DPAM, in refining the classification of challenging protein folds and uncovering distant evolutionary relationships.
The assessment of oligomer targets in the Critical Assessment of Structure Prediction Round 16 (CASP16) suggests that complex structure prediction remains an unsolved challenge. Even the leading groups can only predict slightly more than half of the targets to high accuracy. Most CASP16 groups relied on AlphaFold-Multimer (AFM) or AlphaFold3 (AF3) as their core modeling engines. By optimizing input MSAs, refining modeling constructs (using partial rather than full sequences), and employing massive model sampling and selection, top-performing groups were able to significantly outperform the default AFM/AF3 predictions. CASP16 also introduced two additional challenges: Phase 0, which required predictions without stoichiometry information, and Phase 2, which provided participants with thousands of models generated by MassiveFold (MF) to enable large-scale sampling for resource-limited groups. Across all phases, the MULTICOM series and Kiharalab emerged as top performers based on the quality of their best models. However, these groups did not have a strong advantage in model ranking, and thus their lead over other teams, such as Yang-Multimer and kozakovvajda, was less pronounced when evaluating only the first submitted models. Compared to CASP15, CASP16 showed moderate overall improvement, likely driven by the release of AF3 and the extensive model sampling employed by top groups. Several notable trends highlight frontiers for future development. First, the kozakovvajda group significantly outperformed others on antibody-antigen targets, achieving over a 60% success rate without relying on AFM or AF3 as their primary modeling framework, suggesting that alternative approaches may offer promising solutions for these difficult targets. Second, model ranking and selection continue to be major bottlenecks. The PEZYFoldings group demonstrated a notable advantage in selecting their best models as first models, suggesting that their pipeline for model ranking may offer important insights for the field. Finally, the Phase 0 experiment indicated moderate success in stoichiometry prediction; however, stoichiometry prediction remains challenging for high-order assemblies and targets that differ from available homologous templates. Overall, CASP16 demonstrated steady progress in multimer prediction while emphasizing the need for more effective model ranking strategies, improved stoichiometry prediction, and new modeling methods that extend beyond the current AF-based paradigm.
Homology-based protein domain classification is a powerful tool for gaining biological insights into protein function. This classification process has been significantly enhanced by the availability of experimental structures and high-accuracy structural models generated by advanced tools such as AlphaFold. Our Evolutionary Classification of protein Domains (ECOD) database provides a continuously updated and refined domain classification system. Isolated ("orphan") protein domain families, which have a limited distribution in the protein universe, present a unique challenge in this classification process. These families lack clear or identifiable evolutionary relationships with other sequence families. While some isolated domain families may have emerged through de novo evolution, others potentially share common evolutionary origins with existing domain families but represent difficult cases for traditional classification methods. In this study, we conducted a manual analysis of a set of isolated families of small domains in ECOD. By exploring sequence, structural, and functional evidence, we uncovered distant members and likely homologous relationships between different isolated domain families that were previously unrecognized. Our analysis provides valuable insights into the evolution of isolated domain families and has led to improved classification within ECOD. This work enhances our understanding of protein evolution and underscores the importance of continuous refinement in domain classification systems as new data and analytical methods become available.
Protein-protein interactions (PPIs) are essential for biological function. Coevolutionary analysis and deep-learning (DL)-based protein structure prediction have enabled comprehensive PPI identification in bacteria and yeast, but these approaches have had limited success for the more complex human proteome. We overcame this challenge by enhancing the coevolutionary signals with sevenfold-deeper multiple sequence alignments harvested from 30 petabytes of unassembled genomic data and developing a new DL network trained on augmented datasets of domain-domain interactions from 200 million predicted protein structures. We systematically screened 200 million human protein pairs and predicted 17,849 interactions with an expected precision of 90%, of which 3631 interactions were not identified in previous experimental screens. Three-dimensional models of these predicted interactions provide numerous hypotheses about protein function and mechanisms of human diseases.
Synthetic circuits that regulate protein secretion in human cells could support cell-based therapies by enabling control over local environments. Although protein-level circuits enable such potential clinical applications, featuring orthogonality and compactness, their non-human origin poses a potential immunogenic risk. In this study, we developed Humanized Drug Induced Regulation of Engineered CyTokines (hDIRECT) as a platform to control cytokine activity exclusively using human-derived proteins. We sourced a specific human protease and its FDA-approved inhibitor. We engineered cytokines (IL-2, IL-6 and IL-10) whose activities can be activated and abrogated by proteolytic cleavage. We used species specificity and re-localization strategies to orthogonalize the cytokines and protease from the human context that they would be deployed in. hDIRECT should enable local cytokine activation to support a variety of cell-based therapies, such as muscle regeneration and cancer immunotherapy. Our work offers a proof of concept for the emerging appreciation of humanization in synthetic biology for human health. Engineering of a human-derived protease controlled exogenously by its FDA-approved inhibitor enables control over cytokine activity in cell-based therapies with reduced risk of immunogenicity.
Protein-protein interactions (PPI) are essential for biological function. Recent advances in coevolutionary analysis and Deep Learning (DL) based protein structure prediction have enabled comprehensive PPI identification in bacterial and yeast proteomes, but these approaches have limited success to date for the more complex human proteome. Here, we overcome this challenge by 1) enhancing the coevolutionary signals with 7-fold deeper multiple sequence alignments harvested from 30 petabytes of unassembled genomic data, and 2) developing a new DL network trained on augmented datasets of domain-domain interactions from 200 million predicted protein structures. These advancements allow us to systematically screen through 200 million human protein pairs and predict 18,316 PPIs with an expected precision of 90%, among which 5,578 are novel predictions. 3D models of these predicted PPIs nearly triple the number of human PPIs with accurate structural information, providing numerous insights into protein function and mechanisms of human diseases.
Domain classification of protein predictions released in the AlphaFold Database (AFDB) has been a recent focus of the Evolutionary Classification of protein Domains (ECOD). Although a primary focus of our recent work has been the partition and assignment of domains from these predictions, we here show how these diverse predictions can be used to examine the reference domain set more closely. Using results from DPAM, our AlphaFold-specific domain parsing algorithm, we examine hierarchical groupings that share significant levels of homologous links, both between groups that were not previously assessed to be definitively homologous and between groups that were not previously observed to share significant homologous links. Combined with manual analysis, these large datasets of structural and sequence similarities allow us to merge homologous groups in multiple cases which we detail within. These domains tend to be families of domains from families that are either small, previously had few experimental representatives, or had unknown function. The exception to this is the chromodomains, a large homologous group which were increased from "possibly homologous" to "definitely homologous" to increase the consistency of ECOD based their strong homologous links to the SH3 domains.
The classification of novel protein folds remains a central challenge in structural bioinformatics, particularly as deep learning models like AlphaFold2 dramatically expand the universe of predicted protein structures. In this study, we investigated 664 candidate novel fold (CNF) domains from the TED database that both TED and DPAM methods had classified with low confidence. These CNFs span a structurally diverse and largely non-redundant set of domains, most of which lack clear sequence or structural similarity to known folds. Many CNFs appear as insertions into known transmembrane or enzymatic domains, while others occur in modular architectures, co-occurring with interaction or catalytic folds such as β-barrels, zinc fingers, or Rossmann-like domains. Although some CNFs resemble known folds that have undergone topological rearrangements or circular permutations, others result from errors in domain boundary prediction, often due to truncated sequences or tightly packed domain duplications. Our analyses led to the creation of 190 new Pfam families, many classified as domains of unknown function (DUFs), and revealed intriguing cases of zinc-binding and disulfide-rich architectures that contribute to fold space expansion. A small subset of CNFs helped define new superfamilies by linking previously unclassified but structurally related domains. Taken together, this work underscores the importance of integrating structural, evolutionary, and contextual information to resolve challenging fold assignments and provides a roadmap for extending protein classification frameworks into previously uncharted structural territory.
Targeting Meis1 and Hoxb13 transcriptional activity could be a viable therapeutic strategy for heart regeneration. In this study, we performd an in silico screening to identify FDA-approved drugs that can inhibit Meis1 and Hoxb13 transcriptional activity based on the resolved crystal structure of Meis1 and Hoxb13 bound to DNA. Paromomycin (Paro) and neomycin (Neo) induced proliferation of neonatal rat ventricular myocytes in vitro and displayed dose-dependent inhibition of Meis1 and Hoxb13 transcriptional activity by luciferase assay and disruption of DNA binding by electromobility shift assay. X-ray crystal structure revealed that both Paro and Neo bind to Meis1 near the Hoxb13-interacting domain. Administration of Paro-Neo combination in adult mice and in pigs after cardiac ischemia/reperfusion injury induced cardiomyocyte proliferation, improved left ventricular systolic function and decreased scar formation. Collectively, we identified FDA-approved drugs with therapeutic potential for induction of heart regeneration in mammals.
Identification of bacterial protein-protein interactions and predicting the structures of these complexes could aid in the understanding of pathogenicity mechanisms and developing treatments for infectious diseases. Here we developed RoseTTAFold2-Lite, a rapid deep learning model that leverages residue-residue coevolution and protein structure prediction to systematically identify and structurally characterize protein-protein interactions at the proteome-wide scale. Using this pipeline, we searched through 78 million pairs of proteins across 19 human bacterial pathogens and identified 1,923 confidently predicted complexes involving essential genes and 256 involving virulence factors. Many of these complexes were not previously known; we experimentally tested 12 such predictions, and half of them were validated. The predicted interactions span core metabolic and virulence pathways ranging from post-transcriptional modification to acid neutralization to outer-membrane machinery and should contribute to our understanding of the biology of these important pathogens and the design of drugs to combat them.
Classification of protein domains based on homology and structural similarity serves as a fundamental tool to gain biological insights into protein function. Recent advancements in protein structure prediction, exemplified by AlphaFold, have revolutionized the availability of protein structural data. We focus on classifying about 9000 Pfam families into ECOD (Evolutionary Classification of Domains) by using predicted AlphaFold models and the DPAM (Domain Parser for AlphaFold Models) tool. Our results offer insights into their homologous relationships and domain boundaries. More than half of these Pfam families contain DPAM domains that can be confidently assigned to the ECOD hierarchy. Most assigned domains belong to highly populated folds such as Immunoglobulin-like (IgL), Armadillo (ARM), helix-turn-helix (HTH), and Src homology 3 (SH3). A large fraction of DPAM domains, however, cannot be confidently assigned to ECOD homologous groups. These unassigned domains exhibit statistically different characteristics, including shorter average length, fewer secondary structure elements, and more abundant transmembrane segments. They could potentially define novel families remotely related to domains with known structures or novel superfamilies and folds. Manual scrutiny of a subset of these domains revealed an abundance of internal duplications and recurring structural motifs. Exploring sequence and structural features such as disulfide bond patterns, metal-binding sites, and enzyme active sites helped uncover novel structural folds as well as remote evolutionary relationships. By bridging the gap between sequence-based Pfam and structure-based ECOD domain classifications, our study contributes to a more comprehensive understanding of the protein universe by providing structural and functional insights into previously uncharacterized proteins.
Protein-protein interactions (PPI) are essential for biological function. Recent advances in coevolutionary analysis and Deep Learning (DL) based protein structure prediction have enabled comprehensive PPI identification in bacterial and yeast proteomes, but these approaches have limited success to date for the more complex human proteome. Here, we overcome this challenge by 1) enhancing the coevolutionary signals with 7-fold deeper multiple sequence alignments harvested from 30 petabytes of unassembled genomic data, and 2) developing a new DL network trained on augmented datasets of domain-domain interactions from 200 million predicted protein structures. These advancements allow us to systematically screen through 200 million human protein pairs and predict 18,316 PPIs with an expected precision of 90%, among which 5,578 are novel predictions. 3D models of these predicted PPIs nearly triple the number of human PPIs with accurate structural information, providing numerous insights into protein function and mechanisms of human diseases. ### Competing Interest Statement The authors have declared no competing interest.
Human cytochrome c oxidase (COX) is composed of fourteen subunits, only two of which are catalytic. The functions of a myriad of accessory factors are needed to put together the structural subunits, as well as for the formation of the active centers. The molecular roles of most of these factors are still unknown, even if they are essential for COX function since their defects cause severe mitochondrial disease. COA8 is exclusively found in metazoa and loss-of-function genetic variants cause mitochondrial encephalopathy and COX deficiency in humans. In this work, we have resolved the function of COA8, as a factor that specifically binds to COX10, the heme O synthase. This interaction is necessary not only for the stabilization of human COX10, but also to guarantee its catalytic activity as the first step for the synthesis of the heme A cofactor, which is a crucial component of the COX catalytic center.
A propeptide is removed from a precursor protein to generate its active or mature form. Propeptides play essential roles in protein folding, transportation, and activation and are present in about 2.3% of reviewed proteins in the UniProt database. They are often found in secreted or membrane-bound proteins including proteolytic enzymes, hormones, and toxins. We identified a variety of globular and nonglobular Pfam domains in protein sequences designated as propeptides, some of which form intramolecular interactions with other domains in the mature proteins. Propeptide-containing enzymes mostly function as proteases, as they are depleted in other enzyme classes such as hydrolases acting on DNA and RNA, isomerases, and lyases. We applied AlphaFold to generate structural models for over 7000 proteins with propeptides having no less than 20 residues. Analysis of residue contacts in these models revealed conformational changes for over 300 proteins before and after the cleavage of the propeptide. Examples of conformation change occur in several classes of proteolytic enzymes in the families of subtilisins, trypsins, aspartyl proteases, and thermolysin-like metalloproteases. In most of the observed cases, cleavage of the propeptide releases the constraints imposed by the covalent bond between the propeptide and the mature protein, and cleavage enables stronger interactions between the propeptide and the mature protein. These findings suggest that post-cleavage propeptides could play critical roles in regulating the activity of mature proteins.
Adolescent idiopathic scoliosis (AIS) is a common and progressive spinal deformity in children that exhibits striking sexual dimorphism, with girls at more than five-fold greater risk of severe disease compared to boys. Despite its medical impact, the molecular mechanisms that drive AIS are largely unknown. We previously defined a female-specific AIS genetic risk locus in an enhancer near the PAX1 gene. Here we sought to define the roles of PAX1 and newly-identified AIS-associated genes in the developmental mechanism of AIS. In a genetic study of 10,519 individuals with AIS and 93,238 unaffected controls, significant association was identified with a variant in COL11A1 encoding collagen (α1) XI (rs3753841; NM_080629.2_c.4004C>T; p.(Pro1335Leu); P=7.07e-11, OR=1.118). Using CRISPR mutagenesis we generated Pax1 knockout mice (Pax1-/-). In postnatal spines we found that PAX1 and collagen (α1) XI protein both localize within the intervertebral disc (IVD)-vertebral junction region encompassing the growth plate, with less collagen (α1) XI detected in Pax1-/- spines compared to wildtype. By genetic targeting we found that wildtype Col11a1 expression in costal chondrocytes suppresses expression of Pax1 and of Mmp3, encoding the matrix metalloproteinase 3 enzyme implicated in matrix remodeling. However, this suppression was abrogated in the presence of the AIS-associated COL11A1P1335L mutant. Further, we found that either knockdown of the estrogen receptor gene Esr2, or tamoxifen treatment, significantly altered Col11a1 and Mmp3 expression in chondrocytes. We propose a new molecular model of AIS pathogenesis wherein genetic variation and estrogen signaling increase disease susceptibility by altering a Pax1-Col11a1-Mmp3 signaling axis in spinal chondrocytes.
Identification of bacterial protein-protein interactions and predicting the structures of the complexes could aid in the understanding of pathogenicity mechanisms and developing treatments for infectious diseases. Here, we developed a deep learning-based pipeline that leverages residue-residue coevolution and protein structure prediction to systematically identify and structurally characterize protein-protein interactions at the proteome-wide scale. Using this pipeline, we searched through 78 million pairs of proteins across 19 human bacterial pathogens and identified 1923 confidently predicted complexes involving essential genes and 256 involving virulence factors. Many of these complexes were not previously known; we experimentally tested 12 such predictions, and half of them were validated. The predicted interactions span core metabolic and virulence pathways ranging from post-transcriptional modification to acid neutralization to outer membrane machinery and should contribute to our understanding of the biology of these important pathogens and the design of drugs to combat them.