Establishing a coherent mapping of the relationships among all known proteins is crucial for elucidating processes of protein emergence and evolution. Yet the capacity to fully capture relationships of protein similarity is complicated by the nonstraightforward interplay between sequence and structure; indeed, proteins with unrelated sequences can adopt similar structures, and, conversely, proteins with similar or identical sequences can manifest radically different structures. Here, we introduce Contrastive Learning Sequence-Structure (CLSS), a contrastive protein language model (PLM) trained to coembed sequence and structure information in a self-supervised manner, facilitating a holistic representation of protein relatedness. CLSS represents the structures and sequences of full domains and domain subsequences as vectors in the same high-dimensional latent space. We show that this approach yields meaningful shared representations, which recapitulate the extensive structure- and sequence-based knowledge encoded in human-curated hierarchical protein classification systems (ECOD and CATH). Moreover, the representations generated by CLSS outperform those generated by alternative state-of-the-art PLMs in downstream classification tasks. Notably, we show that even the far larger space of domain subsequences is successfully coembedded, establishing a PLM tailored to these evolutionarily meaningful objects. CLSS embeddings produce informative representations of the protein universe without further downstream processing, as we demonstrate by analyzing preferential associations between protein architectures and ligand types across protein space.
Aminoacyl-tRNA synthetases are the guardians of translational fidelity. Their complex function is mirrored by an elaborate structure, which includes multiple nested domains. While the evolutionary pressures that promoted the emergence of some domains, such as the editing domain, are clear, the pressures acting on other domains, particularly those at the C-terminus, are not. Here, we use a combination of kinetic analysis, X-ray crystallography, and bioinformatics to unveil the history and evolutionary forces that have shaped isoleucyl-tRNA synthetase (IleRS) domain structure. We find that the traditional classification into IleRS1 and IleRS2, based on the C-terminal tRNA-recognition domains, is incomplete, as it fails to capture features of the synthetic domain. Guided by the crystal structure of the Priestia megaterium IleRS2:tRNA complex, we removed key interactions between IleRS2 and its cognate tRNA and characterised their impact on enzyme activity. We found that D-loop interactions with the IleRS2 C-terminal region are non-essential in prokaryotes, and their loss can even increase catalytic turnover. Further, the zinc-binding domain of IleRS1 recognises the anticodon less stringently than the canonical C-terminal domain of IleRS2. Our data suggest that C-terminal evolutionary remodelling of IleRSs is an ongoing process with a historical precedent, consistent with selection for faster aminoacylation rate. ### Competing Interest Statement The authors have declared no competing interest. Croatian Science Foundation, https://ror.org/03n51vw80, IP-2022-10-1400 Swiss National Science Foundation, https://ror.org/00yjd3n13, 310030E_215868 Human Frontier Science Program, RGEC29/2025
Whereas phylogenetic reconstructions are a primary record of protein evolution, it is unknown whether the deep history of enzymes is encoded at higher levels of biological organization. Here, we demonstrate that the emergence and reuse history of enzymatic folds is embedded within the web of metabolite-cofactor-enzyme interdependencies that comprise biosphere-scale metabolic reaction networks. Using a simple network analysis approach, we reconstruct the relative ordering of enzymatic fold emergence and, where possible, the first reaction(s) that each enzymatic fold catalyzed. We find that a large majority of enzymatic folds were sufficient as independent additions to open new avenues for metabolic growth. The resulting network-based histories are broadly concordant with enzyme phyletic distribution in prokaryotes, a proxy for enzyme age. Our results suggest that the earliest enzyme-mediated metabolisms were enriched for α/β proteins, likely due to their strong association with cofactor utilization, and that α-proteins preferentially emerge at later stages. The cradle-loop barrel, a member of the small β-barrel metafold, is predicted to be the founding β-fold, in agreement with analyses of ribosome structure. An examination of how the protein universe responded to the biological production of molecular oxygen reveals that the adaptation of existing enzymatic folds, not novel fold emergence, was the primary driver of metabolic evolution. This work presents a self-consistent model of metabolic and enzyme evolution, key progress toward integrating diverse perspectives into a unified history of protein evolution.
Amino acid sequence dictates the three-dimensional structure and biological function of proteins. Yet, despite decades of research, our understanding of the interplay between sequence and structure is incomplete. To meet this challenge, we introduce Contrastive Learning Sequence-Structure (CLSS), an AI-based contrastive learning model trained to co-embed sequence and structure information in a self-supervised manner. We trained CLSS on large and diverse sets of protein building blocks called domains. CLSS represents both sequences and structures as vectors in the same high-dimensional space, where distance relates to sequence-structure similarity. Thus, CLSS provides a natural way to represent the protein universe, reflecting evolutionary relationships, as well as structural changes. We find that CLSS refines expert knowledge about the global organization of protein space, and highlights transitional forms that resist hierarchical classification. CLSS reveals linkage between domains of seemingly separate lineages, thereby significantly improving our understanding of evolutionary design. ### Competing Interest Statement The authors have declared no competing interest. Israel Science Foundation, https://ror.org/04sazxf24, 1764/21 HSFP, RGEC29/2025
Primitive nucleic acids and peptides likely collaborated in early biochemistry. What forces drove their interactions and how did these forces shape the properties of primitive complexes? We investigated how two model primordial polypeptides associate with DNA. When peptides were coupled to a ferromagnetic substrate, DNA binding depended on the substrate's magnetic moment orientation. Reversing the magnetic field nearly abolished binding despite complementary charges. Inverting the peptide chirality or just the cysteine residue reversed this effect. These results are attributed to the chiral-induced spin selectivity (CISS) effect, where molecular chirality and electron spin alter a protein's electric polarizability. The presence of CISS in simple protein-DNA complexes suggests that it played a significant role in ancient biomolecular interactions. A major consequence of CISS is enhancement of the kinetic stability of protein-nucleic acid complexes. These findings reveal how chirality and spin influence bioassociation, offering insights into primitive biochemical evolution and shaping contemporary protein functions.
Phylogenetic reconstructions are a primary record of protein evolution. But what other records can attest to the deep history of enzymes, and what tools are needed to decode their meaning? Here, we demonstrate that the history of enzyme discovery and reuse is embedded within the web of interdependencies that constitute contemporary, biosphere-scale metabolism. Using a simple network analysis approach, we reconstruct both the relative temporal ordering of enzyme domain emergence and, where possible, the first reactions that they catalyzed. These network-based histories were found to be broadly concordant with phyletic information, suggesting that the two approaches reflect related generative processes. When enzyme emergence is initiated after the discovery of nucleotide cofactors, a predominantly stepwise trajectory of domain discovery is recovered. We find that the earliest enzyme-mediated metabolisms were dominated by α/β domains, likely due to their high discoverability and functional potential under constraint. Finally, we quantify how the protein universe responded to a major transition, the biological production of molecular oxygen, by preferentially reusing pre-existing enzyme domains. This work presents a self-consistent model of metabolic and enzyme evolution, essential progress towards integrating multiple, independent records into a unified history of protein evolution. ### Competing Interest Statement The authors have declared no competing interest. International Human Frontier Science Program Organization (HFSP), Ref.-No: RGEC29/2025 National Aeronautics and Space Administration (NASA), 80NSSC23K1357, 80NSSC25K7873
Recent evidence suggests that peptide-RNA coacervates may have buffered the emergence of folded domains from flexible peptides. As primitive peptides were likely composed of both L- and D-amino acids, we hypothesized that coacervates may have also supported the emergence of chiral control. To test this hypothesis, we compared the coacervation propensities of an isotactic (homochiral) peptide and a syndiotactic (alternating chirality) peptide, both with an identical sequence derived from the ancient helix-hairpin-helix (HhH) motif. Using electron paramagnetic resonance (EPR) spectroscopy and molecular dynamics (MD) simulations, we found that the syndiotactic peptide does not form stable dimers with high α-helicity in solution, unlike the isotactic peptide. However, both peptides do coacervate with RNA, albeit with distinct reentrant phase behaviors. Coacervation in each case is facilitated by oligomer formation, likely dimerization, upon RNA binding that promotes RNA cross-linking. Additionally, RNA cross-linking and coacervation of the syndiotactic peptide seems to involve α-helical conformations. We attribute differences in reentrant phase behavior to differences in dimer flexibility and stability that alter the effectiveness of RNA cross-linking. These results illustrate how RNA-binding and/or coacervation by early protein forms could have promoted the transition of flexible, heterochiral peptides into folded, homochiral domains.
The helix-hairpin-helix (HhH) motif is an ancient and ubiquitous nucleic acid-binding element that has emerged as a model system for studying the evolution of dsDNA-binding domains from simple peptides that phase separate with RNA. We analyzed the entire putative evolutionary trajectory of the HhH motif - from a flexible peptide to a folded domain - for functional robustness to total chiral inversion. Against expectations, functional "ambidexterity" was observed for both the phase separation of HhH peptides with RNA and binding of the duplicated (HhH)2-Fold to dsDNA. Moreover, dissociation kinetics, mutational analysis, and molecular dynamics simulations revealed an overlap between the binding modes adopted by the natural and mirror-image proteins to natural dsDNA. The similarity of several dissociation phases upon chiral inversion may reflect the history of (HhH)2-Fold binding, with the ultimate emergence of a high-affinity binding mode, supported by a bridging metal ion, depopulating but not displacing more primitive (potentially ambidextrous) modes. These data underscore the surprising functional robustness of the HhH protein family and suggest that the veil between worlds with alternative chiral preferences may not be as impenetrable as is often assumed.
At the heart of many nucleoside triphosphatases is a conserved phosphate-binding sequence motif. A current model of early enzyme evolution proposes that this six to eight residue motif could have sparked the emergence of the very first nucleoside triphosphatases-a striking example of evolutionary continuity from simple beginnings, if true. To test this provocative model, seven disembodied Walker A-derived peptides were extensively computationally characterized. Although dynamic flickers of nest-like conformations were observed, significant structural similarity between the situated peptide and its disembodied counterpart was not detected. Simulations suggest that phosphate binding is nonspecific, with a preference for GTP over orthophosphate. Control peptides with the same amino acid composition but different sequences and situated conformations behaved similarly to the Walker A peptides, revealing no indication that the Walker A sequence is privileged as a disembodied peptide. We conclude that the evolutionary history of the P-loop NTPase family is unlikely to have started with a disembodied Walker A peptide in an aqueous environment. The limits of evolutionary continuity for this protein family must be reconsidered. Finally, we argue that motifs such as the Walker A motif may represent incomplete or fragmentary molecular fossils-the true nature of which has been eroded by time.
An unresolved question in the origin and evolution of life is whether a continuous path from geochemical precursors to the majority of molecules in the biosphere can be reconstructed from modern-day biochemistry. Here we identified a feasible path by simulating the evolution of biosphere-scale metabolism, using only known biochemical reactions and models of primitive coenzymes. We find that purine synthesis constitutes a bottleneck for metabolic expansion, which can be alleviated by non-autocatalytic phosphoryl coupling agents. Early phases of the expansion are enriched with enzymes that are metal dependent and structurally symmetric, supporting models of early biochemical evolution. This expansion trajectory suggests distinct hypotheses regarding the tempo, mode and timing of metabolic pathway evolution, including a late appearance of methane metabolisms and oxygenic photosynthesis consistent with the geochemical record. The concordance between biological and geological analyses suggests that this trajectory provides a plausible evolutionary history for the vast majority of core biochemistry.
Homochirality of biopolymers emerged early in the history of life on Earth, nearly 4 billion years ago. Whether the establishment of homochirality was the result of abiotic physical and chemical processes, or biological selection, remains unknown. However, given that significant events in protein evolution predate the last universal common ancestor, the history of homochirality may have been written into some of the oldest protein folds. To test this hypothesis, the evolutionary trajectory of the ancient and ubiquitous helix-hairpin-helix (HhH) protein family was analyzed for functional robustness to total chiral inversion of just one binding partner. Against expectations, functional ‘ambidexterity’ was observed across the entire trajectory, from phase separation of HhH peptides with RNA to dsDNA binding of the duplicated (HhH)2-Fold. Moreover, dissociation kinetics, mutational analysis, and molecular dynamics simulations revealed significant overlap between the binding modes of a natural and a mirror-image protein to natural dsDNA. These data suggest that the veil between worlds with alternative chiral preferences may not be as impenetrable as is often assumed, and that the HhH protein family is an intriguing exception to the dogma of reciprocal chiral substrate specificity proposed by Milton and Kent (Milton et al . Science 1992).### Competing Interest StatementThe authors have declared no competing interest.
As sequence and structure comparison algorithms gain sensitivity, the intrinsic interconnectedness of the protein universe has become increasingly apparent. Despite this general trend, β-trefoils have emerged as an uncommon counterexample: They are an isolated protein lineage for which few, if any, sequence or structure associations to other lineages have been identified. If β-trefoils are, in fact, remote islands in sequence-structure space, it implies that the oligomerizing peptide that founded the β-trefoil lineage itself arose de novo . To better understand β-trefoil evolution, and to probe the limits of fragment sharing across the protein universe, we identified both ‘β-trefoil bridging themes’ (evolutionarily-related sequence segments) and ‘β-trefoil-like motifs’ (structure motifs with a hallmark feature of the β-trefoil architecture) in multiple, ostensibly unrelated, protein lineages. The success of the present approach stems, in part, from considering β-trefoil sequence segments or structure motifs rather than the β-trefoil architecture as a whole, as has been done previously. The newly uncovered inter-lineage connections presented here suggest a novel hypothesis about the origins of the β-trefoil fold itself – namely, that it is a derived fold formed by ‘budding’ from an Immunoglobulin-like β-sandwich protein. These results demonstrate how the emergence of a folded domain from a peptide need not be a signature of antiquity and underpin an emerging truth: few protein lineages escape nature’s sewing table.
Peptide-RNA coacervates can result in the concentration and compartmentalization of simple biopolymers. Given their primordial relevance, peptide-RNA coacervates may have also been a key site of early protein evolution. However, the extent to which such coacervates might promote or suppress the exploration of novel peptide conformations is fundamentally unknown. To this end, we used electron paramagnetic resonance (EPR) spectroscopy to characterize the structure and dynamics of an ancient and ubiquitous nucleic acid binding element, the helix-hairpin-helix (HhH) motif, alone and in the presence of RNA, with which it forms coacervates. Double electron-electron resonance (DEER) spectroscopy applied to singly labeled peptides containing one HhH motif reveals the presence of dimers, even in the absence of RNA, and transient α-helical character. Moreover, dimer formation is promoted upon RNA binding and was detectable within peptide-RNA coacervates. The distance distributions between spin labels are consistent with the symmetric (HhH) 2 -Fold, which is generated upon duplication and fusion of a single HhH motif and traditionally associated with dsDNA binding. These results support the hypothesis that coacervates are a unique testing ground for peptide oligomerization and that phase-separating peptides could have been a resource for the construction of complex protein structures via common evolutionary processes, such as duplication and fusion.
How protein translation evolved from a simple beginning to its complex and accurate contemporary state is unknown. Aminoacyl-tRNA synthetases (AARSs) define the genetic code by activating amino acids and loading them onto cognate tRNAs. As such, their evolutionary history can shed light on early translation. Using structure-based alignments of the conserved core of Class I AARSs, we reconstructed their phylogenetic tree and ancestral states. Unexpectedly, AARSs charging amino acids that are assumed to have emerged later – such as TrpRS and TyrRS or LysRS and CysRS – appear as the earliest splits in the tree; conversely, those AARSs charging abiotic, early-emerging amino acids, e . g . ValRS, seem to have diverged most recently. Furthermore, the inferred Class I ancestor (excluding TrpRS and TyrRS) lacks the residues that mediate selectivity in contemporary AARSs, and appears to be a generalist that could charge a wide range of amino acids. This ancestor subsequently diverged to two clades: “charged” (which gave rise to ArgRS, GluRS, and GlnRS) and “hydrophobics”, which includes CysRS and LysRS as its outgroups. The ancestors of both clades maintain a wide-accepting pocket that could readily diverge to the contemporary, specialized families. Overall, our findings suggest a “generalist-maintaining” model of class I AARS evolution, in which early statistical translation was kept active by a generalist AARS while the evolution of a specialized, accurate translation system took place. Significance Aminoacyl-tRNA synthetases (AARS) define the genetic code by linking amino acids with their cognate tRNAs. While contemporary AARSs leverage exquisite molecular recognition and proofreading to ensure translational fidelity, early translation was likely less stringent and operated on a different pool of amino acids. The co-emergence of translational fidelity and the amino acid alphabet, however, is poorly understood. By inferring the evolutionary history of Class I AARSs we found seemingly conflicting signals: Namely, the oldest AARSs apparently operate on the youngest amino acids. We also observed that the early ancestors had broad amino acid specificities, consistent with a model of statistical translation. Our data suggests that a generalist AARS was actively maintained until complete specialization, thereby resolving the age paradox.
Anthropogenic organophosphorus compounds (AOPCs), such as phosphotriesters, are used extensively as plasticizers, flame retardants, nerve agents, and pesticides. To date, only a handful of soil bacteria bearing a phosphotriesterase (PTE), the key enzyme in the AOPC degradation pathway, have been identified. Therefore, the extent to which bacteria are capable of utilizing AOPCs as a phosphorus source, and how widespread this adaptation may be, remains unclear. Marine environments with phosphorus limitation and increasing levels of pollution by AOPCs may drive the emergence of PTE activity. Here, we report the utilization of diverse AOPCs by four model marine bacteria and 17 bacterial isolates from the Mediterranean Sea and the Red Sea. To unravel the details of AOPC utilization, two PTEs from marine bacteria were isolated and characterized, with one of the enzymes belonging to a protein family that, to our knowledge, has never before been associated with PTE activity. When expressed in Escherichia coli with a phosphodiesterase, a PTE isolated from a marine bacterium enabled growth on a pesticide analog as the sole phosphorus source. Utilization of AOPCs may provide bacteria a source of phosphorus in depleted environments and offers a prospect for the bioremediation of a pervasive class of anthropogenic pollutants.
Nat/Ivy is a diverse and ubiquitous CoA-binding evolutionary lineage that catalyzes acyltransferase reactions, primarily converting thioesters into amides. At the heart of the Nat/Ivy fold is a phosphate-binding loop that bears a striking resemblance to that of P-loop NTPases-both are extended, glycine-rich loops situated between a β-strand and an α-helix. Nat/Ivy, therefore, represents an intriguing intersection between thioester chemistry, a putative primitive energy currency, and an ancient mode of phospho-ligand binding. Current evidence suggests that Nat/Ivy emerged independently of other cofactor-utilizing enzymes, and that the observed structural similarity-particularly of the cofactor binding site-is the product of shared constraints instead of shared ancestry. The reliance of Nat/Ivy on a β-α-β motif for CoA-binding highlights the extent to which this simple structural motif may have been a fundamental evolutionary "nucleus" around which modern cofactor-binding domains condensed, as has been suggested for HUP domains, Rossmanns, and P-loop NTPases. Finally, by dissecting the patterns of conserved interactions between Nat/Ivy families and CoA, the coevolution of the enzyme and the cofactor was analyzed. As with the Rossmann, it appears that the pyrophosphate moiety at the center of the cofactor predates the enzyme, suggesting that Nat/Ivy emerged sometime after the metabolite dephospho-CoA.
Among the enzyme lineages that undoubtedly emerged prior to the last universal common ancestor is the so-called HUP, which includes Class I aminoacyl tRNA synthetases (AARSs) as well as enzymes mediating NAD, FAD, and CoA biosynthesis. Here, we provide a detailed analysis of HUP evolution, from emergence to structural and functional diversification. The HUP is a nucleotide binding domain that uniquely catalyzes adenylation via the release of pyrophosphate. In contrast to other ancient nucleotide binding domains with the αβα sandwich architecture, such as P-loop NTPases, the HUP's most conserved feature is not phosphate binding, but rather ribose binding by backbone interactions to the tips of β1 and/or β4. Indeed, the HUP exhibits unusual evolutionary plasticity and, while ribose binding is conserved, the location and mode of binding to the base and phosphate moieties of the nucleotide, and to the substrate(s) reacting with it, have diverged with time, foremost along the emergence of the AARSs. The HUP also beautifully demonstrates how a well-packed scaffold combined with evolvable surface elements promotes evolutionary innovation. Finally, we offer a scenario for the emergence of the HUP from a seed βαβ fragment, and suggest that despite an identical architecture, the HUP and the Rossmann represent independent emergences.
The P-loop Walker A motif underlies hundreds of essential enzyme families that bind nucleotide triphosphates (NTPs) and mediate phosphoryl transfer (P-loop NTPases), including the earliest DNA/RNA helicases, translocases, and recombinases. What were the primordial precursors of these enzymes? Could these large and complex proteins emerge from simple polypeptides? Previously, we showed that P-loops embedded in simple βα repeat proteins bind NTPs but also, unexpectedly so, ssDNA and RNA. Here, we extend beyond the purely biophysical function of ligand binding to demonstrate rudimentary helicase-like activities. We further constructed simple 40-residue polypeptides comprising just one β-(P-loop)-α element. Despite their simplicity, these P-loop prototypes confer functions such as strand separation and exchange. Foremost, these polypeptides unwind dsDNA, and upon addition of NTPs, or inorganic polyphosphates, release the bound ssDNA strands to allow reformation of dsDNA. Binding kinetics and low-resolution structural analyses indicate that activity is mediated by oligomeric forms spanning from dimers to high-order assemblies. The latter are reminiscent of extant P-loop recombinases such as RecA. Overall, these P-loop prototypes compose a plausible description of the sequence, structure, and function of the earliest P-loop NTPases. They also indicate that multifunctionality and dynamic assembly were key in endowing short polypeptides with elaborate, evolutionarily relevant functions.