The special nature of the fluorine atom imparts remarkable strength and unique physical properties to chemical bonds. Unlike man-made fluorochemicals, fluorinated natural products remain rare due to low bioavailability and toxicity of fluoride. Despite this, defluorinases have evolved in nature to cleave carbon-fluorine bonds, with the hydrolytic fluoroacetate dehalogenase being one of the most well-characterized examples. These enzymes are of fundamental interest and hold unrealized biotechnological potential, yet the scope of this unique chemistry remains underexplored in the biosphere. Here, we trained and applied a machine learning-based framework, termed latent generative landscapes (LGLs), to map the functional sequence space of the α/β-hydrolase superfamily. This approach identified 3014 putative defluorinases that were previously not annotated or plausibly misannotated. Experimental validation of selected candidates led to the reclassification of five novel defluorinases, all exhibiting high thermal stability (T m > 70 °C) and diverse catalytic efficiencies with conserved enantioselectivity on the model substrate 2-fluoro-2-phenylacetate. Notably, the enzyme A0A4Z0BVY8 exhibited 2.7-fold greater defluorination activity than the current state-of-the-art enzyme Q6NAM1. Our results establish that LGL modeling is a powerful strategy to decode cryptic carbon-fluorine bond chemistry in nature, enabling the future discovery and engineering of defluorination biocatalysts.
Efficiency and substrate specificity of proteases in the Potyviridae family have not been comprehensively profiled. Here we develop a model that learns co-evolutionary features to accurately predict and experimentally validate protease performance at single amino-acid resolution. We identify and engineer several proteases that perform better than the commercially available tobacco etch virus protease. To demonstrate the resolving power of our methods, we engineer protease crosstalk to selectively trigger a synthetic cell-death program in human cells.
Antibiotic resistance has become a critical public health problem, rendering many antibiotics ineffective. In particular, the evolution of extended-spectrum β-lactamases (ESBLs) threatens β-lactams, the cornerstone of bacterial infection treatment. We investigated the evolution of Escherichia coli TEM-1 β-lactamase into ESBLs by constructing a combinatorially complete library of all 55,296 TEM-1 variants from 18 clinical mutations across 13 residues. We obtained over 9,000,000 fitness measurements under native (ampicillin) and non-native (aztreonam) selection. Graph-theoretic and epistatic analyses revealed that ampicillin selection produced weak epistasis and predictable evolutionary trajectories, whereas aztreonam selection induced extensive higher-order epistasis, increasing phenotypic unpredictability. Machine learning identified interpretable epistatic rules shaping these landscapes. Evolutionary statistics, including direct coupling analysis and latent generative landscapes, showed that top-performing ESBL variants followed conserved epistatic patterns observed in natural β-lactamases. Our integrated experimental–computational framework provides a foundation for predicting ESBL evolution and quantifying mutational contributions to ESBL variants. Authors investigate the evolution E. coli TEM-1 β-lactamase into ESBLs with a complete library of 55,296 TEM-1 β-lactamase variants, showing that adaptation to a non-native antibiotic is driven by higher-order epistasis, making resistance evolution far less predictable than to the native substrate.
Daptomycin resistance (DAP-R) in enterococci is associated with alterations in the membrane lipid composition. The membrane-bound protein MprF is responsible for the synthesis of amino acid-modified lipids in bacteria, and these modified lipids contribute to DAP-R in some Gram-positive pathogens. In enterococci, MprF synthesizes three lysine-modified lipids: the phospholipid lysyl-phosphatidylglycerol (Lys-PG), and the newly identified cationic glycolipids lysyl-diglucosyl-diacylglycerol (Lys-Glc 2 -DAG) and lysyl-glucosyl-diacylglycerol (Lys-Glc-DAG). Given the recent discovery of cationic glycolipids in enterococci, we re-examined a collection of laboratory-evolved DAP-R E. faecalis to investigate whether these lipids contribute to DAP-R. We found that levels of Lys-Glc 2 -DAG were strikingly reduced in DAP-R variants with high-level resistance. The dramatic alterations in Lys-Glc 2 -DAG levels were temporally coupled with the emergence of loss-of-function mutations in the gene drmA , which encodes a DUF998 family protein of unknown function. DrmA is a membrane protein with six predicted transmembrane helices and is widely distributed among Gram-positive and Gram-negative bacteria, including plant and animal pathogens. Complementation of the DAP-R strains with wild-type E. faecalis drmA significantly lowered their DAP MIC, reversing their trajectory to high-level DAP-R. Using genetic and lipidomic approaches in the natively DAP-sensitive strain OG1RF, we conclusively linked drmA loss-of-function with significantly reduced Lys-Glc 2 -DAG levels as well as a small but significant increase in Lys-PG levels. Yet, drmA inactivation in OG1RF did not alter its DAP MIC. We conclude that drmA loss-of-function confers elevated DAP MIC on the background of preceding mutations in the DAP-R evolutionary trajectory, most likely mutations in cls1 . The recurrence of drmA mutations in multiple studies underscores its importance in DAP-R evolution. Overall, our work identifies a role for the DUF998 family in cellular lipid homeostasis and confirms its significant role in the evolution of DAP-R.
BRCA1/BARD1 is a chromatin-associated E3 ubiquitin ligase that ubiquitylates histone H2A to coordinate DNA damage repair, transcriptional repression, and genome stability. In Caenorhabditis elegans (C. elegans), the orthologous BRC-1/BRD-1 complex performs analogous functions but exhibits structural variation, most notably through an additional 11-residue loop in BRD-1 that is absent from human BARD1. Prior experiments indicate this worm-specific insertion promotes nucleosome engagement and may alter the preferred lysine target for ubiquitylation. Here, we provide a cross-species comparison by integrating computational and experimental investigation to clarify how a discrete structural variation can tune BRCA1-family ligase behavior and, consequently, chromatin regulation. In vitro ubiquitylation assays and mass spectrometry reveal BRC-1/BRD-1 ubiquitylate the C-terminal tail of histone H2A with less specificity than the human homologs. All-atom molecular dynamics simulations of both the C. elegans BRC-1/BRD-1-LET-70-Ubiquitin assembly and the human BRCA1/BARD1-UbcH5c-Ubiquitin complex in the presence of the nucleosome core particle uncover that the BRD-1 loop makes transient contacts with nucleosomal DNA and histone tails, thereby modulating the positioning and conformational flexibility of the bound E2 (ubiquitin-conjugating enzyme). Together, our results suggest that the BRD-1 loop alters the E3-E2 geometry, thereby altering ubiquitylation-site specificity.
The rapid evolution of extended-spectrum β-lactamases (ESBLs) represents a global health threat, undermining the efficacy of β-lactams, the most extensively used antibiotic class. To elucidate the evolutionary dynamics underlying β-lactam resistance, we constructed a comprehensive combinatorial mutant library comprising all 55,296 possible TEM-1 β-lactamase variants integrating 18 clinically observed mutations across 13 key residues. Over eight million empirical fitness measurements were obtained under selection pressure with both a native antibiotic substrate (ampicillin) and a novel antibiotic (aztreonam), establishing the largest experimentally determined fitness landscape for antibiotic resistance to date. Through graph-theoretic and epistatic analyses, we discovered that selection with ampicillin resulted in weak epistasis, with mutants rarely surpassing the fitness of the wild-type enzyme. Conversely, aztreonam selection elicited extensive higher-order epistasis, generating a rugged fitness landscape characterized by increased phenotypic unpredictability. Interpretable machine-learning analyses identified context-dependent epistatic interactions necessary for achieving high-level aztreonam resistance. Further evolutionary statistical analyses, including direct coupling analysis and latent generative landscapes, showed that top-performing TEM-1 variants consistently adhered to conserved epistatic patterns found in naturally occurring β-lactamases. Our findings demonstrate that higher-order epistasis critically shapes fitness landscape ruggedness when enzymes adapt to novel substrates, whereas adaptations to native substrates exhibit predictably smoother landscapes. This integrated experimental and computational framework provides a foundation for predictive evolutionary pharmacology, enabling assessments of newly developed β-lactams or emerging β-lactamase variants for their potential contribution to ESBL evolution. Importantly, incorporating graph-theoretically informed evolutionary constraints can strategically disrupt evolutionary pathways, presenting a viable approach to mitigate the rise of antibiotic resistance.
Microtubule (MT) branch nucleation requires Augmin and NEDD1 proteins, which recruit and activate the gamma-tubulin ring complex (γ-TuRC). Augmin is a fork-shaped assembly of eight coiled-coil subunits, while NEDD1 is a β-propeller protein bridging MTs, Augmin, and γ-TuRC. We reconstitute Arabidopsis thaliana Augmin assemblies and determine 3.7-7.3-Å cryo-EM structures of its V-junction and extended regions using crosslinking mass spectrometry. These structures reveal a complete plant Augmin model showing multi-coiled-coil interfaces stabilizing its 40-nm hetero-octameric fork architecture. The dual calponin homology (CH) domains at the V-junction terminus adopt open and closed conformations for MT binding. A 12-Å cryo-EM structure shows Augmin undergoes anti-parallel dimerization through conserved surfaces on its extended region. We determine the NEDD1 β-propeller structure with Augmin, revealing direct binding inside the V-junction that enhances dimerization. Direct coupling and evolutionary analyses identify co-varying residue pairs validating the eight-subunit model and NEDD1 interface. Cooperativity between dual CH domains and NEDD1 binding may regulate V-junction binding to MT lattices. This V-shaped dual binding anchors Augmin along MTs, creating platforms for γ-TuRC recruitment and branched MT nucleation.
Protein evolution has shaped enzymes that maintain stability and function across diverse thermal environments. While sequence variation, thermal stability and conformational dynamics are known to influence an enzyme's thermal adaptation, how these factors collectively govern stability and function across diverse temperatures remains unresolved. Cytosolic malate dehydrogenase (cMDH), a citric acid cycle enzyme, is an ideal model for studying these mechanisms due to its temperature-sensitive flexibility and broad presence in species from diverse thermal environments. In this study, we employ techniques inspired by deep learning and statistical mechanics to uncover how sequence variation and conformational dynamics shape patterns of cMDH's thermal adaptation. By integrating coevolutionary models with variational autoencoders (VAE), we generate a latent generative landscape (LGL) of the cMDH sequence space, enabling us to explore mutational pathways and predict fitness using direct coupling analysis (DCA). Structure predictions via AlphaFold and molecular dynamics simulations further illuminate how variations in hydrophobic interactions and conformational flexibility contribute to the thermal stability of warm- and cold-adapted cMDH orthologs. Notably, we identify the ratio of hydrophobic contacts between two regions as a predictive order parameter for thermal stability features, providing a quantitative metric for understanding cMDH dynamics across temperatures. The integrative computational framework employed in this study provides mechanistic insights into protein adaptation at both sequence and structural levels, offering unique perspectives on the evolution of thermal stability and creating avenues for the rational design of proteins with optimized thermal properties.
The rapid expansion of protein sequence databases has far outpaced experimental structure determination, leaving many unannotated sequences, particularly the more remote homologs with low sequence identity. Because protein folds are more conserved and functionally informative than sequences alone, structural information offers a powerful lens for analysis. Here, we introduce a generative, structure-aware framework that integrates geometric encoding and coevolutionary constraints to map, cluster, and design protein sequences. Our approach employs the 3D interaction (3Di) alphabet to convert local residue geometries into compact, 20-state discrete representations. Using ProstT5, we enable bidirectional translation between amino acid sequences and 3Di representations, facilitating sensitive homology detection and structure-guided sequence generation. We then augment the latent generative landscape methodology by combining 3Di-based alignments with direct coupling analysis (DCA) and variational autoencoders (VAE), imbuing tasks such as clustering, annotation, and design with structural information. This integrative framework enhances the detection of coevolutionary signals and enables rational sampling of structural variants, even without functional labels. We demonstrate the utility of our method across diverse protein families, including globins, kinases, and malate dehydrogenases, achieving improved contact prediction, homology inference, and sequence generation. Together, our approach offers a quantitative, generative view of protein structure space, advancing protein evolution and design studies.
Design and synthesis of functionally active artificial proteins is challenging, as it requires simultaneous consideration of interconnected factors, such as fold, dynamics, and function. These evolutionary constraints are encoded in protein sequences and can be learned through the latent generative landscape (LGL) framework to predict functional sequences by leveraging evolutionary patterns, enabling exploration of uncharted sequence space. By simulating designed proteins through molecular dynamics (MD), we gain deeper insights into the interdependencies governing structure and dynamics. We present a synergized workflow combining LGL with MD and biochemical characterization, allowing us to explore the sequence space effectively. This approach has been applied to design and characterize two artificial multidomain ATP-driven transmembrane copper transporters, with native-like functionality. This integrative approach proved effective in revealing the intricate relationships between sequence, structure, and function.
The rapid expansion of protein sequence databases has far outpaced experimental structure determination, leaving many unannotated sequences, particularly the more remote homologs with low sequence identity. Because protein folds are more conserved and functionally informative than sequences alone, structural information offers a powerful lens for analysis. Here, we introduce a generative, structure-aware framework that integrates geometric encoding and coevolutionary constraints to map, cluster, and design protein sequences. Our approach employs the 3D interaction (3Di) alphabet to convert local residue geometries into compact, 20-state discrete representations. Using ProstT5, we enable bidirectional translation between amino acid sequences and 3Di representations, facilitating sensitive homology detection and structure-guided sequence generation. We construct a latent sequence landscape by combining 3Di-based alignments with direct coupling analysis (DCA) and variational autoencoders (VAE), unifying tasks such as clustering, annotation, and design. This integrative framework enhances the detection of coevolutionary signals and enables rational sampling of structural variants, even without functional labels. We demonstrate the utility of our method across diverse protein families, including globins, kinases, and malate dehydrogenases, achieving improved contact prediction, homology inference, and sequence generation. Together, our approach offers a quantitative, generative view of protein structure space, advancing protein evolution, and design studies. ### Competing Interest Statement The authors have declared no competing interest.
DNA-transcription factor (TF) interactions are essential for gene regulation. Fully characterizing TF recognition specificities and identifying their genomic binding targets are important to understand TF function and regulatory networks. Recently, high-throughput sequencing technology HT-SELEX (high-throughput systematic evolution of ligands by exponential enrichment) has been used to measure hundreds of TFs, providing massive datasets that comprise TF binding preferences. However, there is a need to develop comprehensive computational modeling to fully extract and characterize critical TF binding preferences and fail to distinguish genome-wide binding targets. In this study, we developed a global pairwise model called DCA-Scapes trained with experimental HT-SELEX data. Our approach uncovered high-resolution TF recognition specificity landscapes, enabled the prediction of in vivo binding sequences, and was validated with ChIP-seq (ChIP sequencing) data. In addition, the DCA-Scapes model was utilized to refine the locations of binding regions and accurately identify the binding sites within the ChIP-seq enriched peaks. Moreover, we extended our model to cover the entire human genome, uncovering potential TF target sites that exhibit tissue-specific TF recognition across various cellular environments.
Inferring the historical and biophysical causes of diversity within protein families is a complex puzzle. A key to unraveling this problem is characterizing the rugged topography of sequence-function adaptive landscapes. Using biochemical data from a 29 = 512 combinatorial library of tobacco 5-epi-aristolochene synthase (TEAS) mutants engineered to make the native major product of Egyptian henbane premnaspirodiene synthase (HPS) and a complementary 512 mutant HPS library, we address the question of how product specificity is controlled. These data sets reveal that HPS is far more robust and resistant to mutations than TEAS, where most mutants are promiscuous. We also combine experimental data with a sequence Potts Hamiltonian model and direct coupling analysis to quantify mutant fitness. Our results demonstrate that the Hamiltonian captures variation in product outputs across both libraries, clusters native family members based on their substrate specificities, and exposes the divergent catalytic roles of couplings between the catalytic and noncatalytic domains of TEAS versus HPS. Specifically, we found that the role of the interdomain connectivities in specifying product output is more important in TEAS than connectivities within the catalytic domain. Despite being 75% identical, this property is not shared by HPS, where connectivities within the catalytic domain are more important for specificity. By solving the X-ray crystal structure of HPS, we assessed structural bases for their interdomain network differences. Last, we calculate the product profile Shannon entropies of the two libraries, which showcases that site-site connectivities also play divergent roles in catalytic accuracy.