Hydrolytic haloalkane dehalogenase enzymes catalyze an SN2 nucleophilic substitution to erase halogen substituents in organohalogen compounds. The acid-base-nucleophile triad secures irreversible SN2 displacement of the halogen for the hydroxyl derived from the water. Catalysis relies on the protonatable imidazole ring of the histidine base, and its substitution with an asparagine traps the enzyme in a covalently bound intermediate state, a principle exploited in the widely used HaloTag technology. By contrast, the histidine-to-phenylalanine substitution triggers reversibility of the SN2 reaction, but the molecular trick by which it reprograms the catalytic pathway remains unknown. Here, we show that the phenylalanine at the site of the histidine base spatially disturbs the adjacent residues, leading to the remodeling of surrounding active-site loops. Consequently, rerouting the access tunnels imparts distinctive kinetic behavior featuring a reversible SN2 chemical step that facilitates transhalogenation reactions. This information is crucial for engineering next-generation biocatalysts for sustainable chemistry.
Cardiovascular diseases, including ischemic stroke, necessitate improved thrombolytic agents. A microbe-encoded plasminogen activator staphylokinase (SAK) is a promising alternative to the widely used tissue plasminogen activator (tPA) due to its high fibrin specificity and low production cost. To overcome potential immunogenicity hampering its use in clinical settings, the low-immunogenic variants SAK SY155 and SAK THR174 were previously engineered. However, the molecular basis underlying their reduced immunogenicity is not understood and requires detailed elucidation. Here, we determine molecular structures and compare flexibility between low-immunogenic and immunogenic SAK variants, using a combination of experimental and computational structural techniques. Our analyses show that all variants share the canonical SAK fold and retain similar plasminogen activation kinetics, despite the number of introduced substitutions. Crucially, the low-immunogenic variants exhibit distinct flexibility profiles, with SAK THR174 showing substantially increased flexibility in the H1 helix and B3 region. SAK SY155 exhibits an increased flexibility in the H1-B3 loop and propensity to homodimerize. These flexibility changes are found in the known immunogenic epitopes. Our multi-scale flexibility analysis provides the molecular explanation for the reduced immunogenicity, altered thermostability, and retained fibrinolytic function of the engineered variants. This information is critical for the design of next-generation thrombolytics. ### Competing Interest Statement The authors have declared no competing interest. Czech Science Foundation, https://ror.org/01pv73b02, GX25-17329X, 25-18233M
Hereditary spastic paraplegia type 56 (SPG56) is a rare autosomal recessive neurodegenerative disorder caused by biallelic variants in the CYP2U1 gene, which encodes a cytochrome P450 enzyme involved in fatty acid metabolism and mitochondrial function. The clinical spectrum includes progressive spasticity of the lower limbs, developmental delay or regression, cognitive impairment, and variable ophthalmological findings. Although several cases have been reported in recent years, the functional characterization of individual variants remains limited. Here we describe a male patient with early-onset SPG56 carrying two CYP2U1 missense variants, NM_183075.3:c.1376 C > T p.(Pro459Leu) and NM_183075.3:c.557G > A p.(Arg186His). Combined genomic, cellular, and in silico analyses confirmed loss of enzymatic activity and protein instability, supporting the pathogenic classification of both variants. Functional validation led to reclassification of the p.(Arg186His) variant from uncertain significance to pathogenic. Further, we link specific CYP2U1 missense changes to convergent molecular defects, thereby refining genotype–phenotype correlations. From a therapeutic perspective, we highlight the relevance of experimental interventions such as folinic acid supplementation and multimodal spasticity management, while emphasizing the future promise of gene therapy for SPG56 patients. Our findings highlight the value of integrating genomic, biochemical, and structural approaches in the diagnostic evaluation of rare neurogenetic disorders, and provide functional evidence that the identified CYP2U1 variants are damaging, consistent with the observed early-onset complex SPG56 phenotype.
Scientific progress has historically depended on collaboration. For the first time, some collaborators may be non-human. Multi-agent artificial intelligence (AI) systems operate as teams of specialised digital scientists addressing shared challenges. Instead of relying on a single model, these systems integrate multiple AI agents, each assigned distinct roles, perspectives, or reasoning strategies. Agents generate proposals, critique one another’s suggestions, and iteratively refine solutions through both collaboration and competition. Advancement arises from the interaction of diverse viewpoints and the exchange of ideas, rather than from a single insight.
Abstract Molecular dynamics simulations provide atomistic views of protein motions, but conventional analyses often struggle with extracting subtle mechanistic insights from complex trajectories. Here, we present an integrated framework, ProtXAI, combining molecular dynamics and explainable artificial intelligence (XAI), to identify residue-level determinants of conformational change across diverse protein systems. By leveraging inter-residue distance dynamics, deep learning, and sequential relevance propagation, the approach captures both local fluctuations and long-range communication pathways within protein structures. We applied this framework to three mechanistically distinct systems: apolipoprotein E4 (ApoE4), staphylokinase (SAK) variants, and an ancestral luciferase. Across these applications, our XAI-based approach recovered experimentally supported dynamic hotspots: ligand-responsive hinges in ApoE4, mutation-dependent flexibility shifts in SAK, and evolutionary redistribution of motions in the luciferase. ProtXAI also revealed additional long-range couplings not accessible to classical analysis. Together, these findings demonstrate that combining molecular dynamics with XAI provides a general and scalable strategy for dissecting protein dynamics and uncovering structural determinants of function, stability, and evolutionary changes without prior bias. This approach thus advances the current methodological repertoire for analysing proteins and their intrinsic properties. Highlights Molecular dynamics simulations are increasingly accessible, yet scalable tools for comparative analysis remain limited. We demonstrate that machine learning coupled with explainable AI can automatically extract structural determinants of protein dynamics from trajectories. ProtXAI identifies key dynamic regions across diverse scenarios, including comparison of protein variants, understanding ligand modulation, and single-trajectory analysis. ProtXAI enables scalable, unbiased interpretation of long trajectories, providing an alternative to manual, time-intensive analysis.
Staphylokinase (SAK) is a highly fibrin-specific plasminogen activator with significant potential as a safe and affordable thrombolytic. Yet, its clinical translation can be limited by potential immunogenicity. To accelerate the development of improved thrombolytics, a critical step is identifying the most suitable molecular template. Therefore, we performed a comparative analysis of biochemical and immunological properties of three engineered low-immunogenic variants (SAK SY155, SAK THR174, and SAK STAR FRIDA) and two wild-types (SAK STAR and SAK 42D), using a newly established panel of assays. All variants retained potent thrombolytic activity, with SAK SY155 displaying the highest catalytic efficiency and fibrin-clot permeability. However, this advantage did not fully translate into improved clot reduction under flow conditions. Comprehensive immunological profiling, including monocyte activation, T lymphocyte proliferation, dendritic cell maturation, and mouse immunization models, revealed no strong immunogenic response in any tested variant. Overall, low-immunogenic SAK SY155 and its background wild-type SAK STAR emerged as the most promising templates for rational engineering of next-generation thrombolytics. ### Competing Interest Statement The authors have declared no competing interest. European Unions Horizon 2020, 857560 Horizon Europe Framework programme, 101136607 the Czech Grant Agency the Czech Health Research Council, NW24-08-00064 the Ministry of Education, Youth and Sports of the Czech Republic, 90254, LM2023055, LM2023069
Abstract The identification of aggregation-prone regions in proteins and their suppression through mutations is a powerful strategy to enhance protein solubility and yield, significantly expanding their potential applications. Here, we developed and experimentally validated a deep neural network-based predictor, AggreProt, that generates a residue-level aggregation profile for protein sequences. The model outperformed or matched current state-of-the-art algorithms, as validated on two independent datasets comprising hexapeptides and full-length proteins with annotated aggregation-prone regions. Importantly, we validated the model experimentally using a set of 34 hexapeptides identified in the model protein haloalkane dehalogenase LinB, along with seven proteins from the AmyPro database. Experimental results agreed with our predictions in 79% of cases and revealed inaccuracies in some database annotations. Finally, the algorithm’s utility was demonstrated by identifying aggregation-prone regions in the LinB enzyme and designing mutations to suppress aggregation in its exposed regions. The resulting variants exhibited reduced aggregation propensity, improved solubility, and up to a 100% increase in yield compared to the wild type. AggreProt is freely available to the scientific community via a user-friendly web server: https://loschmidt.chemi.muni.cz/aggreprot .
Evolution-guided protein design remains one of the most effective strategies for engineering proteins with enhanced stability, activity, or specificity. To make these approaches more accessible, we previously developed FireProtASR-a fully automated pipeline for ancestral sequence reconstruction (ASR). Here, we present FireProtASR 2.0, a significantly enhanced version that extends the design space beyond ancestral inference by integrating a successor sequence predictor (SSP) and a generative model based on variational autoencoders (VAEs). These new modules enable both 'prospective' and 'retrospective' evolutionary design strategies. The SSP module predicts likely future mutations based on site-wise evolutionary trends, and the method was previously validated through in silico benchmarks, demonstrating improvements in thermostability and activity. The VAE module captures global evolutionary constraints in a low-dimensional latent space, from which novel functional ancestral-like variants can be sampled. The VAE-based design strategy was previously validated experimentally on the haloalkane dehalogenase family, yielding variants with enhanced thermostability while maintaining catalytic activity. Both these modules are newly available in FireProtASR in a fully automated pipeline, guiding the users via an interactive graphical user interface. With expanded functionality, modernized user interface, and a more robust backend, FireProtASR 2.0 provides a comprehensive, accessible, and fully automated platform for evolutionary-based protein engineering (https://loschmidt.chemi.muni.cz/fireprotasr/).
Abstract Enzyme cascades enable complex biochemical transformations, but their optimization is resource-intensive, requiring navigation through high-dimensional parameter spaces encompassing reaction conditions, enzyme ratios, and buffer composition. Here we introduce CascadeMAP, an autonomous microfluidic platform for closed-loop optimization of enzyme cascades, integrating high-throughput microfluidics with Bayesian optimization and multi-agent AI system. We demonstrate the platform across two cascades: (i) a glycerol detection pathway monitored by fluorescence and (ii) a 1,2,3-trichloropropane degradation pathway monitored by label-free Raman spectroscopy providing orthogonal detection modalities. Bayesian optimization identified optimal conditions three times faster than Design of Experiments. Multi-agent AI system automated hypothesis generation, processing 11 GB of experimental data, pattern recognition, and insight synthesis. Operating without human intervention for 7 days, CascadeMAP processed ∼220,000 reactions across ∼7,400 different conditions. This capability establishes a generalizable framework for the autonomous optimization of enzyme cascades and metabolic pathways and accelerates the development of biocatalytic and synthetic biological systems.
Enhancing enzymes to improve desired properties remains an expensive and time-consuming process. Scanning databases of known protein sequences to find enzymes with similar catalytic activity and enhanced properties is an efficient and valuable approach. The EnzymeMiner web server has proven integral as an automated, user-friendly tool that identifies enzymes with the desired catalytic activity from provided sequences and essential residues. Here, we introduce EnzymeMiner 2.0 that builds upon its predecessor, retaining its original functionality, while introducing several key improvements: (i) significantly expanded searched protein space; (ii) annotation of discovered sequences with predictions of the melting temperature, optimal pH, catalytic activity and efficiency, and aggregation propensity with state-of-the-art computational tools; and (iii) smart automatic sequence prioritization and filtering based on user-defined goals or a set of predefined scenarios. With all these enhancements, EnzymeMiner 2.0 aims to remain among the leading solutions for efficient discovery of novel enzymes. The server is freely accessible at https://loschmidt.chemi.muni.cz/enzymeminer/.
Staphylokinase (SAK) is a promising third-generation thrombolytic protein, but its clinical potential is limited by immunogenicity and stability concerns. The conformational and colloidal stabilities of four SAK variants-SAK 42D, SAK STAR, and their non-immunogenic derivatives SAK 42D 3A and SAK STAR 3A-were evaluated using differential scanning calorimetry (DSC), dynamic light scattering (DLS), and aggregation kinetics assays. DSC analyses revealed that thermal denaturation of all variants proceeds via two consecutive irreversible steps, with transition parameters strongly dependent on scan rate and protein concentration. SAK STAR variants exhibited markedly exothermic first transitions and reduced scan rate dependence, suggesting stabilization of intermediate states and suppression of aggregation. In contrast, SAK 42D variants exhibited endothermic or weakly exothermic first transitions and a higher aggregation propensity, correlating with reduced conformational stability and formation of less stable dimers. Colloidal stability tests showed that SAK STAR and SAK STAR 3A remained largely aggregation-resistant, whereas SAK 42D and SAK 42D 3A aggregated rapidly at elevated temperatures (>51°C and >38°C, respectively), following apparent second-order kinetics. DLS confirmed concentration-dependent dimerization in all variants, with SAK 42D 3A displaying pronounced polydispersity and instability. We could rationalize this behavior in the context of engineered surface charges. Our results demonstrate that SAK variant stability is shaped by a complex interplay between primary sequence, dimerization behavior, and aggregation propensity, guiding the design of clinically viable thrombolytic agents and their formulations.
Apolipoprotein E4 (ApoE4) is a major genetic risk factor in many neurodegenerative diseases, yet effective therapeutic strategies targeting its associated pathologies remain unresolved. The aggregation of ApoE4, a key pathological feature, is modulated by tramiprosate and its metabolite 3-sulfopropanoic acid. In this study, we provide mechanistic insights into how taurine, a close chemical analogue of tramiprosate, interacts with ApoE4 and may similarly modulate its aggregation behavior. Using an integrated approach, which included molecular dynamics simulations, static light scattering, mass spectrometry, and cerebral organoid models, we investigated the effects of taurine on ApoE4 aggregation. Our results indicate that taurine effectively prevents ApoE4 aggregation and exerts a partial disaggregating effect on pre-formed aggregates. Notably, taurine modulates molecular and cellular features associated with the ApoE4 isoform, shifting them toward patterns observed in the more benign ApoE3 isoform. These observations are consistent with effects similar to those reported for tramiprosate and 3-sulfopropanoic acid and suggest that taurine influences ApoE4-related molecular mechanisms, particularly in the context of the high-risk ApoE4/E4 genotype.
Abstract The nonradiative transport of electronic excitation from one chromophore to another, known as resonance energy transfer, lies at the root of photochemical processes in biology. 1,2 Unlike photosynthesis, bioluminescence converts chemical energy into light through an enzymatic oxygenation of an energy-rich luciferin. 3–5 In glowing cnidarians, the energy is relocated from an excited oxyluciferin to a fluorescent protein, shifting the colour and enhancing the quantum yield of a photogenic reaction. 6,7 How protein-chromophore complexes assemble during this interplay in real space, and what this association entails for function, are unknown. Here, we report co-crystal structures of a 120-kilodalton energy-transfer complex from the luminescent soft coral Renilla reniformis . We find a heterotetrameric 2:2 assembly composed of two coelenteramide-loaded luciferases (RrLuc) docked at opposite sides of a head-to-tail dimer of green fluorescent protein (RrGFP). The edge-to-edge distance between donor and acceptor chromophores is below 3 nm, favouring the Förster-type radiationless energy transfer. Furthermore, RrGFP serves not only as a colour-switchable antenna and luminescence amplifier but also tunes the efficiency of luciferase catalysis by controlling its inherent dynamics. Our results provide detailed spatial information about intermolecular dipole-dipole coupling in Renilla bioluminescence, including the arrangement of donor-acceptor pairs that secure excited-state energy transfer with exquisite precision.
Thermostable proteins are crucial in numerous biomedical and biotechnological applications. However, naturally occurring proteins have evolved to function in mild conditions, and laboratory experiments aiming at improving protein stability have proven laborious and expensive. Computational methods overcome this issue by providing a cheap and scalable alternative. Despite significant progress, their reliability is still hindered by the availability of high-quality data. FireProtDB 2.0 (http://loschmidt.chemi.muni.cz/fireprotdb) is a large-scale database aggregating stability data from multiple sources. The second version builds upon its predecessor, retaining its original functionality while introducing a new approach to data storage and maintenance. The new scheme enables the introduction of both absolute and relative data types connected with measurements of wild-types, mutants, protein domains, and de novo designed proteins. Furthermore, while the original database was limited to single-point mutations, more complex data such as insertions, deletions, and multiple-point mutations are now available. As a result, the inclusion of large-scale mutagenesis has increased the size of the database from 16 000 to almost 5 500 000 experiments. Moreover, the updated abstract scheme is fully expandable with any new measurements and annotations without the need for any restructuring. Finally, the tracking of history together with fixed identifiers is in accordance with the FAIR principles.
Selective hydroxylation of sesquiterpenes remains a major challenge in synthetic chemistry due to their chemical homogeneity and steric complexity. Here, we report the Rieske oxygenase (RO)-catalyzed hydroxylation of sesquiterpenes using cumene dioxygenase (CDO) from Pseudomonas fluorescens IP01. Wild-type CDO catalyzed the monohydroxylation of the sesquiterpene β-bisabolene, producing ring-hydroxylated β-bisabolene-5-ol and tail-hydroxylated (S)-β-bisabolene-13-ol in a 64:21 ratio, with a total product formation of 0.43 ± 0.07 mM. A single-loop deletion variant, L284del, enhanced both total product formation and selectivity toward the ring-hydroxylated product, producing 1.40 ± 0.08 mM and enabling preparative-scale isolation (17 mg, 0.08 mmol, 39% yield). Another single-loop variant, I288del, similarly increased total product formation while shifting regioselectivity from ring- to tail-hydroxylation in a 5:81 ratio, generating 1.07 ± 0.04 mM tail-hydroxylated product and enabling its preparative-scale isolation (24 mg, 0.11 mmol, 55% yield). To understand this clear switch in regioselectivity, we conducted molecular dynamics (MD) and adaptive steered MD simulations. This revealed that I288del increased the substrate tunnel openness and frequency of access, promoting a tail-first orientation of the substrate. This preorientation enhanced the probability of reactive conformations in proximity to the catalytic iron. Binding energy calculations further supported a loss of orientation bias in I288del, giving further insight into the regioselectivity shift. Additional engineering yielded the variant N279T_I288del_A321T, which enhanced the formation of the tail-hydroxylated compound to 1.58 ± 0.26 mM. Further, variant I288del enabled the conversion of α-santalene, a reaction not catalyzed by the wild-type enzyme. Our study provides mechanistic insight into how tunnel dynamics and substrate preorientation govern selectivity in RO-catalyzed C-H functionalization.
Abstract Motivation Protein melting temperature ( T m ) prediction accelerates the discovery of thermostable enzymes which are crucial for industrial biotechnology often requiring harsh reaction conditions. Experimental determination of T m remains labour-intensive and varies across techniques, motivating the development of in silico predictors. Mass-spectrometry datasets such as Meltome Atlas now enable large-scale T m prediction with models based on deep learning, but model generalisation across diverse experimental datasets has not been systematically tested. Results We evaluated the generalisability of state-of-the-art deep learning approaches and explored ESM-based embeddings for T m prediction. To this end, we assembled the ProMelt training dataset (45 441 proteins) and five independent biophysics-based validation datasets. Our analysis revealed substantial differences between proteomics- and biophysics-based T m measurements, highlighting the challenge of cross-domain generalisation. Existing state-of-the-art predictors trained on large-scale proteomics datasets showed reduced performance on biophysics-based validation sets. Our fine-tuned embedding-based models, particularly LoRA-adapted ESM-2 (TmProt 1.0), outperformed state-of-the-art predictors in identifying thermostable proteins ( T m ≥ 60 °C) across heterogeneous datasets, achieving AUC scores of 0.75–0.77. We also demonstrated that the available models could be used efficiently in the sequence prioritization task. Availability The TmProt web server is available at https://loschmidt.chemi.muni.cz/tmprot/ . Source code and data are available at https://github.com/loschmidt/TmProt .
Protein engineering approaches, including rational design and directed evolution, are essential for optimizing protein properties in biotechnology. However, their application is often limited by the need to experimentally characterize large numbers of variants. This is especially true for directed evolution, where selected candidates require heterologous expression and purification before functional testing. To reduce this bottleneck for staphylokinase, a fibrin-specific thrombolytic with therapeutic potential, we developed a rapid screening platform based on cell-free protein synthesis (CFPS) and a chromogenic plasminogen-activation assay. Activity rankings from crude CFPS mixtures closely matched those from purified proteins, showing that the method provides reliable functional readouts without purification. This CFPS-based workflow offers a fast, scalable, and efficient solution for early-stage screening of staphylokinase variants and can accelerate the identification of new thrombolytic candidates. The modularity of this method also allows its facile adjustment for other enzymes.
The α/β-hydrolase (ABH) superfamily is a widespread and functionally versatile protein fold recognized for its ability to adapt to diverse molecular functions across all three domains of life. One such spectacular example of evolutionary adaptation at the ABH fold is an acquisition of oxygenolytic luciferase reaction that occurred within the hydrolytic haloalkane dehalogenase family. The molecular details of this evolution remain puzzling. In this work, we determine crystal structures and explore dynamical behaviour of a bifunctional ancestral ABH-fold enzyme, highlighting molecular features associated with the transition from hydrolytic to oxygenolytic catalysis at this fold. Structures showed a canonical αβα-sandwich shielded with a helical cap domain. The catalytic pocket is voluminous enough to accommodate a bulky substrate. Molecular dynamics simulations demonstrated that coelenterazine entry does not present a major energetic barrier and identified a preferred binding orientation important for oxygenolytic catalysis. Comparisons between ancestral and extant enzymes highlighted specific amino acids and sequence motifs characteristic for oxygenolytic luciferases. Collectively, our results provide an expanded view of the evolutionary transition in which ABH-fold enzymes, originally using water to cleave chemical bonds, adapted to utilize dioxygen for bioluminescence.
Determining why convergent traits use distinct versus shared genetic components is crucial for understanding how evolutionary processes generate and sustain biodiversity. However, the factors dictating the genetic underpinnings of convergent traits remain incompletely understood. Here, we use heterologous protein expression, biochemical assays, and phylogenetic analyses to confirm the origin of a luciferase gene from haloalkane dehalogenases in the brittle star Amphiura filiformis. Through database searches and gene tree analyses, we also show a complex pattern of the presence and absence of haloalkane dehalogenases across organismal genomes. These results first confirm parallel evolution across a vast phylogenetic distance, because octocorals like Renilla also use luciferase derived from haloalkane dehalogenases. This parallel evolution is surprising, even though previously hypothesized, because many organisms that also use coelenterazine as the bioluminescence substrate evolved completely distinct luciferases. The inability to detect haloalkane dehalogenases in the genomes of several bioluminescent groups suggests that the distribution of this gene family influences its recruitment as a luciferase. Together, our findings highlight how biochemical function and genomic availability help determine whether distinct or shared genetic components are used during the convergent evolution of traits like bioluminescence.