A Working Group consisting of the co-authors of this paper was established in 2020 to re-evaluate the standard valence geometry used for the validation of nucleic acid structure models in the Protein Data Bank (PDB). This Working Group re-examined the dependence of Cambridge Structural Database (CSD) derived targets on base and sugar type, sugar pucker, and phosphate and glycosidic conformation, before comparing those targets with the geometry of a quality-filtered reference set of nucleic acid crystal structural models held in the PDB. This revealed that the valence bond and angle mean values are close to the CSD targets, but many parameters have highly non-Gaussian or even multimodal distributions. One explanation is the inconsistency of restraints used over time and by different refinement programs. The Working Group recommends a new validation scheme for use by the PDB. For this purpose, we have developed a new three-tier scale for outlier detection-graded as Preferred, Allowed, and Of Concern intervals-based on a combination of quality-curated reference data from the CSD and the PDB. The proposed approach to validation should lead to improved nucleic acid models in (future) PDB-deposited macromolecular structures.
As the number and complexity of RNA and DNA structures continue to expand, there is a growing need for robust yet accessible tools that support their accurate interpretation, validation, and refinement. We present DNATCO v5.0 (dnatco.datmos.org), an interactive web application for comprehensive structural analysis of nucleic acids. DNATCO integrates the NtC dinucleotide conformational classes and the CANA structural alphabet to provide an intuitive, geometrically complete description of local backbone and base orientations, complemented by interactive visualization of base pairing. The platform performs quantitative validation of conformational similarity and covalent bond lengths and angles, using newly established nucleic-acid valence-geometry standards. Quantitative validation encompasses the confal score and scattergrams mapping the fit between experimental electron density and geometry similarity to the closest NtC class. All outputs are downloadable. Integrated diagnostic tools help users identify unusual or problematic regions, explore alternative conformations, and generate torsion-restraint files for downstream. DNATCO v5.0 is implemented entirely client-side via WebAssembly, ensuring fast performance and preserving data privacy, and supports both PDB and user-provided structural models. By combining a rigorous geometric framework with an approachable interface, DNATCO enables both non-experts and specialists to evaluate nucleic-acid structures with greater confidence and to improve models in ways that support accurate biological interpretation.
Light-responsive proteins are involved in a wide range of essential physiological processes in bacteria, plants, and animals. Engineered light-responsive proteins have also emerged as prospective tools in biotechnology and biomedicine. These proteins are often characterized by short-lived lit states and the need for continuous illumination to reach photostationary states. Therefore, developing methods for studying light-responsive proteins and their interactions under illumination represents an important research goal. Here, we report on a novel front-illuminated surface plasmon resonance (fiSPR) biosensor for monitoring interactions involving light-responsive proteins. The fiSPR biosensor combines the optical platform based on the Kretschmann geometry with advanced transparent microfluidics and an additional light module, enabling in situ illumination of the liquid sample in contact with the SPR chip. We apply the fiSPR biosensor to study the blue light-responsive transcription factor EL222, which recovers to the dark state in a few seconds and plays an important role in the optogenetic control of gene expression. Specifically, we determine the rate and equilibrium constants for EL222 dimerization and DNA binding. The results support the hypothesis that EL222 dimerizes prior to binding DNA. In addition, we provide evidence of the interaction between an interleukin receptor modified with a photocaged tyrosine (IL-20R2-Y70NBY) and its cytokine ligand (IL-24) only upon UV illumination. Overall, this study demonstrates the versatility of the developed fiSPR biosensor for monitoring biomolecular interactions involving both natural and engineered light-responsive proteins, particularly those featuring short lit-state lifetimes.
It remains uncertain whether excited electronic state mixing occurs in the flavin cofactor of the light-oxygen-voltage-sensing (LOV) domain. In this study, we present transient absorption and femtosecond stimulated Raman spectra of both free and EL222 binding flavin mononucleotide (FMN). We observed a change in the shape of the excited-state absorption around 800 nm in the S1 state transient absorption after binding to EL222, alongside a relative intensity increase of the N1-C2 and C2═O2 stretching modes in the S1 state Raman spectra. Based on the previous calculated geometric differences between the ππ* and nπ* states, we propose a probable electronic state mixing in EL222 binding FMN. This mixing is favored by the nonsymmetric hydrogen bonding interaction between the flavin O4 atom and the asparagine residue and fewer hydrogen bonds with the O2 atom in EL222.
The activity of the light-oxygen-voltage/helix-turn-helix (LOV-HTH) photoreceptor EL222 is regulated through protein-protein and protein-DNA interactions, both triggered by photo-excitation of its flavin mononucleotide (FMN) cofactor. To gain molecular-level insight into the photocycle of EL222, we applied complementary methods: macromolecular X-ray crystallography (MX), nuclear magnetic resonance (NMR) spectroscopy, optical spectroscopies (infrared and UV-visible), molecular dynamics/metadynamics (MD/metaD) simulations, and protein engineering using noncanonical amino acids. Kinetic experiments provided evidence for two distinct EL222 conformations (lit1 and lit2) that become sequentially populated under illumination. These two lit states were assigned to covalently bound N5 protonated, and noncovalently bound hydroquinone forms of FMN, respectively. Only subtle structural differences were observed between the monomeric forms of all three EL222 species (dark, lit1, and lit2). While the dark state is largely monomeric, both lit states undergo monomer-dimer exchange. Furthermore, molecular modeling revealed differential dynamics and interdomain separation times arising from the three FMN states (oxidized, adduct, and reduced). Unexpectedly, all three EL222 species can associate with DNA, but only upon blue-light irradiation, a high population of stable complexes is obtained. Overall, we propose a model of EL222 activation where photoinduced changes in the FMN moiety shift the population equilibrium toward an open conformation that favors self-association and DNA-binding.
The activity of the transcription factor EL222 is regulated through protein-chromophore adduct formation, interdomain dynamics, oligomerization and protein-DNA interactions, all triggered by photo-excitation of its flavin mononucleotide (FMN) cofactor. To gain molecular-level insight into the photocycle of EL222, we applied complementary methods: macromolecular X-ray crystallography (MX), nuclear magnetic resonance (NMR) spectroscopy, optical spectroscopies (infrared and UV/visible), molecular dynamics/metadynamics (MD/metaD) simulations, and protein engineering using non-canonical amino acids. The observation of only subtle atomic displacements between crystal structures of EL222 with and without blue-light back-illumination, was confirmed by NMR data indicating no major changes in secondary structure and fold compactness. Kinetic experiments in solution provided evidence for two distinct EL222 conformations (lit1 and lit2) that become sequentially populated under illumination. These two lit states were assigned to covalently-bound N5 protonated, and non-covalently-bound hydroquinone forms of FMN, respectively. Molecular modeling revealed differential dynamics and domain separation times arising from the three FMN states (oxidized, adduct, and reduced). Furthermore, while the dark state is largely monomeric, both lit states undergo slow monomer-dimer exchange. The photoinduced loss of α-helicity, seen by infrared difference spectroscopy, was ascribed to dimeric EL222 species. Unexpectedly, NMR revealed that all three EL222 species (dark, lit1, lit2) can associate with DNA to some extent, but only under illumination a high population of stable complexes is obtained. Overall, we propose a refined model of EL222 photo-activation where photoinduced changes in the oxidation state of FMN and thioadduct formation shift the population equilibrium towards an open conformation that favors self-association and DNA-binding. ### Competing Interest Statement The authors have declared no competing interest.
Here, we present a previously undescribed approach to modify N-terminal sequences of recombinant proteins to increase their production yield in Escherichia coli. Prior research has demonstrated that the nucleotides immediately following the start codon can significantly influence protein expression. However, the impact of these sequences is construct-specific and is not universally applicable to all proteins. Most of the previous research has been limited to selecting from a few rationally designed sequences. In contrast, we used a directed evolution-based methodology, screening large numbers of diversified sequences derived from DNA libraries coding for the N-termini of investigated proteins. To facilitate the identification of cells with increased expression of the target construct, we cloned a GFP gene at the C-terminus of the expressed genes and used fluorescent activated cell sorting (FACS) to separate cells based on their fluorescence. By following this systematic workflow, we successfully elevated the yield of soluble recombinant proteins of multiple constructs up to over 30-fold.
Progress in cytokine engineering is driving therapeutic translation by overcoming these proteins' limitations as drugs. The IL-2 cytokine is a promising immune stimulant for cancer treatment but is limited by its concurrent activation of both pro-inflammatory immune effector cells and antiinflammatory regulatory T cells, toxicity at high doses, and short serum half-life. One approach to improve the selectivity, safety, and longevity of IL-2 is complexing with anti-IL-2 antibodies that bias the cytokine toward immune effector cell activation. Although this strategy shows potential in preclinical models, clinical translation of a cytokine/antibody complex is complicated by challenges in formulating a multiprotein drug and concerns regarding complex stability. Here, we introduced a versatile approach to designing intramolecularly assembled single-agent fusion proteins (immunocytokines, ICs) comprising IL-2 and a biasing anti-IL-2 antibody that directs the cytokine toward immune effector cells. We optimized IC construction and engineered the cytokine/antibody affinity to improve immune bias. We demonstrated that our IC preferentially activates and expands immune effector cells, leading to superior antitumor activity compared with natural IL-2, both alone and combined with immune checkpoint inhibitors. Moreover, therapeutic efficacy was observed without inducing toxicity. This work presents a roadmap for the design and translation of cytokine/antibody fusion proteins.
The EMDataResource Ligand Model Challenge aimed to assess the reliability and reproducibility of modeling ligands bound to protein and protein/nucleic-acid complexes in cryogenic electron microscopy (cryo-EM) maps determined at near-atomic (1.9-2.5 Å) resolution. Three published maps were selected as targets: E. coli beta-galactosidase with inhibitor, SARS-CoV-2 RNA-dependent RNA polymerase with covalently bound nucleotide analog, and SARS-CoV-2 ion channel ORF3a with bound lipid. Sixty-one models were submitted from 17 independent research groups, each with supporting workflow details. We found that (1) the quality of submitted ligand models and surrounding atoms varied, as judged by visual inspection and quantification of local map quality, model-to-map fit, geometry, energetics, and contact scores, and (2) a composite rather than a single score was needed to assess macromolecule+ligand model quality. These observations lead us to recommend best practices for assessing cryo-EM structures of liganded macromolecules reported at near-atomic resolution.
Binder H33 is a small protein binder engineered by ribosome display to bind human interleukin 10. Crystals of binder H33 display severe diffraction anisotropy. A set of data files with correction for diffraction anisotropy based on different local signal-to-noise ratios was prepared. Paired refinement was used to find the optimal anisotropic high-resolution diffraction limit of the data: 3.13-2.47 angstrom. The structure of binder H33 belongs to the 2% of crystal structures with the highest solvent content in the Protein Data Bank.
Human interleukin 24 (IL-24) is a multifunctional cytokine that represents an important target for autoimmune diseases and cancer. Since the biological functions of IL-24 depend on interactions with membrane receptors, on-demand regulation of the affinity between IL-24 and its cognate partners offers exciting possibilities in basic research and may have applications in therapy. As a proof-of-concept, we developed a strategy based on recombinant soluble protein variants and genetic code expansion technology to photocontrol the binding between IL-24 and one of its receptors, IL-20R2. Screening of non-canonical ortho-nitrobenzyl-tyrosine (NBY) residues introduced at several positions in both partners was done by a combination of biophysical and cell signaling assays. We identified one position for installing NBY, tyrosine70 of IL-20R2, which results in clear impairment of heterocomplex assembly in the dark. Irradiation with 365-nm light leads to decaging and reconstitutes the native tyrosine of the receptor that can then associate with IL-24. Photocaged IL-20R2 may be useful for the spatiotemporal control of the JAK/STAT phosphorylation cascade.
Time-resolved femtosecond-stimulated Raman spectroscopy (FSRS) provides valuable information on the structural dynamics of biomolecules. However, FSRS has been applied mainly up to the nanoseconds regime and above 700 cm−1, which covers only part of the spectrum of biologically relevant time scales and Raman shifts. Here we report on a broadband (~200–2200 cm−1) dual transient visible absorption (visTA)/FSRS set-up that can accommodate time delays from a few femtoseconds to several hundreds of microseconds after illumination with an actinic pump. The extended time scale and wavenumber range allowed us to monitor the complete excited-state dynamics of the biological chromophore flavin mononucleotide (FMN), both free in solution and embedded in two variants of the bacterial light-oxygen-voltage (LOV) photoreceptor EL222. The observed lifetimes and intermediate states (singlet, triplet, and adduct) are in agreement with previous time-resolved infrared spectroscopy experiments. Importantly, we found evidence for additional dynamical events, particularly upon analysis of the low-frequency Raman region below 1000 cm−1. We show that fs-to-sub-ms visTA/FSRS with a broad wavenumber range is a useful tool to characterize short-lived conformationally excited states in flavoproteins and potentially other light-responsive proteins.
Unnatural protein side-chains with frequencies within the “transparent window” spectral region (∼1800 - 2700 cm−1), where no native protein vibrations are found, can single out specific residues in vibrational spectra. These moieties contain most often triple bonds, like nitriles (R-C≡N), which can be conveniently attached to proteins in the form of genetically encoded non-canonical amino acids (ncAA) e.g. 4-cyanophenylalanine (CNF). However, the full potential of ncAA as vibrational reporters has not been exploited yet, particularly for Raman spectroscopy. Here we show how CNF residues introduced at multiple positions of the bacterial transcription factor EL222 provide a detailed picture of local changes along the photocycle. Time-resolved infrared spectra in the CNF absorption region disclose two additional kinetic events not sensed by the native probes, one occurring before flavin chromophore relaxation and the other after protein backbone relaxation. In addition, we incorporate an ncAA carrying a diacetylene (R-C≡C-C≡CH) via genetic code expansion technology into EL222 variants. The conjugated diyne features an intense Raman signal which is sensitive to the microenvironment around the triple bonds. We then measure light-induced conformational changes in the flavoprotein EL222 with single-residue precision. Our results suggest that vibrational spectroscopy assisted by ncAA can reveal protein structural dynamics residue-by-residue.
AbstractThe protein structure prediction problem has been solved for many types of proteins by AlphaFold. Recently, there has been considerable excitement to build off the success of AlphaFold and predict the 3D structures of RNAs. RNA prediction methods use a variety of techniques, from physics-based to machine learning approaches. We believe that there are challenges preventing the successful development of deep learning-based methods like AlphaFold for RNA in the short term. Broadly speaking, the challenges are the limited number of structures and alignments making data-hungry deep learning methods unlikely to succeed. Additionally, there are several issues with the existing structure and sequence data, as they are often of insufficient quality, highly biased and missing key information. Here, we discuss these challenges in detail and suggest some steps to remedy the situation. We believe that it is possible to create an accurate RNA structure prediction method, but it will require solving several data quality and volume issues, usage of data beyond simple sequence alignments, or the development of new less data-hungry machine learning methods.
We combined cell‐free ribosome display and cell‐based yeast display selection to build specific protein binders to the extracellular domain of the human interleukin 9 receptor alpha (IL‐9Rα). The target, IL‐9Rα, is the receptor involved in the signalling pathway of IL‐9, a pro‐inflammatory cytokine medically important for its involvement in respiratory diseases. The successive use of modified protocols of ribosome and yeast displays allowed us to combine their strengths—the virtually infinite selection power of ribosome display and the production of (mostly) properly folded and soluble proteins in yeast display. The described experimental protocol is optimized to produce binders highly specific to the target, including selectivity to common proteins such as BSA, and proteins potentially competing for the binder such as receptors of other cytokines. The binders were trained from DNA libraries of two protein scaffolds called 57aBi and 57bBi developed in our laboratory. We show that the described unconventional combination of ribosome and yeast displays is effective in developing selective small protein binders to the medically relevant molecular target.
Photoreceptors containing the light-oxygen-voltage (LOV) domain elicit biological responses upon excitation of their flavin mononucleotide (FMN) chromophore by blue light. The mechanism and kinetics of dark-state recovery are not well understood. Here we incorporated the non-canonical amino acid p-cyanophenylalanine (CNF) by genetic code expansion technology at 45 positions of the bacterial transcription factor EL222. Screening of light-induced changes in infrared (IR) absorption frequency, electric field and hydration of the nitrile groups identified residues CNF31 and CNF35 as reporters of monomer/oligomer and caged/decaged equilibria, respectively. Time-resolved multi-probe UV/visible and IR spectroscopy experiments of the lit-to-dark transition revealed four dynamical events. Predominantly, rearrangements around the A'α helix interface (CNF31 and CNF35) precede FMN-cysteinyl adduct scission, folding of α-helices (amide bands), and relaxation of residue CNF151. This study illustrates the importance of characterizing all parts of a protein and suggests a key role for the N-terminal A'α extension of the LOV domain in controlling EL222 photocycle length.
Human Interleukin 24 (IL-24) is an immunomodulatory cytokine that represents an important target for treatment of autoimmune diseases and cancer. Since the biological functions of IL-24 depend on interactions with its cell membrane receptors, on-demand regulation of the affinity between IL-24 and its cognate partners offers many exciting possibilities in basic research and may have also potential applications in therapy. In particular, control of protein-protein interactions by light has recently found increasing interest. As a proof-of-concept, we developed a strategy based on genetic code expansion technology to photo-control the binding between IL-24 and one of its receptors, IL-20R2. Both these proteins are engineered versions of the native proteins that feature high stability and can be recombinantly expressed in E. coli at adequate yields. Introduction of the photocaged non-canonical amino acid ortho-nitrobenzyl-tyrosine (ONBY) at selected positions of IL-20R2 resulted in severe impairment of heterocomplex assembly as determined by microscale thermophoresis. Irradiation of ONBY-substituted IL-20R2 with UV light (365 nm) reconstitutes the canonical tyrosines and subsequently restores the ability of IL-20R2 to recognize IL-24. We envision that photocontrollable IL-20R2 may be useful for the spatiotemporal regulation of signaling pathways in which IL-24 is involved.