Native amine dehydrogenases offer sustainable access to chiral amines, so the search for scaffolds capable of converting more diverse carbonyl compounds is required to reach the full potential of this alternative to conventional synthetic reductive aminations. Here we report a multidisciplinary strategy combining bioinformatics, chemoinformatics and biocatalysis to extensively screen billions of sequences in silico and to efficiently find native amine dehydrogenases features using computational approaches. In this way, we achieve a comprehensive overview of the initial native amine dehydrogenase family, extending it from 2,011 to 17,959 sequences, and identify native amine dehydrogenases with non-reported substrate spectra, including hindered carbonyls and ethyl ketones, and accepting methylamine and cyclopropylamine as amine donor. We also present preliminary model-based structural information to inform the design of potential (R)-selective amine dehydrogenases, as native amine dehydrogenases are mostly (S)-selective. This integrated strategy paves the way for expanding the resource of other enzyme families and in highlighting enzymes with original features. Sustainable chemistry can benefit from biocatalysis, but a high diversity of enzymes is needed. Here, the authors screen billions of protein sequences to provide an overview of the native amine dehydrogenase family for amine synthesis.
Abstract SulfAtlas (https://sulfatlas.sb-roscoff.fr/) is a knowledge-based resource dedicated to a sequence-based classification of sulfatases. Currently four sulfatase families exist (S1–S4) and the largest family (S1, formylglycine-dependent sulfatases) is divided into subfamilies by a phylogenetic approach, each subfamily corresponding to either a single characterized specificity (or few specificities in some cases) or to unknown substrates. Sequences are linked to their biochemical and structural information according to an expert scrutiny of the available literature. Database browsing was initially made possible both through a keyword search engine and a specific sequence similarity (BLAST) server. In this article, we will briefly summarize the experimental progresses in the sulfatase field in the last 6 years. To improve and speed up the (sub)family assignment of sulfatases in (meta)genomic data, we have developed a new, freely-accessible search engine using Hidden Markov model (HMM) for each (sub)family. This new tool (SulfAtlas HMM) is also a key part of the internal pipeline used to regularly update the database. SulfAtlas resource has indeed significantly grown since its creation in 2016, from 4550 sequences to 162 430 sequences in August 2022.
This repository contains the data specified in the paper entitled "A refined picture of the native Amine Dehydrogenase family revealed by extensive biodiversity screening". It includes: 1. NAD-dependent_enzymes.fa.gz - The library of 20,315,745 sequences of NADPH-dependent enzymes recovered from genomic and metagenomic sequence databases. 2. ref-AmDHs17959_nr.fa.gz - The library of 17,959 ref-AmDH sequences recovered from genomic and metagenomic sequence databases. Considered as the updated nat-AmDH family. 3. NAD_subfams_HMMs.tar.gz - The library of 104,686 Hidden Markov Models (HMMs) of NADPH-dependent protein subfamilies. As described in the paper, those HMMs were obtained by clustering the set uploaded here as NAD-dependent_enzymes.fa.gz and building one HMM per subfamily. 4. ref-AmDHs_HMMs.tar.gz - This repertory includes the HMMs used to update the nat-AmDH family (all_ASMC_no_nad,hmm and all_ASMC_nad_dom.hmm) as well as the ones used to search for distant homologs; HMMs designed for the phylogenetic and structure-based groups built from the set uploaded here as ref-AmDHs17959_nr.fa.gz (asmc_*.hmm and phylo_*.hmm) . 5. ref-AmDH_ASMC_models.tar.gz - The library of 9886 ref-AmDH models built using the ASMC pipeline. 6. 72_ref-AmDH_seqs_representatives.txt.gz - Sequences of the 72 representative ref-AmDHs experimentally tested and found to be active. 7. 17_nat-AmDH_seqs_specific_feature.txt.gz - Sequences of the 17 nat-AmDHs with specific feature that have been heterologously expressed and tested.
Background The growing availability of large genomic datasets presents an opportunity to discover novel metabolic pathways and enzymatic reactions profitable for industrial or synthetic biological applications. Efforts to identify new enzyme functions in this substantial number of sequences cannot be achieved without the help of bioinformatics tools and the development of new strategies. The classical way to assign a function to a gene uses sequence similarity. However, another way is to mine databases to identify conserved gene clusters (i.e. syntenies) as, in prokaryotic genomes, genes involved in the same pathway are frequently encoded in a single locus with an operonic organisation. This Genomic Context (GC) conservation is considered as a reliable indicator of functional relationships, and thus is a promising approach to improve the gene function prediction. Methods Here we present NetSyn (Network Synteny), a tool, which aims to cluster protein sequences according to the similarity of their genomic context rather than their sequence similarity. Starting from a set of protein sequences of interest, NetSyn retrieves neighbouring genes from the corresponding genomes as well as their protein sequence. Homologous protein families are then computed to measure synteny conservation between each pair of input sequences using a GC score. A network is then created where nodes represent the input proteins and edges the fact that two proteins share a common GC. The weight of the edges corresponds to the synteny conservation score. The network is then partitioned into clusters of proteins sharing a high degree of synteny conservation. Results As a proof of concept, we used NetSyn on two different datasets. The first one is made of homologous sequences of an enzyme family (the BKACE family, previously named DUF849) to divide it into sub-families of specific activities. NetSyn was able to go further by providing additional subfamilies in addition to those previously published. The second dataset corresponds to a set of non-homologous proteins consisting of different Glycosyl Hydrolases (GH) with the aim of interconnecting them and finding conserved operon-like genomic structures. NetSyn was able to detect the locus of Cellvibrio japonicus for the degradation of xyloglucan. It contains three non-homologous GH and was found conserved in fourteen bacterial genomes. Discussion NetSyn is able to cluster proteins according to their genomic context which is a way to make functional links between proteins without taking into count their sequence similarity only. We showed that NetSyn is efficient in exploring large protein families to define iso-functional groups. It can also highlight functional interactions between proteins from different families and predicts new conserved genomic structures that have not yet been experimentally characterised. NetSyn can also be useful in pinpointing mis-annotations that have been propagated in databases and in suggesting annotations on proteins currently annotated as “unknown”. NetSyn is freely available at https://github.com/labgem/netsyn.
Macroalgae contribute substantially to primary production in coastal ecosystems. Their biomass, mainly consisting of polysaccharides, is cycled into the environment by marine heterotrophic bacteria using largely uncharacterized mechanisms. Here we describe the complete catabolic pathway for carrageenans, major cell wall polysaccharides of red macroalgae, in the marine heterotrophic bacterium Zobellia galactanivorans. Carrageenan catabolism relies on a multifaceted carrageenan-induced regulon, including a non-canonical polysaccharide utilization locus (PUL) and genes distal to the PUL, including a susCD-like pair. The carrageenan utilization system is well conserved in marine Bacteroidetes but modified in other phyla of marine heterotrophic bacteria. The core system is completed by additional functions that might be assumed by non-orthologous genes in different species. This complex genetic structure may be the result of multiple evolutionary events including gene duplications and horizontal gene transfers. These results allow for an extension on the definition of bacterial PUL-mediated polysaccharide digestion.
Hydroxypyruvate was shown to be a nucleophile for class II pyruvate aldolases isolated from biodiversity, allowing unprecedented stereoselective cross-aldol reactions.
A high-throughput screening for the identification of nitrilases demonstrating activity towards alpha-aminonitriles is reported. A LC–MS assay giving access to both conversion and enantiospecificity was developed. 588 candidate enzymes were screened as cell lysates against six alpha-aminonitriles in 96-well microplates. The candidate enzymes were selected following two criteria, their sequence identity with a set of known nitrilases or their phylogenetic position among the nitrilase superfamily. Five enzymes were identified and found to hydrolyse alpha-aminonitrile into the corresponding alpha-aminoacid. The substrate range was found to be very narrow as only two different alpha-aminonitriles, 2-aminovaleronitrile and 2-amino-2-phenylacetonitrile, were found to be substrates. The biocatalytic capabilities of three enzymes were further investigated and the best result was obtained with an enzyme from Burkholderia xenovorans catalysing the enantiospecific hydrolysis of 2-aminovaleronitrile into (S)-norvaline with excellent conversion and enantiomeric excess.
The emergence of Next Generation Sequencing generates an incredible amount of sequence and great potential for new enzyme discovery. Despite this huge amount of data and the profusion of bioinformatic methods for function prediction, a large part of known enzyme activities is still lacking an associated protein sequence. These particular activities are called "orphan enzymes". The present review proposes an update of previous surveys on orphan enzymes by mining the current content of public databases. While the percentage of orphan enzyme activities has decreased from 38% to 22% in ten years, there are still more than 1,000 orphans among the 5,000 entries of the Enzyme Commission (EC) classification. Taking into account all the reactions present in metabolic databases, this proportion dramatically increases to reach nearly 50% of orphans and many of them are not associated to a known pathway. We extended our survey to "local orphan enzymes" that are activities which have no representative sequence in a given clade, but have at least one in organisms belonging to other clades. We observe an important bias in Archaea and find that in general more than 30% of the EC activities have incomplete sequence information in at least one superkingdom. To estimate if candidate proteins for local orphans could be retrieved by homology search, we applied a simple strategy based on the PRIAM software and noticed that candidates may be proposed for an important fraction of local orphan enzymes. Finally, by studying relation between protein domains and catalyzed activities, it appears that newly discovered enzymes are mostly associated with already known enzyme domains. Thus, the exploration of the promiscuity and the multifunctional aspect of known enzyme families may solve part of the orphan enzyme issue. We conclude this review with a presentation of recent initiatives in finding proteins for orphan enzymes and in extending the enzyme world by the discovery of new activities.
A high-throughput screening of candidate nitrilases against 25 structurally diverse substrates allowed us to create a wide collection of 125 experimentally validated nitrilases. The enzymes were selected by genomic approach from 700 diverse prokaryotic species and one metagenome as representative of the nitrilase family diversity. The enzymatic screening of this collection expands the biocatalytic toolbox for chemical synthesis by providing a large number of tested nitrilases with their assigned substrates. Three examples illustrate the synthetic potential of our enzyme collection. The syntheses of carboxylic acid building blocks, a β-substituted phenylpropanoic acid, a cyclic γ-keto carboxylic acid and a mononitrile monocarboxylic acid, were achieved from the corresponding nitrile substrates, using three new nitrilases (two from Sphingomonas wittichii and one from Syntrophobacter fumaroxidans). Improvements of nitrilase activities through the optimization of reaction parameters and the preparative biocatalytic synthesis are presented for these three examples.
The classification of carbohydrate-active enzymes, presents a number of challenges to conciliate sequence and genome information with protein structure and reaction mechanisms. The principles of the sequence- and structure-based classification according to CAZy, combining structural and mechanistic information with bioinformatics, are described. In this chapter, we give a particular emphasis to the classification and description of glycosyltransferases relying on phospho-activated sugars. Issued from a limited number of protein folds, the corresponding diverse set of enzyme families was subdivided according to the nature of the phospho-activated sugar (mono- vs. diphospho-activated sugars; axial- vs. equatorial-linked glycosidic moieties; nature of activator) and the type of catalytic mechanism (inverting vs. retaining) to show that these features are usually strictly conserved at family level. Given the practical difficulties to determine glycosyltransferase activity, existing knowledge of the common features found at family level should be used to limit the range of activities to assess experimentally.
Family GH13, also known as the alpha-amylase family, is the largest sequence-based family of glycoside hydrolases and groups together a number of different enzyme activities and substrate specificities acting on alpha-glycosidic bonds. This polyspecificity results in the fact that the simple membership of this family cannot be used for the prediction of gene function based on sequence alone. In order to establish robust groups that show an improved correlation between sequence and enzymatic specificity, we have performed a large-scale analysis of 1691 family GH13 sequences by combining clustering, similarity search and phylogenetic methods. About 80% of the sequences could be reliably classified into 35 subfamilies. Most subfamilies appear monofunctional (i.e. contain enzymes with the same substrate and the same product). The close examination of the other, apparently polyspecific, subfamilies revealed that they actually group together enzymes with strongly related (or even sometimes virtually identical) activities. Overall our subfamily assignment allows to set the limits for genomic function prediction on this large family of biologically and industrially important enzymes.
Because of the fast accumulation of sequences derived from genome sequencing efforts, the sampling of the sequence space in glycosidase and related enzyme families is such that sensitive sequence similarity detection methods like PSI-BLAST are now able to reveal distant, but clear, structural and evolutionary relations between glycosidases acting on α- and β-bonds. We have observed this trend within groups of glycosidases with completely different folds. We postulate that the evolutionary interconversion between α- and β-acting glycosidases was greatly facilitated by the fact that both types share a similar axial orientation of the glycosidic bond in the reactive bound substrate. Glycosides in the β anomeric configuration, require a sugar ring distortion, resulting in an axial orientation of the glycosidic bond, equivalent to that of an α glycosidic bond, prior to displacement by nucleophilic substitution.
Wood formation is a fundamental biological process with significant economic interest. While lignin biosynthesis is currently relatively well understood, the pathways leading to the synthesis of the key structural carbohydrates in wood fibers remain obscure. We have used a functional genomics approach to identify enzymes involved in carbohydrate biosynthesis and remodeling during xylem development in the hybrid aspen Populus tremula x tremuloides. Microarrays containing cDNA clones from different tissue-specific libraries were hybridized with probes obtained from narrow tissue sections prepared by cryosectioning of the developing xylem. Bioinformatic analyses using the sensitive tools developed for carbohydrate-active enzymes allowed the identification of 25 xylem-specific glycosyltransferases belonging to the Carbohydrate-Active EnZYme families GT2, GT8, GT14, GT31, GT43, GT47, and GT61 and nine glycosidases (or transglycosidases) belonging to the Carbohydrate-Active EnZYme families GH9, GH10, GH16, GH17, GH19, GH28, GH35, and GH51. While no genes encoding either polysaccharide lyases or carbohydrate esterases were found among the secondary wall-specific genes, one putative O-acetyltransferase was identified. These wood-specific enzyme genes constitute a valuable resource for future development of engineered fibers with improved performance in different applications.