Terrestrial plants emerged from the water about 500 million years ago. Thereafter, they have diversified and now inhabit most of the Earth's surface. More recently, some species have re-adapted to an aquatic lifestyle, both in fresh and salt water, and fully or partially submerged. The mechanisms enabling these adaptations between terrestrial and aquatic life are extremely numerous, making it difficult to have a comprehensive overview of the phenomenon. Here, we performed a series of intraspecific measurements of the selection pressure affecting orthologous genes in eight aquatic and four terrestrial plants. Our analyses showed that aquatic plants have a relaxed selection pressure on nutrient assimilation mechanisms, probably linked to a greater bioavailability, as well as stronger adaptations to oxidative stress, while terrestrial plants evolution is linked to environment perception. Inter-species analyses have also highlighted a different evolution of chloroplast proteins between these two types of plants, suggesting adaptations to gas availability.
Formerly considered as part of “junk DNA”, pseudogenes are nowadays known for their role in the post-transcriptional regulation of functional genes. Their identification also contributes to a better understanding of gene evolution, particularly in relation to adaptive responses and the evolution of multigene families. Despite this, there is, to our knowledge, no fully automatic pipeline allowing annotation of the pseudogenes on a whole genome. Here, we propose a new software named Pseudo-Gene Retriever (P-GRe). This is a completely automated pseudogene prediction tool requiring only a genome sequence, its corresponding GFF annotation file, and a protein sequences file. The aligner miniprot has been integrated in our pipeline, because of its high speed and sensitivity. With several filtering and post-analysis steps P-GRe outperforms existing software, while being more sensitive and bringing the new capacity of annotating unitary pseudogenes.
Climate change is expected to intensify the occurrence of abiotic stress in plants, such as hypoxia and salt stresses, leading to the production of reactive oxygen species (ROS), which need to be effectively managed by various oxido-reductases encoded by the so-called ROS gene network. Here, we studied six oxido-reductases families in three Brassicaceae species, Arabidopsis thaliana as well as Nasturtium officinale and Eutrema salsugineum, which are adapted to hypoxia and salt stress, respectively. Using available and new genomic data, we performed a phylogenomic analysis and compared RNA-seq data to study genomic and transcriptomic adaptations. This comprehensive approach allowed for the gaining of insights into the impact of the adaptation to saline or hypoxia conditions on genome organization (gene gains and losses) and transcriptional regulation. Notably, the comparison of the N. officinale and E. salsugineum genomes to that of A. thaliana highlighted changes in the distribution of ohnologs and homologs, particularly affecting class III peroxidase genes (CIII Prxs). These changes were specific to each gene, to gene families subjected to duplication events and to each species, suggesting distinct evolutionary responses. The analysis of transcriptomic data has allowed for the identification of genes related to stress responses in A. thaliana, and, conversely, to adaptation in N. officinale and E. salsugineum.
Formerly considered as part of “junk DNA”, pseudogenes are nowadays known for their role in the post-transcriptional regulation of functional genes. In addition, their identification allows a better understanding of gene evolution in the frame of multigenic families. Despite this, there is, to our knowledge, no fully automatic user-friendly software allowing the annotation of pseudogenes on a whole genome. Here, we present Pseudo-Gene Retriever (P-GRe), a fully automated pseudogene prediction software requiring only a genome sequence and its corresponding GFF annotation file. P-GRe detects the sequences of the pseudogenes on a whole genome and returns to the user all their genomic sequences and their pseudo-coding sequences. The ability of P-GRe to finely reconstruct the structure of pseudogenes also allow to obtain a set of proteins virtually encoded by the predicted pseudogenes. We show here that in 70% of the cases, virtual proteins constructed by P-GRe from Arabidopsis thaliana proteome and genome aligned better to their parent protein than their annotated counterpart. ### Competing Interest Statement The authors have declared no competing interest.
In the last few decades, the explosion of genomic projects has produced huge sets of predicted genes and annotated sequences. The prediction of a gene structure can be defined as the capacity to determine the start and the stop of the gene as well as the positions of introns, if present. Despite the number of performant gene prediction programs combining ab initio and homology-based approaches (Mathe et al., 2002; Hoff and Stanke, 2015), the rate of mis-predicted genes is not negligible and can be due to several factors (Scalzitti et al., 2020). For example, unusually long introns, short exons or long genes can generate incomplete or partially predicted gene structure; short intergenic regions can lead to gene fusion; DNA sequencing errors (nucleotide deletions or insertions) introducing frameshifts can affect predictions; non-canonical splice sites, overlapping genes and genes located within introns are also a source of erroneous predictions. Due to high sequence identity and duplication rate, the risks of mis-prediction are exacerbated in the case of multigenic families (Figure 1, Fawal et al., 2014). In addition, protein annotation or function assignment, based on the presence of a hypothetical protein domain or on homology with known proteins, can also lead to an inappropriate annotation. The risk of mis-annotations is high for proteins containing multiple domains or small domain(s) common to several classes of proteins. For example, the PFAM domain PF07992 (Pyridine nucleotide-disulphide oxidoreductase) is detected in MonoDehydroAscorbate Reductases (MDARs), Glutathione Reductases (GRs), and in the Thioredoxin family (Trx) but does not discriminate between these three different families (Table 1). Mis-annotations are also observed for proteins belonging to superfamilies with conserved domain and large number of protein families and classes. As an example, 198 genes of the MYB superfamily have been detected in Arabidopsis thaliana (Yanhui et al., 2006), but the PFAM domain PF00249 (Myb_DNA-binding) does not discriminate between the R2R3-MYB, the R1R2R3-MYB, the MYB-related, and the atypical MYB families. In addition, the PF00249 entry also contains the SANT domain, which has a strong structural similarity to the Myb domain but is functionally divergent. Therefore, using this PFAM entry to extract MYB proteins returns many false positives (total of 326 sequences from A. thaliana).
Ligninolytic peroxidases are microbial enzymes involved in depolymerisation of lignin, a plant cell wall polymer found in land plants. Among fungi, only Dikarya were found to degrade lignin. The increase of available fungal genomes allows performing an expert annotation of lignin-degrading peroxidase encoding sequences with a particular focus on Class II peroxidases (CII Prx). In addition to the previously described LiP, MnP and VP classes, based on sequence similarity, six new sub-classes have been defined: three found in plant pathogen ascomycetes and three in basidiomycetes. The presence of CII Prxs could be related to fungal life style. Typically, necrotrophic or hemibiotrophic fungi, either ascomycetes or basidiomycetes, possess CII Prxs while symbiotic, endophytic or biotrophic fungi do not. CII Prxs from ascomycetes are rarely subjected to duplications unlike those from basidiomycetes, which can form large recent duplicated families. Even if these CII Prxs classes form two well distinct clusters with divergent gene structures (intron numbers and positions), they share the same key catalytic residues suggesting that they evolved independently from similar ancestral sequences with few or no introns. The lack of CII Prxs encoding sequences in early diverging fungi, together with the absence of duplicated class I peroxidase (CcP) in fungi containing CII Prxs, suggests the potential emergence of an ancestral CII Prx sequence from the duplicated CcP after the separation between ascomycetes and basidiomycetes. As some ascomycetes and basidiomycetes did not possess CII Prx, late gene loss could have occurred.
We present a new database, specifically devoted to ROS homeostasis regulated proteins. This database replaced our previous database, the PeroxiBase, which was focused only on various peroxidase families. The addition of 20 new protein families related with ROS homeostasis justifies the new name for this more complex and comprehensive database as RedoxiBase. Besides enlarging the focus of the database, new analysis tools and functionalities have been developed and integrated through the web interface, with which the users can now directly access to orthologous sequences and see the chromosomal localization of sequences when available. OrthoMCL tool, completed with a post-treatment process, provides precise predictions of orthologous gene groups for the sequences present in this database. In order to explore and analyse orthogroups results, taxonomic visualization of organisms containing sequence of a specific orthogroup as well as chromosomal distribution of the orthogroup with one or two organisms have been included.
Plant organisms contain a large number of genes belonging to numerous multigenic families whose evolution size reflects some functional constraints. Sequences from eight multigenic families, involved in biotic and abiotic responses, have been analyzed in Eucalyptus grandis and compared with Arabidopsis thaliana. Two transcription factor families APETALA 2 (AP2)/ethylene responsive factor and GRAS, two auxin transporter families PIN-FORMED and AUX/LAX, two oxidoreductase families (ascorbate peroxidases [APx] and Class III peroxidases [CIII Prx]), and two families of protective molecules late embryogenesis abundant (LEA) and DNAj were annotated in expert and exhaustive manner. Many recent tandem duplications leading to the emergence of species-specific gene clusters and the explosion of the gene numbers have been observed for the AP2, GRAS, LEA, PIN, and CIII Prx in E. grandis, while the APx, the AUX/LAX and DNAj are conserved between species. Although no direct evidence has yet demonstrated the roles of these recent duplicated genes observed in E. grandis, this could indicate their putative implications in the morphological and physiological characteristics of E. grandis, and be the key factor for the survival of this nondormant species. Global analysis of key families would be a good criterion to evaluate the capabilities of some organisms to adapt to environmental variations.
A major challenge facing bioinformatics today is the efficient annotation of the exponential flow of genomic data. This has led to an increasing dependence on automatic annotation procedures, despite the relatively high error rates of these programs, particularly for multigenic families. We discuss here the errors and biases introduced by automatic genome annotations, focusing on issues with structural annotations of gene families, and suggest ways to overcome these limitations.
SUMMARY GECA is a fast, user-friendly and freely-available tool for representing gene exon/intron organization and highlighting changes in gene structure among members of a gene family. It relies on protein alignment, completed with the identification of common introns in the corresponding genes using CIWOG. GECA produces a main graphical representation showing the resulting aligned set of gene structures, where exons are to scale. The important and original feature of GECA is that it combines these gene structures with a symbolic display highlighting sequence similarity between subsequent genes. It is worth noting that this combination of gene structure with the indications of similarities between related genes allows rapid identification of possible events of gain or loss of introns, or points to erroneous structural annotations. The output image is generated in a portable network graphics format which can be used for scientific publications.
The PeroxiBase (http://peroxibase.toulouse.inra.fr/) is a specialized database devoted to peroxidases' families, which are major actors of stress responses. In addition to the increasing number of sequences and the complete modification of the Web interface, new analysis tools and functionalities have been developed since the previous publication in the NAR database issue. Nucleotide sequences and graphical representation of the gene structure can now be included for entries containing genomic cross-references. An expert semi-automatic annotation strategy is being developed to generate new entries from genomic sequences and from EST libraries. Plus, new internal and automatic controls have been included to improve the quality of the entries. To compare gene structure organization among families' members, two new tools are available, CIWOG to detect common introns and GECA to visualize gene structure overlaid with sequence conservation. The multicriteria search tool was greatly improved to allow simple and combined queries. After such requests or a BLAST search, different analysis processes are suggested, such as multiple alignments with ClustalW or MAFFT, a platform for phylogenetic analysis and GECA's display in association with a phylogenetic tree. Finally, we updated our family specific profiles implemented in the PeroxiScan tool and made new profiles to consider new sub-families.
Phylogenetic, genomic and functional analyses have allowed the identification of a new class of putative heme peroxidases, so called APx-R (APx-Related). These new class, mainly present in the green lineage (including green algae and land plants), can also be detected in other unicellular chloroplastic organisms. Except for recent polyploid organisms, only single-copy of APx-R gene was detected in each genome, suggesting that the majority of the APx-R extra-copies were lost after chromosomal or segmental duplications. In a similar way, most APx-R co-expressed genes in Arabidopsis genome do not have conserved extra-copies after chromosomal duplications and are predicted to be localized in organelles, as are the APx-R. The member of this gene network can be considered as unique gene, well conserved through the evolution due to a strong negative selection pressure and a low evolution rate.
The transition from water to land was a major evolutionary step for the green lineage. Based on fossil data, this event probably occurred some 480-430 million years ago, during the Ordovician and the early Silurian and initiated the explosive evolution that led to the modern diversity of photosynthetic organisms living on Earth. The chronological steps are still puzzling, but the great advances in genetics have allowed some of them to be positioned on the time axis.Chloroplastic organisms evolving towards terrestrialization have had to solve many problems: limited water supply, scarcity of mineral and especially phosphorus, harmful effect of ultraviolet and cosmic rays, pronounced fluctuations of temperature and attacks from new and diversified microbes. Many adaptations, such as the modification of the life cycle (sporophytes, seeds), organ diversification (root and leaves), the appearance of complex phenolic compounds (lignin, flavonoids), vascularization, the accumulation of new compounds (cutin, suberin), the development of specialized cells and the establishment of symbiotic interactions, have all played major roles during the transition from water to land and have resulted in the rich plant biodiversity of today. Some molecular and biochemical aspects putatively associated with land plant emergence are summarized here. (C) 2011 Elsevier GmbH. All rights reserved.
Class III peroxidases are members of a large multigenic family, only detected in the plant kingdom and absent from green algae sensu stricto (chlorophyte algae or Chlorophyta). Their evolution is thought to be related to the emergence of the land plants. However class III peroxidases are present in a lower copy number in some basal Streptophytes (Charapyceae), which predate land colonization. Gene structures are variable among organisms and within species with respect to the number of introns, but their positions are highly conserved. Their high copy number, as well as their conservation could be related to plant complexity and adaptation to increasing stresses. No specific function has been assigned to respective isoforms, but in large multigenic families, particular structure–function relations can be expected. Plant peroxidase sequences contain highly conserved residues and motifs, variable domains surrounded by conserved residues and present a low identity level among their promoter regions, further suggesting the existence of sub-functionalization of the different isoforms.
A pathosystem between Aphanomyces euteiches, the causal agent of pea root rot disease, and the model legume Medicago truncatula was developed to gain insights into mechanisms involved in resistance to this oomycete. The F83005.5 French accession and the A17-Jemalong reference line, susceptible and partially resistant, respectively, to A. euteiches, were selected for further cytological and genetic analyses. Microscopy analyses of thin root sections revealed that a major difference between the two inoculated lines occurred in the root stele, which remained pathogen free in A17. Striking features were observed in A17 roots only, including i) frequent pericycle cell divisions, ii) lignin deposition around the pericycle, and iii) accumulation of soluble phenolic compounds. Genetic analysis of resistance was performed on an F7 population of 139 recombinant inbred lines and identified a major quantitative trait locus (QTL) near the top of chromosome 3. A second study, with near-isogenic line responses to A. euteiches confirmed the role of this QTL in expression of resistance. Fine-mapping allowed the identification of a 135-kb sequenced genomic DNA region rich in proteasome-related genes. Most of these genes were shown to be induced only in inoculated A17. Novel mechanisms possibly involved in the observed partial resistance are proposed.
Aphanomyces euteiches is an oomycete pathogen that causes seedling blight and root rot of legumes, such as alfalfa and pea. The genus Aphanomyces is phylogenically distinct from well-studied oomycetes such as Phytophthora sp., and contains species pathogenic on plants and aquatic animals. To provide the first foray into gene diversity of A. euteiches, two cDNA libraries were constructed using mRNA extracted from mycelium grown in an artificial liquid medium or in contact to plant roots. A unigene set of 7,977 sequences was obtained from 18,864 high-quality expressed sequenced tags (ESTs) and characterized for potential functions. Comparisons with oomycete proteomes revealed major differences between the gene content of A. euteiches and those of Phytophthora species, leading to the identification of biosynthetic pathways absent in Phytophthora, of new putative pathogenicity genes and of expansion of gene families encoding extracellular proteins, notably different classes of proteases. Among the genes specific of A. euteiches are members of a new family of extracellular proteins putatively involved in adhesion, containing up to four protein domains similar to fungal cellulose binding domains. Comparison of A. euteiches sequences with proteomes of fully sequenced eukaryotic pathogens, including fungi, apicomplexa and trypanosomatids, allowed the identification of A. euteiches genes with close orthologs in these microorganisms but absent in other oomycetes sequenced so far, notably transporters and non-ribosomal peptide synthetases, and suggests the presence of a defense mechanism against oxidative stress which was initially characterized in the pathogenic trypanosomatids.
In this era of whole genome sequencing, reliable genome annotations ( identification of functional regions) are the cornerstones for many subsequent analyses. Not only is careful annotation important for studying the gene and gene family content of a genome and its host, but also for wide-scale transcriptome and proteome analyses attempting to describe a certain biological process or to get a global picture of a cell's behavior. Although the number of sequenced genomes is increasing thanks to the application of new technologies, genome-wide analyses will critically depend on the quality of the genome annotations. However, the annotation process is more complicated in the plant field than in the animal field because of the limited funding that leads to much fewer experimental data and less annotation expertise. This situation calls for highly automated annotation platforms that can make the best use of all available data, experimental or not. We discuss how the gene prediction (the process of predicting protein gene structures in genomic sequences) research field increasingly shifts from methods that typically exploited one or two types of data to more integrative approaches that simultaneously deal with various experimental, statistical, or other in silico evidence. We illustrate the importance of integrative approaches for producing high-quality automatic annotations of genomes of plants and algae as well as of fungi that live in close association with plants using the platform EuGene as an example.
BACKGROUND:The Oomycete genus Aphanomyces comprises devastating plant and animal pathogens. However, little is known about the molecular mechanisms underlying pathogenicity of Aphanomyces species. In this study, we report on the development of a public database called AphanoDB which is dedicated to Aphanomyces genomic data. As a first step, a large collection of Expressed Sequence Tags was obtained from the legume pathogen A. euteiches, which was then processed and collected into AphanoDB.DESCRIPTION:Two cDNA libraries of A. euteiches were created: one from mycelium growing on synthetic medium and one from mycelium grown in contact to root tissues of the model legume Medicago truncatula. From these libraries, 18,684 expressed sequence tags were obtained and assembled into 7,977 unigenes which were compared to public databases for annotation. Queries on AphanoDB allow the users to retrieve information for each unigene including similarity to known protein sequences, protein domains and Gene Ontology classification. Statistical analysis of EST frequency from the two different growth conditions was also added to the database.CONCLUSION:AphanoDB is a public database with a user-friendly web interface. The sequence report pages are the main web interface which provides all annotation details for each unigene. These interactive sequence report pages are easily available through text, BLAST, Gene Ontology and expression profile search utilities. AphanoDB is available from URL: http://www.polebio.scsv.ups-tlse.fr/aphano/.
We recently cloned a novel human nuclear factor (designated THAP1) from postcapillary venule endothelial cells (ECs) that contains a DNA-binding THAP domain, shared with zebrafish E2F6 and several Caenorhabditis elegans proteins interacting genetically with retinoblastoma gene product (pRB). Here, we show that THAP1 is a physiologic regulator of EC proliferation and cell-cycle progression, 2 essential processes for angiogenesis. Retroviral-mediated gene transfer of THAP1 into primary human ECs inhibited proliferation, and large-scale expression profiling with microarrays revealed that THAP1-mediated growth inhibition is due to coordinated repression of pRB/E2F cell-cycle target genes. Silencing of endogenous THAP1 through RNA interference similarly inhibited EC proliferation and G1/S cell-cycle progression, and resulted in down-regulation of several pRB/E2F cell-cycle target genes, including RRM1, a gene required for S-phase DNA synthesis. Chromatin immunoprecipitation assays in proliferating ECs showed that endogenous THAP1 associates in vivo with a consensus THAP1-binding site found in the RRM1 promoter, indicating that RRM1 is a direct transcriptional target of THAP1. The similar phenotypes observed after THAP1 overexpression and silencing suggest that an optimal range of THAP1 expression is essential for EC proliferation. Together, these data provide the first links in mammals among THAP proteins, cell proliferation, and pRB/E2F cell-cycle pathways.
We have recently described an evolutionarily conserved protein motif, designated the THAP domain, which defines a previously uncharacterized family of cellular factors (THAP proteins). The THAP domain exhibits similarities to the site-specific DNA-binding domain of Drosophila P element transposase, including a putative metal-coordinating C2CH signature (CX(2-4)CX(35-53)CX(2)H). In this article, we report a comprehensive list of approximately 100 distinct THAP proteins in model animal organisms, including human nuclear proapoptotic factors THAP1 and DAP4/THAP0, transcriptional repressor THAP7, zebrafish orthologue of cell cycle regulator E2F6, and Caenorhabditis elegans chromatin-associated protein HIM-17 and cell-cycle regulators LIN-36 and LIN-15B. In addition, we demonstrate the biochemical function of the THAP domain as a zinc-dependent sequence-specific DNA-binding domain belonging to the zinc-finger superfamily. In vitro binding-site selection allowed us to identify an 11-nucleotide consensus DNA-binding sequence specifically recognized by the THAP domain of human THAP1. Mutations of single nucleotide positions in this sequence abrogated THAP-domain binding. Experiments with the zinc chelator 1,10-o-phenanthroline revealed that the THAP domain is a zinc-dependent DNA-binding domain. Site-directed mutagenesis of single cysteine or histidine residues supported a role for the C2CH motif in zinc coordination and DNA-binding activity. The four other conserved residues (P, W, F, and P), which define the THAP consensus sequence, were also found to be required for DNA binding. Together with previous genetic data obtained in C. elegans, our results suggest that cellular THAP proteins may function as zinc-dependent sequence-specific DNA-binding factors with roles in proliferation, apoptosis, cell cycle, chromosome segregation, chromatin modification, and transcriptional regulation.
Pierre Rouzé合作论文数Ghent University
Bioinformatics & Evolutionary Genomics;Laboratoire Associ?e l'INRA;VIB4