Whole-genome sequencing of the protozoan pathogen Trypanosoma cruzi revealed that the diploid genome contains a predicted 22,570 proteins encoded by genes, of which 12,570 represent allelic pairs. Over 50% of the genome consists of repeated sequences, such as retrotransposons and genes for large families of surface molecules, which include trans-sialidases, mucins, gp63s, and a large novel family (>1300 copies) of mucin-associated surface protein (MASP) genes. Analyses of the T. cruzi, T. brucei, and Leishmania major (Tritryp) genomes imply differences from other eukaryotes in DNA repair and initiation of replication and reflect their unusual mitochondrial DNA. Although the Tritryp lack several classes of signaling molecules, their kinomes contain a large and diverse set of protein kinases and phosphatases; their size and diversity imply previously unknown interactions and regulatory processes, which may be targets for intervention.
Protein tyrosine kinases and phosphatases play important roles in the regulation of cell growth, development, and differentiation. We report here the identification in Trypanosoma cruzi of a gene (TcPRL-1) encoding a protein tyrosine phosphatase. The predicted protein (TcPRL-1) shares ca. 35% identity with the mammalian protein tyrosine phosphatase known as phosphatase of regenerating liver 1 (PRL-1). Four copies of this protein tyrosine phosphatase are present in the T. cruzi genome, and Northern blot assays showed a transcript of approximately 750 bases. TcPRL-1 was detected by Western blot analysis only in amastigote extracts as a 21-kDa protein. TcPRL-1 was expressed in Escherichia coli, and its phosphatase activity was determined by using p-nitrophenylphosphate and a phosphorylated protein as substrates. In contrast to other PRLs, TcPRL-1 activity was not affected by pentamidine, and it was inhibited by very low concentrations of o-vanadate. TcPRL-1 has a C-terminal CAAX motif (CAVM) and is farnesylated in vitro by T. cruzi epimastigote extracts and in vivo according to the transfection results. After transfection of T. cruzi with a vector that expresses TcPRL-1 as a C-terminal fusion to green fluorescent protein, GFP-TcPRL-1 was detected in the endocytic pathway of epimastigotes, amastigotes, and trypomastigotes by colocalization with cruzipain and concanavalin A. Interestingly, a mutant form without the CAAX motif localized to the cytoplasm, in contrast to its mammalian counterparts that localize to the nucleus. The results of these studies on TcPRL-1 reveal that, even though the animal and parasite PRLs share similar kinetic properties, their susceptibilities to inhibitors, as well as their localization, are distinct, implying that they may be involved in different cellular processes.
Adrenal corticosteroids influence the function of the hippocampus, the brain structure in which the highest expression of glucocorticoid receptors is found. Chronic high levels of cortisol elicited by stress or through exogenous administration can cause irreversible damage and cognitive deficits. In this study, we searched for genes expressed in the hippocampal formation after chronic cortisol treatment in male tree shrews. Animals were treated orally with cortisol for 28 days. At the end of the experiments, we generated two subtractive hippocampal hybridization libraries from which we sequenced 2,246 expressed sequenced tags (ESTs) potentially regulated by cortisol. To validate this approach further, we selected some of the candidate clones to measure mRNA expression levels in hippocampus using real-time PCR. We found that 66% of the sequences tested (10 of 15) were differentially represented between cortisol-treated and control animals. The complete set of clones was subjected to a bioinformatic analysis, which allowed classification of the ESTs into four different main categories: 1) known proteins or genes (approximately 28%), 2) ESTs previously published in the database (approximately 16%), 3) novel ESTs matching only the reference human or mouse genome (approximately 5%), and 4) sequences that do not match any public database (50%). Interestingly, the last category was the most abundant. Hybridization assays revealed that several of these clones are indeed expressed in hippocampal tissue from tree shrew, human, and/or rat. Therefore, we discovered an extensive inventory of new molecular targets in the hippocampus that serves as a reference for hippocampal transcriptional responses under various conditions. Finally, a detailed analysis of the genomic localization in human and mouse genomes revealed a survey of putative novel splicing variants for several genes of the nervous system.
The surface of Trypanosoma cruzi is covered by mucin-type glycoproteins involved in parasite protection, attachment and immunoevasion. The gene family coding for the mucins expressed by the parasite in the vertebrate host, named TcMUC, is composed of several hundred members and presents high variability. The genes encoding mucins expressed in the insect-dwelling parasite stages are part of a much more homogeneous family, named TcSMUG. Here, we addressed the organization and evolution of physically linked T. cruzi mucin genes by sequencing large chromosomal fragments containing these genes. Specific accumulation of mutations was restricted to particular domains of TcMUC genes, showing that these regions have, or have had, an accelerated evolution rate. Sequence analysis of several TcMUC genes allowed for the identification of members sharing features of TcMUC I and II, thus evidencing that one group of genes was generated from the other. The highly conserved intergenic regions of both TcMUC and TcSMUG families contained TG-rich microsatellites that were not present in unrelated genes in the cosmids, suggesting a role for homologous recombination in shuffling and/or amplification of T. cruzi mucin genes. The comparison of putative homologous TcMUC II genes from different strains of T. cruzi showed that their central variable domains are conserved. This conservation was always higher at the DNA level suggesting positive selection in these particular regions of TcMUC II genes.
We have generated 2771 expressed sequence tags (ESTs) from two cDNA libraries of Trypanosoma cruzi CL-Brener. The libraries were constructed from trypomastigote and amastigotes, using a spliced leader primer to synthesize the cDNA second strand, thus selecting for full-length cDNAs. Since the libraries were not normalized nor pre-screened, we compared the representation of transcripts between the two using a statistical test and identify a subset of transcripts that show apparent differential representation. A non-redundant set of 1619 reconstructed transcripts was generated by sequence clustering. This dataset was used to perform similarity searches against protein and nucleotide databases. Based on these searches, 339 sequences could be assigned a putative identity. One thousand one-hundred and sixteen sequences in the non-redundant clustered dataset (68.8%) are new expression tags, not represented in the T. cruzi epimastigote ESTs that are in the public databases. Additional information is provided online at http://genoma.unsam.edu.ar/projects/tram. To the best of our knowledge these are the first ESTs reported for the life cycle stages of T. cruzi that occur in the vertebrate host.
ABSTRACTgp63 is a highly abundant glycosylphosphatidylinositol (GPI)-anchored membrane protein expressed predominantly in the promastigote but also in the amastigote stage ofLeishmaniaspecies. InLeishmaniaspp., gp63 has been implicated in a number of steps in establishment of infection. Here we demonstrate thatTrypanosoma cruzi, the etiological agent of Chagas' disease, has a family of gp63 genes composed of multiple groups. Two of these groups,Tcgp63-Iand-II, are present as high-copy-number genes. The genomic organization and mRNA expression pattern were specific for each group.Tcgp63-Iwas widely expressed, while theTcgp63-IIgroup was scarcely detected in Northern blots, even though it is well represented in theT. cruzigenome. Western blots using sera directed against a synthetic peptide indicated that theTcgp63-Igroup produced proteins of ∼78 kDa, differentially expressed during the life cycle. Immunofluorescence staining and phosphatidylinositol-specific phospholipase C digestion confirmed that Tcgp63-I group members are surface proteins bound to the membrane by a GPI anchor. We also demonstrate the presence of metalloprotease activity which is attributable, at least in part, to Tcgp63-I group. Since antibodies against Tcgp63-I partially blocked infection of Vero cells by trypomastigotes, a possible role for this group in infection is suggested.
Chagas' disease is a chronic, debilitating, multisystemic disorder that affects millions of people in Latin America. The protozoan parasite Trypanosoma cruzi, the etiological agent of Chagas' disease, has a large number of O-glycosylated Thr/Ser/Pro-rich mucin molecules on its surface (TcMuc). These mucins are the main acceptors of sialic acid and have been suggested to play a role on various host-parasite interactions, such as adhesion to macrophages, protection from complement lysis, and immunomodulation of the immune response mounted by the host. To observe the immunologic effect obtained by the heterologous expression of a TcMuc gene in higher eukaryotic cells exposed to xenogeneic lymphocytes, we developed a strategy based on the transfection of a known T. cruzi mucin gene (TcMuc-e2) into Vero cells. In contrast to the brisk proliferation and activation of human lymphocytes observed at 3, 4, and 5 days induced by normal Vero cells, neither proliferation nor signicant activation of human lymphocytes was observed with TcMuc-e2-transfected Vero cells. This TcMuc-e2 mucin-induced suppression of T cell response can be reversed by the addition of exogenous IL-2. In addition it was demonstrated that the immunosuppressive reaction was not related to the induction of an important degree of apoptosis in human lymphocytes. Posttranslational modification are required for the inhibitory effect that TcMuc-e2 exerts when transfected to Vero cells. O-glycosylation and sialylation are required to obtain the immunomodulatory effect as assessed by O-sialoglycoprotease and neuraminidase treatments. These results are consistent with other studies showing that surface glycoconjugates from T. cruzi and mammalian cells can induce an inhibition of the immune response.
ABSTRACT A total of 1,921 expressed sequence tags (ESTs) were obtained from bloodstream trypomastigotes of Trypanosoma carassii , a parasite of economic importance due to its high prevalence in fish farms. Analysis of the data set allowed us to identify a trans -sialidase (TS)-like gene and three ESTs coding for putative mucin-like genes. TS activity was detected in cell extracts of bloodstream trypomastigotes. We have also used the sequence information obtained to identify genes that have not been previously described in trypanosomatids. (Additional information on these ESTs can be found at http://genoma.unsam.edu.ar/projects/tca .)
ABSTRACT Brucella abortus is the etiological agent of brucellosis, a disease that affects bovines and human. We generated DNA random sequences from the genome of B. abortus strain 2308 in order to characterize molecular targets that might be useful for developing immunological or chemotherapeutic strategies against this pathogen. The partial sequencing of 1,899 clones allowed the identification of 1,199 genomic sequence surveys (GSSs) with high homology (BLAST expect value < 10−5) to sequences deposited in the GenBank databases. Among them, 925 represent putative novel genes for the Brucella genus. Out of 925 nonredundant GSSs, 470 were classified in 15 categories based on cellular function. Seven hundred GSSs showed no significant database matches and remain available for further studies in order to identify their function. A high number of GSSs with homology toAgrobacterium tumefaciens and Rhizobium meliloti proteins were observed, thus confirming their close phylogenetic relationship. Among them, several GSSs showed high similarity with genes related to nodule nitrogen fixation, synthesis of nod factors, nodulation protein symbiotic plasmid, and nodule bacteroid differentiation. We have also identified severalB. abortus homologs of virulence and pathogenesis genes from other pathogens, including a homolog to both the Shda gene fromSalmonella enterica serovar Typhimurium and the AidA-1 gene from Escherichia coli. Other GSSs displayed significant homologies to genes encoding components of the type III and type IV secretion machineries, suggesting that Brucella might also have an active type III secretion machinery.
A random sequence survey of the genome of Trypanosoma cruzi, the agent of Chagas disease, was performed and 11,459 genomic sequences were obtained, resulting in approximately 4.3 Mb of readable sequences or approximately 10% of the parasite haploid genome. The estimated total GC content was 50.9%, with a high representation of A and T di- and trinucleotide repeats. Out of the estimated 5000 parasite genes, 947 putative new genes were identified. Another 1723 sequences corresponded to genes detected previously in T. cruzi through expression sequence tag analysis. 7735 sequences had no matches in the database, but the presence of open reading frames that passed Fickett's test suggests that some might contain coding DNA. The survey was highly redundant, with approximately 35% of the sequences included in a few large sequence families. Some of them code for protein families present in dozens of copies, including proteins essential for parasite survival and retrotransposons. Other sequence families include repetitive DNA present in thousands of copies per haploid genome. Some families in the latter group are new, parasite-specific, repetitive DNAs. These results suggest that T. cruzi could constitute an interesting model to analyze gene and genome evolution due to its plasticity in terms of sequence amplification and divergence. Additional information can be found at http://www.iib.unsam.edu.ar/tcruzi.gss. html.
Trypanosoma cruzi has a complex mucin gene family of 500 members with hypervariable regions expressed preferentially in vertebrate associated stages of the parasite. In this work, a novel mucin-type gene family is reported, composed of two groups of genes organized in independent tandems and having very short open reading frames. The structures of deduced proteins share the N and C termini but differ in central regions. One group has repeats with the consensus Lys-Asn-Thr7-Ser-Thr3-Ser(Ser/Lys)-Ala-Pro and the other a Thr-rich sequence of the type Asp-Gln-Thr17–20-Asn-Ala-Pro-Ala-Lys-Asp-Thr5–7-Asn-Ala-Pro-Ala-Lys. In both cases, expected mature core proteins are around 7 kDa. Both groups, named L and S, respectively, differ in the structure of genomic loci and mRNA, with differential blocks in the 3′-untranslated region. The highest mRNA level for S and L groups are in the epimastigote stage but they show distinct developmentally regulated patterns. Transcripts are short lived and their steady-state abundance is regulated post-transcriptionally with increased mRNA stability in insect stage epimastigote. AU-rich sequences, similar to ARE motives known to cause mRNA instability in higher eukaryotes, are present in the 3′-untranslated region of the transcripts. In transfection experiments this sequence is shown to be functional for the L group destabilizing its mRNA in a stage-specific manner. Furthermore, an effect of this AU-rich region on translation efficiency is shown. To our knowledge, this is the first time that a functional ARE sequence-dependent post-transcriptional regulation mechanism is reported in a lower eukaryote.
Five years ago the Special Programme for Research and Training in Tropical Diseases (TDR) from the World Health Organization (WHO) launched the Parasite Genome Project. The aims were to obtain information on genome organization and gene discovery in five parasites, namely, Schistosoma, Filaria, Leishmania and Trypanosomas brucei and cruzi, Organization of research networks for each parasite under study, promotion of international collaboration and training of researchers in developing countries, were also main objectives of the programme. After five years, a large amount of information has been obtained, which is now available to researchers in the field.
trans-sialidase is a unique sialidase in that, instead of hydrolizing sialic acid, it preferentially transfers the monosaccharide to a terminal beta-galactose in glycoproteins and glycolipids. This enzyme, originally identified in Trypanosoma cruzi, belongs to a large family of proteins. Some members of the family lack the enzymatic activity. No function has been yet assigned to them. In this work, the gene copy number and the possible function of inactive members of the trans -sialidase family was studied. It is shown that genes encoding inactive members are not a few, but rather, are present in the same copy number (60-80 per haploid genome) as those encoding active trans -sialidases. Recombinant inactive proteins were purified and assayed for sialic acid and galactose binding activity in agglutination tests. The enzymatically inactive trans -sialidases were found to agglutinate de-sialylated erythrocytes but not untreated red blood cells. Assays made with mouse and rabbit red blood cells suggest that inactive trans -sialidases bind to beta, rather than alpha, terminal galactoses, the same specificity required by active trans -sialidases. A recombinant molecule that was made enzymatically inactive through a mutation in a single amino acid also retained the galactose binding activity. The binding was competed by lactose and was dependent on conservation of the protein native conformation. Therefore, at least some molecules in the trans -sialidase family that have lost their enzymatic function still retain their Gal-binding properties and might have a function as lectins in the parasite-host interaction.
In previous works we have identified genes in the protozoan parasite Trypanosoma cruzi whose structure resemble those of mammalian mucin genes. Indirect evidence suggested that these genes might encode the core protein of parasite mucins, glycoproteins that were proposed to be involved in the interaction with, and invasion of, mammalian host cells. We now show that the mucin gene family from T. cruzi is much larger and diverse than expected. A minimal number of 484 mucin genes per haploid genome is calculated for a parasite clone. Most, if not all, genes are transcribed, as deduced from cDNA analysis. Comparison of the cDNA sequences showed evidences of a high mutation rate in localized regions of the genes. Sequence conservation among members of the family is much higher in the untranslated (UTR) regions than in the sequences encoding the mature mucin core protein. Transcription units can be classified into two main subfamilies according to the sequence homologies in the 5'-UTR, whereas the 3'-UTR is highly conserved in all clones analyzed. The common origin of members of this gene family as well as their relationships can be defined by sequence comparison of different domains in the transcription units. The regions encoding the N and C termini, supposed to correspond to the leader peptide and membrane-anchoring signal, respectively, (Di Noia, J. M., Sanchez, D. O., and Frasch, A. C. C. (1995) J. Biol. Chem. 270, 24146-24149) are highly conserved. Conversely, the central regions are highly variable. These regions encode the target sites for O-glycosylation and are made of a variable number of repetitive units rich in Thr and Pro residues or are nonrepetitive but still rich in Thr/Ser and Pro residues. The region putatively coding for the N-terminal domain of the mature core protein is hypervariable, being different in most of the transcripts sequenced. Nonrepetitive central domains are unique to each gene. Gene-specific probes show that the relative abundance of different mRNAs varies greatly within the same parasite clone.
Analysis of expressed sequence tags (ESTs) constitutes a useful approach for gene identification that, in the case of human pathogens, might result in the identification of new targets for chemotherapy and vaccine development. As part of the Trypanosoma cruzi genome project, we have partially sequenced the 5' ends of 1,949 clones to generate ESTs, The clones were randomly selected from a normalized CL Brener epimastigote cDNA library. A total of 14.6% of the clones were homologous to previously identified T. cruzi genes, while 18.4% had significant matches to genes from other organisms in the database. A total of 67% of the ESTs had no matches in the database, and thus, some of them might be T. cruzi-specific genes. Functional groups of those sequences with matches in the database were constructed according to their putative biological functions, The two largest categories were protein synthesis (23.3%) and cell surface molecules (10.8%). The information reported in this paper should be useful for researchers in the field to analyze genes and proteins of their own interest.
Mucins are highly O-glycosylated molecules which in mammalian cells accomplish essential functions, like cytoprotection and cell-cell interactions. In the protozoan parasite Trypanosoma cruzi, mucin-related glycoproteins have been shown to play a relevant role in the interaction with and invasion of host cells. We have previously reported a family of mucin-like genes in T. cruzi whose overall structure resembled that of mammalian mucin genes. We have now analyzed the relationship between these genes and mucin proteins. A monoclonal antibody specific for a mucin sugar epitope and a polyclonal serum directed to peptide epitopes in a MUC gene-encoded recombinant protein, detected identical bands in three out of seven strains of T. cruzi. Immunoprecipitation experiments confirmed these results. When expressed in eukaryotic cells, the MUC gene product is post-translationally modified, most likely, through extensive O-glycosylation. Gene sequencing showed that the central domains encoding the repeated sequences with the consensus T8KP2, varies in number from 1 to 10, and the number of Thr residues in each repeat could be 7, 8, or 10. A run of 16 to 18 Thr residues was present in some, but not all, MUC gene-derived sequences. Direct compositional analysis of mucin core proteins showed that Thr residues are much more frequent than Ser residues. The same fact occurs in MUC gene-derived protein sequences. Molecular mass determinations of the 35-kDa glycoproteins further extend the heterogeneity of the family to the natural mucin molecules. Difficulties in assigning each of the several MUC genes identified to a mucin product arise from the high diversity and partial sequence conservation of the members of this family.
In order to generate contiguous cosmid coverage of the genome of the protozoan parasite Trypanosoma cruzi for large-scale sequence analysis, a cosmid library of 36864 individual, primary clones was generated. Total genomic DNA of the reference strain CL Brener was fragmented both by partial digestion with MboI and by physical shearing. For cloning, a modified cosmid vector was used that simplifies analyses such as restriction mapping. The library's representation is about 25 genome equivalents, assuming a size of 55 Mb per haploid genome. No chimerism of inserts in the clones could be detected. The colinearity between cosmid inserts and genomic DNA was verified. Also, hybridizations to the gel-separated karyotype of the organism were carried out as a quality check. Gridded onto two nylon filters, the library was analyzed with a variety of probes. Apart from being used for combined physical and transcriptional mapping of the genome, library filters and clones are also available to interested parties.
A full-length DNA clone encoding a putative pyruvate dehydrogenase alpha subunit (E1 alpha) gene was isolated from a Trypanosoma cruzi (RA strain) DNA library. Sequencing of this clone revealed it to encode a 378 amino acid protein (M(r) 42774) with high sequence similarity to E1 alpha obtained from different sources. The highest score is obtained with human E1 alpha: 43,3% similarity. Southern blot analysis is consistent with the existence of a single copy of this putative T. cruzi E1 alpha gene per haploid genome in different parasite strains. Expression of this gene was demonstrated by Northern blot analysis and its trans-splicing acceptor site was identified by Polymerase Chain Reaction-mediated amplification of its cDNA.
Several genes encode members of the Trypanosoma cruzi (Tc) trans-sialidase (TS) family. These proteins contain an enzymatic domain on the N terminus, the only one required for TS activity, and an antigenic domain (SAPA (shed acute phase antigen) amino acid (aa) repeats) on the C terminus. Only some members of this glycoprotein family are enzymatically active. The complete sequence of two clones encoding the enzymatic domain of active and inactive protein from each of two Tc strains has now been obtained. Comparison of these sequences showed a limited divergence among them: 20 out of the 642 deduced aa in the enzymatic domain were found to differ. From these 20 aa, only one was found to be essential for enzymatic activity. A Tyr342 residue is deduced in both active proteins while a His342 is present in both inactive ones. This naturally occurring Tyr342→ His substitution completely abolished the TS activity. In addition to Tyr342, a second deduced aa, Pro231, was found to be necessary for full enzymatic TS activity; a Pro231 → Ala change rendered the TS protein partially active. Fourteen aa residues, including Tyr342, out of the 16 aa in the active site of a sialidase from Salmonella typhimurium are present at the same or very similar positions in the Tc TS.