The population structure of Toxoplasma gondii includes three highly prevalent clonal lineages referred to as types I, II, and III, which differ greatly in virulence in the mouse model. Previous studies have implicated a family of serine/threonine protein kinases found in rhoptries (ROPs) as important in mediating virulence differences between strain types. Here, we explored the genetic basis of differences in virulence between the highly virulent type I lineage and moderately virulent type II based on successful genetic cross between these lineages. Genome-wide association revealed that a single quantitative trait locus controls the dramatic difference in lethality between these strain types. Neither ROP16 nor ROP18, previously implicated in virulence of T. gondii , was found to contribute to differences between types I and II. Instead, the major virulence locus contained a tandem cluster of polymorphic alleles of ROP5, which showed similar protein expression between strains. ROP5 contains a conserved serine/threonine protein kinase domain that includes only part of the catalytic triad, and hence, all members are considered to be pseudokinases. Genetic disruption of the entire ROP5 locus in the type I lineage led to complete attenuation of acute virulence, and complementation with ROP5 restored lethality to WT levels. These findings reveal that a locus of polymorphic pseudokinases plays an important role in pathogenesis of toxoplasmosis in the mouse model.
Plasmodium yoelii is an excellent model for studying malaria pathogenesis that is often intractable to investigate using human parasites; however, genetic studies of the parasite have been hindered by lack of genome-wide linkage resources. Here, we performed 14 genetic crosses between three pairs of P. yoelii clones/subspecies, isolated 75 independent recombinant progeny from the crosses, and constructed a high-resolution linkage map for this parasite. Microsatellite genotypes from the progeny formed 14 linkage groups belonging to the 14 parasite chromosomes, allowing assignment of sequence contigs to chromosomes. Growth-related virulent phenotypes from 25 progeny of one of the crosses were significantly associated with a major locus on chromosome 13 and with two secondary loci on chromosomes 7 and 10. The chromosome 10 and 13 loci are both linked to day 5 parasitemia, and their effects on parasite growth rate are independent but additive. The locus on chromosome 7 is associated with day 10 parasitemia. The chromosome 13 locus spans ∼220 kb of DNA containing 51 predicted genes, including the P. yoelii erythrocyte binding ligand, in which a C741Y substitution in the R6 domain is implicated in the change of growth rate. Similarly, the chromosome 10 locus spans ∼234 kb with 71 candidate genes, containing a member of the 235-kDa rhoptry proteins (Py235) that can bind to the erythrocyte surface membrane. Atypical virulent phenotypes among the progeny were also observed. This study provides critical tools and information for genetic investigations of virulence and biology of P. yoelii.
Background Apicomplexan parasites replicate by varied and unusual processes where the typically eukaryotic expansion of cellular components and chromosome cycle are coordinated with the biosynthesis of parasite-specific structures essential for transmission. Methodology/Principal Findings Here we describe the global cell cycle transcriptome of the tachyzoite stage of Toxoplasma gondii. In dividing tachyzoites, more than a third of the mRNAs exhibit significant cyclical profiles whose timing correlates with biosynthetic events that unfold during daughter parasite formation. These 2,833 mRNAs have a bimodal organization with peak expression occurring in one of two transcriptional waves that are bounded by the transition into S phase and cell cycle exit following cytokinesis. The G1-subtranscriptome is enriched for genes required for basal biosynthetic and metabolic functions, similar to most eukaryotes, while the S/M-subtranscriptome is characterized by the uniquely apicomplexan requirements of parasite maturation, development of specialized organelles, and egress of infectious daughter cells. Two dozen AP2 transcription factors form a series through the tachyzoite cycle with successive sharp peaks of protein expression in the same timeframes as their mRNA patterns, indicating that the mechanisms responsible for the timing of protein delivery might be mediated by AP2 domains with different promoter recognition specificities. Conclusion/Significance Underlying each of the major events in apicomplexan cell cycles, and many more subordinate actions, are dynamic changes in parasite gene expression. The mechanisms responsible for cyclical gene expression timing are likely crucial to the efficiency of parasite replication and may provide new avenues for interfering with parasite growth.
Most pairwise and multiple sequence alignment programs seek alignments with optimal scores. Central to defining such scores is selecting a set of substitution scores for aligned amino acids or nucleotides. For local pairwise alignment, substitution scores are implicitly of log-odds form. We now extend the log-odds formalism to multiple alignments, using Bayesian methods to construct “BILD” (“Bayesian Integral Log-odds”) substitution scores from prior distributions describing columns of related letters. This approach has been used previously only to define scores for aligning individual sequences to sequence profiles, but it has much broader applicability. We describe how to calculate BILD scores efficiently, and illustrate their uses in Gibbs sampling optimization procedures, gapped alignment, and the construction of hidden Markov model profiles. BILD scores enable automated selection of optimal motif and domain model widths, and can inform the decision of whether to include a sequence in a multiple alignment, and the selection of insertion and deletion locations. Other applications include the classification of related sequences into subfamilies, and the definition of profile-profile alignment scores. Although a fully realized multiple alignment program must rely upon more than substitution scores, many existing multiple alignment programs can be modified to employ BILD scores. We illustrate how simple BILD score based strategies can enhance the recognition of DNA binding domains, including the Api-AP2 domain in Toxoplasma gondii and Plasmodium falciparum.
Toxoplasma gondii strains differ dramatically in virulence despite being genetically very similar. Genetic mapping revealed two closely adjacent quantitative trait loci on parasite chromosome VIIa that control the extreme virulence of the type I lineage. Positional cloning identified the candidate virulence gene ROP18, a highly polymorphic serine-threonine kinase that was secreted into the host cell during parasite invasion. Transfection of the virulent ROP18 allele into a nonpathogenic type III strain increased growth and enhanced mortality by 4 to 5 logs. These attributes of ROP18 required kinase activity, which revealed that secretion of effectors is a major component of parasite virulence.
Toxoplasma gondii is a highly successful protozoan parasite in the phylum Apicomplexa, which contains numerous animal and human pathogens. T.gondii is amenable to cellular, biochemical, molecular and genetic studies, making it a model for the biology of this important group of parasites. To facilitate forward genetic analysis, we have developed a high-resolution genetic linkage map for T.gondii. The genetic map was used to assemble the scaffolds from a 10X shotgun whole genome sequence, thus defining 14 chromosomes with markers spaced at approximately 300 kb intervals across the genome. Fourteen chromosomes were identified comprising a total genetic size of approximately 592 cM and an average map unit of approximately 104 kb/cM. Analysis of the genetic parameters in T.gondii revealed a high frequency of closely adjacent, apparent double crossover events that may represent gene conversions. In addition, we detected large regions of genetic homogeneity among the archetypal clonal lineages, reflecting the relatively few genetic outbreeding events that have occurred since their recent origin. Despite these unusual features, linkage analysis proved to be effective in mapping the loci determining several drug resistances. The resulting genome map provides a framework for analysis of complex traits such as virulence and transmission, and for comparative population genetic studies.
1. 0. Diels and K. Alder, Justus Liebigs Ann. Chem. 460, 98 (1928); 0. Diels, J. H. Blom, W. Koll, ibid. 443, 242 (1925); B. M. Trost, I. Fleming, L. A. Paquette, Comprehensive Organic Synthesis (Pergamon, Oxford, 1991), vol. 5, pp. 316; F. Friguelli and A. Tatichi, Dienes in the Diels-Alder Reaction (Wiley, New York, 1990), and references therein; W. Carruthers, Cycloaddition Reactions in Organic Synthesis (Pergamon, Oxford, 1990), and references therein; G. Desimoni, G. Tacconi, A. Barco, G. P. Pollini, ACS Monograph No. 180 (American Chemical Society, Washington, DC, 1983), and references cited therein. 2. For recent reviews, see U. Pindur, G. Lutz, C. Otto, Chem. Rev. 93, 741 (1993); H. B. Kagan and 0. Riant, ibid. 92, 1007 (1992), and references therein. 3. D. Hilvert, K. W. Hill, K. D. Nared, M-T. M. Auditor, J. Am Chem. Soc. 111, 9261 (1989). 4. A. C. Braisted and P. G. Schultz, ibid. 112, 7430 (1 990). 5. C. J. Suckling, M. C. Tedford, L. M. Bence, J. I. Irvine, W. H. Stimson, Biorg. Med. Chem. Lett. 2, 49 (1992). 6. K. D. Janda, C. G. Sheviin, R. A. Lerner, Science 259, 490 (1993). 7. L. E. Overman, G. F. Taylor, C. B. Petty, P. J. Jessup, J. Org. Chem. 43, 2164 (1978). 8. For example: d-pumiliotoxin, C. L. E. Overman and P. J. Jessup, Tetrahedron Lett. 1253 (1977); J. Am. Chem. Soc. 100, 5179 (1978); d-perhydrogephyrotoxin, L. E. Overman and C. Fukaya, ibid. 102, 1454 (1980); dIisogabaculine, S. Danishefsky and F. M. Hershenson, J. Org. Chem. 44, 1180 (1979); tylidine, L. E. Overman, C. B. Petty, R. J. Doedens, ibid., p. 4183; amaryllidaceae alkaloids, L. E. Overman et al., J. Am. Chem. Soc. 103, 2816 (1981); gephirotoxine, L. E. Overman, D. Lesuisse, M. Hashimoto, ibid. 105, 16, 5373 (1983). 9. Diene 1b was prepared by neutralizing the corresponding carboxylic acid suspended in PBS (sodium phosphate buffer) with one equivalent of NaOH. 10. All new adducts were characterized by NMR and mass spectrometry. 11. In the crude mixture of the cycloaddition, we could not detect any trace of the meta regioisomer by 1H NMR (>98% ortho regioselective). 12. M. J. Frisch et al., Gaussian 92 (Gaussian, Pittsburgh, PA, 1992). 13. R. J. Loncharich, F. K. Brown, K. N. Houk, J. Org. Chem. 54, 1129 (1989). All structures were fully optimized with the 3-21G basis set. Reactants were also fully optimized with the 6-31G* basis set. Vibrational frequencies for all transition structures were calculated at the RHF/3-21 G level and shown to have one imaginary frequency. In addition, the energies were evaluated by single-point calculations with the 6-31G* basis set on 3-21G geometries. 14. K. N. Houk, R. J. Loncharich, J. Blake, W. Jorgensen, J. Am. Chem. Soc. 111, 9172 (1989). 15. L. E. Overman, G. F. Taylor, K. N. Houk, I. N. Domelsmith, ibid. 100, 3182 (1978). 16. M. I. Page and W. P. Jencks, Proc. Nati. Acad. Sci. U.S.A. 68, 1678 (1971). 17. F. Bernardi et al., J. Am. Chem. Soc. 110, 3050 (1989); F. K. Brown and K. N. Houk, Tetrahedron Lett. 25, 4609 (1984). 18. This strategy to avoid product inhibition was first described by Braisted and Schultz in (4). 19. The stereochemistry was established on the basis of 'H NMR analysis. For example: 5-exo, BH6 = 2.65 ppm and 4J6 7a= 3.5 Hz; 5-endo, &H6 = 3.12 ppm, and 6 7=a 0 Hz. 20. G. Kohler and C. Milstein, Nature 256, 495 (1975); the KLH conjugate was prepared by slowly adding 2 mg of the hapten dissolved in 100 &l of dimethylformamide to 2 mg of KLH in 900 gl of 0.01 M sodium phosphate buffer, pH 7.4. Four 8-week-old 129GIX+ mice each received 2 intraperitoneal (IP) injections of 100 gg of the hapten conjugated to KLH and RIBI adjuvant (MPL and TDM emulsion) 2 weeks apart. A 50 >g IP
Summary The Plasmodium falciparum genome sequence has boosted hopes for a new era of malaria research and for the application of comprehensive molecular knowledge to disease control, but formidable obstacles remain: ≈ 60% of the predicted P. falciparum proteins have no known functions or homologues, and most life cycle stages of this haploid eukaryotic parasite are relatively intractable to cultivation and biochemical manipulation. Genetic mapping based on high‐resolution maps saturated with single‐nucleotide polymorphisms or microsatellites is now providing effective strategies for discovering candidate genes determining important parasite phenotypes. Here we review classical linkage studies using laboratory crosses and population associations that are now amenable to genome‐wide approaches and are revealing multiple candidate genes involved in complex drug responses. Moreover, mapping by linkage disequilibrium is practicable in cases where chromosomal segments flanking drug‐selected genes have been preserved in populations during relatively recent P. falciparum evolution. We discuss the advantages and limitations of these various genetic mapping strategies, results from which offer complementary insights to those emerging from gene knockout experiments and/or high‐throughput genomic technologies.
Mutations and/or overexpression of various transporters are known to confer drug resistance in a variety of organisms. In the malaria parasite Plasmodium falciparum, a homologue of P-glycoprotein, PfMDR1, has been implicated in responses to chloroquine (CQ), quinine (QN) and other drugs, and a putative transporter, PfCRT, was recently demonstrated to be the key molecule in CQ resistance. However, other unknown molecules are probably involved, as different parasite clones carrying the same pfcrt and pfmdr1 alleles show a wide range of quantitative responses to CQ and QN. Such molecules may contribute to increasing incidences of QN treatment failure, the molecular basis of which is not understood. To identify additional genes involved in parasite CQ and QN responses, we assayed the in vitro susceptibilities of 97 culture-adapted cloned isolates to CQ and QN and searched for single nucleotide polymorphisms (SNPs) in DNA encoding 49 putative transporters (total 113 kb) and in 39 housekeeping genes that acted as negative controls. SNPs in 11 of the putative transporter genes, including pfcrt and pfmdr1, showed significant associations with decreased sensitivity to CQ and/or QN in P. falciparum. Significant linkage disequilibria within and between these genes were also detected, suggesting interactions among the transporter genes. This study provides specific leads for better understanding of complex drug resistances in malaria parasites.
Let A denote an alphabet consisting of n types of letters. Given a sequence S of length L with v(i) letters of type i on A, to describe the compositional properties and combinatorial structure of S, we propose a new complexity function of S, called the reciprocal complexity of S, as C(S) = (i=1) product operator (n) (L/nv(i))(vi) Based on this complexity measure, an efficient algorithm is developed for classifying and analyzing simple segments of protein and nucleotide sequence databases associated with scoring schemes. The running time of the algorithm is nearly proportional to the sequence length. The program DSR corresponding to the algorithm was written in C++, associated with two parameters (window length and cutoff value) and a scoring matrix. Some examples regarding protein sequences illustrate how the method can be used to find regions. The first application of DSR is the masking of simple sequences for searching databases. Queries masked by DSR returned a manageable set of hits below the E-value cutoff score, which contained all true positive homologues. The second application is to study simple regions detected by the DSR program corresponding to known structural features of proteins. An extensive computational analysis has been made of protein sequences with known, physicochemically defined nonglobular segments. For the SWISS-PROT amino acid sequence database (Release 40.2 of 02-Nov-2001), we determine that the best parameters and the best BLOSUM matrix are, respectively, for automatic segmentation of amino acid sequences into nonglobular and globular regions by the DSR program: Window length k = 35, cutoff value b = 0.46, and the BLOSUM 62.5 matrix. The average "agreement accuracy (sensitivity)" of DSR segmentation for the SWISS-PROT database is 97.3%.
Let Pl,n denote the partition lattice of l with n parts, ordered by Hardy–Littlewood–Polya majorization. For any two comparable elements x and y of Pl,n, we denote by M(x,y), m(x,y), f(x,y), and F(x,y), respectively, the sizes of four typical chains between x and y: the longest chain, the shortest chain, the lexicographic chain, and the counter-lexicographic chain. The covers u=(u1,…,un)≻v=(v1,…,vn) in Pl,n are of two types: N-shift (nearby shift) where vi=ui−1, vi+1=ui+1+1 for some i; and D-shift (distant shift) where ui−1=vi=vi+1=⋯=vj=uj+1 for some i and j. An N-shift (a D-shift) is pure if it is not a D-shift (an N-shift). We develop linear algorithms for calculating M(x,y), m(x,y), f(x,y), and F(x,y), using the leftmost pure N-shift first search, the rightmost pure D-shift first search, the leftmost N-shift first search, and the rightmost D-shift first search, respectively. Those algorithms have significant applications in complexity analysis of biological sequences.
Widespread use of antimalarial agents can profoundly influence the evolution of the human malaria parasite Plasmodium falciparum. Recent selective sweeps for drug-resistant genotypes may have restricted the genetic diversity of this parasite, resembling effects attributed in current debates1,2,3,4 to a historic population bottleneck. Chloroquine-resistant (CQR) parasites were initially reported about 45 years ago from two foci in southeast Asia and South America5, but the number of CQR founder mutations and the impact of chlorquine on parasite genomes worldwide have been difficult to evaluate. Using 342 highly polymorphic microsatellite markers from a genetic map6, here we show that the level of genetic diversity varies substantially among different regions of the parasite genome, revealing extensive linkage disequilibrium surrounding the key CQR gene pfcrt7 and at least four CQR founder events. This disequilibrium and its decay rate in the pfcrt-flanking region are consistent with strong directional selective sweeps occurring over only ∼20–80 sexual generations, especially a single resistant pfcrt haplotype spreading to very high frequencies throughout most of Asia and Africa. The presence of linkage disequilibrium provides a basis for mapping genes under drug selection in P. falciparum.
We have investigated intragenic recombination in Block 2 of the merozoite surface protein-1 (MSP-1), where three allele-specific families: K1, Mad20, and RO33 were previously known. Using parasites from western Kenya, we have found a fourth Block 2 allele type, which is a recombinant between Mad20 and RO33 alleles. These recombinant alleles, which we have termed MR, contain sequence from the 5′ region of Mad20 and the 3′ region of RO33. The results of this study provide new data on the complexity of the MSP-1 antigen gene, which is a candidate vaccine antigen, and further support the importance of intragenic recombination in generating genetic variability in Plasmodium falciparum parasites in nature.
Several protein sequence analysis algorithms are based on properties of amino acid composition and repetitiveness. These include methods for prediction of secondary structure elements, coiled-coils, transmembrane segments or signal peptides, and for assignment of low-complexity, nonglobular, or intrinsically unstructured regions. The quality of such analyses can be greatly enhanced by graphical software tools that present predicted sequence features together in context and allow judgment to be focused simultaneously on several different types of supporting information. For these purposes, we describe the SFINX package, which allows many different sets of segmental or continuous-curve sequence feature data, generated by individual external programs, to be viewed in combination alongside a sequence dot-plot or a multiple alignment of database matches. The implementation is currently based on extensions to the graphical viewers Dotter and Blixem and scripts that convert data from external programs to a simple generic data definition format called SFS. We describe applications in which dot-plots and flanking database matches provide valuable contextual information for analyses based on compositional and repetitive sequence features. The system is also useful for comparing results from algorithms run with a range of parameters to determine appropriate values for defaults or cutoffs for large-scale genomic analyses.
Chloroquine (CQ)-resistant Plasmodium vivax malaria was first reported 12 years ago, nearly 30 years after the recognition of CQ-resistant P. falciparum. Loss of CQ efficacy now poses a severe problem for the prevention and treatment of both diseases. Mutations in a digestive vacuole protein encoded by a 13-exon gene, pfcrt, were shown recently to have a central role in the CQ resistance (CQR) of P. falciparum. Whether mutations in pfcrt orthologues of other Plasmodium species are involved in CQR remains an open question. This report describes pfcrt homologues from P. vivax, P. knowlesi, P. berghei, and Dictyostelium discoideum. Synteny between the P. falciparum and P. vivax genes is demonstrated. However, a survey of patient isolates and monkey-adapted lines has shown no association between in vivo CQR and codon mutations in the P. vivax gene. This is evidence that the molecular events underlying P. vivax CQR differ from those in P. falciparum.