Most cancer genomes are characterized by the gain or loss of copies of some sequences through deletion, amplification or unbalanced translocations. Delineating and quantifying these changes is important in understanding the initiation and progression of cancer, in identifying novel therapeutic targets, and in the diagnosis and prognosis of individual patients. Conventional methods for measuring copy‐number are limited in their ability to analyse large numbers of loci, in their dynamic range and accuracy, or in their ability to analyse small or degraded samples. This latter limitation makes it difficult to access the wealth of fixed, archived material present in clinical collections, and also impairs our ability to analyse small numbers of selected cells from biopsies. Molecular copy‐number counting (MCC), a digital PCR technique, has been used to delineate a non‐reciprocal translocation using good quality DNA from a renal carcinoma cell line. We now demonstrate µMCC, an adaptation of MCC which allows the precise assessment of copy number variation over a significant dynamic range, in template DNA extracted from formalin‐fixed paraffin‐embedded clinical biopsies. Further, µMCC can accurately measure copy number variation at multiple loci, even when applied to picogram quantities of grossly degraded DNA extracted after laser capture microdissection of fixed specimens. Finally, we demonstrate the power of µMCC to precisely interrogate cancer genomes, in a way not currently feasible with other methodologies, by defining the position of a junction between an amplified and non‐amplified genomic segment in a bronchial carcinoma. This has tremendous potential for the exploitation of archived resources for high‐resolution targeted cancer genomics and in the future for interrogating multiple loci in cancer diagnostics or prognostics. Copyright © 2008 Pathological Society of Great Britain and Ireland. Published by John Wiley & Sons, Ltd.
The social amoebae are exceptional in their ability to alternate between unicellular and multicellular forms. Here we describe the genome of the best-studied member of this group, Dictyostelium discoideum. The gene-dense chromosomes of this organism encode approximately 12,500 predicted proteins, a high proportion of which have long, repetitive amino acid tracts. There are many genes for polyketide synthases and ABC transporters, suggesting an extensive secondary metabolism for producing and exporting small molecules. The genome is rich in complex repeats, one class of which is clustered and may serve as centromeres. Partial copies of the extrachromosomal ribosomal DNA (rDNA) element are found at the ends of each chromosome, suggesting a novel telomere structure and the use of a common mechanism to maintain both the rDNA and chromosomal termini. A proteome-based phylogeny shows that the amoebozoa diverged from the animal-fungal lineage after the plant-animal split, but Dictyostelium seems to have retained more of the diversity of the ancestral genome than have plants, animals or fungi.
The apicomplexan Cryptosporidium parvum is an intestinal parasite that affects healthy humans and animals, and causes an unrelenting infection in immunocompromised individuals such as AIDS patients. We report the complete genome sequence of C. parvum , type II isolate. Genome analysis identifies extremely streamlined metabolic pathways and a reliance on the host for nutrients. In contrast to Plasmodium and Toxoplasma , the parasite lacks an apicoplast and its genome, and possesses a degenerate mitochondrion that has lost its genome. Several novel classes of cell-surface and secreted proteins with a potential role in host interactions and pathogenesis were also detected. Elucidation of the core metabolism, including enzymes with high similarities to bacterial and plant counterparts, opens new avenues for drug development.
The apicomplexan Cryptosporidium parvum is one of the most prevalent protozoan parasites of humans. We report the physical mapping of the genome of the Iowa isolate, sequencing and analysis of chromosome 6, and approximately 0.9 Mbp of sequence sampled from the remainder of the genome. To construct a robust physical map, we devised a novel and general strategy, enabling accurate placement of clones regardless of clone artefacts. Analysis reveals a compact genome, unusually rich in membrane proteins. As in Plasmodium falciparum, the mean size of the predicted proteins is larger than that in other sequenced eukaryotes. We find several predicted proteins of interest as potential therapeutic targets, including one exhibiting similarity to the chloroquine resistance protein of Plasmodium. Coding sequence analysis argues against the conventional phylogenetic position of Cryptosporidium and supports an earlier suggestion that this genus arose from an early branching within the Apicomplexa. In agreement with this, we find no significant synteny and surprisingly little protein similarity with Plasmodium. Finally, we find two unusual and abundant repeats throughout the genome. Among sequenced genomes, one motif is abundant only in C. parvum, whereas the other is shared with (but has previously gone unnoticed in) all known genomes of the Coccidia and Haemosporida. These motifs appear to be unique in their structure, distribution and sequences.
When the Human Genome Mapping Project began in earnest, the allocation of large funding and the opening of dedicated research facilities promised much for the rapid improvement of DNA sequencing methods and, as a consequence, sequencing rates and throughput. Today, several years later, some novel approaches to the problem are being developed, such as sequencing by hybridization, and sequence detection using mass spectrometry and biochips, but these are far from in general use.
We have made a high-resolution HAPPY map of chromosome 6 of Dictyostelium discoideum consisting of 300 sequence-tagged sites with an average spacing of 14 kb along the approximately 4-Mb chromosome. The majority of the marker sequences were derived from randomly chosen clones from four different chromosome 6-enriched plasmid libraries or from subclones of YACs previously mapped to chromosome 6. The map appears to span the entire chromosome, although marker density is greater in some regions than in others and is lowest within the telomeric region. Our map largely supports previous gene-based maps of this chromosome but reveals a number of errors in the physical map. In addition, we find that a high proportion of the plasmid sequences derived from gel-enriched chromosome 6 (and that form the basis of a chromosome-specific sequencing project) originates from other chromosomes.
We have mapped 1001 novel sequence-tagged sites on human chromosome 14. The mean spacing between markers is approximately 90 kb, most markers are mapped with a resolution of better than 100 kb, and physical distances are determined. The map was produced using HAPPY mapping, a simple and widely applicable in vitro approach that is analogous to linkage or to radiation hybrid mapping, but that circumvents many of the difficulties and potential artifacts associated with these methods. We show also that the map serves as a robust scaffold for building physical maps using large-insert clones.
Cryptosporidium parvum proteases have been associated with release of infective sporozoites from oocysts, and their specific inhibition blocks parasite excystation in vitro. Additionally, proteases have been implicated in the processing of parasite adhesion molecules found on the surface of sporozoites and merozoites. In this study, we cloned and expressed the C. parvum aminopeptidase N gene by screening a large insert, P1 artificial chromosome library with a probe identified from a Cryptosporidium genome survey-sequencing project. Analysis of the predicted protein encoded by the 2.3 kb gene demonstrated a high degree of homology with prokaryotic and eukaryotic aminopeptidases. The 783 amino acid sequence predicted a Mr of ∼89,000. The active site sequence was found to be highly conserved when compared with other Apicomplexan aminopeptidases. Motifs commonly found in aminopeptidases of this class and a unique single Arg-Gly-Asp (RGD) tripeptide motif predictive of cell adhesion were identified. The aminopeptidase N mRNA was expressed in infective sporozoites and during the infection of human HCT-8 enterocytes as revealed by reverse transcription PCR.
We have localized the gene encoding human RNase k6 to within approximately 120 kb on the long (q) arm of chromosome 14 by HAPPY mapping. With this information, the relative positions of the six human RNase A ribonucleases that have been mapped to this locus can be inferred. To further our understanding of the individual lineages comprising the RNase A superfamily, we have isolated and characterized 10 novel genes orthologous to that encoding human RNase k6 from Great Ape, Old World, and New World monkey genomes. Each gene encodes a complete ORF with no less than 86% amino acid sequence identity to human RNase k6 with the eight cysteines and catalytic histidines (H15 and H123) and lysine (K38) typically observed among members of the RNase A superfamily. Interesting trends include an unusually low number of synonymous substitutions (Ks) observed among the New World monkey RNase k6 genes. When considering nonsilent mutations, RNase k6 is a relatively stable lineage, with a nonsynonymous substitution rate of 0.40 x 10(-9) nonsynonymous substitutions/nonsynonymous site/year (ns/ns/yr). These results stand in contrast to those determined for the primate orthologs of the two closely related ribonucleases, the eosinophil-derived neurotoxin (EDN) and eosinophil cationic protein (ECP), which have incorporated nonsilent mutations at very rapid rates (1.9 x 10(-9) and 2.0 x 10(-9) ns/ns/yr, respectively). The uneventful trends observed for RNase k6 serve to spotlight the unique nature of EDN and ECP and the unusual evolutionary constraints to which these two ribonuclease genes must be responding. [The sequence data described in this paper have been submitted to the GenBank data library under accession nos. AF037081-AF037090.]
We have constructed a HAPPY map of the apicomplexan parasite Cryptosporidium parvum. We have placed 204 markers on the 10.4-Mb genome, giving an average marker spacing of approximately 50 kb, with an effective resolution of approximately 40 kb. HAPPY mapping (an in vitro linkage technique based on screening approximately haploid amounts of DNA by the polymerase chain reaction) is fast and accurate and is not subject to the distortions inherent in cloning, meiotic recombination, or hybrid cell formation. In addition, little genomic DNA is needed as a substrate, and the AT content of the genome is largely immaterial, making it an ideal method for mapping otherwise intractable parasite genomes. The map, covering all eight chromosomes, consists of 10 linkage groups, each of which has been chromosomally assigned. We have verified the accuracy of the map by several methods, including the construction of a >140-kb PAC contig on chromosome VI. Less than 1% of our markers detect non-rDNA duplicated sequences.
This directory was made possible by a unique international collaboration between the 633 scientists whose names appear below. It represents both the first published description of the complete sequence of most chromsomes from Saccharomyces cerevisiae , and the first published overview of the entire sequence. As such, the authors would like future papers referring to the entire sequence and/or its contents to cite this directory; future papers referring to the sequence of individual chromosomes should refer to the papers listed at the head of page 9. The authors’ affiliations appear in the papers describing the individual chromosomes.
The growth and purification of M13 DNA from small volume (1.5mL) cultures is a rapid and easily performed procedure (). The samples can be processed in microcentrifuges in disposable polypropylene tubes and yield sufficient, pure single-stranded DNA (4 μM) for five or more sequencing experiments. Even when several cultures are to be grown and purified simultaneously, up to 100 can be processed to completion in a day. The handling of this number of samples is tedious, however, and much time is spent opening and closing tubes and transferring tubes in and out of microcentrifuges. One hundred small volume phenol extractions and ethanol precipitations tasks even the more dedicated sequencer.
The underlying principle of DNA sequencing by either the Sanger () or Maxam and Gilbert method (), is the ability to fractionate and resolve long, single-stranded DNA molecules that differ in length by only one nucleotide. Denaturing polyacrylamide gels have been reported to give interpretable separation of molecules up to 0.6 kb in length, on 1-m long gels ().
This chapter describes the methods for the preparation and fluorescent sequencing of M13 clones. The shotgun sequencing strategy—in which randomly generated DNA subfragments are cloned into bacteriophage M13 vectors and sequenced, using dideoxynucleotide chain terminators—is a proven and reliable scheme. A major rate-limiting step is the generation of single-stranded M13 DNA templates. The M13 library of the DNA to be sequenced is generated by sonication. Sonication conditions are adjusted to give different size fractions as required; for example, libraries used for fluorescent sequencing contain inserts in the 500- to 2500-bp size range. Sonication time course are carried out to determine the minimum conditions that produce DNA fragments in the desired size range. The chapter also describes the self-ligation of DNA fragments prior to sonication, fragment end repair, size fractionation, ligation in M13, and transformation procedures in detail. The preparation of M13 DNA templates in 96-well microtiter plates can be carried out using a standard multichannel pipette, or it can be semiautomated by the use of a robot pipetting device.