The region Xq21.3/Yp11.1 represents the largest segment of homology between the sex chromosomes in humans, though no recombination occurs in male meiosis. It presumably arose as a transposition from the X to the Y chromosome; the present-day organization in the latter chromosome indicates a paracentric inversion that disrupted its continuity. Moreover, an X-specific block (defined by the marker DXS214) is embedded in the region. Previously, no hypotheses about the length, origin, or evolution of this X-specific segment have been proposed. Here we report on the refinement of the size and the sequence of the distal boundary of the X-specific block. Furthermore, we have tracked by FISH experiments the evolution of this region in primates. This further clarifies the multistep mechanism of origin for the XY homology region, by demonstrating that the X-specific block was deleted from the Y chromosome after the initial transfer from the X chromosome.
GPC3,the gene modified in the Simpson–Golabi–Behmel gigantism/overgrowth syndrome (SGBS), is shown to span more than 500 kb of genomic sequence, with the transcript beginning 197 bp 5′ of the translational start site. The Xq26.1 region containingGPC3as the only known gene has been extended to >900 kb by sequence analysis of flanking BAC clones. Two GC isochores (40.6 and 42.6% GC) are observed at the 5′ and 3′ ends of the locus, with a large repertoire of repetitive sequences that includes an unusual cluster of four L1 elements >92% identical over 2.8 kb. Eight exons, accounting for the full 2.4-kbGPC3cDNA, have been sequenced along with neighboring intronic regions. PCR assays have been developed to amplify each exon and exon/intron junction sequence, to help discriminate instances of SGBS among individuals with overgrowth syndromes and to facilitate mutational analysis of lesions in the gene.
So far we have used OSS strategy to sequence over 2 megabases DNA in large-insert clones from regions of human X chromosome with different characteristic levels of GC content. The method starts by randomly fragmenting a BAG, YAC or PAC to 8-12 kb pieces and subcloning those into lambda phage. Insert-ends of these clones are sequenced and overlapped to create a partial map. Complete sequencing is then done on a minimal tiling path of selected subclones, recursively focusing on those at the edges of contigs to facilitate mergers of clones across the entire target. To reduce manual labor, PCR processes have been adapted to prepare sequencing templates throughout the entire operation. The streamlined process can thus lend itself to further automation.The OSS approach is suitable for large-scale genomic sequencing, providing considerable flexibility in the choice of subclones or regions for more or less intensive sequencing. For example, subclones containing contaminating host cell DNA or cloning vector can be recognized and ignored with minimal sequencing effort; regions overlapping a neighboring clone already sequenced need not be redone; and segments containing tandem repeats or long repetitive sequences can be spotted early on and targeted for additional attention.
DNA comprising 219 447 bp was sequenced in nine cosmids and verified at > 99.9% precision. Of the standard repetitive elements, 187 Alus make up 20.6% of the sequence, but there were only 27 MERs (2.9%) and 17 L1 fragments (1.6%). This may be characteristic of such high GC (57%) regions. The sequence also includes an 11.3 kb tract duplicated with 99.2% identity at a distance of 38 kb. The region is 80-90% transcribed and 12.5% translated. Thirteen known genes and their exon-intron borders are all accurately predicted at least in part by GRAIL programs, as are six additional genes. From centromere to telomere, the orientation of transcription varies among the first eight genes, then runs centromeric to telomeric for the next five, and is in the opposite sense for the last six. Eighteen of the 19 genes are associated with CpG islands. Two islands are exact copies in the 11.3 kb repeat units, and could thus give rise to double dosage levels of an X-linked gene. Another island is associated with two genes transcribed in opposite directions. From the sequence data, three genes and their exon structure are inferred. One of them, previously associated with HEX2, is shown to be a different gene unrelated to hexokinases; a second gene, previously known by an EST, is plexin, from its 65.5% identity with the Xenopus analog; and a third is a subunit of a vacuolar H-ATPase, and is named VATPS1.
Ectodermal dysplasias comprise over 150 syndromes of unknown pathogenesis. X-linked anhidrotic ectodermal dysplasia (EDA) is characterized by abnormal hair, teeth and sweat glands. We now describe the positional cloning of the gene mutated in EDA. Two exons, separated by a 200-kilobase intron, encode a predicted 135-residue transmembrane protein. The gene is disrupted in six patients with X;autosome translocations or submicroscopic deletions; nine patients had point mutations. The gene is expressed in keratinocytes, hair follicles, and sweat glands, and in other adult and fetal tissues. The predicted EDA protein may belong to a novel class with a role in epithelial-mesenchymal signalling.
Simpson-Golabi-Behmel syndrome (SGBS) is an X-linked condition characterized by pre- and postnatal overgrowth with visceral and skeletal anomalies, To identify the causative gene, breakpoints in two female patients with X;autosome translocations were identified. The breakpoints occur near the 5' and 3' ends of a gene, GPC3, that spans more than 500 kilobases in Xq26; in three families, different microdeletions encompassing exons cosegregate with SGBS. GPC3 encodes a putative extracellular proteoglycan, glypican 3, that is inferred to play an important role in growth control in embryonic mesodermal tissues in which it is selectively expressed. Initial western- and ligand-blotting experiments suggest that glypican 3 forms a complex with insulin-like growth factor 2 (IGF2), and might thereby modulate IGF2 action.
The humanZFX,humanZFY,and mouseZfxgenes have CpG islands near their 5′ ends. These islands are typical in that they span about 1.5 kb, contain transcription initiation sites, and encompass some 5′ untranslated exons and introns. However, comparative nucleotide sequencing of these human and mouse islands provided evidence of evolutionary conservation to a degree unprecedented among mammalian 5′ CpG islands. In one stretch of 165 nucleotides containing 19 CpGs, mouseZfxand humanZFXare identical to each other and differ from humanZFYat only 9 nucleotides. In contrast, we found no evidence of homologous CpG islands in the mouseZfygenes, whose transcription is more circumscribed than that of humanZFX,humanZFY,and mouseZfx.Using the isoschizomersHpaII andMspI to examine a highly conserved segment of theZFXCpG island, we detected methylation on inactive mouse X chromosomes but not on inactive human X chromosomes. These observations parallel the previous findings that mouseZfxundergoes X inactivation while humanZFXescapes it.
A cosmid containing 36.4 kb of high GC human DNA centromeric to the G6PD gene has been analyzed. The sequence was 99.9% precise, based on the comparison of 4.3 kb that overlaps an earlier analysis of 20.1 kb containing G6PD. Properties of the entire 52 kb region that may be characteristic of high GC portions of the genome include a very high density of sixty-two half or full Alu sequences, or 1.2/kb, and an absence of L1 sequences. Other highly repetitive sequences include 11 MER sequences, one of them interrupted at two positions by groups of 3 Alu elements. In segments of unique sequence, computer-aided analysis predicted three possible genes, one of which has thus far been confirmed by the recovery of a corresponding cDNA, both by a direct hybridization method and by a PCR-based method based on a primer pair inferred from the genomic sequence. The cDNA has been sequenced, and is completely concordant with counterpart genomic sequence; it has no resemblance to any previously described gene.
Cytogenetic analysis of a pediatric patient with T-cell acute lymphoblastic leukemia (T-ALL) revealed a mosaic karyotype, 47,XX,+17,t(11;14)(p13;q11)/47,XX,+17,t(9;22)(q34;q11),t(11;14) (p13;q11). DNA blot analysis was used to examine the break-point within the BCR gene on chromosome 22 and showed that the breakpoint occurred within the 20-kb minor breakpoint cluster region (m-bcr) located within the first intron of the BCR gene. Immunoprecipitation analysis demonstrated that the leukemic cells expressed the P185 BCR-ABL protein tyrosine kinase. P185 BCR-ABL has previously been shown to be expressed in most cases of Ph+ acute leukemia of myeloid and B-progenitor origin. Here, we demonstrate for the first time that P185 can also be expressed in the T-cell lineage. DNA blot hybridization was also used to characterize the t(11;14) translocation. This showed rearrangement on chromosome 11 within the T-ALLbcr region, upstream of the RBTN-2 gene. Polymerase chain reaction revealed the presence of RBTN-2 transcripts in the leukemic cells. Finally, comparison of the T-ALLbcr, BCR-ABL, IGH, TCR beta and gamma gene rearrangements in leukemic cells obtained at the time of diagnosis and at first relapse showed that relapse occurred in a leukemic clone indistinguishable from the major Ph+ clone involved at diagnosis. Together, these data support a multistep pathogenesis in which the Philadelphia (Ph) chromosome translocation appeared subsequent to the +17 and t(11;14) and imparted a growth advantage over the Ph-negative cells that carried these abnormalities.
Cytogenetic analysis of a pediatric patient with T-cell acute lymphoblastic leukemia (T-ALL) revealed a mosaic karyotype, 47,XX,+17,t(11;14)(p13;q11)/47,XX,+17,t(9;22)(q34;q11),t(11;14) (p13;q11). DNA blot analysis was used to examine the breakpoint within the BCR gene on chromosome 22 and showed that the breakpoint occurred within the 20-kb minor breakpoint cluster region (m-bcr) located within the first intron of the BCR gene. Immunoprecipitation analysis demonstrated that the leukemic cells expressed the P185 BCR-ABL protein tyrosine kinase. P185 BCR-ABL has previously been shown to be expressed in most cases of Ph+ acute leukemia of myeloid and B-progenitor origin. Here, we demonstrate for the first time that P185 can also be expressed in the T-cell lineage. DNA blot hybridization was also used to characterize the t(11;14) translocation. This showed rearrangement on chromosome 11 within the T-ALL(bcr) region, upstream of the RBTN-2 gene. Polymerase chain reaction revealed the presence of RBTN-2 transcripts in the leukemic cells. Finally, comparison of the T-ALL(bcr), BCR-ABL, IGH, TCR beta and gamma gene rearrangements in leukemic cells obtained at the time of diagnosis and at first relapse showed that relapse occurred in a leukemic clone indistinguishable from the major Ph+ clone involved at diagnosis. Together, these data support a multistep pathogenesis in which the Philadelphia (Ph) chromosome translocation appeared subsequent to the +17 and t(11,14) and imparted a growth advantage over the Ph-negative cells that carried these abnormalities.
A temperature-sensitive mutation (act1-1) in the essential actin gene of Saccharomyces cerevisiae can be suppressed by mutations in the SAC2 gene. A cloned genomic DNA fragment that complements the cold-sensitive growth phenotype associated with such a suppressor mutation (sac2-1) was sequenced. The fragment contained an open reading frame that encodes a 641 amino acid predicted hydrophilic protein with a molecular weight of 74,445. No sequences with significant similarity to SAC2 were found in the GenBank and EMBL databases. A SAC2 disruption mutation was constructed which had phenotypes similar to the sac2-1 point mutation. A haploid SAC2 disruption strain failed to grow at low temperature and the disruption allele suppressed the temperature-sensitive act1-1 growth defect. The suppression phenotype was dependent on the strain background.
A full-length mouse glucose-6-phosphate dehydrogenase (G6PD) cDNA has been isolated and sequenced, and the evolutionary conservation of many portions of the sequence has been verified by comparison with that of human and other sources.
Ordered shotgun sequencing proposes to organize the mapping and sequencing of YACs with a hierarchical strategy that incorporates a feedback loop. Building on current protocols, a YAC is subcloned into plasmids, plasmid insert ends are sequenced, and the sequences are overlapped to create a partial map. Complete sequencing then starts with plasmids whose end-sequence tracts have overlapped, but to a minimal extent. The next plasmids to be sequenced are again selected for least overlap, striking out progressively to span the YAC with minimal directed gap-filling. Simulations support its feasibility and indicate that during the generation of the complete sequence, the approach facilitates the early choice of regions for selective sequencing, for example, for coding units. The sequencing of plasmids would also require less redundancy, and discriminate repetitive sequences more easily, than random sequencing across larger clones. The overall effort scales with YAC size and can be further reduced by additional mapping information.
We isolated a gene encoding a 218 kDa myosin-like protein from Saccharomyces cerevisiae using a monoclonal antibody directed against human platelet myosin as a probe. The protein sequence encoded by the MLP1 gene (for myosin-like protein) contains extensive stretches of a heptad-repeat pattern suggesting that the protein can form coiled coils typical of myosins. Immunolocalization experiments using affinity-purified antibodies raised against a TrpE-MLP1 fusion protein showed a dot-like structure adjacent to the nucleus in yeast cells bearing the MLP1 gene on a multicopy plasmid. In mouse epithelial cells the yeast anti-MLP1 antibodies stained the nucleus. Mutants bearing disruptions of the MLP1 gene were viable, but more sensitive to ultraviolet light than wild-type strains, suggesting an involvement of MLP1 in DNA repair. The MLP1 gene was mapped to chromosome 11, 25 cM from met1.
We have developed a general method for screening randomly mutagenized expression libraries in mammalian cells by using fluorescence-activated cell sorting (FACS). The cDNA sequence of a secreted protein is randomly mutagenized by PCR under conditions of reduced Taq polymerase fidelity. The mutated DNA is inserted into an expression vector encoding the membrane glycophospholipid anchor sequence of decay-accelerating factor (DAF) fused to the C terminus of the secreted protein. This results in expression of the protein on the cell surface in transiently transfected mammalian cells, which can then be screened by FACS. This method was used to isolate mutants in the kringle 1 (K1) domain of tissue plasminogen activator (t-PA) that would no longer be recognized by a specific monoclonal antibody (mAb387) that inhibits binding of t-PA to its clearance receptor. DNA sequence analysis of the mutants and localization of the mutated residues on a three-dimensional model of the K1 domain identified three key discontinuous amino acid residues that are essential for mAb387 binding. Mutants with changes in any of these three residues were found to have reduced binding to the t-PA receptor on human hepatoma HepG2 cells but to retain full clot lysis activity.
The sequence of 20,114 bp of DNA including the human glucose-6-phosphate dehydrogenase (G6PD) gene was determined. The region included a prominent CpG island, starting about 680 nucleotides upstream of the transcription start site, extending about 1050 nucleotides downstream of the start site, and ending just at the start of the first intron. The transcribed region from the start site to the poly (A) addition site covers 15,860 bp. The sequence of the 13 exons agreed with published cDNA sequence and for the 11 exons tested, with the corresponding sequence in a yeast artificial chromosome (YAC). The latter confirms YAC cloning fidelity at the DNA sequence level. Sixteen Alu sequences constitute 24% of the total sequence tract. Four were outside the borders of the mRNA transcript of the gene; all the others were found in a large (9858 bp) intron between exons 2 and 3. Two Alu clusters each contain Alus lying between the monomers of another.