
ABSTRACT Recent Escherichia coli genome sequencing efforts in the 76–81.5 min and 92.8–0.1 min regions have revealed three new operons or gene clusters concerned with sugar metabolism, which we designate sga (6 ORFs), sgb (4 ORFs), and sgc (6 ORFs). All of these operons encode proteins homologous to pentose-phosphate-4-epimerases ( sga and sgb ) or pentose-phosphate-3-epimerases ( sgc ), indicating that these operons are involved in the metabolism of pentoses or pentitols. sga and sgc (but not sgb ) encode also proteins of the bacterial phosphotransferase system (PTS), whereas sga and sgb encode three non-PTS proteins (one of which is the pentose-phosphate-4-epimerase homologue) that are homologous to each other. The two PTS protein homologues in the sga operon are members of (1) the family of mannitol-specific and fructose-specific IIA proteins and (2) the family of lactose-specific and cellobiose-specific IIB proteins. A IIC PTS homologue was not found in the sga operon, but a permease-like protein designated SgaT with 12 putative transmembrane helical segments may provide this transport function. Although the sgb operon lacks ORF homologous to PTS proteins, it contains a cryptic gene encoding L -xylulose kinase. The sgc gene cluster encodes two PTS proteins homologous to (1) the mannitol-specific and fructose-specific IIA proteins and (2) the galactitol IIC protein. It also encodes a transcriptional regulator of the DeoR family, a pentose-phosphatase-3-epimerase homologue, and an ORF homologous to a protein encoded in the recently described frv operon. Computer analyses of the DNA sequences and the encoded protein sequences are presented, and the potential roles of these operons in carbohydrate metabolism are discussed.
ABSTRACT The expressed sequence tag (EST) strategy is a useful approach for describing gene diversity and identifying changes in expression patterns between different populations of cells. We constructed two directionally cloned cDNA libraries from human Jurkat T cell leukemia cells synchronized in either the G 1 or the S phase of the cell cycle. Sequence analysis of 3398 randomly selected clones from G 1 phase and 3428 randomly selected clones from S phase was carried out for the purpose of comparing gene expression profiles. EST sequencing identified 1728 distinct transcripts from the G 1 phase library and 1623 distinct transcripts from the S phase library. Approximately 13% of these transcripts appeared to be differentially expressed between the G1 and S phases of the cell cycle. Among the differentially expressed genes were alpha k-1 tubulin, alpha enolase, CDC p55, and ubiquitin. The human homologue of cyclin B2 is also differentially expressed in the G 1 phase vs the S phase of the cell cycle. Northern blot analysis of mRNA levels in cells from either the G 1 or S phase of the cell cycle correlated well with the quantitation of mRNA from EST profile analysis.
Genome Science and TechnologyVol. 1, No. 4 The Eighth International Genome Sequencing and Analysis Conference, Hilton Head Island, October 5-8, 1996: SummaryDARRELL J. DOYLEDARRELL J. DOYLESearch for more papers by this authorPublished Online:27 Apr 2009https://doi.org/10.1089/mcg.1996.1.339AboutSectionsPDF/EPUB Permissions & CitationsPermissionsDownload CitationsTrack CitationsAdd to favorites Back To Publication ShareShare onFacebookTwitterLinked InRedditEmail FiguresReferencesRelatedDetails Volume 1Issue 4Jan 1996 To cite this article:DARRELL J. DOYLE.The Eighth International Genome Sequencing and Analysis Conference, Hilton Head Island, October 5-8, 1996: Summary.Genome Science and Technology.Jan 1996.339-342.http://doi.org/10.1089/mcg.1996.1.339Published in Volume: 1 Issue 4: April 27, 2009PDF download
ABSTRACT Current methods of forensic DNA identification at our facilities use PCR for typing single nucleotide and tandem repeat polymorphisms. Unfortunately, these PCR-based methods are relatively expensive and time consuming and are not well suited for total automation. The ligase detection reaction (LDR) when used in conjunction with PCR offers distinct advantages. In LDR, two adjacent primers hybridize to the target and are ligated only when there is perfect complementarity at the junction. Since Taq DNA ligase used in LDR is thermostable, several rounds of thermal cycling may be used to unambiguously distinguish any single nucleotide polymorphisms. We have developed a coupled multiplex PCR-LDR assay to type single base variations at 12 biallelic loci, giving a power of discrimination of 1.12 × 10 5 . The 12 loci are PCR amplified in a single reaction using a unique two-step method that produces similar amounts of multiplexed products without the need to carefully adjust primer concentrations or PCR conditions. Following PCR, these products are used in a single LDR to generate products that are resolved and typed on an Applied Biosystems 373 DNA sequencer, creating an LDR profile. Our ability to easily generate similar amounts of product in a 12-locus multiplex PCR amplification is the basis for expanding the assay to type 30 biallelic loci to give a theoretical power to discriminate one individual in 10 12 .
ABSTRACT Pyrococcus furiosus is a marine microorganism with the unusual ability to grow optimally at 100°C. It is classified as a member of the domain Archaea (Archaebacteria). We present studies on the genome of P. furiosus , consisting of a precise determination of the size of the chromosome, 2.05 mb, and an analysis of sequence data from cDNA and genomic libraries. The sequence analysis included a total of 1176 sequences, representing 14.7% of the genome, which were compared with the current databases using the BLAST X algorithm. Expressed sequence tag (EST) analysis is skewed toward the repeated detection of highly expressed genes, for example, ribosomal RNA, in P. furiosus . The relatively small size of the genome implies a high coding density, and randomly chosen genomic sequences provide good performance in gene discovery with relatively fewer identical hits compared with the EST database. The recovery of database matches with P(n) ≤ 1.0e −05 was 30% of the total sequences tested. We estimate that the genome contains approximately 1800 genes, and this study has provided evidence for 309 genes. The average G+C content of the sequences obtained was 41.2%, and no repetitive sequences were detected. Genes from Eukarya and Bacteria are almost equally represented in the matches obtained in this study. The homologs representing central metabolism exhibit preferential similarity to their homologs from bacteria, despite the availability of eukaryotic homologs in the databases. Several characteristically eukaryotic protein homologs were found with functions in transcription, translation, and membrane transport. We propose that the archaeal genome is a mosaic, consisting of ancient gene sequences related to eukaryal homologs and more recently acquired genes, many of which are related to bacterial metabolic functions. Further sequencing and phylogenetic studies are needed to confirm this hypothesis, which predicts that there was extensive lateral genetic transfer among Bacteria, Eukarya, and Archaea during a period when the original gene pool was expanding into the early lineages of life.
ABSTRACT Identification of quantitative alterations in the gene expression during small cell lung cancer (SCLC), if sufficiently characterized, may result in novel molecular markers that may be useful in the diagnosis and treatment of human small cell lung cancer. The recently developed mRNA differential display technique has been used to identify differentially expressed sequence tags (EST), short complementary DNA fragments corresponding to mRNA that are differentially expressed in SCLC as compared with normal human bronchial epithelial cell lines as control. DNA sequencing followed by computer search against sequences present in Genbank and EMBL DNA databases indicated that one tag was novel and two had high homology with reported genes: the lissencephaly-1 gene involved in Miller-Dieker syndrome and human peptide-binding protein, which is a new member of the heat shock protein 70 (hsp70) family. Lissencephaly-1 gene maps to a region of chromosome 17p13.3, which has been found to be repeatedly deleted in SCLC. This is the first report, to our knowledge, on the tags of the genes differentially expressed between normal human bronchial epithelial and SCLC cells. Loss of their expression in SCLC could contribute to tumor formation or progression or both.
The complete sequence of the Mycoplasma genitalium chromosome has recently been determined. We here report analyses of the genes encoding proteins of the phosphoenolpyruvatersugar phosphotransferase system, PTS. These genes encode (1) Enzyme I, (2) HPr, (3) a glucose-specific Enzyme IICBA, (4) an inactive glucose-specific Enzyme IIB, lacking the active site cysteyl residue, and (5) a fructose-specific Enzyme IIABC. Some of the unique features of these genes and their enzyme products are as follows. (1) Each of the genes is encoded within a distinct operon. (2) Both Enzyme I and HPr have basic isoelectric points. (3) The glucose-specific Enzyme IIC bears a centrally located, hydrophilic, 200 amino acyl residue insert that lacks sequence similarity with any protein in the current database. (4) The fructose-specific Enzyme II has a domain order (IIABC), different from those of previously characterized fructose permeases, and its IIA domain more closely resembles the IIANtr protein of Escherichia coli than other fructose-specific IIA domains. The potential significance of these novel features is discussed.
We describe here the first general survey of the genomic content and the coding capacity of the 1.1 Mb genome of Rickettsia prowazekii based on an analysis of a total of 200 kb of unique sequence data collected in a random manner. Based on nucleotide distribution profiles, we estimate that the R. prowazekii genome may have a coding density of 60%-70% and that it may contain a total of circa 800 genes. Here, we have tentatively identified and classified 173 of these genes. Our analysis suggests that the R. prowazekii genome is a highly derived, reduced genome that has lost many genes involved in amino acid biosynthetic pathways and regulatory functions. Furthermore, the R. prowazekii genome seems to lack glycolytic genes, but it does contain genes encoding components of the tricarboxylic acid cycle as well as of the electron transport system. We have also encountered a family of homologous genes coding for ATP/ADP translocases, as observed in several mitochondrial genomes. We relate these findings to previous phylogenetic studies that suggest that Rickettsia and mitochondria share a common ancestor.
A pyrimidine-rich element (PyRE), present in the 21st intron of the PKD1 gene, posed a significant obstacle in determining the primary structure of the gene. Only cycle sequencing of nested, single-stranded phage templates of the CT-rich strand enabled complete and accurate sequence data. Similar attempts on the GA-rich strand were unsuccessful. The resulting primary structure showed the 3 kb 21st intron to contain a 2.5 kb PyRE, whose sense-strand is 97% C + T. The PKD1 PyRE does not appear to be polymorphic based on RFLP analysis of DNA from 6 unrelated individuals digested with 9 different restriction enzymes. This is the largest pyrimidine tract sequenced to date, being over twice as large as those previously identified and shows little homology to other polypyrimidine tracts. Additional analysis of this PyRE revealed the presence of 23 mirror repeats with stem lengths of at least 10 nucleotides. The 23 H-DNA-forming sequences in the PKD1 PyRE exceed the cumulative total of 22 found in 157 human genes that have been completely sequenced. The mirror repeats confer this region of the PKD1 gene with a strong probability of forming H-DNA or triplex structures under appropriate conditions. Based on studies with PyRE found in other eukaryotic genes, the PKD1 PyRE may play a role in regulating PKD1 expression, and its potential for forming an extended triplex structure may explain some of the observed instability in the PKD1 locus.
A large-scale sequencing project requires a tool to control the quality of the input data because a sizable number of trace data may be of low quality. If these data are allowed to enter the sequence assembly pipeline, harm will be done. Hence, it is important to detect such data as soon as possible. MTT (Move-Track-Trim) is a software package analyzing the quality of the lanes. It subjects each lane to a series of tests, and if a lane does not pass all tests, it is flagged as a "bad" lane. The use has a chance to examine both the "good" and the "bad" lanes and reclassify a "bad" lane as "good," or vice versa. Alternatively, the user may decide to retrack the gel or get rid of some lanes altogether. As a by-product of the analysis, MTT performs other useful functions. It trims the lanes and compresses the lane files and moves them to the directories where assembly is carried out. It also generates some useful statistics describing the quality of the gel.
We describe a computer program, named DNA-Protein Search (DPS), for comparing a megabase DNA sequence with a protein sequence database. The DPS program addresses the problems of frameshifts and introns in the DNA sequence. The DPS program was used to compare each of the following sequences with the Swiss-Prot database: the 1.8-megabase sequence of the Haemophilus influenzae Rd genome, the 0.58-megabase sequence of the Mycoplasma genitalium genome, and the 0.56-megabase sequence of Saccharomyces cerevisiae chromosome VIII. The comparisons found new regions that are similar to protein sequences. The sensitivity of DPS was evaluated using as test data the known coding regions of the three DNA sequences. The results demonstrate that the DPS program is a useful tool for finding the coding regions of the DNA sequence. The DPS program uses an order of magnitude less computer memory and is several times faster than the BLASTX program.
With the advent of megabase genome sequencing, the need for computational analyses increases exponentially. Sequencing errors must be corrected, encoded proteins must be identified, functions must be assigned to these proteins, and distant phylogenetic relationships must be recognized in order to maximize the yield of information obtainable from genome sequencing projects. Both the computer and the human brain have their limitations, but using them in combination, the biologist can vastly extend his or her analytic capabilities. Computer techniques can be used to estimate protein structure, function, biogenesis, and evolution. In this review, the application of available computer programs to several protein families, particularly transport, receptor, and transcriptional regulatory protein families, illustrate our current capabilities and limitations. Although some multidomain protein families are evolutionarily homogeneous, others have mosaic origins. Evidence concerning the nature and frequency of occurrence of domain shuffling, splicing, fusion, deletion, and duplication during evolution of specific protein families is evaluated. It is shown that specific families of enzymes, receptors, transport proteins, and transcriptional regulatory proteins share a common evolutionary origin, frequently diverging in function because of domain splicing and ligation. Some large families arose gradually over evolutionary time, whereas others developed suddenly, due to bursts of intragenic or intergenic (or both) duplication events occurring over relatively short periods of time. It is argued that energy coupling to transport was a late occurrence, superimposed on preexisting mechanisms of solute facilitation. It is also shown that several transport protein families have evolved independently of each other, employing different routes, at different times in evolutionary history, to give topologically similar transmembrane protein complexes.
As genomic research proliferates, DNA banking will become more common. In research, samples will be banked largely in an effort to find and clone genes that predispose to disease. Commercially oriented banks, those that offer services to families, may also become more common. These entities will hold sensitive information. DNA banking is not yet regulated. We argue here that new laws are not needed at this time to regulate DNA banking. We suggest an approach that relies on a professional code of conduct and draws on principles of disclosure inherent to the process used in obtaining informed consent. In addition to suggesting 12 specific recommendations for the code of conduct, we suggest that items should be included in depositor's agreements. We offer a rationale for our suggestions.
Analysis of genomic sequences is necessarily an ongoing process. Initial gene assignments tend (wisely) to be on the conservative side (Venter, 1996). The analysis of the genome then grows in an iterative fashion as additional data and more sophisticated algorithms are brought to bear on the data. The present report is an emendation of the original gene list of Methanococcus jannaschii (Bult et al., 1996). By using a somewhat more updated database and more relaxed (and operator-intensive) pattern matching methods, we were able to add significantly to, and in a few cases amend, the gene identification table originally published by Bult et al. (1996).
ABSTRACT A new approach to assembling large, random shotgun sequencing projects has been developed. The TIGR Assembler overcomes several major obstacles to assembling such projects: the large number of pairwise comparisons required, the presence of repeat regions, chimeras introduced in the cloning process, and sequencing errors. A fast initial comparison of fragments based on oligonucleotide content is used to eliminate the need for a more sensitive comparison between most fragment pairs, thus greatly reducing computer search time. Potential repeat regions are recognized by determining which fragments have more potential overlaps than expected given a random distribution of fragments. Repeat regions are dealt with by increasing the match criteria stringency and by assembling these regions last so that maximum information from nonrepeat regions can be used. The algorithm also incorporates a number of constraints, such as clone length and the placement of sequences from the opposite ends of a clone. TIGR Assembler has been used to assemble the complete 1.8 Mbp Haemophilus influenzae (Fleischmann et al., 1995) and 0.58 Mbp Mycoplasma genitalium (Fraser et al., 1995) genomes.
ABSTRACT A procedure using a 96-well format for growth, template preparation, and quantification was optimized to facilitate high-throughput automated sequencing of double stranded plasmid DNA templates. Modification of the 96-well Miniprep Kit [Advanced Genetic Technologies (AGTC), Gaithersburg, MD] protocol combined with the DNA quantification with a Millipore Cytofluor 2350 (Millipore, Bedford, MA) provided high-quality DNA plasmid templates rapidly, efficiently, and inexpensively. We utilized this procedure to prepare more than 130,000 templates for a number of large-scale cDNA and genomic projects. Four specific projects were: human expressed sequence tag (EST) project (Adams et al., 1994) (70,080 templates), Haemophilus influenzae (Fleischmann et al., 1995) (22,656 templates), Mycoplasma genitalium (Fraser et al., 1995) (5760 templates), and Methanococcus jannaschii (22,175 templates). Data from these projects support a sequencing success rate of 89%–95%, with an average read length of 432 bases with <1% ambiguous bases.
ABSTRACT The first wave of practical products from the international effort to produce systematic maps of the human genome and improve DNA sequencing technologies is taking the form of new tools to predict the risk of specific diseases in individual patients. DNA-based tests for molecular mutations associated with clinical syndromes increasingly allow clinicians to detect disease processes and health risks before clinical problems occur, sometimes making prevention possible. The increasing predictive power of medical diagnostics, however, also poses serious health policy challenges at both the professional and societal levels. At the professional level, DNA-based health risk assessments challenge traditional ethical commitments to confidentially, informed consent, and nondirective genetic counseling. At the societal level, these tests challenge institutions and governments to clarify their policies regarding access to opportunities by defining fair uses of genetic health risk information about individuals outside the clinical setting. Underlying all these issues are basic questions about how we interpret the meaning of DNA-based diagnostic findings.