ABSTRACT A cornerstone of bacterial molecular biology is the ability to genetically manipulate the microbe under study. Many bacteria are difficult to manipulate genetically, a phenotype due in part to robust removal of newly acquired DNA, for example, by restriction-modification (R-M) systems. Here, we report approaches that dramatically improve bacterial transformation efficiency, piloted using a microbe that is challenging to transform due to expression of many R-M systems, Helicobacter pylori . Initially, we identified conditions that dampened expression of several R-M systems and concomitantly enhanced transformation efficiency. We then identified an approach that would broadly protect newly acquired DNA. We computationally predicted under-represented short DNA sequences in the H. pylori genome, with the idea that these sequences reflect targets of sequence-based surveillance such as R-M systems. We then used this information to modify and eliminate such sites in antibiotic resistance cassettes, creating a “stealth” version. Modifying antibiotic resistance cassettes in this way resulted in significantly higher transformation efficiency compared to non-modified cassettes, a response that was genomic loci independent. Our results suggest that avoiding R-M systems, via modification of under-represented DNA sequences or transformation conditions, is a powerful method to enhance DNA transformation. Our approach to identify under-represented sequences is applicable to any microbe with a sequenced genome. IMPORTANCE Manipulating the genomes of bacteria is critical to many fields. Such manipulations are made by genetic engineering, which often requires new pieces of DNA to be added to the genome. Bacteria have robust systems for identifying and degrading new DNA, some of which rely on restriction enzymes. These enzymes cut DNA at specific sequences. We identified a set of DNA sequences that are missing normally from a bacterium’s genome, more than would be expected by chance. Eliminating these sequences from a new piece of DNA allowed it to be incorporated into the bacterial genome at a higher frequency than new DNA containing the sequences. Removing such sequences appears to allow the new DNA to fly under the bacterial radar in “stealth” mode. This transformation improvement approach is straightforward to apply and likely broadly applicable.
ABSTRACT Many bacterial genomes are highly variable but nonetheless are typically published as a single assembled genome. Experiments tracking bacterial genome evolution have not looked at the variation present at a given point in time. Here, we analyzed the mouse-passaged Helicobacter pylori strain SS1 and its parent PMSS1 to assess intra- and intergenomic variability. Using high sequence coverage depth and experimental validation, we detected extensive genome plasticity within these H. pylori isolates, including movement of the transposable element IS607, large and small inversions, multiple single nucleotide polymorphisms, and variation in cagA copy number. The cagA gene was found as 1 to 4 tandem copies located off the cag island in both SS1 and PMSS1; this copy number variation correlated with protein expression. To gain insight into the changes that occurred during mouse adaptation, we also compared SS1 and PMSS1 and observed 46 differences that were distinct from the within-genome variation. The most substantial was an insertion in cagY, which encodes a protein required for a type IV secretion system function. We detected modifications in genes coding for two proteins known to affect mouse colonization, the HpaA neuraminyllactose-binding protein and the FutB α-1,3 lipopolysaccharide (LPS) fucosyltransferase, as well as genes predicted to modulate diverse properties. In sum, our work suggests that data from consensus genome assemblies from single colonies may be misleading by failing to represent the variability present. Furthermore, we show that high-depth genomic sequencing data of a population can be analyzed to gain insight into the normal variation within bacterial strains. IMPORTANCE Although it is well known that many bacterial genomes are highly variable, it is nonetheless traditional to refer to, analyze, and publish “the genome” of a bacterial strain. Variability is usually reduced (“only sequence from a single colony”), ignored (“just publish the consensus”), or placed in the “too-hard” basket (“analysis of raw read data is more robust”). Now that whole-genome sequences are regularly used to assess virulence and track outbreaks, a better understanding of the baseline genomic variation present within single strains is needed. Here, we describe the variability seen in typical working stocks and colonies of pathogen Helicobacter pylori model strains SS1 and PMSS1 as revealed by use of high-coverage mate pair next-generation sequencing (NGS) and confirmed by traditional laboratory techniques. This work demonstrates that reliance on a consensus assembly as “the genome” of a bacterial strain may be misleading.
Significance Modification of cytosine bases in DNA can determine when genes are turned on in biological cells. These modifications are important during cell differentiation, embryogenesis, and aberrant cell growth in cancer. Here, we present a nanopore technique that permits direct detection of cytosine, 5-hydroxymethylcytosine, and 5-methylcytosine on individual synthetic DNA strands of known sequence. This technique focuses on three ionic current amplitudes that occur as an enzyme motor pulls the cytosine on a captured DNA strand through the nanopore. For genomic DNA, we predict that a given strand must be read 5–19 times to achieve cytosine methylation calls that are sufficiently accurate for epigenetic studies.
A key obstacle to sequencing DNA as it passes through a nanopore is that the translocation rate is too fast to resolve individual bases. Cherf et al. solve this problem with an improved method for ratcheting DNA forward and backward through the nanopore using a DNA polymerase. An emerging DNA sequencing technique uses protein or solid-state pores to analyze individual strands as they are driven in single-file order past a nanoscale sensor1,2,3. However, uncontrolled electrophoresis of DNA through these nanopores is too fast for accurate base reads4. Here, we describe forward and reverse ratcheting of DNA templates through the α-hemolysin nanopore controlled by phi29 DNA polymerase without the need for active voltage control. DNA strands were ratcheted through the pore at median rates of 2.5–40 nucleotides per second and were examined at one nucleotide spatial precision in real time. Up to 500 molecules were processed at ∼130 molecules per hour through one pore. The probability of a registry error (an insertion or deletion) at individual positions during one pass along the template strand ranged from 10% to 24.5% without optimization. This strategy facilitates multiple reads of individual strands and is transferable to other nanopore devices for implementation of DNA sequence analysis.
Multiple Sequence Alignments are fundamental to many sequence analysis methods. Most alignments are computed using the Progressive Alignment heuristic. These methods are starting to become a bottleneck in some analysis pipelines when faced with data-sets of the size of many thousands of sequences. Some methods allow computation of larger datasets while sacrificing quality, and others produce high quality alignments, but scale badly with the number of sequences. In this paper, we describe a new program called Clustal Omega which can align virtually any number of protein sequences quickly and that delivers accurate alignments. The accuracy of the package on smaller test-cases is similar to that of the high-quality aligners. On larger data-sets Clustal Omega outperforms other packages in terms of execution time and quality. Clustal Omega also has powerful features for adding sequences to and exploiting information in existing alignments, making use of the vast amount of precomputed information in public databases like Pfam.
Motivation: Our focus has been on detecting topological properties that are rare in real proteins, but occur more frequently in models generated by protein structure prediction methods such as Rosetta. We previously created the Knotfind algorithm, successfully decreasing the frequency of knotted Rosetta models during CASP6. We observed an additional class of knot-like loops that appeared to be equally un-protein-like and yet do not contain a mathematical knot. These topological features are commonly referred to as slip-knots and are caused by the same mechanisms that result in knotted models. Slip-knots are undetectable by the original Knotfind algorithm. We have generalized our algorithm to detect them, and analyzed CASP6 models built using the Rosetta loop modeling method. Results: After analyzing known protein structures in the PDB, we found that slip-knots do occur in certain proteins, but are rare and fall into a small number of specific classes. Our group used this new Pokefind algorithm to distinguish between these rare real slip-knots and the numerous classes of slip-knots that we discovered in Rosetta models and models submitted by the various CASP7 servers. The goal of this work is to improve future models created by protein structure prediction methods. Both algorithms are able to detect un-protein-like features that current metrics such as GDT are unable to identify, so these topological filters can also be used as additional assessment tools. Contact: firas@u.washington.edu
Residue burial, which describes a protein residue's exposure to solvent and neighboring atoms, is key to protein structure prediction, modeling, and analysis. We assessed 21 alphabets representing residue burial, according to their predictability from amino acid sequence, conservation in structural alignments, and utility in one fold‐recognition scenario. This follows upon our previous work in assessing nine representations of backbone geometry. 1 The alphabet found to be most effective overall has seven states and is based on a count of Cβ atoms within a 14 Å‐radius sphere centered at the Cβ of a residue of interest. When incorporated into a hidden Markov model (HMM), this alphabet gave us a 38% performance boost in fold recognition and 23% in alignment quality. Proteins 2004. © 2004 Wiley‐Liss, Inc.
Local structure alphabets are discrete encodings of one or more properties of local protein structure that cluster residues with similar properties into the same state. They allow us to represent a protein's structure, in simplified form, as a one-dimensional string. I explore whether there are preferred ways to encode local protein structure to best recognize relationships between distantly related proteins. To identify the most informative alphabets of local protein structure, I have developed an evaluation protocol and applied it to 48 candidate alphabets. The evaluation includes many new alphabets, as well as some taken from the literature, covering descriptions of backbone geometry, residue burial, and side-chain orientations. The main criteria sought in a local structure alphabet are predictability, conservation within a collection of structurally-similar proteins, and improvement in fold recognition and alignment quality. An important problem in computational biology is predicting the structure of the large number of putative proteins discovered by genome sequencing projects. Knowledge-based methods attempt to solve the problem by relating the target proteins to known structures, searching for template proteins homologous to the target. Distant homologs, which may have significant structural similarity, are often not detectable by sequence similarities alone. These weak relationships may be recognized by combining sequence information and evolutionary information (about the target and template), structural information (about the template), and predicted information about the target's local structure. All these kinds of information can be incorporated into threading algorithms [39, 175], hidden Markov models (HMMs) [66, 109] or profiles [113]. When compared to the baseline fold-recognition and alignment performance of a HMM that uses only amino-acid information, HMM s enhanced with a secondary track of local-structure-alphabet emissions show a substantial improvement when judiciously selected alphabets are used. A simple three-state classification of secondary structure and a seven-state description of residue burial, based on a count of neighboring C β atoms within a 14Å-radius spherical cutoff, are most useful for fold recognition. A six-state secondary-structure alphabet and a fourteen-state secondary-structure alphabet that includes classifications of beta-strand orientation are most useful for improving alignments. The best fold recognition alphabet contributes a 40% improvement to HMM performance and the best alignment alphabet contributes a 62% improvement.
This textbook brings together the collective expertise of distinguished laboratorians in a comprehensive volume on hematology for students of laboratory sciences. Its organization is particularly well suited for promoting active learning. Each chapter begins with a list of learning objectives, a case study to be considered, and concludes with a list of important teaching points and several review questions. Answers to review questions and a brief case discussion are featured in the Appendix. In addition to covering topics traditionally included under the broad umbrella of hematology (including hematopoiesis, disorders of red blood cells, white blood cells, and platelets, as well as sections on hemostasis and thrombosis), the textbook considers important, but often neglected, topics such as laboratory safety, practical guides to specimen collection, and quality control, as well as quality assurance. Also noteworthy is a discussion on laboratory management and related issues, to which technologists often have limited exposure. This section offers an introduction to important concepts such as laboratory organization, staffing and scheduling, developing human resources, financial planning, purchasing and inventory control, and analysis of service operations. Other practical features of this textbook include overviews of routine laboratory evaluation of blood cells, hematology and coagulation instrumentation, and body-fluid analysis. A section on special studies is particularly notable for its discussion of molecular diagnostic techniques in the clinical laboratory. Another practical feature of this textbook is the easily accessible guide to reference intervals on the inside front and back covers. Throughout the text, material is presented in a clear and simple manner, targeted at a level of complexity appropriate for the beginning student of hematology. Illustrations and diagrams are particularly well done and clarify as well as complement textual information. The chapters on hemoglobinopathies and thalassemias are sufficiently detailed and comprehensive to serve as a guide for the identification and classification of abnormal hemoglobins in specialty laboratories. The section on erythrocyte disorders is notable for a subsection entitled laboratory diagnosis, which serves to integrate clinical laboratory evaluation with relevant pathophysiology. Sections on hemostasis and thrombosis provide a broad overview of complex processes and pathways, simplified by the use of figures and tables. An introduction to therapeutic strategies is provided, but often the mechanisms of action are not sufficiently detailed for the beginning student. Moreover, some sections on the evaluation of platelet function are not as up to date as corresponding sections on qualitative platelet disorders. However, advanced methods for identification and diagnosis are extremely thorough and well described. Chapters on leukocytes cover nonmalignant alterations, as well as the major neoplasms, in a cogent manner. Accompanying photomicrographs are of very high quality. Particular attention is given to chromosomal abnormalities associated with different hematologic malignancies. Overall, Hematology: Clinical Principles and Applications succeeds in providing a comprehensive yet userfriendly approach to laboratory hematology. The integration of basic and practical information will foster the development of effective laboratory consultants.
Kestrel is a programmable linear array processordesigned for sequence analysis. Among other features, Kestrelincludes an 8-bit word, a single-cycle add-and-minimizeinstruction, a multiplier and efficient communication usingshared registers. This paper describes Kestrel‘s functionalunits in detail, and examines each of their effects on systemperformance. With functional prototype chips completed, we willassemble a full single-board Kestrel array, with 512 processingelements on eight chips, in early 1998.
This paper presents a semi-systolic architecture for decoding cyclic linear error-correcting codes at high speed. The architecture implements a variant of Tanner's Algorithm B, modified for simpler and faster implementation. The main features of the architecture are low computational complexity, a simple, regular arrangement of cells for easy layout, short critical paths, and a high clock rate.A prototype chip has been designed to decode a 73-bit perfect difference set code. This 4600-mu-m x 6800-mu-m chip should achieve 25MHz decoding in 2-mu-m n-well cMOS.The success of the implementation illustrates the value of using technology dependent constraints and cost measures to guide the design of algorithms and architectures.