Pan-genome ortholog clustering tool (PanOCT) is a tool for pan-genomic analysis of closely related prokaryotic species or strains. PanOCT uses conserved gene neighborhood information to separate recently diverged paralogs into orthologous clusters where homology-only clustering methods cannot. The results from PanOCT and three commonly used graph-based ortholog-finding programs were compared using a set of four publicly available strains of the same bacterial species. All four methods agreed on ∼70% of the clusters and ∼86% of the proteins. The clusters that did not agree were inspected for evidence of correctness resulting in 85 high-confidence manually curated clusters that were used to compare all four methods.
Generation of syntactically correct and unambiguous names for proteins is a challenging, yet vital task for functional annotation processes. Proteins are often named based on homology to known proteins, many of which have problematic names. To address the need to generate high-quality protein names, and capture our significant experience correcting protein names manually, we have developed the Protein Naming Utility (PNU, http://www.jcvi.org/pn-utility). The PNU is a web-based database for storing and applying naming rules to identify and correct syntactically incorrect protein names, or to replace synonyms with their preferred name. The PNU allows users to generate and manage collections of naming rules, optionally building upon the growing body of rules generated at the J. Craig Venter Institute (JCVI). Since communities often enforce disparate conventions for naming proteins, the PNU supports grouping rules into user-managed collections. Users can check their protein names against a selected PNU rule collection, generating both statistics and corrected names. The PNU can also be used to correct GenBank table files prior to submission to GenBank. Currently, the database features 3080 manual rules that have been entered by JCVI Bioinformatics Analysts as well as 7458 automatically imported names.
Viruses are the most abundant biological entities on our planet. Interactions between viruses and their hosts impact several important biological processes in the world's oceans such as horizontal gene transfer, microbial diversity and biogeochemical cycling. Interrogation of microbial metagenomic sequence data collected as part of the Sorcerer II Global Ocean Expedition (GOS) revealed a high abundance of viral sequences, representing approximately 3% of the total predicted proteins. Cluster analyses of the viral sequences revealed hundreds to thousands of viral genes encoding various metabolic and cellular functions. Quantitative analyses of viral genes of host origin performed on the viral fraction of aquatic samples confirmed the viral nature of these sequences and suggested that significant portions of aquatic viral communities behave as reservoirs of such genetic material. Distributional and phylogenetic analyses of these host-derived viral sequences also suggested that viral acquisition of environmentally relevant genes of host origin is a more abundant and widespread phenomenon than previously appreciated. The predominant viral sequences identified within microbial fractions originated from tailed bacteriophages and exhibited varying global distributions according to viral family. Recruitment of GOS viral sequence fragments against 27 complete aquatic viral genomes revealed that only one reference bacteriophage genome was highly abundant and was closely related, but not identical, to the cyanomyovirus P-SSM4. The co-distribution across all sampling sites of P-SSM4-like sequences with the dominant ecotype of its host, Prochlorococcus supports the classification of the viral sequences as P-SSM4-like and suggests that this virus may influence the abundance, distribution and diversity of one of the most dominant components of picophytoplankton in oligotrophic oceans. In summary, the abundance and broad geographical distribution of viral sequences within microbial fractions, the prevalence of genes among viral sequences that encode microbial physiological function and their distinct phylogenetic distribution lend strong support to the notion that viral-mediated gene acquisition is a common and ongoing mechanism for generating microbial diversity in the marine environment.
MOTIVATION:Many genomes are sequenced by a collaboration of several centers, and then each center produces an assembly using their own assembly software. The collaborators then pick the draft assembly that they judge to be the best and the information contained in the other assemblies is usually not used.METHODS:We have developed a technique that we call assembly reconciliation that can merge draft genome assemblies. It takes one draft assembly, detects apparent errors, and, when possible, patches the problem areas using pieces from alternative draft assemblies. It also closes gaps in places where one of the alternative assemblies has spanned the gap correctly.RESULTS:Using the Assembly Reconciliation technique, we produced reconciled assemblies of six Drosophila species in collaboration with Agencourt Bioscience and The J. Craig Venter Institute. These assemblies are now the official (CAF1) assemblies used for analysis. We also produced a reconciled assembly of Rhesus Macaque genome, and this assembly is available from our website http://www.genome.umd.edu.AVAILABILITY:The reconciliation software is available for download from http://www.genome.umd.edu/software.htm
Marvin E. Frazier,Douglas B. Rusch, Aaron L. Halpern, Karla B. Heidelberg, Granger Sutton, Shannon Williamson, Shibu Yooseph, Dongying Wu, Jonathan A. Eisen, Jeff Hoffman, Charles H. Howard, Cyrus Foote, Brooke A. Dill, Karin Remington, Karen Beeson, Bao Tran, Hamilton Smith, Holly Baden-Tillson, Clare Stewart, Joyce Thorpe, Jason Freemen, Cindy Pfannkoch, Joseph E. Venter, John Heidelberg, Terry Utterback, Yu-Hui Rogers, Shaojie Zhang, Vineet Bafna, Luisa Falcon, Valeria Souza,German Bonilla, Luis E. Eguiarte , David M. Karl, Ken Nealson, Shubha Sathyendranath, Trevor Platt, Eldredge Bermingham, Victor Gallardo, Giselle Tamayo, Robert Friedman, Robert Strausberg, J. Craig Venter 1 J. Craig Venter Institute, Rockville, Maryland, United States Of America 2 The Institute For Genomic Research, Rockville, Maryland, United States Of America 3 Department of Computer Science, University of California San Diego 4 Instituto de Ecologia Dept. Ecologia Evolutiva, National Autonomous University of Mexico Mexico City, 04510 Distrito Federal, Mexico 5 University of Hawaii, Honolulu, United States of America 6 Dept. of Earth Sciences, University of Southern California, Los Angeles, California, United States of America 7 Dalhousie University, Halifax, Nova Scotia, Canada 8 Smithsonian Tropical Research Institute, Balboa, Ancon, Republic of Panama 9 University of Concepcion, Concepcion, Chile 10 University of Costa Rica, San Pedro, San Jose, Republic of Costa Rica
Accurate annotated assemblies of the mouse and human genomes enable a detailed comparison of the organization and evolution of the two genomes. We have completed several assemblies of both the mouse, with and without public data, and human genomes. Analysis of these assemblies suggests the mouse genome is about 10% smaller than the human genome primarily because of a difference in the content of repetitive DNA between the two genomes. More than 300,000 positions in these two genomes can be aligned with one another based on short segments of sequence similarity. These conserved segments significantly enhance the resolution of the resultant comparative maps and can be used to divide the genomes into regions of conserved-shared synteny. The genes found in such regions are highly conserved as is their relative order and orientation.