Nature 496, 498–503 (2013); doi:10.1038/nature12111 In this Letter, five authors were inadvertently omitted: Sharmin Begum and Christine Lloyd from the Wellcome Trust Sanger Institute, and Christa Lanz, Günter Raddatz and Stephan C. Schuster from the Max Planck Institute for Developmental Biology. David Elliot was incorrectly listed as David Eliot, Beverley Mortimore was incorrectly listed as Beverly Mortimer, and James D.
The mission of the Universal Protein Resource (UniProt) (http://www.uniprot.org) is to support biological research by providing a freely accessible, stable, comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase. It integrates, interprets and standardizes data from numerous resources to achieve the most comprehensive catalogue of protein sequences and functional annotation. UniProt comprises four major components, each optimized for different uses, the UniProt Archive, the UniProt Knowledgebase, the UniProt Reference Clusters and the UniProt Metagenomic and Environmental Sequence Database. UniProt is produced by the UniProt Consortium, which consists of groups from the European Bioinformatics Institute (EBI), the SIB Swiss Institute of Bioinformatics (SIB) and the Protein Information Resource (PIR). UniProt is updated and distributed every 4 weeks and can be accessed online for searches or downloads.
The GO annotation dataset provided by the UniProt Consortium (GOA: http://www.ebi.ac.uk/GOA) is a comprehensive set of evidenced-based associations between terms from the Gene Ontology resource and UniProtKB proteins. Currently supplying over 100 million annotations to 11 million proteins in more than 360,000 taxa, this resource has increased 2-fold over the last 2 years and has benefited from a wealth of checks to improve annotation correctness and consistency as well as now supplying a greater information content enabled by GO Consortium annotation format developments. Detailed, manual GO annotations obtained from the curation of peer-reviewed papers are directly contributed by all UniProt curators and supplemented with manual and electronic annotations from 36 model organism and domain-focused scientific resources. The inclusion of high-quality, automatic annotation predictions ensures the UniProt GO annotation dataset supplies functional information to a wide range of proteins, including those from poorly characterized, non-model organism species. UniProt GO annotations are freely available in a range of formats accessible by both file downloads and web-based views. In addition, the introduction of a new, normalized file format in 2010 has made for easier handling of the complete UniProt-GOA data set.
The primary mission of Universal Protein Resource (UniProt) is to support biological research by maintaining a stable, comprehensive, fully classified, richly and accurately annotated protein sequence knowledgebase, with extensive cross-references and querying interfaces freely accessible to the scientific community. UniProt is produced by the UniProt Consortium which consists of groups from the European Bioinformatics Institute (EBI), the Swiss Institute of Bioinformatics (SIB) and the Protein Information Resource (PIR). UniProt is comprised of four major components, each optimized for different uses: the UniProt Archive, the UniProt Knowledgebase, the UniProt Reference Clusters and the UniProt Metagenomic and Environmental Sequence Database. UniProt is updated and distributed every 4 weeks and can be accessed online for searches or download at http://www.uniprot.org.
Manual annotation (the "museum" model of annotation) relies on a small group of specialized curators to catalogue and classify genes according to their functional roles. This is both costly and time consuming and therefore is used only for model organisms with sufficient funding. Smaller research communities often have to rely on other models of annotation, mainly automated annotation (the "factory" model, e.g. Ensembl), and the "jamboree" model (in which a group of leading biologists from the community and bioinformaticians come together for a short intensive annotation workshop). At the Wellcome Trust Sanger Institute (WTSI), the Havana team provides high quality manual annotation of finished vertebrate genome sequences, namely human, mouse and zebrafish. We also perform the curation of specific finished regions such as the MHC in dog, cow and pig, whose whole genomes have been assembled from unfinished BACs or from whole genome shotgun sequences. In addition, we at Havana have also hosted annotation jamborees for the cow (Bos taurus) and pig (Sus scrofa) genomes. During those sessions, the research community had the opportunity to annotate their genes of interest under expert guidance using the custom written publicly available Otterlace annotation system, and the unified manual annotation guidelines. By making use of the tools and skills acquired during the cow and pig jamborees, the delegates can continue annotating their genomes remotely. For the pig genome, a highly contiguous physical map has been generated by an international effort of four laboratories (available in Pre!Ensembl) and is being used as a substrate for the swine genome sequencing project. Upcoming vertebrate genomes will be sequenced to a high depth coverage with the next generation sequencing technologies (e.g. Illumina, 454, SOLiD) but will have the drawback of not being manually finished. Manual annotation will be more accurate than the automated predictions at coping with any assembly problems derived from these high coverage but unfinished (or automatic pre-finished) genomes. Once these inherent assembly errors are corrected and the gene structures are accurately identified with manual annotation, the curated genes will be incorporated and merged with the predicted gene models in Ensembl to provide a unified view of the landscape of vertebrate genomes. I will present an introduction to our manual annotation system and our experience using it for annotation jamborees at the WTSI.
There are two main classes of natural killer (NK) cell receptors in mammals, the killer cell immunoglobulin-like receptors (KIR) and the structurally unrelated killer cell lectin-like receptors (KLR). While KIR represent the most diverse group of NK receptors in all primates studied to date, including humans, apes, and Old and New World monkeys, KLR represent the functional equivalent in rodents. Here, we report a first digression from this rule in lemurs, where the KLR (CD94/NKG2) rather than KIR constitute the most diverse group of NK cell receptors. We demonstrate that natural selection contributed to such diversification in lemurs and particularly targeted KLR residues interacting with the peptide presented by MHC class I ligands. We further show that lemurs lack a strict ortholog or functional equivalent of MHC-E, the ligands of non-polymorphic KLR in "higher" primates. Our data support the existence of a hitherto unknown system of polymorphic and diverse NK cell receptors in primates and of combinatorial diversity as a novel mechanism to increase NK cell receptor repertoire.
Background The domestic pig is being increasingly exploited as a system for modeling human disease. It also has substantial economic importance for meat-based protein production. Physical clone maps have underpinned large-scale genomic sequencing and enabled focused cloning efforts for many genomes. Comparative genetic maps indicate that there is more structural similarity between pig and human than, for example, mouse and human, and we have used this close relationship between human and pig as a way of facilitating map construction. Results Here we report the construction of the most highly continuous bacterial artificial chromosome (BAC) map of any mammalian genome, for the pig ( Sus scrofa domestica ) genome. The map provides a template for the generation and assembly of high-quality anchored sequence across the genome. The physical map integrates previous landmark maps with restriction fingerprints and BAC end sequences from over 260,000 BACs derived from 4 BAC libraries and takes advantage of alignments to the human genome to improve the continuity and local ordering of the clone contigs. We estimate that over 98% of the euchromatin of the 18 pig autosomes and the X chromosome along with localized coverage on Y is represented in 172 contigs, with chromosome 13 (218 Mb) represented by a single contig. The map is accessible through pre-Ensembl, where links to marker and sequence data can be found. Conclusion The map will enable immediate electronic positional cloning of genes, benefiting the pig research community and further facilitating use of the pig as an alternative animal model for human disease. The clone map and BAC end sequence data can also help to support the assembly of maps and genome sequences of other artiodactyls.
Human killer immunoglobulin-like receptors (KIR) are expressed on natural killer (NK) cells and are involved in their immunoreactivity. While KIR with a long cytoplasmic tail deliver an inhibitory signal when bound to their respective major histocompatibility complex class I ligands, KIR with a short cytoplasmic tail can activate NK responses. The expansion of the KIR gene family originally appeared to be a phenomenon restricted to primates (human, apes, and monkeys) in comparison to rodents, which via convergent evolution have numerous C-type lectin-like Ly49 molecules that function analogously. Further studies have shown that multiple KIR are also present in cow and horse. In this study, we have identified by comparative genomics the first and possibly only KIR gene, named KIR2DL1, in the domesticated pig (Sus scrofa) allowing further evolutionary comparisons to be made. It encodes a protein with two extracellular immunoglobulin domains (D0 + D2), and a long cytoplasmic tail containing two inhibitory motifs. We have mapped the pig KIR2DL1 gene to chromosome 6q. Flanked by LILRa, LILRb, and LILRc, members of the leukocyte immunoglobulin-like receptor (LILR) family, on the centromeric end, and FCAR, NCR1, NALP7, NALP2, and GP6 on the telomeric end, pig demonstrates conservation of synteny with the human leukocyte receptor complex (LRC). Both the porcine KIR and LILR genes have diverged sufficiently to no longer be clearly orthologous with known human LRC family members.
The human X chromosome has a unique biology that was shaped by its evolution as the sex chromosome shared by males and females. We have determined 99.3% of the euchromatic sequence of the X chromosome. Our analysis illustrates the autosomal origin of the mammalian sex chromosomes, the stepwise process that led to the progressive loss of recombination between X and Y, and the extent of subsequent degradation of the Y chromosome. LINE1 repeat elements cover one-third of the X chromosome, with a distribution that is consistent with their proposed role as way stations in the process of X-chromosome inactivation. We found 1,098 genes in the sequence, of which 99 encode proteins expressed in testis and in various tumour types. A disproportionately high number of mendelian diseases are documented for the X chromosome. Of this number, 168 have been explained by mutations in 113 X-linked genes, which in many cases were characterized with the aid of the DNA sequence.