Abstract COSMIC (Catalogue Of Somatic Mutations in Cancer) has evolved from an initial catalogue to the world's most comprehensive knowledgebase of somatic variants in cancer, built upon a foundation of continuous, expert curation. COSMIC currently aggregates over 29 million unique somatic variants carefully curated from 1.5 million samples, establishing an essential resource for studying the cancer genome. This rich dataset is the result of extensive, dedicated work, drawing on insights from more than 30,000 scientific publications and major studies. COSMIC is structured into a suite of specialized modules that collectively transform raw genomic variants into biologically and clinically meaningful insight. These include the Cancer Gene Census (CGC), which systematically classifies causal cancer genes; the Cancer Mutation Census (CMC), which distinguishes driver from passenger mutations through computational and evidence-based annotation; Mutational Signatures, which captures genome-wide mutagenic processes; COSMIC 3D, which contextualizes variants within protein structures; and the Actionability and Resistance resources, which map genomic alterations to therapeutic response and resistance mechanisms. Together, these modules provide a framework essential for interpreting somatic variant landscapes in precision oncology. We highlight ongoing research on the next iteration of the Cancer Mutation Census (CMC v2), designed to enhance COSMIC’s ability to extract biologically meaningful signals from large-scale somatic datasets. CMC v2 applies refined background models to identify mutation hotspots at the amino acid level across cancer genes in a pan-cancer context, focusing on positions exhibiting statistically significant enrichment of somatic variants. This approach isolates non-random, spatially coherent clusters of mutations that represent strong candidates for driver activity. Although currently under development, these analyses demonstrate the potential of CMC v2 to provide higher-resolution insights into cancer gene dysregulation and support more nuanced interpretation of tumor evolution. The development of CMC v2 marks a key step in COSMIC’s evolution toward more data-driven, biologically grounded interpretation of cancer variants. By integrating statistical modeling with expert insight, CMC v2 refines our capacity to distinguish meaningful mutational patterns from background noise across diverse tumor contexts. These advances exemplify COSMIC’s ongoing commitment to translating large-scale genomics into actionable biological knowledge. As the knowledgebase continues to expand and engage with the global cancer research community, COSMIC remains an indispensable foundation for understanding cancer gene function, refining biomarker discovery, and supporting precision oncology. Citation Format: Madhumita Madhumita, Madiha Ahmed, Joanna Argasinska, David Armstrong, Nidhi Bindal Dhir, Denise Carvalho-Silva, Lucie Chadelle, Patrick Dao, Stephen Duke, Giovanna Fasanella, Muhammad Fouzan, Abishekraj Gnanasambandam, Avirup Neogi, Susan Haller, Bhavana Harsha, Balazs Hetenyi, Leonie Hodges, Steven Jupe, Rachel Lyne, Thomas Maurel, Karen McLaren, Thomas Mutimer, Sumodh Nair, Hanna Najgebauer, Helder Pedro, Sophie Poole, Amaia Sangrador-Vegas, Zoe Sheard, Manpreet Singh Chawla, Michael Starkey, Rebecca Steele, Sari Ward, Ellen Wiedemann, Jennifer Wilding, Siew Yit Yong, Jon Teague. COSMIC: Advancing the cancer genomics knowledgebase of somatic mutations [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 56.
The Catalogue of Somatic Mutations in Cancer (COSMIC) is a vital resource for cancer genomics, offering extensive data on somatic mutations, cell lines, and mutation signatures. While the existing COSMIC dataset provides wealth of diverse, high-quality information, accessing and fully utilising it requires significant processing and expertise in data analysis. To address this, we are developing a new suite of tools to enhance COSMIC data integration, usability and exploration.The Cell Line Explorer serves as a starting point for this suite, enabling users to easily access and analyse COSMIC data related to specific cell lines. This tool integrates genetic variants with annotations and reference datasets, providing dynamic insights into clinical significance. Beyond the cell lines mutation catalogue, the platform connects to key COSMIC resources, including the Cancer Gene Census, Actionability, Cancer Mutation Census, Mutational Signatures, and COSMIC3D. This integration allows users to derive deeper insights, and easily correlate mutation data with functional impact and potential clinical significance.The tools are designed for flexibility and scalability, meeting the needs of diverse users, from computational researchers, clinicians seeking actionable insights, educators incorporating bioinformatics into their teaching, and individuals with limited bioinformatics expertise who require intuitive solutions. With its intuitive interface and robust integration of COSMIC resources, this platform lowers the barriers to access and streamline complex analyses, providing researchers with powerful tools to explore and analyse cancer genomic data at unprecedented depth. Helder Pedro, Zbyslaw Sondka. Transforming cancer genomics research: A platform for integrated exploration of COSMIC resources [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 1080.
COSMIC, the Catalogue of Somatic Mutations in Cancer ( http://cancer.sanger.ac.uk ), is the world’s largest source of expert manually curated somatic mutation information relating to human cancers. Data is curated from many sources including whole genome studies, large next generation sequencing panels and case reports from many different journals. The first step in data curation is the identification of relevant papers from the literature. In COSMIC we focus searches on known cancer genes together with more focussed searches for tumour types. However, with hundreds of papers published in the field of cancer research every week, identifying those that contain somatic mutation data at the patient or sample level is a formidable task. To this end we have tested and applied Artificial Intelligence (AI) models to analyse papers returned from PubMed searches and select those that contain curatable data. A number of AI models were tested and trained using a large corpus of historical data spanning 20 years of papers tagged in COSMIC as either ‘rejected’ or ‘curated’. Microsoft's PubMedBERT was found to give the best predictions and has been incorporated into a pipeline to run weekly PubMed searches and return a report of papers categorised as either ‘curatable’ or ‘to be rejected’. Subsequent testing has indicated that the predictions are over 85% accurate. The application of AI to a repetitive and time consuming curator task has produced a pipeline that saves many hours of work. The increased capacity for paper triage enables a much broader view of cancer research to be scanned, potentially allowing new breakthroughs and gaps in data to be identified. The development, testing and future improvements and applications of the work will be discussed.
Abstract In 2004, COSMIC was one of the first initiatives to integrate global data on somatic mutations in cancer. At the time it was explicit that the fragmentation of genetic datasets was a major obstacle to understand the processes driving cancer. A team of expert curators and bioinformaticians was tasked with identifying and cataloguing data related to somatic variants, as well as relevant demographic, clinical, and patient information from published studies and making these data easily accessible to the research community. Over the last two decades we have witnessed the incredible progress in cancer genomics enabling whole genome studies, resulting in the exponential growth of data generated by cancer research. During its first year, COSMIC integrated data from 1672 scientific papers, cataloguing 1755 unique mutations across 21 genes. At present, twenty years later it is common for a single publication to describe tens of thousands of genome-wide mutations and now COSMIC includes 24 million genomic variants collected from more than 1.5 million patient samples acquired from over 29,000 scientific publications. Today, an important part of COSMIC's mission is to help translate this information for improving cancer treatment and patient care. To achieve this, the main catalogue of somatic mutations is supported by 6 accompanying resources that focus on different aspects of molecular oncology. The Cancer Gene Census and Cancer Mutation Census describe the roles of genes and mutations in oncogenesis, based on literature curation and analysis of the core mutation catalogue. COSMIC 3-D visualises mutation frequency in the context of the protein structure, and Mutational Signatures catalogues the mutagenic processes behind the nature of mutations at the genome level. To help target molecular alterations in clinical practice, Actionability and the catalogue of mutations causing drug resistance inform the availability of therapeutic options for cancer and how their efficacy is influenced by the genetic profile of the cancer. The ongoing development of sequencing techniques and new diagnostic methods as well as computational approaches, including deep learning and use of generative AI, give us hope that the next 20 years will be equally revolutionary for data-driven oncology. However, the abundance of new sources of diverse datasets makes data fragmentation an evolving challenge for the whole research community. Developing and adopting common standards for data formats, management and usage will be critical to assure inclusive, efficient, and effective translation of genomic research into a clinical practice. Citation Format: Zbyslaw Sondka, Madiha Ahmed, Joanna Argasinska, David Beare, Nidhi Bindal Dhir, Denise Carvalho-Silva, Manpreet Singh Chawla, Stephen Duke, Ilaria Fasanella, Muhammed Fouzan, Avirup Guha Neogi, Susan Haller, Bhavana Harsha, Balazs Hetenyi, Leonie Hodges, Alex Holmes, Steven Jupe, Rachel Lyne, Madhumita Madhumita, Thomas Maurel, Karen McLaren, Sumodh Nair, Helder Pedro, Amaia Sangrador-Vegas, Helen Schuilenburg, Zoe Sheard, Michael Starkey, Rebecca Steele, Sari Ward, Jennifer Wilding, Siew Yit Yong, Jon Teague. COSMIC: two decades of curating somatic variants in cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 3555.
The Catalogue Of Somatic Mutations In Cancer (COSMIC), https://cancer.sanger.ac.uk/cosmic, is an expert-curated knowledgebase providing data on somatic variants in cancer, supported by a comprehensive suite of tools for interpreting genomic data, discerning the impact of somatic alterations on disease, and facilitating translational research. The catalogue is accessed and used by thousands of cancer researchers and clinicians daily, allowing them to quickly access information from an immense pool of data curated from over 29 thousand scientific publications and large studies. Within the last 4 years, COSMIC has substantially expanded its utility by adding new resources: the Mutational Signatures catalogue, the Cancer Mutation Census, and Actionability. To improve data accessibility and interoperability, somatic variants have received stable genomic identifiers that are associated with their genomic coordinates in GRCh37 and GRCh38, and new export files with reduced data redundancy have been made available for download.
Somatic mutations accumulate in cells throughout their life. Most of them do not bring any negative effect. However, certain mutations change protein behaviour, structure, or level of expression. More importantly, some mutations are known to initiate and drive oncogenic transformation. These mutations often make good therapeutic targets but recognising this small subset in a cancer sample is a major challenge. The average cancer cell carries a life-long baggage of somatic mutations, and the mutational process is sped up in these cells through genomic instability (one of the hallmarks of cancer). As a result, there are hundreds of thousands of variants of unknown significance identified through sequencing of cancer DNA. COSMIC Cancer Mutation Census (CMC) answers this challenge by identifying coding mutations with a potential to drive cancer. This is achieved by combining manually curated information regarding cancer genes and genetic variants with data on variant frequencies in cancer and non-cancer populations, and algorithmic evaluation of variant significance. It applies a simple and transparent set of rules to the whole set of coding mutations in COSMIC to identify variants with the highest potential of clinical relevance. In current version (v95, November 2021) the CMC describes 4.7 million somatic variants and segregates them into four tiers. Tier 1 is the highest confidence set. This set includes 1558 mutations that are found in Cancer Gene Census genes and are also described as pathogenic in cancer by ClinVar. Tiers 2 and 3 contain variants with less extensive evidence of involvement in carcinogenesis. The dN/dS algorithm is used to include variants that are under positive selection in cancer cells. Finally, mutations without evidence for driving cancer are classified as Tier 4. In addition to this classification, CMC integrates and presents the information used to prioritise variants, including their frequencies in various cancer types (COSMIC), germline frequencies (gnomAD), ClinVar annotations, dN/dS analysis results, and nucleotide and amino acid conservation. Data can be accessed and scrutinised through a dedicated website at https://cancer.sanger.ac.uk/cmc. Citation Format: Zbyslaw Sondka, Bhavana Harsha, Helder Pedro, Nidhi Bindal Dhir, Charlie Hathaway, Sumodh Nair, Doron Sondheimer, Simon A. Forbes. COSMIC cancer mutation census: Classifying somatic coding variants by their potential to drive cancer [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 1200.
Ensembl Genomes (https://www.ensemblgenomes.org) provides access to non-vertebrate genomes and analysis complementing vertebrate resources developed by the Ensembl project (https://www.ensembl.org). The two resources collectively present genome annotation through a consistent set of interfaces spanning the tree of life presenting genome sequence, annotation, variation, transcriptomic data and comparative analysis. Here, we present our largest increase in plant, metazoan and fungal genomes since the project's inception creating one of the world's most comprehensive genomic resources and describe our efforts to reduce genome redundancy in our Bacteria portal. We detail our new efforts in gene annotation, our emerging support for pangenome analysis, our efforts to accelerate data dissemination through the Ensembl Rapid Release resource and our new AlphaFold visualization. Finally, we present details of our future plans including updates on our integration with Ensembl, and how we plan to improve our support for the microbial research community. Software and data are made available without restriction via our website, online tools platform and programmatic interfaces (available under an Apache 2.0 license). Data updates are synchronised with Ensembl's release cycle.
Abstract Since 2005, the Pathogen–Host Interactions Database (PHI-base) has manually curated experimentally verified pathogenicity, virulence and effector genes from fungal, bacterial and protist pathogens, which infect animal, plant, fish, insect and/or fungal hosts. PHI-base (www.phi-base.org) is devoted to the identification and presentation of phenotype information on pathogenicity and effector genes and their host interactions. Specific gene alterations that did not alter the in host interaction phenotype are also presented. PHI-base is invaluable for comparative analyses and for the discovery of candidate targets in medically and agronomically important species for intervention. Version 4.12 (September 2021) contains 4387 references, and provides information on 8411 genes from 279 pathogens, tested on 228 hosts in 18, 190 interactions. This provides a 24% increase in gene content since Version 4.8 (September 2019). Bacterial and fungal pathogens represent the majority of the interaction data, with a 54:46 split of entries, whilst protists, protozoa, nematodes and insects represent 3.6% of entries. Host species consist of approximately 54% plants and 46% others of medical, veterinary and/or environmental importance. PHI-base data is disseminated to UniProtKB, FungiDB and Ensembl Genomes. PHI-base will migrate to a new gene-centric version (version 5.0) in early 2022. This major development is briefly described.
Abstract The pathogen–host interactions database (PHI-base) is available at www.phi-base.org. PHI-base contains expertly curated molecular and biological information on genes proven to affect the outcome of pathogen–host interactions reported in peer reviewed research articles. PHI-base also curates literature describing specific gene alterations that did not affect the disease interaction phenotype, in order to provide complete datasets for comparative purposes. Viruses are not included, due to their extensive coverage in other databases. In this article, we describe the increased data content of PHI-base, plus new database features and further integration with complementary databases. The release of PHI-base version 4.8 (September 2019) contains 3454 manually curated references, and provides information on 6780 genes from 268 pathogens, tested on 210 hosts in 13,801 interactions. Prokaryotic and eukaryotic pathogens are represented in almost equal numbers. Host species consist of approximately 60% plants (split 50:50 between cereal and non-cereal plants), and 40% other species of medical and/or environmental importance. The information available on pathogen effectors has risen by more than a third, and the entries for pathogens that infect crop species of global importance has dramatically increased in this release. We also briefly describe the future direction of the PHI-base project, and some existing problems with the PHI-base curation process.
Ensembl Genomes (http://www.ensemblgenomes.org) is an integrating resource for genome-scale data from non-vertebrate species, complementing the resources for vertebrate genomics developed in the context of the Ensembl project (http://www.ensembl.org). Together, the two resources provide a consistent set of interfaces to genomic data across the tree of life, including reference genome sequence, gene models, transcriptional data, genetic variation and comparative analysis. Data may be accessed via our website, online tools platform and programmatic interfaces, with updates made four times per year (in synchrony with Ensembl). Here, we provide an overview of Ensembl Genomes, with a focus on recent developments. These include the continued growth, more robust and reproducible sets of orthologues and paralogues, and enriched views of gene expression and gene function in plants. Finally, we report on our continued deeper integration with the Ensembl project, which forms a key part of our future strategy for dealing with the increasing quantity of available genome-scale data across the tree of life.
Plant pathogens continue to threaten food security and have a significant economic impact. For this, an understanding of gene function and pathogen-host interactions is critical; underpinned by accurate and comprehensive annotation of genomic sequence. The pathogenic species within Ensembl (bacterial, fungal and protists) have continued to grow, accumulating genomic, transcriptomic, variation, pathogen-host interaction and comparative data. However, it was obvious that some key plant pathogens had incomplete and inconsistent gene sets. Through PhytoPath, a BBSRC-funded project in collaboration with Rothamsted Research, we were able to initiate successful community annotation projects to rectify this problem. In the first instance, we facilitated the collaboration of over 40 members of the Botrytis cinerea community spread across 8 countries. Using the gene editing tool Apollo, web-based training and infrastructure support from Ensembl Fungi, this group systematically reviewed theentire gene set. This new gene set was checked and disseminated via Ensembl Fungi. Following this, a similar project was completed for another important pathogen,Blumeria graminis, also leading to a refined gene set and publication. Currently, we are engaging with the Zymoseptoria tritici (a growing threat to wheat) community to enable them to achieve the same. While the tangible outcome of these projects appears to be a better gene set, it is apparent that the inherent agreement and dialogue among community members as they undergo this curation process is also consequential to the acceleration of the field. With increasing volumes of easily-accessible data, the role of unifying both the divergent data sets and the associated research teams is paramount. We anticipate that by continuing to be a hub for community-driven efforts, the microbial portals within Ensembl can pave the way for a new kind of collaboration among geographically- dispersed pathogen research communities.
Accurate and comprehensive annotation of genomic sequences underpins advances in managing plant disease. However, important plant pathogens still have incomplete and inconsistent gene sets and lack dedicated funding or teams to improve this annotation. This paper describes a collaborative approach to gene curation to address this shortcoming. In the first instance, over 40 members of the Botrytis cinerea community from eight countries, with training and infrastructural support from Ensembl Fungi, used the gene editing tool Apollo to systematically review the entire gene set (11,707 protein coding genes) in 6-7 months. This has subsequently been checked and disseminated. Following this, a similar project for another pathogen, Blumeria graminis f. sp. hordei, also led to a completely redefined gene set. Currently, we are working with the Zymoseptoria tritici community to enable them to achieve the same. While the tangible outcome of these projects is improved gene sets, it is apparent that the inherent agreement and ownership of a single gene set by research teams as they undergo this curation process are consequential to the acceleration of research in the field. With the generation of large data sets increasingly affordable, there is value in unifying both the divergent data sets and their associated research teams, pooling time, expertise, and resources. Community-driven annotation efforts can pave the way for a new kind of collaboration among pathogen research communities to generate well-annotated reference data sets, beneficial not just for the genome being examined but for related species and the refinement of automatic gene prediction tools.
Ensembl Genomes (http://www.ensemblgenomes.org) is an integrating resource for genome-scale data from non-vertebrate species, complementing the resources for vertebrate genomics developed in the Ensembl project (http://www.ensembl.org). Together, the two resources provide a consistent set of programmatic and interactive interfaces to a rich range of data including genome sequence, gene models, transcript sequence, genetic variation, and comparative analysis. This paper provides an update to the previous publications about the resource, with a focus on recent developments and expansions. These include the incorporation of almost 20 000 additional genome sequences and over 35 000 tracks of RNA-Seq data, which have been aligned to genomic sequence and made available for visualization. Other advances since 2015 include the release of the database in Resource Description Framework (RDF) format, a large increase in community-derived curation, a new high-performance protein sequence search, additional cross-references, improved annotation of non-protein-coding genes, and the launch of pre-release and archival sites. Collectively, these changes are part of a continuing response to the increasing quantity of publicly-available genome-scale data, and the consequent need to archive, integrate, annotate and disseminate these using automated, scalable methods.
Abstract: The pathogen-host interactions database PHI-base (www.phi-base.org) is a knowledge database. It contains expertly curated molecular and biological information on genes proven to affect the outcome of pathogen-host interactions reported in peer reviewed research articles. Genes not affecting the disease interaction phenotype are also curated. Viruses are not included. Here we describe a revised PHI-base Version 4 data platform with improved search, filtering and extended data display functions. Also a BLAST search function is now provided. The database links to PHI-Canto, a new multi-species author self-curation tool adapted from PomBase-Canto. The recent release of PHI-base version 4 has an increased data content containing information from >2000 manually curated references. The data provide information on 4460 genes from 264 pathogens tested on 176 hosts in 8046 interactions. Pro- and eukaryotic pathogens are represented in almost equal numbers. Host species belong ~70% to plants and 30% to other species of medical and/or environmental importance. Additional data types included into PHI-base 4 are the direct targets of pathogen effector proteins in experimental and natural host organisms. The different use types and the future directions of PHI-base as a community database are discussed. Urban et al. , (2017) PHI-base: a new interface and further additions for the multi-species pathogen-host interactions database. Nucleic Acids Research (04 January 2017) 45 (D1): D604-D610 doi: 10.1093/nar/gkw1089 This work is supported by the UK Biotechnology and Biological Sciences Research Council (BBSRC) (BB/I/001077/1, BB/K020056/1). PHI-base receives additional support from the BBSRC as a National Capability (BB/J/004383/1).
The pathogen–host interactions database (PHI-base) is available at www.phi-base.org. PHI-base contains expertly curated molecular and biological information on genes proven to affect the outcome of pathogen–host interactions reported in peer reviewed research articles. In addition, literature that indicates specific gene alterations that did not affect the disease interaction phenotype are curated to provide complete datasets for comparative purposes. Viruses are not included. Here we describe a revised PHI-base Version 4 data platform with improved search, filtering and extended data display functions. A PHIB-BLAST search function is provided and a link to PHI-Canto, a tool for authors to directly curate their own published data into PHI-base. The new release of PHI-base Version 4.2 (October 2016) has an increased data content containing information from 2219 manually curated references. The data provide information on 4460 genes from 264 pathogens tested on 176 hosts in 8046 interactions. Prokaryotic and eukaryotic pathogens are represented in almost equal numbers. Host species belong ∼70% to plants and 30% to other species of medical and/or environmental importance. Additional data types included into PHI-base 4 are the direct targets of pathogen effector proteins in experimental and natural host organisms. The curation problems encountered and the future directions of the PHI-base project are briefly discussed.
Fusarium culmorum is a soilborne fungal plant pathogen that causes foot and root rot and Fusarium head blight on small-grain cereals, in particular on wheat and barley. We report herein the draft genome sequence of a 1998 field strain called FcUK99 adapted to the temperate climate found in England.
Ensembl Genomes (http://www.ensemblgenomes.org) is an integrating resource for genome-scale data from non-vertebrate species, complementing the resources for vertebrate genomics developed in the context of the Ensembl project (http://www.ensembl.org). Together, the two resources provide a consistent set of programmatic and interactive interfaces to a rich range of data including reference sequence, gene models, transcriptional data, genetic variation and comparative analysis. This paper provides an update to the previous publications about the resource, with a focus on recent developments. These include the development of new analyses and views to represent polyploid genomes (of which bread wheat is the primary exemplar); and the continued up-scaling of the resource, which now includes over 23 000 bacterial genomes, 400 fungal genomes and 100 protist genomes, in addition to 55 genomes from invertebrate metazoa and 39 genomes from plants. This dramatic increase in the number of included genomes is one part of a broader effort to automate the integration of archival data (genome sequence, but also associated RNA sequence data and variant calls) within the context of reference genomes and make it available through the Ensembl user interfaces.
PhytoPath (www.phytopathdb.org) is a resource for genomic and phenotypic data from plant pathogen species, that integrates phenotypic data for genes from PHI-base, an expertly curated catalog of genes with experimentally verified pathogenicity, with the Ensembl tools for data visualization and analysis. The resource is focused on fungi, protists (oomycetes) and bacterial plant pathogens that have genomes that have been sequenced and annotated. Genes with associated PHI-base data can be easily identified across all plant pathogen species using a BioMart-based query tool and visualized in their genomic context on the Ensembl genome browser. The PhytoPath resource contains data for 135 genomic sequences from 87 plant pathogen species, and 1364 genes curated for their role in pathogenicity and as targets for chemical intervention. Support for community annotation of gene models is provided using the WebApollo online gene editor, and we are working with interested communities to improve reference annotation for selected species.
Ensembl Genomes (http://www.ensemblgenomes.org) is an integrating resource for genome-scale data from non-vertebrate species. The project exploits and extends technologies for genome annotation, analysis and dissemination, developed in the context of the vertebrate-focused Ensembl project, and provides a complementary set of resources for non-vertebrate species through a consistent set of programmatic and interactive interfaces. These provide access to data including reference sequence, gene models, transcriptional data, polymorphisms and comparative analysis. This article provides an update to the previous publications about the resource, with a focus on recent developments. These include the addition of important new genomes (and related data sets) including crop plants, vectors of human disease and eukaryotic pathogens. In addition, the resource has scaled up its representation of bacterial genomes, and now includes the genomes of over 9000 bacteria. Specific extensions to the web and programmatic interfaces have been developed to support users in navigating these large data sets. Looking forward, analytic tools to allow targeted selection of data for visualization and download are likely to become increasingly important in future as the number of available genomes increases within all domains of life, and some of the challenges faced in representing bacterial data are likely to become commonplace for eukaryotes in future.