Advancements in immunogenomics and immuno-oncology have enabled the development of personalized cancer vaccines (PCVs) that target cancer cell-specific somatic variants. A subset of these variants produce neoantigens that, when presented on tumor cells by MHC molecules, have the potential to elicit a robust and specific immune response. To date, there are over one hundred interventional studies listed on clinicaltrials.gov that explore the use of PCVs. We have supported a number of these trials through the creation of bioinformatic pipelines, tools, and procedures for the identification of patient-specific neoantigen candidates. While many of these steps have been automated, the final selection of neoantigen candidates often relies on expert manual review, creating a bottleneck that limits scalability and full automation of PCV workflows. Addressing this challenge, we introduce NEAT (Neoantigen Evaluation & Automated Triage), a machine learning-based approach that enables automated neoantigen candidate prioritization and supports the transition toward more scalable and reproducible PCV design. We implemented a prediction model trained and tested on existing vaccine design results from 33 patients and 1,943 peptides, across 3 clinical trials, including 439 peptides prioritized for PCV inclusion. This model uses features such as tumor variant allele frequency, RNA expression, driver gene status, binding/presentation scores, and transcript support level to automatically predict whether a peptide will be accepted, rejected, or require further human review before inclusion in a vaccine. The model achieved a sensitivity of 0.847 and specificity of 0.924, with an area under the curve of 0.955. The model predictions have been incorporated in pVACtools v7.0.0. By integrating this model into the vaccine development pipeline, we foresee a significant reduction in the time required to transition from patient sample collection to vaccine manufacturing, thereby enhancing the efficiency and scalability of PCV production.
BACKGROUND:Optimal consolidation therapy for patients with intermediate-risk acute myeloid leukemia (AML) in first complete remission (CR1) is controversial. Retrospective studies have suggested that the clearance of leukemia-associated mutations (LAMs) in CR1 may predict lower relapse risk and better outcomes with high-dose cytarabine (HiDAC) consolidation. We tested this hypothesis prospectively. METHODS:We performed a phase II, multicenter study of intermediate-risk, transplant-eligible, de novo AML in patients 18-60 years of age who achieved a complete remission (CR) or CR with incomplete count recovery (CRi) after induction therapy. Tumor and normal whole-exome sequencing was performed at presentation to identify somatic LAMs (median ∼30 LAMs/patient). In remission marrow samples, LAM variant allele frequencies (VAFs) were then remeasured using a VAF cutoff of less than 2.5% to define clearance. Patients who met this LAM clearance threshold received HiDAC consolidation, whereas those with persistent LAMs (VAF ≥2.5%) were recommended to undergo allogeneic hematopoietic cell transplantation. The primary endpoint compared relapse-free survival (RFS) of intermediate-risk patients with complete LAM clearance to historical cohorts with intermediate-risk AML who received HiDAC-based regimens in CR1. To account for an unplanned interim assessment, the significance threshold for the primary analysis was 0.01. RESULTS:Among 100 patients who were evaluated, intermediate-risk patients who cleared all LAMs in CR1 (n=33) had a median RFS of 33.1 months (95% confidence interval, 11.7-NA) compared to a median RFS of 11.7 months in the historical cohort (n=239; 95% confidence interval, 9.9-15.6, P=0.015). CONCLUSIONS:Among patients with intermediate-risk AML, clearance of LAMs after induction, followed by HiDAC consolidation in CR1, was associated with longer RFS compared with similarly treated historical controls. Although this result did not meet the prespecified threshold for statistical significance, the reported association sets the stage for a randomized trial to further evaluate this strategy. (ClinicalTrials.gov number, NCT02756962.).
Personalized neoantigen vaccines represent a promising immunotherapy approach that harnesses tumor-specific antigens to stimulate anti-tumor immune responses. However, the design of these vaccines requires sophisticated computational workflows to predict and prioritize neoantigen candidates from patient sequencing data, coupled with rigorous review to ensure candidate quality. While numerous computational tools exist for neoantigen prediction, to our knowledge, there are no established protocols detailing the complete process from raw sequencing data through systematic candidate selection. Here, we present ImmunoNX (Immunogenomics Neoantigen eXplorer), an end-to-end protocol for neoantigen prediction and vaccine design that has supported over 185 patients across 11 clinical trials. The workflow integrates tumor DNA/RNA and matched normal DNA sequencing data through a computational pipeline built with Workflow Definition Language (WDL) and executed via Cromwell on Google Cloud Platform. ImmunoNX employs consensus-based variant calling, in-silico HLA typing, and pVACtools for neoantigen prediction. Additionally, we describe a two-stage immunogenomics review process with prioritization of neoantigen candidates, enabled by pVACview, followed by manual assessment of variants using the Integrative Genomics Viewer (IGV). This workflow enables vaccine design in under three months. We demonstrate the protocol using the HCC1395 breast cancer cell line dataset, identifying 78 high-confidence neoantigen candidates from 322 initial predictions. Although demonstrated here for vaccine development, this workflow can be adapted for diverse neoantigen therapies and experiments. Therefore, this protocol provides the research community with a reproducible, version-controlled framework for designing personalized neoantigen vaccines, supported by detailed documentation, example datasets, and open-source code.
Introduction Ph-like B-ALL is a poor-risk B-ALL subtype defined by gene expression signatures that resemble the BCR::ABL1-associated transcriptional program in the absence of the BCR-ABL1 oncoprotein. Recent studies have demonstrated that up to 80% of Ph-like B-ALLs have recurrent gene fusions involving cytokine receptors or tyrosine kinases that can be identified via cytogenetics, fluorescence in situ hybridization (FISH) or RNA sequencing. However, the inability to rapidly and accurately identify these alterations in the clinical setting is a major impediment to developing therapeutic strategies that may have efficacy in this B-ALL subtype. Here we used 88 patient (pt) samples to validate a rapid, streamlined clinical WGS assay (ChromoSeq) for B-ALL and assessed the diagnostic yield for Ph-like-associated SVs in 136 pts enrolled in the Alliance A041501 clinical trial. Methods ChromoSeq is a clinical WGS assay for hematologic malignancies that uses rapid laboratory workflows and focused analysis methods to identify clinically relevant SVs, copy number alterations and gene-level mutations. We updated the ChromoSeq pipeline for use in B-ALL pts and validated it using 88 retrospective B-ALL samples (80 adult/young adult and 8 pediatric from Washington University School of Medicine and St. Louis Children's Hospital), 54 of which were analyzed via RNAseq for gene fusions and a 6-gene Ph-like B-ALL signature (CA6, SPATS2L, MUC4, JCHAIN, BMPR1B, ADGRF1, NRXN3 and CRLF2). The “ChromoSeqV2” assay was then used to analyze 136 pretreatment blood or bone marrow samples from the Alliance A041501 (trial NCT03150693). These results were compared to data from RNAseq obtained for research purposes (N=35), clinical cytogenetics (N=69) and Ph-like classification via low-density microarray (LDA) card analysis (all pts). Results ChromoSeqV2 was validated in a CLIA-licensed setting using cohort of 88 retrospective B-ALL samples, 45 of which had ICC and WHO class-defining SVs based on G-banded karyotyping, FISH or RNAseq. ChromoSeqV2 detected 100% of the non-Ph-like SVs (37/37), including BCR::ABL1 (N=20), KMT2Ar (N=9), ETV6::RUNX1 (N=3), TCF3::PBX1 (N=1), and rearrangements involving MYC (N=2) or MEF2D (N=2). Fifteen patients met criteria for Ph-like B-ALL via cytogenetics or RNAseq either due to rearrangements of CRLF2 (N=5), JAK2 (N=2) or ABL2 (N=1) or high expression of a 6-gene Ph-like signature (N=7). ChromoSeqV2 detected a Ph-like SV in 93% of these cases (14/15) with 100% specificity in samples negative by RNAseq (N=23). ChromoSeqV2 was then used to profile 136 samples from the Alliance A041501 trial, where it identified ICC and WHO class-defining SVs in 87 pts (64%; 95% CI: 55-72%), including Ph-like SVs (56/136, 41%), KMT2A rearrangements (N=5), TCF3::PBX1 (N=3) and SVs involving MYC (N=3), MEF2D (N=2), DUX4 (N=2), and ZNF384 (N=1). Ph-like SVs were dominated by CRLF2 rearrangements (N=45) and included SVs involving EPOR (N=4), JAK2 (N=3), ABL1/ABL2 (N=2) and PDGFRB (N=1). IGH was the most common CRLF2 partner (N=32), followed by P2RY8::CRLF2 (N=7) and other genes (N=5). Multiple CRLF2 SVs were detected in 4 pts, and 33% (15/45) of pts with a CRLF2 rearrangement had a co-occurring CRLF2 gene mutation. The yield of ChromoSeqV2 for SVs in pts with the Ph-like expression signature via LDA card analysis was 66% (54/82; 95% CI 55%-76%) compared to 24% (11/45; 95% CI 13%-40%) by cytogenetics when these studies were successful and 70% (12/17; 95% CI 44%-90%) vs RNAseq in pts with available data. The sensitivity of ChromoSeqV2 for Ph-like SVs identified by either cytogenetics, FISH or RNAseq was 100% (21/21; 95% CI 84%-100%) and the positive predictive value of Ph-like SVs for the Ph-like expression signature by the LDA card assay was 96% (54/56; 95% CI 88%-99%). Conclusions Genomic assessment of 136 B-ALL pts from the Alliance A041501 trial using the ChromoSeqV2 rapid clinical WGS assay was 100% sensitive for Ph-like SVs identified by other molecular methods and had a 96% PPV for the Ph-like expression signature via LDA card analysis. The diagnostic yield of SV pts with the Ph-like signature was 66% using ChromoSeqV2, which was similar to RNAseq and superior to cytogenetics. These results demonstrate that ChromoSeqV2 is an effective approach for identifying potentially targetable rearrangements in B-ALL pts suspected of having the Ph-like phenotype. Support: U10CA180821, U10CA180882; . Servier; Pfizer.
Advancements in immunogenomics and immuno-oncology have enabled the development of neoantigen vaccines, offering personalized cancer therapies by targeting cancer cell-specific somatic mutations. These mutations produce neoantigens that, when presented on tumor cells by MHC molecules, can elicit a robust and specific immune response. To date, there are 108 interventional studies listed on clinicaltrials.gov that explore the use of cancer vaccines. We have supported a number of these trials through the creation of bioinformatic pipelines, tools and procedures for the identification of patient-specific neoantigen candidates. Final prioritization of neoantigen candidates relies on manual review by an Immunogenomics Tumor Board (ITB) that meets weekly, increasing turnaround time and presenting a barrier to scaling.Addressing this challenge, we introduce a machine learning-based approach to automate the selection of neoantigens peptides. We implemented a random forest model to train and test on existing ITB results from 21 patients and 1,324 peptides, including 297 peptides prioritized for personalized vaccine inclusion. This model aims to use features such as mutation position, driver gene status, tumor variant allele frequency, RNA expression, and other features to automatically predict whether a peptide will be accepted, rejected, or require further review for the vaccine. The model achieved an 88.89% sensitivity and 86.4% specificity, with an area under the curve of 0.933. By integrating this model into the vaccine development pipeline, we foresee a significant reduction in the time required to transition from patient sample collection to vaccine manufacturing, thereby enhancing the efficiency and scalability of personalized cancer vaccine production.
Personalized cancer vaccines (PCVs) leverage immunogenomics strategies to combat cancer. Somatic mutations in tumor cells generate neoantigens that may get presented on the tumor cell's surface by MHC molecules. Immunotherapies target neoantigens to stimulate tumor-specific immune responses. Our bioinformatics workflow has designed vaccines for over 170 patients across 11 of the 180 neoantigen vaccine trials on clinicaltrials.gov.Despite the rise in PCV-related interventions, gaps in established protocols addressing the complexities associated with the design of PCVs still remain. Here, we summarize our bioinformatics pipeline and describe measures taken to ensure robust support for clinical trials at Washington University. Our Google Cloud immunotherapy pipeline (open MIT license) to predict neoantigen epitopes is implemented in Workflow Definition Language and containerized using Docker to ensure portability and reliability. The pVACtools software suite (pvactools.org) that carries out neoantigen identification and prioritization, is developed and updated following industry best practices including version control (Git), formal code review, automated unit and integration tests, and benchmark tests. The final steps of the bioinformatics workflow generate files recording the analysis parameters and QC results tailored to the FDA's requests. Candidates generated by the pipeline are reviewed at an Immunogenomics Tumor Board using the pVACview tool. Prioritized candidates undergo a rigorous examination of data QC metrics, variant support at genomic and transcriptomic levels, MHC binding prediction algorithms, and HLA allele concordance between the clinical data and in-silico prediction tools. Finally, a long-peptide order form generated by the pipeline is sent to the vaccine manufacturer for synthesis.
BACKGROUND:Genomic analysis is essential for risk stratification in patients with acute myeloid leukemia (AML) or myelodysplastic syndromes (MDS). Whole-genome sequencing is a potential replacement for conventional cytogenetic and sequencing approaches, but its accuracy, feasibility, and clinical utility have not been demonstrated.METHODS:We used a streamlined whole-genome sequencing approach to obtain genomic profiles for 263 patients with myeloid cancers, including 235 patients who had undergone successful cytogenetic analysis. We adapted sample preparation, sequencing, and analysis to detect mutations for risk stratification using existing European Leukemia Network (ELN) guidelines and to minimize turnaround time. We analyzed the performance of whole-genome sequencing by comparing our results with findings from cytogenetic analysis and targeted sequencing.RESULTS:Whole-genome sequencing detected all 40 recurrent translocations and 91 copy-number alterations that had been identified by cytogenetic analysis. In addition, we identified new clinically reportable genomic events in 40 of 235 patients (17.0%). Prospective sequencing of samples obtained from 117 consecutive patients was performed in a median of 5 days and provided new genetic information in 29 patients (24.8%), which changed the risk category for 19 patients (16.2%). Standard AML risk groups, as defined by sequencing results instead of cytogenetic analysis, correlated with clinical outcomes. Whole-genome sequencing was also used to stratify patients who had inconclusive results by cytogenetic analysis into risk groups in which clinical outcomes were measurably different.CONCLUSIONS:In our study, we found that whole-genome sequencing provided rapid and accurate genomic profiling in patients with AML or MDS. Such sequencing also provided a greater diagnostic yield than conventional cytogenetic analysis and more efficient risk stratification on the basis of standard risk categories. (Funded by the Siteman Cancer Research Fund and others.).
In this work, we present the Genome Modeling System (GMS), an analysis information management system capable of executing automated genome analysis pipelines at a massive scale. The GMS framework provides detailed tracking of samples and data coupled with reliable and repeatable analysis pipelines. The GMS also serves as a platform for bioinformatics development, allowing a large team to collaborate on data analysis, or an individual researcher to leverage the work of others effectively within its data management system. Rather than separating ad-hoc analysis from rigorous, reproducible pipelines, the GMS promotes systematic integration between the two. As a demonstration of the GMS, we performed an integrated analysis of whole genome, exome and transcriptome sequencing data from a breast cancer cell line (HCC1395) and matched lymphoblastoid line (HCC1395BL). These data are available for users to test the software, complete tutorials and develop novel GMS pipeline configurations. The GMS is available at https://github.com/genome/gms.
Abstract Background: Estrogen receptors are over-expressed in around 70% of breast cancer cases. The genetic changes that occur during aromatase inhibitor (AI) treatment are not well understood and may differ depending upon the patient's response phenotype. Methods: We performed whole genome sequencing (WGS) of matched blood, pre-treatment, and post-treatment biopsy samples from 22 estrogen receptor positive breast cancer patients treated with neoadjuvant aromatase inhibitors. For 5 cases, we performed the whole genome sequencing (WGS) on patients’ matched normal, two pre AI-treatment, and two post AI-treatment DNA isolates from biopsy samples. We validated all putative coding and non-coding somatic mutations using deep sequencing. By comparing the validated somatic mutations from pre- and post- AI treatment biopsy samples, we were able to determine the alterations in the tumor genomes. In every case we defined the clonal architecture of each pair of pre-treatment and post-treatment biopsy samples by comparing the variant allele frequencies from thousands of validated somatic mutations. Results: Comparisons of the two pre AI-treatment biopsy samples from the same patient indicates that the variant allele frequencies of mutations showed high concordances in all 5 cases, 0.74 to 0.95 range of correlation coefficient. Only a small percentage of somatic mutations were detected in one pre-treatment sample and not the other (4.65% overall). In comparing the somatic variations between pre-treatment and matched post-treatment biopsy samples in 22 cases, we found that patients with good clinical response to AI treatment retained known driver mutations only in their pre-treatment tumors. Conversely, those patients with poor clinical response presented new driver mutations in their post-treatment samples. Furthermore, the variant allele frequency for most mutated genes decreased in post AI treatment samples for patients with good AI treatment response; on the contrary, the variant allele frequency increased for patients with poor clinical response. Conclusions: From WGS of matched normal, pre-treatment, and post-treatment biopsy samples, we identified new driver genes mutated in patients with poor clinical response, while patients with good clinical response had lost mutated driver genes in their post-treatment biopsy samples. The genetic landscape revealed by WGS of pre-treatment and post-treatment biopsy samples reveals mutational repertoires are remodeled by AI therapy. This finding suggests deep sequencing of AI treated samples will be necessary to reveal the complete complement of mutations present in a patient's tumor. Citation Format: {Authors}. {Abstract title} [abstract]. In: Proceedings of the 103rd Annual Meeting of the American Association for Cancer Research; 2012 Mar 31-Apr 4; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2012;72(8 Suppl):Abstract nr LB-423. doi:1538-7445.AM2012-LB-423
To correlate the variable clinical features of oestrogen-receptor-positive breast cancer with somatic alterations, we studied pretreatment tumour biopsies accrued from patients in two studies of neoadjuvant aromatase inhibitor therapy by massively parallel sequencing and analysis. Eighteen significantly mutated genes were identified, including five genes (RUNX1, CBFB, MYH9, MLL3 and SF3B1) previously linked to haematopoietic disorders. Mutant MAP3K1 was associated with luminal A status, low-grade histology and low proliferation rates, whereas mutant TP53 was associated with the opposite pattern. Moreover, mutant GATA3 correlated with suppression of proliferation upon aromatase inhibitor treatment. Pathway analysis demonstrated that mutations in MAP2K4, a MAP3K1 substrate, produced similar perturbations as MAP3K1 loss. Distinct phenotypes in oestrogen-receptor-positive breast cancer are associated with specific patterns of somatic mutations that map into cellular pathways linked to tumour biology, but most recurrent mutations are relatively infrequent. Prospective clinical trials based on these findings will require comprehensive genome sequencing.
Massively parallel DNA sequencing technologies provide an unprecedented ability to screen entire genomes for genetic changes associated with tumour progression. Here we describe the genomic analyses of four DNA samples from an African-American patient with basal-like breast cancer: peripheral blood, the primary tumour, a brain metastasis and a xenograft derived from the primary tumour. The metastasis contained two de novo mutations and a large deletion not present in the primary tumour, and was significantly enriched for 20 shared mutations. The xenograft retained all primary tumour mutations and displayed a mutation enrichment pattern that resembled the metastasis. Two overlapping large deletions, encompassing CTNNA1 , were present in all three tumour samples. The differential mutation frequencies and structural variation patterns in metastasis and xenograft compared with the primary tumour indicate that secondary tumours may arise from a minority of cells within the primary tumour.
BACKGROUND:The full complement of DNA mutations that are responsible for the pathogenesis of acute myeloid leukemia (AML) is not yet known.METHODS:We used massively parallel DNA sequencing to obtain a very high level of coverage (approximately 98%) of a primary, cytogenetically normal, de novo genome for AML with minimal maturation (AML-M1) and a matched normal skin genome.RESULTS:We identified 12 acquired (somatic) mutations within the coding sequences of genes and 52 somatic point mutations in conserved or regulatory portions of the genome. All mutations appeared to be heterozygous and present in nearly all cells in the tumor sample. Four of the 64 mutations occurred in at least 1 additional AML sample in 188 samples that were tested. Mutations in NRAS and NPM1 had been identified previously in patients with AML, but two other mutations had not been identified. One of these mutations, in the IDH1 gene, was present in 15 of 187 additional AML genomes tested and was strongly associated with normal cytogenetic status; it was present in 13 of 80 cytogenetically normal samples (16%). The other was a nongenic mutation in a genomic region with regulatory potential and conservation in higher mammals; we detected it in one additional AML tumor. The AML genome that we sequenced contains approximately 750 point mutations, of which only a small fraction are likely to be relevant to pathogenesis.CONCLUSIONS:By comparing the sequences of tumor and skin genomes of a patient with AML-M1, we have identified recurring mutations that may be relevant for pathogenesis.
We report an improved draft nucleotide sequence of the 2.3-gigabase genome of maize, an important crop plant and model for biological research. Over 32,000 genes were predicted, of which 99.8% were placed on reference chromosomes. Nearly 85% of the genome is composed of hundreds of families of transposable elements, dispersed nonuniformly across the genome. These were responsible for the capture and amplification of numerous gene fragments and affect the composition, sizes, and positions of centromeres. We also report on the correlation of methylation-poor regions with Mu transposon insertions and recombination, and copy number variants with insertions and/or deletions, as well as how uneven gene losses between duplicated regions were involved in returning an ancient allotetraploid to a genetically diploid state. These analyses inform and set the stage for further investigations to improve our understanding of the domestication and agricultural improvements of maize.
We have sequenced five distinct mitochondrial genomes in maize: two fertile cytotypes (NA and the previously reported NB) and three cytoplasmic-male-sterile cytotypes (CMS-C, CMS-S, and CMS-T). Their genome sizes range from 535,825 bp in CMS-T to 739,719 bp in CMS-C. Large duplications (0.5–120 kb) account for most of the size increases. Plastid DNA accounts for 2.3–4.6% of each mitochondrial genome. The genomes share a minimum set of 51 genes for 33 conserved proteins, three ribosomal RNAs, and 15 transfer RNAs. Numbers of duplicate genes and plastid-derived tRNAs vary among cytotypes. A high level of sequence conservation exists both within and outside of genes (1.65–7.04 substitutions/10 kb in pairwise comparisons). However, sequence losses and gains are common: integrated plastid and plasmid sequences, as well as noncoding “native” mitochondrial sequences, can be lost with no phenotypic consequence. The organization of the different maize mitochondrial genomes varies dramatically; even between the two fertile cytotypes, there are 16 rearrangements. Comparing the finished shotgun sequences of multiple mitochondrial genomes from the same species suggests which genes and open reading frames are potentially functional, including which chimeric ORFs are candidate genes for cytoplasmic male sterility. This method identified the known CMS-associated ORFs in CMS-S and CMS-T, but not in CMS-C.
Human chromosome 2 is unique to the human lineage in being the product of a head-to-head fusion of two intermediate-sized ancestral chromosomes. Chromosome 4 has received attention primarily related to the search for the Huntington's disease gene, but also for genes associated with Wolf-Hirschhorn syndrome, polycystic kidney disease and a form of muscular dystrophy. Here we present approximately 237 million base pairs of sequence for chromosome 2, and 186 million base pairs for chromosome 4, representing more than 99.6% of their euchromatic sequences. Our initial analyses have identified 1,346 protein-coding genes and 1,239 pseudogenes on chromosome 2, and 796 protein-coding genes and 778 pseudogenes on chromosome 4. Extensive analyses confirm the underlying construction of the sequence, and expand our understanding of the structure and evolution of mammalian chromosomes, including gene deserts, segmental duplications and highly variant regions.
Salmonella enterica serovars often have a broad host range, and some cause both gastrointestinal and systemic disease. But the serovars Paratyphi A and Typhi are restricted to humans and cause only systemic disease. It has been estimated that Typhi arose in the last few thousand years. The sequence and microarray analysis of the Paratyphi A genome indicates that it is similar to the Typhi genome but suggests that it has a more recent evolutionary origin. Both genomes have independently accumulated many pseudogenes among their ∼4,400 protein coding sequences: 173 in Paratyphi A and ∼210 in Typhi. The recent convergence of these two similar genomes on a similar phenotype is subtly reflected in their genotypes: only 30 genes are degraded in both serovars. Nevertheless, these 30 genes include three known to be important in gastroenteritis, which does not occur in these serovars, and four for Salmonella-translocated effectors, which are normally secreted into host cells to subvert host functions. Loss of function also occurs by mutation in different genes in the same pathway (e.g., in chemotaxis and in the production of fimbriae).
Human chromosome 7 has historically received prominent attention in the human genetics community, primarily related to the search for the cystic fibrosis gene and the frequent cytogenetic changes associated with various forms of cancer. Here we present more than 153 million base pairs representing 99.4% of the euchromatic sequence of chromosome 7, the first metacentric chromosome completed so far. The sequence has excellent concordance with previously established physical and genetic maps, and it exhibits an unusual amount of segmentally duplicated sequence (8.2%), with marked differences between the two arms. Our initial analyses have identified 1,150 protein-coding genes, 605 of which have been confirmed by complementary DNA sequences, and an additional 941 pseudogenes. Of genes confirmed by transcript sequences, some are polymorphic for mutations that disrupt the reading frame.
Salmonella enterica subspecies I, serovar Typhimurium ( S. typhimurium ), is a leading cause of human gastroenteritis, and is used as a mouse model of human typhoid fever 1 . The incidence of non-typhoid salmonellosis is increasing worldwide 2 , 3 , 4 , causing millions of infections and many deaths in the human population each year. Here we sequenced the 4,857-kilobase (kb) chromosome and 94-kb virulence plasmid of S. typhimurium strain LT2. The distribution of close homologues of S. typhimurium LT2 genes in eight related enterobacteria was determined using previously completed genomes of three related bacteria, sample sequencing of both S. enterica serovar Paratyphi A ( S. paratyphi A) and Klebsiella pneumoniae , and hybridization of three unsequenced genomes to a microarray of S. typhimurium LT2 genes. Lateral transfer of genes is frequent, with 11% of the S. typhimurium LT2 genes missing from S. enterica serovar Typhi ( S. typhi ), and 29% missing from Escherichia coli K12. The 352 gene homologues of S. typhimurium LT2 confined to subspecies I of S. enterica —containing most mammalian and bird pathogens 5 —are useful for studies of epidemiology, host specificity and pathogenesis. Most of these homologues were previously unknown, and 50 may be exported to the periplasm or outer membrane, rendering them accessible as therapeutic or vaccine targets.