Abstract Introduction Sequencing the Adaptive Immune Receptor Repertoire (AIRR) allows the characterization of immune states in health and disease, including infectious diseases, (auto)immune diseases, and cancer. Although numerous tools exist for reconstructing B and T cell receptor (BCR and TCR) sequences and inferring clonal relationships from AIRR sequencing (AIRR-seq) data, many lack scalability, efficient sample-level parallelization, or portability to high-performance computing environments. We addressed these limitations with nf-core/airrflow (https://nf-co.re/airrflow), a scalable and reproducible Nextflow-based workflow for processing bulk and single-cell AIRR-seq data. Since its implementation, we have expanded the workflow with new functionality including BCR and TCR sequence embedding using large-language models (LLM), immunoglobulin (IG) loci genotyping and support for the AIRR community germline references. Methods nf-core/airrflow integrates tools from the Immcantation Framework (immcantation.org) following BCR and TCR data analysis best practices. We recently expanded the workflow to include LLM sequence embedding with AMULETY, as well as IG loci genotyping and novel allele detection using TIgGER. We additionally provide support for the newly released AIRR Community germline reference datasets hosted in the Open Germline Receptor Database (OGRDB). Results We demonstrate the applicability of nf-core/airrflow by genotyping and generating embeddings of publicly available BCR sequencing datasets from individuals with autoimmune diseases, including systemic lupus erythematosus, type 1 diabetes and rheumatoid arthritis. Conclusion nf-core/airrflow is a comprehensive and scalable workflow for AIRR-seq data analysis, enabling a wide range of applications in immune mediated and infectious disease research and supporting the reproducible analysis of increasingly large AIRR-seq datasets. Funding Source This work was supported by the National Institutes of Health National Institute for Allergy and Infectious Diseases grant U01AI184647 to G.G. Topic Categories Computational and Systems Immunology (COMP)
Adaptive Immune Receptor Repertoire sequencing (AIRR-seq) is a valuable experimental tool to study the immune state in health and following immune challenges such as infectious diseases, (auto)immune diseases, and cancer. Several tools have been developed to reconstruct B cell and T cell receptor sequences from AIRR-seq data and infer B and T cell clonal relationships. However, currently available tools offer limited parallelization across samples, scalability or portability to high-performance computing infrastructures. To address this need, we developed nf-core/airrflow, an end-to-end bulk and single-cell AIRR-seq processing workflow which integrates the Immcantation Framework following BCR and TCR sequencing data analysis best practices. The Immcantation Framework is a comprehensive toolset, which allows the processing of bulk and single-cell AIRR-seq data from raw read processing to clonal inference. nf-core/airrflow is written in Nextflow and is part of the nf-core project, which collects community contributed and curated Nextflow workflows for a wide variety of analysis tasks. We assessed the performance of nf-core/airrflow on simulated sequencing data with sequencing errors and show example results with real datasets. To demonstrate the applicability of nf-core/airrflow to the high-throughput processing of large AIRR-seq datasets, we validated and extended previously reported findings of convergent antibody responses to SARS-CoV-2 by analyzing 97 COVID-19 infected individuals and 99 healthy controls, including a mixture of bulk and single-cell sequencing datasets. Using this dataset, we extended the convergence findings to 20 additional subjects, highlighting the applicability of nf-core/airrflow to validate findings in small in-house cohorts with reanalysis of large publicly available AIRR datasets.
Antibiotic resistance is a persistent problem in health care and a factor that may help maintain bacterial diversity in natural environments. Bacteriophages (“phages”) are viruses that specifically infect bacteria.
Congenital hydrocephalus (CH), featuring markedly enlarged brain ventricles, is thought to arise from failed cerebrospinal fluid (CSF) homeostasis and is treated with lifelong surgical CSF shunting with substantial morbidity. CH pathogenesis is poorly understood. Exome sequencing of 125 CH trios and 52 additional probands identified three genes with significant burden of rare damaging de novo or transmitted mutations: TRIM71 (p = 2.15 x 10(-7)), SMARCC1 (p = 8.15 x 10(-10)), and PTCH1 (p = 1.06 x 10(-6)). Additionally, two de novo duplications were identified at the SHH locus, encoding the PTCH1 ligand (p = 1.2 x 10(-4)). Together, these probands account for similar to 10% of studied cases. Strikingly, all four genes are required for neural tube development and regulate ventricular zone neural stem cell fate. These results implicate impaired neurogenesis (rather than active CSF accumulation) in the pathogenesis of a subset of CH patients, with potential diagnostic, prognostic, and therapeutic ramifications.
Genetic variation affecting gene regulation is a driver of phenotypic differences between individuals and can be used to uncover how biological processes are organized in a cell. Although detecting cis -eQTLs is now routine, trans -eQTLs have proven more challenging to find due to the modest variance explained and the multiple testing burden when comparing millions of SNPs for association to thousands of transcripts. Here, we provide evidence for the existence of trans -eQTLs by looking for SNPs associated with the expression of multiple genes simultaneously. We find substantial evidence of trans -eQTLs, with an 1.8-fold enrichment in nominally significant markers in all three populations and significant overlap between results across the populations. These trans -eQTLs target the same genes and show the same direction of effect across populations. We define a high-confidence set of eight independent trans -eQTLs which are associated to multiple transcripts in all three populations, and affect the same targets in all three populations with the same direction of effect. We then show that target transcripts of trans -eQTLs encode proteins that interact more frequently than expected by chance, and are enriched for pathway annotations indicative of roles in basic cell homeostasis. Thus, we have demonstrated that trans -eQTLs can be accurately identified even in studies of limited sample size.
Despite efforts to interrogate human genome variation through large-scale databases, systematic preference toward populations of Caucasian descendants has resulted in unintended reduction of power in studying non-Caucasians. Here we report a compilation of coding variants from 1,055 healthy Korean individuals (KOVA; Korean Variant Archive). The samples were sequenced to a mean depth of 75x, yielding 101 singleton variants per individual. Population genetics analysis demonstrates that the Korean population is a distinct ethnic group comparable to other discrete ethnic groups in Africa and Europe, providing a rationale for such independent genomic datasets. Indeed, KOVA conferred 22.8% increased variant filtering power in addition to Exome Aggregation Consortium (ExAC) when used on Korean exomes. Functional assessment of nonsynonymous variant supported the presence of purifying selection in Koreans. Analysis of copy number variants detected 5.2 deletions and 10.3 amplifications per individual with an increased fraction of novel variants among smaller and rarer copy number variable segments. We also report a list of germline variants that are associated with increased tumor susceptibility. This catalog can function as a critical addition to the pre-existing variant databases in pursuing genetic studies of Korean individuals.
Congenital heart disease (CHD) is the leading cause of mortality from birth defects. Here, exome sequencing of a single cohort of 2,871 CHD probands, including 2,645 parent-offspring trios, implicated rare inherited mutations in 1.8%, including a recessive founder mutation in GDF1 accounting for ∼5% of severe CHD in Ashkenazim, recessive genotypes in MYH6 accounting for ∼11% of Shone complex, and dominant FLT4 mutations accounting for 2.3% of Tetralogy of Fallot. De novo mutations (DNMs) accounted for 8% of cases, including ∼3% of isolated CHD patients and ∼28% with both neurodevelopmental and extra-cardiac congenital anomalies. Seven genes surpassed thresholds for genome-wide significance, and 12 genes not previously implicated in CHD had >70% probability of being disease related. DNMs in ∼440 genes were inferred to contribute to CHD. Striking overlap between genes with damaging DNMs in probands with CHD and autism was also found.
We report a significantly-enhanced bioinformatics suite and database for proteomics research called Yale Protein Expression Database (YPED) that is used by investigators at more than 300 institutions worldwide. YPED meets the data management, archival, and analysis needs of a high-throughput mass spectrometry-based proteomics research ranging from a single laboratory, group of laboratories within and beyond an institution, to the entire proteomics community. The current version is a significant improvement over the first version in that it contains new modules for liquid chromatography-tandem mass spectrometry (LC-MS/MS) database search results, label and label-free quantitative proteomic analysis, and several scoring outputs for phosphopeptide site localization. In addition, we have added both peptide and protein comparative analysis tools to enable pairwise analysis of distinct peptides/proteins in each sample and of overlapping peptides/proteins between all samples in multiple datasets. We have also implemented a targeted proteomics module for automated multiple reaction monitoring (MRM)/selective reaction monitoring (SRM) assay development. We have linked YPED's database search results and both label-based and label-free fold-change analysis to the Skyline Panorama repository for online spectra visualization. In addition, we have built enhanced functionality to curate peptide identifications into an MS/MS peptide spectral library for all of our protein database search identification results.
Richard Lifton and colleagues report a genomic analysis of cutaneous T cell lymphoma (CTCL). Their results implicate several pathways in CTCL pathogenesis, including genes involved in T cell activation and apoptosis, NF-κB signaling, chromatin remodeling and DNA damage response. Cutaneous T cell lymphoma (CTCL) is a non-Hodgkin lymphoma of skin-homing T lymphocytes. We performed exome and whole-genome DNA sequencing and RNA sequencing on purified CTCL and matched normal cells. The results implicate mutations in 17 genes in CTCL pathogenesis, including genes involved in T cell activation and apoptosis, NF-κB signaling, chromatin remodeling and DNA damage response. CTCL is distinctive in that somatic copy number variants (SCNVs) comprise 92% of all driver mutations (mean of 11.8 pathogenic SCNVs versus 1.0 somatic single-nucleotide variant per CTCL). These findings have implications for new therapeutics.
The Trypanosoma brucei complex contains a number of subspecies with exceptionally variable life histories, including zoonotic subspecies, which are causative agents of human African trypanosomiasis (HAT) in sub-Saharan Africa.Paradoxically, genomic variation between taxa is extremely low.We analyzed the whole-genome sequences of 39 isolates across the T. brucei complex from diverse hosts and regions, identifying 608,501 single nucleotide polymorphisms that represent 2.33% of the nuclear genome.We show that human pathogenicity occurs across a wide range of parasite genotypes, and taxonomic designation does not reflect genetic variation across the group, as previous studies have suggested based on a small number of genes.This genome-wide study allowed the identification of significant host and geographic location associations.Strong purifying selection was detected in genomic regions associated with cytoskeleton structure, and regulatory genes associated with antigenic variation, suggesting conservation of these regions in African trypanosomes.In agreement with expectations drawn from meiotic reciprocal recombination, differences in average linkage disequilibrium between chromosomes in T. brucei correlate positively with chromosome size.In addition to insights into the life history of a diverse group of eukaryotic parasites, the documentation of genomic variation across the T. brucei complex and its association with specific hosts and geographic localities will aid in the development of comprehensive monitoring tools crucial to the proposed elimination of HAT by 2020, and on a shorter term, for monitoring the feared merger between the two human infective parasites, T. brucei rhodesiense and T. b. gambiense, in northern Uganda.
Efficient DNA double-strand break (DSB) repair is a critical determinant of cell survival in response to DNA damaging agents, and it plays a key role in the maintenance of genomic integrity. Homologous recombination (HR) and non-homologous end-joining (NHEJ) represent the two major pathways by which DSBs are repaired in mammalian cells. We now understand that HR and NHEJ repair are composed of multiple sub-pathways, some of which still remain poorly understood. As such, there is great interest in the development of novel assays to interrogate these key pathways, which could lead to the development of novel therapeutics, and a better understanding of how DSBs are repaired. Furthermore, assays which can measure repair specifically at endogenous chromosomal loci are of particular interest, because of an emerging understanding that chromatin interactions heavily influence DSB repair pathway choice. Here, we present the design and validation of a novel, next-generation sequencing-based approach to study DSB repair at chromosomal loci in cells. We demonstrate that NHEJ repair "fingerprints" can be identified using our assay, which are dependent on the status of key DSB repair proteins. In addition, we have validated that our system can be used to detect dynamic shifts in DSB repair activity in response to specific perturbations. This approach represents a unique alternative to many currently available DSB repair assays, which typical rely on the expression of reporter genes as an indirect read-out for repair. As such, we believe this tool will be useful for DNA repair researchers to study NHEJ repair in a high-throughput and sensitive manner, with the capacity to detect subtle changes in DSB repair patterns that was not possible previously.
Genetic variation affecting gene regulation is a central driver of phenotypic differences between individuals and can be used to uncover how biological processes are organized in a cell. Although detecting cis-eQTLs is now routine, trans-eQTLs have proven more challenging to find due to the modest variance explained and the multiple tests burden of testing millions of SNPs for association to thousands of transcripts. Here, we successfully map trans-eQTLs with the complementary approach of looking for SNPs associated to the expression of multiple genes simultaneously. We find 732 trans- eQTLs that replicate across two continental populations; each trans-eQTL controls large groups of target transcripts (regulons), which are part of interacting networks controlled by transcription factors. We are thus able to uncover co-regulated gene sets and begin describing the cell circuitry of gene regulation.
BACKGROUND:Current research suggests that a small set of "driver" mutations are responsible for tumorigenesis while a larger body of "passenger" mutations occur in the tumor but do not progress the disease. Due to recent pharmacological successes in treating cancers caused by driver mutations, a variety of methodologies that attempt to identify such mutations have been developed. Based on the hypothesis that driver mutations tend to cluster in key regions of the protein, the development of cluster identification algorithms has become critical.RESULTS:We have developed a novel methodology, SpacePAC (Spatial Protein Amino acid Clustering), that identifies mutational clustering by considering the protein tertiary structure directly in 3D space. By combining the mutational data in the Catalogue of Somatic Mutations in Cancer (COSMIC) and the spatial information in the Protein Data Bank (PDB), SpacePAC is able to identify novel mutation clusters in many proteins such as FGFR3 and CHRM2. In addition, SpacePAC is better able to localize the most significant mutational hotspots as demonstrated in the cases of BRAF and ALK. The R package is available on Bioconductor at: http://www.bioconductor.org/packages/release/bioc/html/SpacePAC.html.CONCLUSION:SpacePAC adds a valuable tool to the identification of mutational clusters while considering protein tertiary structure.
BRAF inhibitors improve melanoma patient survival, but resistance invariably develops. Here we report the discovery of a novel BRAF mutation that confers resistance to PLX4032 employing whole-exome sequencing of drug-resistant BRAF(V600K) melanoma cells. We further describe a new screening approach, a genome-wide piggyBac mutagenesis screen that revealed clinically relevant aberrations (N-terminal BRAF truncations and CRAF overexpression). The novel BRAF mutation, a Leu505 to His substitution (BRAF(L505H)), is the first resistance-conferring second-site mutation identified in BRAF mutant cells. The mutation replaces a small nonpolar amino acid at the BRAF-PLX4032 interface with a larger polar residue. Moreover, we show that BRAF(L505H), found in human prostate cancer, is itself a MAPK-activating, PLX4032-resistant oncogenic mutation. Lastly, we demonstrate that the PLX4032-resistant melanoma cells are sensitive to novel, next-generation BRAF inhibitors, especially the paradox-blocker' PLX8394, supporting its use in clinical trials for treatment of melanoma patients with BRAF-mutations.
Exome sequencing of patients with congenital heart disease (CHD) and their unaffected parents reveals an excess of strong-effect, protein-altering de novo mutations in genes expressed in the developing heart, many of which regulate chromatin modification in key developmental genes; collectively, these mutations are predicted to account for approximately 10% of severe CHD cases. This paper demonstrates that de novo mutations with large effect have a role in the pathogenesis of at least 10% of cases of congenital heart disease (CHD). Using exome sequence analysis in parent–offspring trios Richard Lifton and colleagues compared the frequency of de novo mutations, identified by exome sequencing, in 362 CHD parent–offspring trios and 264 control trios. Gene ontology analysis demonstrated significant enrichment of de novo protein-altering mutation of genes involved in chromatin modification, notably a marked enrichment of genes involved in the production, removal and reading of methylation of histone H3K4 and H3K27. Congenital heart disease (CHD) is the most frequent birth defect, affecting 0.8% of live births1. Many cases occur sporadically and impair reproductive fitness, suggesting a role for de novo mutations. Here we compare the incidence of de novo mutations in 362 severe CHD cases and 264 controls by analysing exome sequencing of parent–offspring trios. CHD cases show a significant excess of protein-altering de novo mutations in genes expressed in the developing heart, with an odds ratio of 7.5 for damaging (premature termination, frameshift, splice site) mutations. Similar odds ratios are seen across the main classes of severe CHD. We find a marked excess of de novo mutations in genes involved in the production, removal or reading of histone 3 lysine 4 (H3K4) methylation, or ubiquitination of H2BK120, which is required for H3K4 methylation2,3,4. There are also two de novo mutations in SMAD2, which regulates H3K27 methylation in the embryonic left–right organizer5. The combination of both activating (H3K4 methylation) and inactivating (H3K27 methylation) chromatin marks characterizes ‘poised’ promoters and enhancers, which regulate expression of key developmental genes6. These findings implicate de novo point mutations in several hundreds of genes that collectively contribute to approximately 10% of severe CHD.
Despite considerable efforts to sequence hypermutated cancers such as melanoma, distinguishing cancer-driving genes from thousands of recurrently mutated genes remains a significant challenge. To circumvent the problematic background mutation rates and identify new melanoma driver genes, we carried out a low-copy piggyBac transposon mutagenesis screen in mice. We induced eleven melanomas with mutation burdens that were 100-fold lower relative to human melanomas. Thirty-eight implicated genes, including two known drivers of human melanoma, were classified into three groups based on high, low, or background-level mutation frequencies in human melanomas, and we further explored the functional significance of genes in each group. For two genes overlooked by prevailing discovery methods, we found that loss of membrane associated guanylate kinase, WW and PDZ domain containing 2 and protein tyrosine phosphatase, receptor type, O cooperated with the v-raf murine sarcoma viral oncogene homolog B (BRAF) recurrent V600E mutation to promote cellular transformation. Moreover, for infrequently mutated genes often disregarded by current methods, we discovered recurrent mitogen-activated protein kinase kinase kinase 1 (Map3k1)-activating insertions in our screen, mirroring recurrent MAP3K1 up-regulation in human melanomas. Aberrant expression of Map3k1 enabled growth factor-autonomous proliferation and drove BRAF-independent ERK signaling, thus shedding light on alternative means of activating this prominent signaling pathway in melanoma. In summary, our study contributes several previously undescribed genes involved in melanoma and establishes an important proof-of-principle for the utility of the low-copy transposon mutagenesis approach for identifying cancer-driving genes, especially those masked by hypermutation.
Exome sequencing identifies mutations in kelch-like 3 and cullin 3 as causes of a syndrome featuring high blood pressure and electrolyte abnormalities. Exome sequencing in a family with pseudohypoaldosteronism type II, a rare Mendelian syndrome featuring hypertension, has identified mutations in kelch-like 3 (KLHL3) and cullin 3 (CUL3). This implicates a specific ubiquitin ligase pathway in the regulation of blood pressure and electrolyte homeostasis. This study also demonstrates the value of exome sequencing — a cheaper alternative to whole genome sequencing — in the identification of disease-associated genes. Hypertension affects one billion people and is a principal reversible risk factor for cardiovascular disease. Pseudohypoaldosteronism type II (PHAII), a rare Mendelian syndrome featuring hypertension, hyperkalaemia and metabolic acidosis, has revealed previously unrecognized physiology orchestrating the balance between renal salt reabsorption and K+ and H+ excretion1. Here we used exome sequencing to identify mutations in kelch-like 3 (KLHL3) or cullin 3 (CUL3) in PHAII patients from 41 unrelated families. KLHL3 mutations are either recessive or dominant, whereas CUL3 mutations are dominant and predominantly de novo. CUL3 and BTB-domain-containing kelch proteins such as KLHL3 are components of cullin–RING E3 ligase complexes that ubiquitinate substrates bound to kelch propeller domains2,3,4,5,6,7,8. Dominant KLHL3 mutations are clustered in short segments within the kelch propeller and BTB domains implicated in substrate9 and cullin5 binding, respectively. Diverse CUL3 mutations all result in skipping of exon 9, producing an in-frame deletion. Because dominant KLHL3 and CUL3 mutations both phenocopy recessive loss-of-function KLHL3 mutations, they may abrogate ubiquitination of KLHL3 substrates. Disease features are reversed by thiazide diuretics, which inhibit the Na–Cl cotransporter in the distal nephron of the kidney; KLHL3 and CUL3 are expressed in this location, suggesting a mechanistic link between KLHL3 and CUL3 mutations, increased Na–Cl reabsorption, and disease pathogenesis. These findings demonstrate the utility of exome sequencing in disease gene identification despite the combined complexities of locus heterogeneity, mixed models of transmission and frequent de novo mutation, and establish a fundamental role for KLHL3 and CUL3 in blood pressure, K+ and pH homeostasis.
Rare de novo single nucleotide variants in brain-expressed genes are found to be associated with autism spectrum disorders and to carry large effects. Although it is well accepted that genetics makes a strong contribution to autism spectrum disorder, most of the underlying causes of the condition remain unknown. Three groups present large-scale exome-sequencing studies of individuals with sporadic autism spectrum disorder, including many parent–child trios and unaffected siblings. The overall message from the three papers is that there is extreme locus heterogeneity among autistic individuals, with hundreds of genes involved in the condition, and with no single gene contributing to more than a small fraction of cases. Sanders et al. report the association of the gene SCN2A, previously identified in epilepsy syndromes, with the risk of autism. Neale et al. find strong evidence that CHD8 and KATNAL2 are autism risk factors. O'Roak et al. observe that a large proportion of the mutated proteins have crucial roles in fundamental developmental pathways, including β-catenin and p53 signalling. Multiple studies have confirmed the contribution of rare de novo copy number variations to the risk for autism spectrum disorders1,2,3. But whereas de novo single nucleotide variants have been identified in affected individuals4, their contribution to risk has yet to be clarified. Specifically, the frequency and distribution of these mutations have not been well characterized in matched unaffected controls, and such data are vital to the interpretation of de novo coding mutations observed in probands. Here we show, using whole-exome sequencing of 928 individuals, including 200 phenotypically discordant sibling pairs, that highly disruptive (nonsense and splice-site) de novo mutations in brain-expressed genes are associated with autism spectrum disorders and carry large effects. On the basis of mutation rates in unaffected individuals, we demonstrate that multiple independent de novo single nucleotide variants in the same gene among unrelated probands reliably identifies risk alleles, providing a clear path forward for gene discovery. Among a total of 279 identified de novo coding mutations, there is a single instance in probands, and none in siblings, in which two independent nonsense variants disrupt the same gene, SCN2A (sodium channel, voltage-gated, type II, α subunit), a result that is highly unlikely by chance.
ABSTRACT Ancient endosymbionts have been associated with extreme genome structural stability with little differentiation in gene inventory between sister species. Tsetse flies (Diptera: Glossinidae) harbor an obligate endosymbiont, Wigglesworthia, which has coevolved with the Glossina radiation. We report on the ~720-kb Wigglesworthia genome and its associated plasmid from Glossina morsitans morsitans and compare them to those of the symbiont from Glossina brevipalpis. While there was overall high synteny between the two genomes, a large inversion was noted. Furthermore, symbiont transcriptional analyses demonstrated host tissue and development-specific gene expression supporting robust transcriptional regulation in Wigglesworthia, an unprecedented observation in other obligate mutualist endosymbionts. Expression and immunohistochemistry confirmed the role of flagella during the vertical transmission process from mother to intrauterine progeny. The expression of nutrient provisioning genes (thiC and hemH) suggests that Wigglesworthia may function in dietary supplementation tailored toward host development. Furthermore, despite extensive conservation, unique genes were identified within both symbiont genomes that may result in distinct metabolomes impacting host physiology. One of these differences involves the chorismate, phenylalanine, and folate biosynthetic pathways, which are uniquely present in Wigglesworthia morsitans. Interestingly, African trypanosomes are auxotrophs for phenylalanine and folate and salvage both exogenously. It is possible that W. morsitans contributes to the higher parasite susceptibility of its host species. IMPORTANCE Genomic stasis has historically been associated with obligate endosymbionts and their sister species. Here we characterize the Wigglesworthia genome of the tsetse fly species Glossina morsitans and compare it to its sister genome within G. brevipalpis. The similarity and variation between the genomes enabled specific hypotheses regarding functional biology. Expression analyses indicate significant levels of transcriptional regulation and support development- and tissue-specific functional roles for the symbiosis previously not observed in obligate mutualist symbionts. Retention of the genetically expensive flagella within these small genomes was demonstrated to be significant in symbiont transmission and tailored to the unique tsetse fly reproductive biology. Distinctions in metabolomes were also observed. We speculate an additional role for Wigglesworthia symbiosis where infections with pathogenic trypanosomes may depend upon symbiont species-specific metabolic products and thus influence the vector competence traits of different tsetse fly host species.