The human and mouse antibody repertoires are formed by identical processes, but like all small animals, mice only have sufficient lymphocytes to express a small part of the potential antibody repertoire. In this study, we determined how the heavy chain repertoires of two mouse strains are generated. Analysis of IgM- and IgG-associated VDJ rearrangements generated by high-throughput sequencing confirmed the presence of 99 functional immunoglobulin heavy chain variable (IGHV) genes in the C57BL/6 genome, and inferred the presence of 164 IGHV genes in the BALB/c genome. Remarkably, only five IGHV sequences were common to both strains. Compared with humans, little N nucleotide addition was seen in the junctions of mouse VDJ genes. Germline human IgG-associated IGHV genes are rare, but many murine IgG-associated IGHV genes were unmutated. Together these results suggest that the expressed mouse repertoire is more germline-focused than the human repertoire. The apparently divergent germline repertoires of the mouse strains are discussed with reference to reports that inbred mouse strains carry blocks of genes derived from each of the three subspecies of the house mouse. We hypothesize that the germline genes of BALB/c and C57BL/6 mice may originally have evolved to generate distinct germline-focused antibody repertoires in the different mouse subspecies.
Antigen selection of B cells within the germinal center reaction generally leads to the accumulation of replacement mutations in the complementarity-determining regions (CDRs) of immunoglobulin genes. Studies of mutations in IgE-associated VDJ gene sequences have cast doubt on the role of antigen selection in the evolution of the human IgE response, and it may be that selection for high affinity antibodies is a feature of some but not all allergic diseases. The severity of IgE-mediated anaphylaxis is such that it could result from higher affinity IgE antibodies. We therefore investigated IGHV mutations in IgE-associated sequences derived from ten individuals with a history of anaphylactic reactions to bee or wasp venom or peanut allergens. IgG sequences, which more certainly experience antigen selection, served as a control dataset. A total of 6025 unique IgE and 5396 unique IgG sequences were generated using high throughput 454 pyrosequencing. The proportion of replacement mutations seen in the CDRs of the IgG dataset was significantly higher than that of the IgE dataset, and the IgE sequences showed little evidence of antigen selection. To exclude the possibility that 454 errors had compromised analysis, rigorous filtering of the datasets led to datasets of 90 core IgE sequences and 411 IgG sequences. These sequences were present as both forward and reverse reads, and so were most unlikely to include sequencing errors. The filtered datasets confirmed that antigen selection plays a greater role in the evolution of IgG sequences than of IgE sequences derived from the study participants.
Somatic point mutations provide glimpses into B‐cell histories, and mutation numbers generally correlate with antibody affinity. We recently proposed a model of human isotype function, based in part on mutation analysis, in which the dominant pathway of isotype switching involves B cells moving sequentially through the four immunoglobulin (Ig) G subclasses. This should result in predictable differences in affinity between isotypes, and this helps explain how different isotypes work together. The model built on analysis of rearranged immunoglobulin heavy chain sequences amplified from Papua New Guinean villagers, which showed highly significant differences in the mean number of V‐REGION mutations in sequences, associated with the different IgG subclasses. To determine whether this relationship between mutation levels and isotypes is a more general phenomenon, the present study was conducted in healthy, urban residents of Sydney, Australia. VDJ sequences were generated from eight individuals, using 454 pyrosequencing, from cells expressing all isotypes except IgD and IgE. This resulted in 35 118 unique, productive VDJ sequences for the study. The data confirm that VDJ genes associated with progressively more 3′ Ig heavy chain gamma (IGHG) constant region genes show increasing levels of point mutation. Mean V‐REGION mutations in IgA1 and IgA2 sequences were similar. Patterns of mutations also differed between isotypes. Despite their association with T‐independent responses, IgG2 sequences showed significantly more mutational evidence of antigen selection than other IgG isotypes. Antigen selection was also significantly higher in IgA2 than in IgA1 sequences, raising the possibility of a preferential switch pathway from IGHG2 to IGHA2.
Immunoglobulin heavy chain rearrangements were sequences from PBMCs of 8 individuals using IGHC primers for all isotypes bar IgE. Each indivdual has a unique label 'Subject' identifier of the format NSXXX. Sequencing was carried out on the 454 pyrosequencing platform and rearranged immunoglobulin heavy chain sequences were partitioned against germline IGHV, IGHD and IGHJ using iHMMune-align. The datasets for each individual has been filtered to remove non-productive rearrangements (frame shift of the IGHJ or stop codons), those with greater than 45 IGHV mutations, sequences which contained indels in the V or J, and those with ambiguities. Within each data only unique sequences were reatained (100% identical) with a 'read copy no.' assigned to track the underlying copies of a sequence. Clonally-related sequences within each subject's dataset were also identified. Following germline reversion of all non-CDR3 nucleotide positions, sequence were clustered using the vmatch package (http://www.vmatch.de/) allowing up to 3 differences. A single representative for each cluster (representing a likely clonal set) was retained within the dataset. Representative sequences were selected based on the sequence with the highest copy number, or in cases where sequences shared the highest copy number, the least mutated sequence. If a cluster spanned multiple isotypes then a representative for each isotype was kept. Clusters are identified by a label that includes the subject in which the cluster was identified. Sequences that weren't part of a cluster as labelled with the subject identifier and "na". Replacement (R) and silent (S) mutation counts for complementarity determining regions (CDRs) and framework regions (FR) were determined for each sequence. CDR definitions that take the outer limits of the Kabat and IMGT systems were used. Where more than one mutation occurred within a single codon, all independent pathways to the final mutational outcome were considered and the R and S counts were weighted accordingly. For sequences with an Mv (total V-REGION mutation from codon 25 through 104, that is, excluding the CDR3 contribution of the V-REGION) greater than zero the proportion of total mutations that represent R in CDRs (CDR1 and CDR2) was calculated. These values were used in the analysis of antigen selection between sequences of different isotypes.
Somatic point mutations provide glimpses into B-cell histories, and mutation numbers generally correlate with antibody affinity. We recently proposed a model of human isotype function, based in part on mutation analysis, in which the dominant pathway of isotype switching involves B cells moving sequentially through the four immunoglobulin (Ig) G subclasses. This should result in predictable differences in affinity between isotypes, and this helps explain how different isotypes work together. The model built on analysis of rearranged immunoglobulin heavy chain sequences amplified from Papua New Guinean villagers, which showed highly significant differences in the mean number of V-REGION mutations in sequences, associated with the different IgG subclasses. To determine whether this relationship between mutation levels and isotypes is a more general phenomenon, the present study was conducted in healthy, urban residents of Sydney, Australia. VDJ sequences were generated from eight individuals, using 454 pyrosequencing, from cells expressing all isotypes except IgD and IgE. This resulted in 35 118 unique, productive VDJ sequences for the study. The data confirm that VDJ genes associated with progressively more 3′ Ig heavy chain gamma (IGHG) constant region genes show increasing levels of point mutation. Mean V-REGION mutations in IgA1 and IgA2 sequences were similar. Patterns of mutations also differed between isotypes. Despite their association with T-independent responses, IgG2 sequences showed significantly more mutational evidence of antigen selection than other IgG isotypes. Antigen selection was also significantly higher in IgA2 than in IgA1 sequences, raising the possibility of a preferential switch pathway from IGHG2 to IGHA2 .
Both the B cell receptor (BCR) and the T cell receptor (TCR) repertoires are generated through essentially identical processes of V(D)J recombination, exonuclease trimming of germline genes, and the random addition of non-template encoded nucleotides. The naïve TCR repertoire is constrained by thymic selection, and TCR repertoire studies have therefore focused strongly on the diversity of MHC-binding complementarity determining region (CDR) CDR3. The process of somatic point mutations has given B cell studies a major focus on variable (IGHV, IGLV, and IGKV) genes. This in turn has influenced how both the naïve and memory BCR repertoires have been studied. Diversity (D) genes are also more easily identified in BCR VDJ rearrangements than in TCR VDJ rearrangements, and this has allowed the processes and elements that contribute to the incredible diversity of the immunoglobulin heavy chain CDR3 to be analyzed in detail. This diversity can be contrasted with that of the light chain where a small number of polypeptide sequences dominate the repertoire. Biases in the use of different germline genes, in gene processing, and in the addition of non-template encoded nucleotides appear to be intrinsic to the recombination process, imparting “shape” to the repertoire of rearranged genes as a result of differences spanning many orders of magnitude in the probabilities that different BCRs will be generated. This may function to increase the precursor frequency of naïve B cells with important specificities, and the likely emergence of such B cell lineages upon antigen exposure is discussed with reference to public and private T cell clonotypes.
Scleractinian corals occur in symbiosis with a range of organisms including the dinoflagellate alga, Symbiodinium, an association that is mutualistic. However, not all symbionts benefit the host. In particular, many organisms within the microbial mucus layer that covers the coral epithelium can cause disease and death. Other organisms in symbiosis with corals include the recently described Chromera velia, a photosynthetic relative of the apicomplexan parasites that shares a common ancestor with Symbiodinium. To explore the nature of the association between C. velia and corals we first isolated C. velia from the coral Montipora digitata and then exposed aposymbiotic Acropora digitifera and A. tenuis larvae to these cultures. Three C. velia cultures were isolated, and symbiosis was established in coral larvae of both these species exposed to all three clones. Histology verified that C. velia was located in the larval endoderm and ectoderm. These results indicate that C. velia has the potential to be endosymbiotic with coral larvae.
We have analysed the transcribed immunoglobulin kappa (IGK) repertoire of peripheral blood B cells from four individuals from two genetically distinct populations, Papua New Guinean and Australian, using high-throughput DNA sequencing. The depth of sequencing data for each individual averaged 5,548 high-quality IGK reads, and permitted genotyping of the inferred IGKV and IGKJ germline gene segments for each individual. All individuals were homozygous at each IGKJ locus and had highly similar inferred IGKV genotypes. Preferential gene usage was seen at both the IGKV and IGKJ loci, but only IGKV segment usage varied significantly between individuals. Despite the differences in IGKV gene utilisation, the rearranged IGK repertoires showed extensive identity at the amino acid level. Public rearrangements (those shared by two or more individuals) made up 60.2% of the total sequenced IGK rearrangements. The total diversity of IGK rearrangements of each individual was estimated to range from just 340 to 549 unique amino acid sequences. Thus, the repertoire of unique expressed IGK rearrangements is dramatically less than previous theoretical estimates of IGK diversity, and the majority of expressed IGK rearrangements are likely to be extensively shared in individual human beings.
The existence of many highly similar genes in the lymphocyte receptor gene loci makes them difficult to investigate, and the determination of phased "haplotypes" has been particularly problematic. However, V(D)J gene rearrangements provide an opportunity to infer the association of Ig genes along the chromosomes. The chromosomal distribution of H chain genes in an Ig genotype can be inferred through analysis of VDJ rearrangements in individuals who are heterozygous at points within the IGH locus. We analyzed VDJ rearrangements from 44 individuals for whom sufficient unique rearrangements were available to allow comprehensive genotyping. Nine individuals were identified who were heterozygous at the IGHJ6 locus and for whom sufficient suitable VDJ rearrangements were available to allow comprehensive haplotyping. Each of the 18 resulting IGHV│IGHD│IGHJ haplotypes was unique. Apparent deletion polymorphisms were seen that involved as many as four contiguous, functional IGHV genes. Two deletion polymorphisms involving multiple contiguous IGHD genes were also inferred. Three previously unidentified gene duplications were detected, where two sequences recognized as allelic variants of a single gene were both inferred to be on a single chromosome. Phased genomic data brings clarity to the study of the contribution of each gene to the available repertoire of rearranged VDJ genes. Analysis of rearrangement frequencies suggests that particular genes may have substantially different yet predictable propensities for rearrangement within different haplotypes. Together with data highlighting the extent of haplotypic variation within the population, this suggests that there may be substantial variability in the available Ab repertoires of different individuals.
Complete and accurate knowledge of the genes and allelic variants of the human immunoglobulin gene loci is critical for studies of B cell repertoire development and somatic point mutation, but evidence from studies of VDJ rearrangements suggests that our knowledge of the available immunoglobulin gene repertoire is far from complete. The reported repertoire has changed little over the last 15 years. This is, in part, a consequence of the inefficiencies involved in searching for new members of large, multigenic gene families by cloning and sequencing. The advent of high-throughput sequencing provides a new avenue by which the germline repertoire can be explored. In this report, we describe pyrosequencing studies of the heavy chain IGHV1, IGHV3 and IGHV4 gene subgroups in ten Papua New Guineans. Thousands of 454 reads aligned with complete identity to 51 previously reported functional IGHV genes and allelic variants. A new gene, IGHV3-NL1*01, was identified, which differs from the nearest previously reported gene by 15 nucleotides. Sixteen new IGHV alleles were also identified, 15 of which varied from previously reported functional IGHV genes by between one and four nucleotides, while one sequence appears to be a functional variant of the pseudogene IGHV3-25. BLAST searches suggest that at least six of these new genes are carried within the relatively well-studied populations of North America, Europe or Asia. This study substantially expands the known immunoglobulin gene repertoire and demonstrates that genetic variation of immunoglobulin genes can now be efficiently explored in different human populations using high-throughput pyrosequencing.
Patterns of somatic mutation in IgE genes from allergic individuals have been a focus of study for many years, but IgE sequences have never been reported from parasitized individuals. To study the role of antigen selection in the evolution of the anti-parasite response, we therefore generated 118 IgE sequences from donors living in Papua New Guinea (PNG), an area of endemic parasitism. For comparison, we also generated IgG1, IgG2, IgG3 and IgG4 sequences from these donors, as well as IgG1 sequences from Australian donors. IgE sequences had, on average, 23.0 mutations. PNG IgG sequences had average mutation levels that varied from 17.7 (IgG3) to 27.1 (IgG4). Mean mutation levels correlated significantly with the position of their genes in the constant region gene locus (IgG3 < IgG1 < IgG2 < IgG4). Interestingly, given the heavy, life-long antigen burden experienced by PNG villagers, average mutation levels in IgG sequences were little different to that seen in Australian IgG1 sequences (19.2). Patterns of mutation provide clear evidence of antigen selection in many IgG sequences. The percentage of IgG sequences that showed significant accumulations of replacement mutations in the complementarity determining regions ranged from 22% of IgG3 sequences to 39% of IgG2 sequences. By contrast, only 12% of IgE sequences had such evidence of antigen selection, and this was significantly less than in PNG IgG1, IgG2 and IgG4 subclass sequences (P < 0.01). The anti-parasite IgE response therefore has the reduced evidence of antigen selection that has previously been reported in studies of IgE sequences from allergic individuals.
BACKGROUND:Clonal expansion of B lymphocytes coupled with somatic mutation and antigen selection allow the mammalian humoral immune system to generate highly specific immunoglobulins (IG) or antibodies against invading bacteria, viruses and toxins. The availability of high-throughput DNA sequencing methods is providing new avenues for studying this clonal expansion and identifying the factors guiding the generation of antibodies. The identification of groups of rearranged immunoglobulin gene sequences descended from the same rearrangement (clonally-related sets) in very large sets of sequences is facilitated by the availability of immunoglobulin gene sequence alignment and partitioning software that can accurately predict component germline gene, but has required painstaking visual inspection and analysis of sequences.RESULTS:We have developed and implemented an algorithm for identifying sets of clonally-related sequences in large human immunoglobulin heavy chain gene variable region sequence sets. The program processes sequences that have been partitioned using iHMMune-align, and uses pairwise comparisons of CDR3 sequences and similarity in IGHV and IGHJ germline gene assignments to construct a distance matrix. Agglomerative hierarchical clustering is then used to identify likely groups of clonally-related sequences. The program is available for download from http://www.cse.unsw.edu.au/~ihmmune/ClonalRelate/ClonalRelate.zip.CONCLUSIONS:The method was evaluated on several benchmark datasets and provided a more accurate and considerably faster identification of clonally-related immunoglobulin gene sequences than visual inspection by domain experts.
We describe a bioinformatic analysis of germline and rearranged immunoglobulin kappa chain (IGK) gene sequences, performed in order to assess the completeness and reliability of the reported IGK repertoire. In contrast to the reported heavy-chain gene repertoire, which includes many dubious sequences, only five IGK variable gene (IGKV) alleles appear to have been reported in error. There was, however, insufficient evidence to justify removing these IGKV genes from the germline repertoire. Bioinformatic analysis of apparent mismatches between reported germline genes and 1,863 expressed IGK sequences suggested the existence of two unreported IGKV polymorphisms. Genomic screening of 12 individuals led to the confirmation of both of these polymorphisms, IGKV1-16*02 and IGKV2-30*02. We also show that in contrast to the heavy chain, the IGK repertoire is dominated by sequences that use just a handful of kappa variable (IGKV) and junction (IGKJ) gene pairs. There is also little modification of IGKV and IGKJ genes by the processes of exonuclease removal and N nucleotide addition. The expressed IGK repertoire therefore lacks diversity and the junction region is particularly constrained. Remarkably, the analysis of a dataset of 435 relatively unmutated rearranged kappa genes showed that ten amino acid sequences account for almost 10% of the rearrangements, with identical sequences being derived from as many as seven independent sources. Such dominant sequences are likely to have important roles in the operation of the humoral immune response.
The identification of the genes that make up rearranged immunoglobulin genes is critical to many studies. For example, the enumeration of mutations in immunoglobulin genes is important for the prognosis of chronic lymphocytic leukemia, and this requires the accurate identification of the germline genes from which a particular sequence is derived. The immunoglobulin heavy-chain variable (IGHV) gene repertoire is generally considered to be highly polymorphic. In this report, we describe a bioinformatic analysis of germline and rearranged immunoglobulin gene sequences which casts doubt on the existence of a substantial proportion of reported germline polymorphisms. We report a five-level classification system for IGHV genes, which indicates the likelihood that the genes have been reported accurately. The classification scheme also reflects the likelihood that germline genes could be incorrectly identified in mutated VDJ rearrangements, because of similarities to other alleles. Of the 226 IGHV alleles that have previously been reported, our analysis suggests that 104 of these alleles almost certainly include sequence errors, and should be removed from the available repertoire. The analysis also highlights the presence of common mismatches, with respect to the germline, in many rearranged heavy-chain sequences, suggesting the existence of twelve previously unreported alleles. Sequencing of IGHV genes from six individuals in this study confirmed the existence of three of these alleles, which we designate IGHV3-49 * 04, IGHV3-49 * 05 and IGHV4-39 * 07. We therefore present a revised repertoire of expressed IGHV genes, which should substantially improve the accuracy of immunoglobulin gene analysis.
The identification of the genes that make up rearranged immunoglobulin genes is critical to many studies. For example, the enumeration of mutations in immunoglobulin genes is important for the prognosis of chronic lymphocytic leukemia, and this requires the accurate identification of the germline genes from which a particular sequence is derived. The immunoglobulin heavy-chain variable (IGHV) gene repertoire is generally considered to be highly polymorphic. In this report, we describe a bioinformatic analysis of germline and rearranged immunoglobulin gene sequences which casts doubt on the existence of a substantial proportion of reported germline polymorphisms. We report a five-level classification system for IGHV genes, which indicates the likelihood that the genes have been reported accurately. The classification scheme also reflects the likelihood that germline genes could be incorrectly identified in mutated VDJ rearrangements, because of similarities to other alleles. Of the 226 IGHV alleles that have previously been reported, our analysis suggests that 104 of these alleles almost certainly include sequence errors, and should be removed from the available repertoire. The analysis also highlights the presence of common mismatches, with respect to the germline, in many rearranged heavy-chain sequences, suggesting the existence of twelve previously unreported alleles. Sequencing of IGHV genes from six individuals in this study confirmed the existence of three of these alleles, which we designate IGHV3-49*04, IGHV3-49*05 and IGHV4-39*07. We therefore present a revised repertoire of expressed IGHV genes, which should substantially improve the accuracy of immunoglobulin gene analysis.