MICA is a stress-induced ligand of the NKG2D receptor that stimulates NK and T cell responses and was identified as a key determinant of anti-tumor immunity. The MICA gene is located inside the MHC complex and is in strong linkage disequilibrium with HLA-B. While an HLA-B*48-linked MICA deletion-haplotype was previously described in Asian populations, little is known about other MICA copy number variations. Here, we report the genotyping of more than two million individuals revealing high frequencies of MICA duplications (1%) and MICA deletions (0.4%). Their prevalence differs between ethnic groups and can rise to 2.8% (Croatia) and 9.2% (Mexico), respectively. Targeted sequencing of more than 70 samples indicates that these copy number variations originate from independent nonallelic homologous recombination events between segmental duplications upstream of MICA and MICB. Overall, our data warrant further investigation of disease associations and consideration of MICA copy number data in oncological study protocols.
The Genotype List (GL) String grammar for reporting HLA and Killer‐cell Immunoglobulin‐like Receptor (KIR) genotypes in a text string was described in 2013. Since this initial description, GL Strings have been used to describe HLA and KIR genotypes for more than 40 million subjects, allowing these data to be recorded, stored and transmitted in an easily parsed, text‐based format. After a decade of working with HLA and KIR data in GL String format, with advances in HLA and KIR genotyping technologies that have fostered the generation of full‐gene sequence data, the need for an extension of the GL String system has become clear. Here, we introduce the new GL String delimiter “?,” which addresses the need to describe ambiguity in assigning a gene sequence to gene paralogs. GL Strings that do not include a “?” delimiter continue to be interpreted as originally described. This extension represents version 1.1 of the GL String grammar.
A catalog of common, intermediate and well-documented (CIWD) HLA-A, -B, -C, -DRB1, -DRB3, -DRB4, -DRB5, -DQB1 and -DPB1 alleles has been compiled from over 8 million individuals using data from 20 unrelated hematopoietic stem cell volunteer donor registries. Individuals are divided into seven geographic/ancestral/ethnic groups and data are summarized for each group and for the total population. P (two-field) and G group assignments are divided into one of four frequency categories: common (≥1 in 10 000), intermediate (≥1 in 100 000), well-documented (≥5 occurrences) or not-CIWD. Overall 26% of alleles in IPD-IMGT/HLA version 3.31.0 at P group resolution fall into the three CIWD categories. The two-field catalog includes 18% (n = 545) common, 17% (n = 513) intermediate, and 65% (n = 1997) well-documented alleles. Full-field allele frequency data are provided but are limited in value by the variations in resolution used by the registries. A recommended CIWD list is based on the most frequent category in the total or any of the seven geographic/ancestral/ethnic groups. Data are also provided so users can compile a catalog specific to the population groups that they serve. Comparisons are made to three previous CWD reports representing more limited population groups. This catalog, CIWD version 3.0.0, is a step closer to the collection of global HLA frequencies and to a clearer view of HLA diversity in the human population as a whole.
The impact of the highly polymorphic Killer-cell immunoglobulin-like receptor (KIR) gene cluster on the outcome of hematopoietic stem cell transplantation (HCST) is subject of current research. To further understand the involvement of this gene family into Natural Killer (NK) cell-mediated graft-versus-leukemia reactions, knowledge of haplotype structures, and allelic linkage is of importance. In this analysis, we estimate population-specific KIR haplotype frequencies at allele group resolution in a cohort of n = 458 German families. We addressed the polymorphism of the KIR gene complex and phasing ambiguities by a combined approach. Haplotype inference within first-degree family relations allowed us to limit the number of possible diplotypes. Structural restriction to a pattern set of 92 previously described KIR copy number haplotypes further reduced ambiguities. KIR haplotype frequency estimation was finally accomplished by means of an expectation-maximization algorithm. Applying a resolution threshold of ½ n, we were able to identify a set of 551 KIR allele group haplotypes, representing 21 KIR copy number haplotypes. The haplotype frequencies allow studying linkage disequilibrium in two-locus as well as in multi-locus analyses. Our study reveals associations between KIR haplotype structures and allele group frequencies, thereby broadening our understanding of the KIR gene complex.
MICA and MICB are ligands of the NKG2D receptor and thereby influence NK and T cell activity. MICA/B gene polymorphisms, expression levels and the amount of soluble MICA/B in the serum have been linked to autoimmune diseases, infections, and cancer. In hematopoietic stem cell transplantation, MICA matching between donor and patient has been correlated with reduced acute and chronic graft-vs.-host disease and improved survival. Hence, we developed an extremely cost-efficient high-throughput workflow for genotyping MICA/B for newly registered potential stem cell donors. Since mid-2017, we have genotyped over two million samples using NGS amplicon sequencing for MICA/B exons 2–5. In donors of German origin, MICA*008 is the most common MICA allele with a frequency of 42.3%. It is followed by MICA*002 (11.7%) and MICA*009 (8.8%). The three most common MICB alleles are MICB*005 (43.9%), MICB*004 (21.7%), and MICB*002 (18.9%). In general, MICB is less diverse than MICA and only 6 alleles, instead of 15, account for a cumulative allele frequency of 99.5%. In 0.5% of the samples we observed at least one allele of MICA or MICB which has so far not been reported to the IPD/IMGT-HLA database. By providing MICA/B typed voluntary donors, clinicians become empowered to include MICA/B into their donor selection process to further improve unrelated hematopoietic stem cell transplantation.
HLA-E has been reported to impact the outcome of hematopoietic stem cell transplantations (HSCT). Starting July 2017, DKMS added HLA-E to the default typing profile for newly enrolled potential donors for HSCT. Since then, DKMS Germany has typed more than 390,000 samples for HLA-E together with the HLA loci A, B, C, DRB1, DQB1, and DPB1. In this work, we present 7-locus haplotype and allele frequencies for donors from Germany (N = 325,955), Turkey (N = 10,649), Poland (N = 4144), Russia (N = 2896) and Italy (N = 2035). HLA-E is typed by our well-establish high-throughput approach relying on an amplicon spanning exons 2 and 3. Amplicons are sequenced by Illumina MiSeq or HiSeq 2500 instruments without fragmentation, preserving phase information. Reads are analysed by our in–house typing software neXtype. Typing results comprise the high-resolved alleles HLA-E*01:01:02, 01:03:03, 01:04, 01:05, 01:07 and 01:08 N and three G-groups: HLA-E*01:01:01G, 01:03:01G and 01:03:02G. Upon enrollment with DKMS Germany, potential stem cell donors provide information on their self-assessed ethnic background by statement of their respective country of origin. Using our open-source software Hapl-o-Mat, we computed frequencies on g-group resolution level for both, alleles and 7-loci haplotypes. The most common allele group in all five populations is HLA-E*01:01 g with a fairly similar frequency (fGermany = 55.5%, fTurkey = 53.4%, fPoland = 56.9%, fRussia = 57.0%, fItaly = 53.0%). Frequencies for HLA-E*01:03 g are (fGermany = 44.4%, fTurkey = 46.6%, fPoland = 43.1%, fRussia = 43.0%, fItaly = 46.9%). In all populations, except Russia, HLA-E*01:05 is the third most common allele. In the Russian population only the two dominant alleles were observed. Our approach provided reliable HLA-E typing data for a large number of samples. In all considered populations, identical allele groups are dominant with similar frequencies.
The MICA and MICB molecules serve as ligands of the activating NKG2D receptor expressed on natural killer (NK) cells. Recent studies indicate effects of MIC genotypes on hematopoietic stem cell transplantation (HSCT) outcome. To provide MIC allele information for donor selection, we developed an NGS-based high-throughput genotyping workflow for MICA and MICB and applied it for genotyping registry donor samples. Exons 2 and 3, as well as most of exons 4 and 5 of MICA and MICB are amplified in a multiplexed PCR reaction. The PCR products are sequenced on Illumina HiSeq or MiSeq instruments. The data are processed by an updated neXtype software version to provide allele-level genotyping information. Using this NGS based workflow, we genotyped 350,000 donors registered in Germany and report on the observed MICA allele frequencies. Due to the restricted sequence coverage, 9 alleles encoding distinct proteins cannot be resolved. However, the unprecedented depth of the study allowed us to estimate allele frequencies for 49 of the 84 described MICA alleles distinguished at the protein coding level. In addition we identified novel alleles in 0.2% of the samples. The 13 (31) most abundant alleles account for a cumulative allele frequency of 99% (99.99%). This newly developed NGS-based genotyping approach offers the opportunity to analyze the genetic diversity of MICA and MICB in large cohorts at high-resolution.
The killer-cell immunoglobulin-like receptor (KIR) genes regulate natural killer cell activity, influencing predisposition to immune mediated disease, and affecting hematopoietic stem cell transplantation (HSCT) outcome. Owing to the complexity of the KIR locus, with extensive gene copy number variation (CNV) and allelic diversity, high-resolution characterization of KIR has so far been applied only to relatively small cohorts. Here, we present a comprehensive high-throughput KIR genotyping approach based on next generation sequencing. Through PCR amplification of specific exons, our approach delivers both copy numbers of the individual genes and allelic information for every KIR gene. Ten-fold replicate analysis of a set of 190 samples revealed a precision of 99.9%. Genotyping of an independent set of 360 samples resulted in an accuracy of more than 99% taking into account consistent copy number prediction. We applied the workflow to genotype 1.8 million stem cell donor registry samples. We report on the observed KIR allele diversity and relative abundance of alleles based on a subset of more than 300,000 samples. Furthermore, we identified more than 2,000 previously unreported KIR variants repeatedly in independent samples, underscoring the large diversity of the KIR region that awaits discovery. This cost-efficient high-resolution KIR genotyping approach is now applied to samples of volunteers registering as potential donors for HSCT. This will facilitate the utilization of KIR as additional selection criterion to improve unrelated donor stem cell transplantation outcome. In addition, the approach may serve studies requiring high-resolution KIR genotyping, like population genetics and disease association studies.
The killer cell immunoglobulin-like receptor (KIR) gene family is seen to play an important role in unrelated hematopoietic stem cell transplantations. Allele level KIR genotyping may improve the donor selection process. The IPD-KIR database currently comprises 753 alleles. With our newly established high-throughput workflow for KIR allele-level typing, we discovered a vast number of previously unknown variations of KIR alleles. We perform cost efficient KIR allele level genotyping based on a paired-end short-amplicon approach by NGS using Illumina HiSeq 2500 instruments. The amplicons cover exons 3, 4, 5, 7, 8 and 9 of each KIR gene. Our neXtype software was adapted to report KIR genotyping results at allelic Level. Starting in October 2016, we added KIR allele-level typing to the DKMS standard genotyping profile for new potential donors and already applied it to more than 500,000 samples. Since then, we have detected more than 5000 distinct novel SNP sequences, see Fig. 1, i.e. sequences that differ from the reported alleles by at least one nucleotide. About half of the novel sequences exhibit new protein sequences. In addition, we identified many alleles with so far unreported combinations of described exon sequences. We analyzed the distribution of new variants according to ethnic self-assessment of recruited donors. The ability to genotype KIR genes at allelic level in a high-throughput framework gives the opportunity to detect a huge number of unknown sequences on a rather short time scale. Consequently, the database of currently named alleles can be extended significantly and thus gives input to studies around the complex KIR gene Family.
The human killer-cell immunoglobulin-like receptor (KIR) family of genes is a key regulator of natural killer cell activity. Several studies have indicated that KIR genotypes affect hematopoietic stem cell transplantation outcome and the evidence is rising that the extensive allelic diversity discovered in KIR genes may be an important component to consider. However, high-resolution characterization of KIR genotypes has so far been applied only to smaller cohorts. Therefore, we developed a KIR genotyping approach based on NGS that would be cost-effective and amendable for high-throughput application. We amplified KIR exons 3, 4, 5, 7, 8 and 9, targeting one to two exons per PCR reaction but multiplexed across all KIR genes. Amplicons of 2 × 3760 samples were combined for sequencing on Illumina HiSeq 2500 instruments together with amplicons for HLA and blood groups. The reads were mapped against the described KIR alleles (IPD-KIR Release 2.6). Based on the read coverage, we estimated the genomic copy number of each sequence feature. These sequence copy calculations enabled precise allele calling and in addition empowered us to determine gene copy numbers. After successful validation, we applied this workflow to genotype 500,000 registry samples. Despite the focus on throughput and cost efficiency, we achieved mostly allotype (3 digit) resolution with the exception of certain allotypes differing only in the short non-targeted exon regions. In addition, certain phasing ambiguities remained. We report the observed allele diversity and the relative abundance of alleles for a mainly Caucasian cohort. We demonstrate that high-resolution KIR genotyping is feasible using a very cost-efficient workflow. This approach enables KIR genotyping for population genetics, disease association studies and other applications requiring large cohorts. As of October 2016 we have been applying this workflow to the analysis of all volunteer samples registering with DKMS as potential donors for hematopoietic stem cell transplantation. High-resolution KIR genotyping results will thus become available to search coordinators and may soon be used as an additional selection criterion to improve transplantation outcome.
In 2013, we implemented a short amplicon based Next Generation Sequencing (NGS) high-throughput workflow for HLA donor registry typing. Data analysis is performed by the in-house typing software neXtype. Currently, over 30,000 samples per week can be typed for 6 HLA loci, ABO, RhD, CCR5, and KIR with this cost-efficient workflow in our lab. As HLA-E has been reported to impact the outcome of stem cell transplantations in certain settings, we evaluated the feasibility of low-cost high-throughput HLA-E typing. On a clinically relevant antigen recognition domain (ARD) level, there are currently 9 groups of HLA-E alleles known (Fig. 1). HLA-E is amplified by PCR with two primers spanning exons 2 and 3. This amplification product is subject to sequencing on Illumina MiSeq or HiSeq 2500 instruments without fragmentation. Thereby, information about the phasing between exon 2 and exon 3 is preserved, which is an important advantage of this method since it allows genotyping at ARD-level resolution. As shown in Fig. 1, alleles HLA-E∗01:01:02, 01:03:03, 01:04, 01:05, 01:07 and 01:08N can be resolved to the allele level. The remaining alleles fall into three distinguishable G-groups, namely HLA-E∗01:01:01G, 01:03:01G and 01:03:02G. The primer set was tested using 384 samples. All successfully sequenced samples corresponded to the G-groups HLA-E∗01:01:01G, 01:03:01G and 01:03:02G. Validation of the workflow against pre-typed samples is projected. Sequencing HLA-E on allelic level has been show feasible in a cost efficient high throughput workflow. As the impact of HLA-E on the transplantation outcome becomes more settled, it could be considered for inclusion into the recruitment typing profile.Download : Download high-res image (244KB)Download : Download full-size image