Disparities in survival after allogeneic hematopoietic cell transplantation have been reported for some race and ethnic groups, despite comparable HLA matching. Individuals' ethnic and race groups, as reported through self-identification, can change over time because of multiple sociological factors. We studied the effect of 2 measures of genetic similarity in 1378 recipients who underwent myeloablative first allogeneic hematopoietic cell transplantation between 1995 and 2011 and their unrelated 10 of 10 HLA-A, -B, -C, -DRB1, and-DQB1- matched donors. The studied factors were as follows (1) donor and recipient genetic ancestral admixture and (2) pairwise donor/recipient genetic distance. Increased African genetic admixture for either transplant recipients or donors was associated with increased risk of overall mortality (hazard ratio [HR], 2.26; P = .005 and HR, 3.09; P = .0002, respectively) and transplant-related mortality (HR, 3.3; P = .0003 and HR, 3.86; P = .0001, respectively) and decreased disease-free survival (HR, 1.9; P = .02 and HR, 2.46; P = .002 respectively). The observed effect, albeit statistically significant, was relevant to a small subset of the studied population and was notably correlated with self-reported African-American race. We were not able to control for other nongenetic factors, such as access to health care or other socioeconomic factors; however, the results suggest the influence of a genetic driver. Our findings confirm what has been previously reported for African-American recipients and show similar results for donors. No significant association was found with donor/recipient genetic distance.
Genetic interactions have been reported to underlie phenotypes in a variety of systems, but the extent to which they contribute to complex disease in humans remains unclear. In principle, genome-wide association studies (GWAS) provide a platform for detecting genetic interactions, but existing methods for identifying them from GWAS data tend to focus on testing individual locus pairs, which undermines statistical power. Importantly, a global genetic network mapped for a model eukaryotic organism revealed that genetic interactions often connect genes between compensatory functional modules in a highly coherent manner. Taking advantage of this expected structure, we developed a computational approach called BridGE that identifies pathways connected by genetic interactions from GWAS data. Applying BridGE broadly, we discover significant interactions in Parkinson’s disease, schizophrenia, hypertension, prostate cancer, breast cancer, and type 2 diabetes. Our novel approach provides a general framework for mapping complex genetic networks underlying human disease from genome-wide genotype data.
Four single nucleotide polymorphism (SNP)-based human leukocyte antigen (HLA) imputation methods (e-HLA, HIBAG, HLA*IMP:02 and MAGPrediction) were trained using 1000 Genomes SNP and HLA genotypes and assessed for their ability to accurately impute molecular HLA-A, -B, -C and -DRB1 genotypes in the Human Genome Diversity Project cell panel. Imputation concordance was high (>89%) across all methods for both HLA-A and HLA-C, but HLA-B and HLA-DRB1 proved generally difficult to impute. Overall, < 27.8% of subjects were correctly imputed for all HLA loci by any method. Concordance across all loci was not enhanced via the application of confidence thresholds; reliance on confidence scores across methods only led to noticeable improvement (+3.2%) for HLA-DRB1. As the HLA complex is highly relevant to the study of human health and disease, a standardized assessment of SNP-based HLA imputation methods is crucial for advancing genomic research. Considerable room remains for the improvement of HLA-B and especially HLA-DRB1 imputation methods, and no imputation method is as accurate as molecular genotyping. The application of large, ancestrally diverse HLA and SNP reference data sets and multiple imputation methods has the potential to make SNP-based HLA imputation methods a tractable option for determining HLA genotypes.
Disparities in survival after allogeneic hematopoietic cell transplantation have been reported for some race and ethnic groups, despite comparable HLA matching. Individuals' ethnic and race groups, as reported through self-identification, can change over time because of multiple sociological factors. We studied the effect of 2 measures of genetic similarity in 1378 recipients who underwent myeloablative first allogeneic hematopoietic cell transplantation between 1995 and 2011 and their unrelated 10 of 10 HLA-A, -B, -C, -DRB1, and-DQB1-matched donors. The studied factors were as follows (1) donor and recipient genetic ancestral admixture and (2) pairwise donor/recipient genetic distance. Increased African genetic admixture for either transplant recipients or donors was associated with increased risk of overall mortality (hazard ratio [HR], 2.26; P=.005 and HR, 3.09; P=.0002, respectively) and transplant-related mortality (HR, 3.3; P=.0003 and HR, 3.86; P=.0001, respectively) and decreased disease-free survival (HR, 1.9; P=.02 and HR, 2.46; P=.002 respectively). The observed effect, albeit statistically significant, was relevant to a small subset of the studied population and was notably correlated with self-reported African-American race. We were not able to control for other nongenetic factors, such as access to health care or other socioeconomic factors; however, the results suggest the influence of a genetic driver. Our findings confirm what has been previously reported for African-American recipients and show similar results for donors. No significant association was found with donor/recipient genetic distance. (C) 2017 American Society for Blood and Marrow Transplantation.
Unrelated stem cell registries have been collecting HLA typing of volunteer bone marrow donors for over 25years. Donor selection for hematopoietic stem cell transplantation is based primarily on matching the alleles of donors and patients at five polymorphic HLA loci. As HLA typing technologies have continually advanced since the beginnings of stem cell transplantation, registries have accrued typings of varied HLA typing ambiguity. We present a new typing resolution score (TRS), based on the likelihood of self-match, that allows the systematic comparison of HLA typings across different methods, data sets and populations. We apply the TRS to chart improvement in HLA typing within the Be The Match Registry of the United States from the initiation of DNA-based HLA typing to the current state of high-resolution typing using next-generation sequencing technologies. In addition, we present a publicly available online tool for evaluation of any given HLA typing. This TRS objectively evaluates HLA typing methods and can help define standards for acceptable recruitment HLA typing.
Standard measures of linkage disequilibrium (LD) provide an incomplete description of the correlation between two loci. Recently, Thomson and Single (2014) described a new asymmetric pair of LD measures (ALD) that give a more complete description of LD. The ALD measures are symmetric and equivalent to the correlation coefficient r when both loci are bi-allelic. When the numbers of alleles at the two loci differ, the ALD measures capture this asymmetry and provide additional detail about the LD structure. In disease association studies the ALD measures are useful for identifying additional disease genes in a genetic region, by conditioning on known effects. In evolutionary genetic studies ALD measures provide insight into selection acting on individual amino acids of specific genes, or other loci in high LD (see Thomson and Single (2014) for these examples). Here we describe new software for computing and visualizing ALD. We demonstrate the utility of this software using haplotype frequency data from the National Marrow Donor Program (NMDP). This enhances our understanding of LD patterns in the NMDP data by quantifying the degree to which LD is asymmetric and also quantifies this effect for individual alleles.
Aim Survival after hematopoietic cell transplantation (HCT) is dependent on donor/recipient (D/R) HLA matching. However disparities in survival were reported for some ethnicities despite comparable HLA matching. Individual ethnicities/races, as reported through self-identification, can change over time. Most studies have shown that African-American recipients (AAFA race) experience worse survival. Another way to investigate ancestry is to use Ancestry Informative Marker SNPs (AIMs), providing ancestral admixture. We hypothesized that information on donor and recipient genetic admixture may be used to evaluate D/R genetic disparity and that this may be associated with outcomes of HLA matched unrelated donor HCTs. Methods Study population included 1295 10/10 HLA matched D/R pairs receiving HCT for AML, ALL, CML and MDS between 1995 and 2011. Samples were genotyped for 500 AIMs. We estimated African (AFR), European (EUR), Asian (ASI) and South European/Amerindian (SE/A) admixtures for donors and recipients using STRUCTURE at K=4 clusters. Tables 1 and 2 show the admixture distributions in each of the self-identified race groups. To model D/R genetic disparity we ran principal components analysis (PCA) on the D/R genotypes, then calculated the pairwise Euclidean distance between a subset of PCA eigenvectors of each D/R duo. Multivariate analyses were performed using Cox proportional hazards models for overall survival (OS), disease free survival (DFS), relapse, transplant related mortality (TRM), acute and chronic graft versus host disease (GVHD) for admixture and genetic disparity. Results For transplant recipients , increasing AFR admixture was associated with worse OS and TRM at p donors, increasing AFR admixture was associated with worse OS, DFS and TRM at p donor ASI, EUR and SE/A. We tested for a cut point for AFR admixture that best associated with survival. For recipients, the optimal cut point was > 14% AFR admixture. This only included 2.8% of the population (N=34 recipients) but 90% of the African-American self-identified recipients. For donors the cut point was > 23% AFR admixture which only included 1.9% of the population (N=24 donors) but 89% of African-American self-identified donors. Recipients and donors with high AFR admixture were highly confounded and numbers were too small to explore whether recipient or donor AFR admixture was more important. No significant associations were observed for D/R pairwise genetic distance and clinical outcomes. Conclusion While no significant associations were found for the D/R genetic distance used here, increasing AFR admixture in recipient and donor associated with increased risk of overall mortality, TRM and decreased DFS. The observed effect was attributed to 2.8% of the recipient and 1.9% of donor population with higher AFR admixture, groups that contained >89% of the self-identified African-Americans. Our findings are consistent with studies showing that self-identified African-Americans are associated with suboptimal HCT outcomes. It is still unclear whether the deleterious effect of AFR admixture is genetic in nature or due to other non-genetic factors such as access to healthcare or other socio-economic factors mostly pertinent to recipients. Disclosures Majhail: Gamida Cell Ltd.: Consultancy; Anthem Inc.: Consultancy. Lee: Kadmon: Consultancy; Bristol-Myers Squibb: Consultancy.
As clinical matching definitions for transplantation evolve there is a need to quantify the degree of uncertainty in a typing result relative to a particular resolution target. We aimed to develop such a measure and then applied it to quantifying the improvement in typing resolution of the Be The Match® Registry since 1993. We describe and apply a new typing ambiguity score, based on the likelihood of self-match which allows for comparison of HLA typings across different methods, data sets and populations. In order to compute the typing ambiguity score, we perform HLA genotype imputation on 14 million donors using high-resolution haplotype frequencies generated from unrelated donors from the National Marrow Donor Program database for 5 population categories. For this experiment the resolution target was 5-locus ARS exons with equivalent amino acid sequence. The Registry has seen an increasing trend in the score over time for all populations and all HLA loci, with the overall scores for recruitment typing rising from 0.31 in 1993 to 0.96 in 2015. Many discontinuities in scores coincide with changes in recruitment typing policy, such as the transition of HLA-A and B typing from serology to DNA, the start of sequence based typing, the inclusion of HLA-C and DQB1 at recruitment, etc. We find evidence that oligo-based kits were tuned to reduce typing ambiguity primarily for majority race/ethnic groups as European American donors generally had the most rapid increase in scores, while African American donors have the lowest scores and a slower increase in scores. Finally, we show that new recruitment HLA typing performed for the US registry today has very little ambiguity under the current standard of matching at HLA-A, C, B, DRB1, and DQB1, using the current laboratory methods that employ next-generation sequencing or sequence-based typing with panels of group-specific sequencing primers. Our typing ambiguity score objectively measured the improvement in HLA typing within the US registry from 1993 to the current state of high-resolution typing. We next aim to assess ambiguity among global registries in BMDW. This method is general and can be applied to other loci (e.g. DPB1, DPA1, DQA1) other systems (KIR) and other definitions of allele (all-exons or full-gene).
The Human Leukocyte Antigen (HLA) genes are some of the most studied genes on the genome. This is due to their importance in bone marrow and solid organ transplantation, as well as their strong associations with many autoimmune, infectious, and inflammatory diseases. As such, they can be a highly valuable asset to clinicians and researchers for elucidating biological mechanism that may drive those diseases. The extraordinary genetic polymorphism that exists in this region makes it very challenging to type. Therefore, several approaches were proposed for prediction of HLA genes from widely available genome-wide single nucleotide polymorphism (SNP) data sets in the attempt to reduce cost and utilize existing data. These methods use SNPs and high-resolution training HLA data to build models for prediction of HLA genes in new samples. However, most of the existing HLA data sets are not available in high-resolution (exact allele assignment) but contain allelic ambiguities (inexact allele assignments). This is a result of existing typing methodologies not always being able to distinguish between several possible alleles at a given gene and produce ambiguous allele as a result. Current approaches for prediction of HLA genes from SNP data do not accommodate learning from ambiguous HLA data and, as such, miss the potential for an increased sample size and consequently improvements in prediction performance. In this paper, we propose Amb-EM, a novel algorithm for SNP-based prediction of HLA genes that utilizes ambiguities in the HLA data and predicts high-resolution alleles using ambiguous HLA alleles for building the model. Additionally, we measure the impact that the uncertainty in the training data has on the prediction accuracy, and evaluate it on a real world data set. Our results show that the prediction from ambiguous HLA data outperforms the alternative approach which first imputes the ambiguous data into high-resolution HLA alleles and uses it to build the model.
Aim Using Shannon’s entropy, we objectively measured the HLA typing ambiguity of various commercial Sequence-Specific Oligonucleotide (SSO) protocols (groups of SSO kits used in tandem). We previously compared the HLA typing ambiguity obtained by serology, allele family level DNA-based typing, SSO, and single-pass sequence based typing (SBT). However, commercial companies have made significant advances in SSO technology in recent years, and those results only evaluated early generic versions of SSO kits. Methods For several US populations, and for 5 HLA loci relevant to transplantation (HLA-A, -B, -C, -DRB1, -DQB1), we identify a set of commonly used SSO protocols. We use population-level haplotype frequency data to generate a cohort of simulated subjects, followed by 5-locus genotype imputation to generate a list of genotypes that would result from SSO typing each simulated subject. The imputation step computes the relative likelihood of each genotype, which is then used for entropy calculation. Results Distribution of entropies across populations and for each locus is generally similar, while the magnitudes differ. For example, entropy for locus HLA-A in African American population reaches 0.19, while the maximum entropy in Caucasians is 0.055. We generally observed that protocols comprised of newer kits had lower ambiguity across all populations. In certain populations, some newer kits had higher entropy than older ones. This may indicate those kits were designed to detect alleles uncommonly found in those populations, but resolving ambiguity in other populations. Conclusions Ambiguity in SSO typing kits has reduced substantially since their introduction in the mid-1990s. Results also indicate a European focus toward selection of additional probes, which is suboptimal for typing diverse populations. We hope to use Shannon’s entropy next to evaluate group-specific sequence primer strategies with SBT. These results will guide selection of HLA typing methods by researchers and laboratories.
The Human Leukocyte Antigen (HLA) gene system plays a crucial role in hematopoietic stem cell transplantation, where patients and donors are matched with respect to their HLA genes in order to maximize the chances of a successful transplant. It is the most polymorphic region of the human genome with some of the strongest associations with autoimmune, infectious, and inflammatory diseases. The availability of HLA data is, therefore, of high importance to clinicians and researchers. However, due to its high polymorphism, obtaining it is time- and cost-prohibitive. We previously described a method for the prediction of HLA genes from widely available Single Nucleotide Polymorphism (SNP) data. In this paper we show that using HLA gene dependency information improves prediction performance on multiple real-world data sets. More specifically, we propose and evaluate different approaches for integrating HLA gene dependency into the prediction process. The results from experiments on two real data sets show that adding dependency information is a valuable asset for HLA gene prediction, particularly for smaller data sets.
Discriminative pattern mining seeks patterns that are more prevalent in one class than another and provide good classification accuracy for the objects in which the patterns occur. A number of approaches have been proposed for finding such patterns, which are also known under a variety of names, e.g., contrast sets and emerging patterns. However, fundamental questions about the nature and limits of such patterns remain unanswered. For instance, a discriminative pattern is only interesting if it provides better discriminative power than any of its subpatterns, but it is not obvious, for example, how much additional discriminative power can be provided by a pattern over and above the discriminative power of its subpatterns. Also, what do the patterns that provide the most additional discrimination look like? And, what is the relationship of different measures for discrimination (e.g., mutual information and DiffSup, the difference of the supports in the two classes) In previous work, we made an initial attempt at analyzing the first two questions. In this paper we present several new developments. Specifically, we present a more elegant and efficient formulation of the problem of determining the best discriminative pattern that can be obtained for a particular number of variables. We also explore for the first time the limits of patterns that go beyond the ‘and’ logic of traditional pattern mining, e.g., patterns based on the logic of ‘or’, ‘n of k’ or ‘majority wins’. We show that the discriminative advantage of ‘and’ based patterns over their subpatterns is more limited than that of some of the other patterns, and hence, these patterns may represent a potential area for future development of discriminative pattern mining. Finally, we explore the relationship of various measures of discriminative pattern mining. We show that our results, although based on one of the measures (DiffSup) have implications for mutual information. More generally, we identify a potential avenue of exploration, which although challenging, may offer the opportunity for making a more definitive and general statement about a certain class of discriminative measures.
There has been a dramatic increase in the quantity, quality, and types of advanced biomedical information available to individuals and their medical providers. These types of data include, but are not limited to, cell process information provided by DNA microarrays and RNA seq, genetic information in the form of Single Nucleotide Polymorphisms (SNPs), metabolomics data in terms of proteins and other metabolites, and structural and functional brain data from magnetic resonance imaging (MRI). Together with the increasing availability of clinical data from electronic medical records, this abundance of data has created the very real possibility of personalized medicine, i.e., using detailed biomedical, clinical, and environmental information about a person for a customized and more effective approach to patient care [11], [16], [3]. Achieving this goal requires identifying those features of the data that can distinguish not …