HLA-G encodes an immune checkpoint molecule with restricted expression in immune-privileged tissues and pathological conditions. It exhibits limited coding diversity but substantial regulatory-region variation influencing expression levels. Strong linkage disequilibrium across HLA-G creates a structured genetic architecture in which regulatory and coding variants co-segregate into well-defined haplotypes, enabling the imputation of complete HLA-G haplotypes from partial genomic data. We developed imputation models to predict HLA-G 4-field alleles, promoter, and 3'UTR haplotypes from whole-exome sequencing and SNP array data using HIBAG. Multi-ethnic reference panels were constructed from 5,347 individuals from three diverse cohorts (1000 Genomes, Human Genome Diversity Project, and Brazilian SABE cohort). Models were validated through cross-validation and independent datasets. Exome-based imputation achieved high accuracy (>99%) for common alleles (frequency > 1%), with mean posterior probabilities exceeding 0.95. SNP array-based models showed slightly lower but still robust performance (>95% accuracy). Our approach enables simultaneous prediction of coding and regulatory sequences, providing comprehensive functional information from datasets that do not capture the complete HLA-G diversity. These models facilitate HLA-G analysis in widely available genomic datasets lacking introns and regulatory regions (tumor exomes, SNP arrays), enabling investigation of HLA-G's role in immune regulation, transplantation, cancer, and pregnancy complications without full-gene sequencing.
Genomic studies of autism spectrum disorder (ASD) have largely excluded admixed populations. To address this gap, we characterized the genomic landscape of ASD in Brazil by combining a systematic literature review with whole-exome sequencing analysis of 441 Brazilian individuals and their families. Our analysis revealed a conclusive molecular diagnosis in 13.1% of probands. The diagnostic yield was higher among individuals with clinical features, particularly comorbid signs of intellectual disability, hypotonia, and seizures, providing a basis for prioritizing genetic testing. The sample presented a diverse ancestry, with major European, African, and Native American contributions. Notably, more than half of the identified rare risk variants were located on non-European haplotypes. Both de novo and inherited variants contributed to ASD risk, and we reinforce NPAS3 as a candidate ASD risk gene. This study provides the first comprehensive genomic overview of ASD in a large Brazilian cohort, reinforcing the critical need to include diversely admixed populations in genomic research to expand the understanding of ASD architecture and improve diagnostic strategies in resource-limited settings.
Killer cell immunoglobulin-like receptors (KIRs) regulate natural killer (NK) cell responses by activating or inhibiting their functions. Genotyping KIR genes from short-read second-generation sequencing data remains challenging as cross-alignments among genes and alignment failure arise from gene similarities and extreme polymorphism. Several bioinformatics pipelines and programs, including PING and T1K, have been developed to analyse KIR diversity. We found discordant results among tools in a systematic comparison using the same dataset. Additionally, they do not provide SNPs in the context of the reference genome, making them unsuitable for whole-genome association studies. Here, we present kir-mapper, a toolkit to analyse KIR genes from short-read sequencing, focusing on detecting KIR alleles, copy number variation, as well as SNPs and InDels in the context of the hg38 reference genome. kir-mapper can be used with whole-genome sequencing (WGS), whole-exome sequencing (WES) and sequencing data generated after probe-based capture methods. It presents strategies for phasing SNPs and InDels within and among genes, reducing the number of ambiguities reported by other methods. We have applied kir-mapper and other tools to data from various sources (WGS, WES) in worldwide samples and compared the results. Using long-read data as a truth set, we found that WGS kir-mapper analyses provided more accurate genotype calls than PING and T1K. For WES, kir-mapper provides more accurate genotype calls than T1K for some genes, particularly highly polymorphic ones (KIR3DL3 and KIR3DL2). This comparison highlights that the choice of method has to be considered as a function of the available data type and the targeted genes. kir-mapper is available at the GitHub repository (https://github.com/erickcastelli/kir-mapper/).
As genomics initiatives have spread around the world–often in the name of genetic diversity and inclusion–they have not only invoked promises of a medical revolution, but also revived categories of human difference that resemble erstwhile racial classifications. This is despite the fact that geneticists broadly dismissed racial categories as obsolete and unfounded after the Human Genome Project was completed in 2003. In fact, contemporary genomics initiatives have often ended up reinforcing ethnocentric and nativist conceptions of difference, drawing intense criticism from activists and critical social scientists. This roundtable brings leading population geneticists grappling with the question of genetic identity and ancestry, especially in the global South, together with some of the most prominent scholars of race in genomics. The result is an engaging and insightful dialogue on questions that have vexed the field for decades. How do we—indeed “can” we reconcile the boundaries of biological and social difference? How do notions of “genetic ancestry” and “biogeographical ancestry differ from erstwhile racial and ethnic categories? Can racial categories ever be shorn of their colonial and oppressive legacies? Here we scrutinise the methodological and epistemological frameworks in contemporary genomics that work to define populations and shape our understanding of biology, society, health, and disease. We seek to clarify perspectives across the disciplinary divide, and to advance constructive and grounded critiques that contend with the question of justice in genomics.
The MHC class I region contains crucial genes for the innate and adaptive immune response, playing a key role in susceptibility to many autoimmune and infectious diseases. Genome-wide association studies have identified numerous disease-associated SNPs within this region. However, these associations do not fully capture the immune-biological relevance of specific HLA alleles. HLA imputation techniques may leverage available SNP arrays by predicting allele genotypes based on the linkage disequilibrium between SNPs and specific HLA alleles. Successful imputation requires diverse and large reference panels, especially for admixed populations. This study employed a bioinformatics approach to call SNPs and HLA alleles in multi-ethnic samples from the 1000 genomes (1KG) dataset and admixed individuals from Brazil (SABE), utilising 30X whole-genome sequencing data. Using HIBAG, we created three reference panels: 1KG (n = 2504), SABE (n = 1171), and the full model (n = 3675) encompassing all samples. In extensive cross-validation of these reference panels, the multi-ethnic 1KG reference exhibited overall superior performance than the reference with only Brazilian samples. However, the best results were achieved with the full model. Additionally, we expanded the scope of imputation by developing reference panels for non-classical, MICA, MICB and HLA-H genes, previously unavailable for multi-ethnic populations. Validation in an independent Brazilian dataset showcased the superiority of our reference panels over the Michigan Imputation Server, particularly in predicting HLA-B alleles among Brazilians. Our investigations underscored the need to enhance or adapt reference panels to encompass the target population's genetic diversity, emphasising the significance of multiethnic references for accurate imputation across different populations.
Human leukocyte antigen (HLA) class I and II loci are essential elements of innate and acquired immunity. Their functions include antigen presentation to T cells leading to cellular and humoral immune responses, and modulation of NK cells. Their exceptional influence on disease outcome has now been made clear by genome-wide association studies. The exons encoding the peptide-binding groove have been the main focus for determining HLA effects on disease susceptibility/pathogenesis. However, HLA expression levels have also been implicated in disease outcome, adding another dimension to the extreme diversity of HLA that impacts variability in immune responses across individuals. To estimate HLA expression, immunogenetic studies traditionally rely on quantitative PCR (qPCR). Adoption of alternative high-throughput technologies such as RNA-seq has been hampered by technical issues due to the extreme polymorphism at HLA genes. Recently, however, multiple bioinformatic methods have been developed to accurately estimate HLA expression from RNA-seq data. This opens an exciting opportunity to quantify HLA expression in large datasets but also brings questions on whether RNA-seq results are comparable to those by qPCR. In this study, we analyze three classes of expression data for HLA class I genes for a matched set of individuals: (a) RNA-seq, (b) qPCR, and (c) cell surface HLA-C expression. We observed a moderate correlation between expression estimates from qPCR and RNA-seq for HLA-A, -B, and -C (0.2 ≤ rho ≤ 0.53). We discuss technical and biological factors which need to be accounted for when comparing quantifications for different molecular phenotypes or using different techniques.
The SNP-HLA Reference Consortium (SHLARC), a component of the 18th International HLA and Immunogenetics Workshop, is aimed at collecting diverse and extensive human leukocyte antigen (HLA) data to create custom reference panels and enhance HLA imputation techniques. Genome-wide association studies (GWAS) have significantly contributed to identifying genetic associations with various diseases. The HLA genomic region has emerged as the top locus in GWAS, particularly in immune-related disorders. However, the limited information provided by single nucleotide polymorphisms (SNPs), the hallmark of GWAS, poses challenges, especially in the HLA region, where strong linkage disequilibrium (LD) spans several megabases. HLA imputation techniques have been developed using statistical inference in response to these challenges. These techniques enable the prediction of HLA alleles from genotyped GWAS SNPs. Here we present the SHLARC activities, a collaborative effort to create extensive, and multi-ethnic reference panels to enhance HLA imputation accuracy.
The identification of genomic regions and genes that have evolved under natural selection is a fundamental objective in the field of evolutionary genetics. While various approaches have been established for the detection of targets of positive selection, methods for identifying targets of balancing selection, a form of natural selection that preserves genetic and phenotypic diversity within populations, have yet to be fully developed. Despite this, balancing selection is increasingly acknowledged as a significant driver of diversity within populations, and the identification of its signatures in genomes is essential for understanding its role in evolution. In recent years, a plethora of sophisticated methods has been developed for the detection of patterns of linked variation produced by balancing selection, such as high levels of polymorphism, altered allele-frequency distributions, and polymorphism sharing across divergent populations. In this review, we provide a comprehensive overview of classical and contemporary methods, offer guidance on the choice of appropriate methods, and discuss the importance of avoiding artifacts and of considering alternative evolutionary processes. The increasing availability of genome-scale datasets holds the potential to assist in the identification of new targets and the quantification of the prevalence of balancing selection, thus enhancing our understanding of its role in natural populations.
In his 1972 paper ‘The apportionment of human diversity’, Lewontin showed that, when averaged over loci, genetic diversity is predominantly attributable to differences among individuals within populations. However, selection can alter the apportionment of diversity of specific genes or genomic regions. We examine genetic diversity at the human leucocyte antigen (HLA) loci, located within the major histocompatibility complex (MHC) region. HLA genes code for proteins that are critical to adaptive immunity and are well-documented targets of balancing selection. The single-nucleotide polymorphisms (SNPs) within HLA genes show strong signatures of balancing selection on large timescales and are broadly shared among populations, displaying low FST values. However, when we analyse haplotypes defined by these SNPs (which define ‘HLA alleles’), we find marked differences in frequencies between geographic regions. These differences are not reflected in the FST values because of the extreme polymorphism at HLA loci, illustrating challenges in interpreting FST. Differences in the frequency of HLA alleles among geographic regions are relevant to bone-marrow transplantation, which requires genetic identity at HLA loci between patient and donor. We discuss the case of Brazil's bone marrow registry, where a deficit of enrolled volunteers with African ancestry reduces the chance of finding donors for individuals with an MHC region of African ancestry. This article is part of the theme issue ‘Celebrating 50 years since Lewontin's apportionment of human diversity’.
Background Although aging correlates with a worse prognosis for Covid-19, super elderly still unvaccinated individuals presenting mild or no symptoms have been reported worldwide. Most of the reported genetic variants responsible for increased disease susceptibility are associated with immune response, involving type I IFN immunity and modulation; HLA cluster genes; inflammasome activation; genes of interleukins; and chemokines receptors. On the other hand, little is known about the resistance mechanisms against SARS-CoV-2 infection. Here, we addressed polymorphisms in the MHC region associated with Covid-19 outcome in super elderly resilient patients as compared to younger patients with a severe outcome. Methods SARS-CoV-2 infection was confirmed by RT-PCR test. Aiming to identify candidate genes associated with host resistance, we investigated 87 individuals older than 90 years who recovered from Covid-19 with mild symptoms or who remained asymptomatic following positive test for SARS-CoV-2 as compared to 55 individuals younger than 60 years who had a severe disease or died due to Covid-19, as well as to the general elderly population from the same city. Whole-exome sequencing and an in-depth analysis of the MHC region was performed. All samples were collected in early 2020 and before the local vaccination programs started. Results We found that the resilient super elderly group displayed a higher frequency of some missense variants in the MUC22 gene (a member of the mucins’ family) as one of the strongest signals in the MHC region as compared to the severe Covid-19 group and the general elderly control population. For example, the missense variant rs62399430 at MUC22 is two times more frequent among the resilient super elderly (p = 0.00002, OR = 2.24). Conclusion Since the pro-inflammatory basal state in the elderly may enhance the susceptibility to severe Covid-19, we hypothesized that MUC22 might play an important protective role against severe Covid-19, by reducing overactive immune responses in the senior population.
As whole-genome sequencing (WGS) becomes the gold standard tool for studying population genomics and medical applications, data on diverse non-European and admixed individuals are still scarce. Here, we present a high-coverage WGS dataset of 1,171 highly admixed elderly Brazilians from a census-based cohort, providing over 76 million variants, of which ~2 million are absent from large public databases. WGS enables identification of ~2,000 previously undescribed mobile element insertions without previous description, nearly 5 Mb of genomic segments absent from the human genome reference, and over 140 alleles from HLA genes absent from public resources. We reclassify and curate pathogenicity assertions for nearly four hundred variants in genes associated with dominantly-inherited Mendelian disorders and calculate the incidence for selected recessive disorders, demonstrating the clinical usefulness of the present study. Finally, we observe that whole-genome and HLA imputation could be significantly improved compared to available datasets since rare variation represents the largest proportion of input from WGS. These results demonstrate that even smaller sample sizes of underrepresented populations bring relevant data for genomic studies, especially when exploring analyses allowed only by WGS.
We recently described a novel missense variant [c.2090T>G:p.(Leu697Trp)] in the MYO3A gene, found in two Brazilian families with late-onset autosomal dominant nonsyndromic hearing loss (ADNSHL). Since then, with the objective of evaluating its contribution to ADNSHL in Brazil, the variant was screened in additional 101 pedigrees with probable ADNSHL without conclusive molecular diagnosis. The variant was found in three additional families, explaining 3/101 (~3%) of cases with ADNSHL in our Brazilian pedigree collection. In order to identify the origin of the variant, 21 individuals from the five families were genotyped with a high-density SNP array (~600 K SNPs— Axiom Human Origins; ThermoFisher). The identity by descent (IBD) approach revealed that many pairs of individuals from the different families have a kinship coefficient equivalent to that of second cousins, and all share a minimum haplotype of ~607 kb which includes the c.2090T>G variant suggesting it probably arose in a common ancestor. We inferred that the mutation occurred in a chromosomal segment of European ancestry and the time since the most common ancestor was estimated in 1100 years (CI = 775–1425). This variant was also reported in a Dutch family, which shares a 87,121 bp haplotype with the Brazilian samples, suggesting that Dutch colonists may have brought it to Northeastern Brazil in the 17th century. Therefore, the present study opens new avenues to investigate this variant not only in Brazilians but also in European families with ADNSHL.
Diagnosis of individuals affected by monogenic disorders was significantly improved by next-generation sequencing targeting clinically relevant genes. Whole exomes yield a large number of variants that require several filtering steps, prioritization, and pathogenicity classification. Among the criteria recommended by ACMG, those that rely on population databases critically affect analyses of individuals with underrepresented ancestries. Population-specific allelic frequencies need consideration when characterizing potential deleteriousness of variants. An orthogonal input for classification is annotation of variants previously classified as pathogenic as a criterion that provide supporting evidence widely sourced at ClinVar. We used a whole-genome dataset from a census-based cohort of 1,171 elderly individuals from São Paulo, Brazil, highly admixed, and unaffected by severe monogenic disorders, to investigate if pathogenic assertions in ClinVar are enriched with higher proportions of European ancestry, indicating bias. Potential loss of function (pLOF) variants were filtered from 4,250 genes associated with Mendelian disorders and annotated with ClinVar assertions. Over 1,800 single nucleotide pLOF variants were included, 381 had non-benign assertions. Among carriers (N = 463), average European ancestry was significantly higher than noncarriers (N = 708; p = .011). pLOFs in genomic contexts of non-European local ancestries were nearly three times less likely to have any ClinVar entry (OR = 0.353; p <.0001). Independent pathogenicity assertions are useful for variant classification in molecular diagnosis. However, European overrepresentation of assertions can promote distortions when classifying variants in non-European individuals, even in admixed samples with a relatively high proportion of European ancestry. The investigation and deposit of clinically relevant findings of diverse populations is fundamental improve this scenario.