The Depression Genetics in Africa (DepGenAfrica) study seeks to address the underrepresentation of African populations in psychiatric genomics by investigating the genetic basis of major depressive disorder in three African countries: Ethiopia, Malawi and Nigeria. This Editorial reflects on lessons from project set-up and offers recommendations, highlighting trust, communication, ethical oversight, data tools, workforce training and resource governance.
COVID-19 and its long-term consequences remain a major global health challenge, with host genetic factors contributing to inter-individual variation in susceptibility. Apolipoprotein E (APOE) genotypes, particularly those including the ε4 allele, have been associated with an increased risk of COVID-19 infection in Europeans, but evidence from continental Africans remains limited. We investigated the association between APOE genotypes (ε2/ε2, ε2/ε3, ε2/ε4, ε3/ε3, ε3/ε4 and ε4/ε4) and COVID-19 positivity in 921 Bantu-speaking South Africans. COVID-19 positivity was determined serologically, while APOE genotypes were identified using TaqMan™ SNP genotyping of rs429358 and rs7412. Associations were assessed using multinomial logistic regression across unadjusted, partially adjusted (age, sex, education), and fully adjusted models (age, sex, education, wealth index, hypertension, diabetes, episodic memory, HIV), with ε3/ε3 as the reference. The most frequent genotype was ε3/ε3 (36%), followed by ε3/ε4 (29%) and ε2/ε3 (19%). None of the genotypes were significantly associated with COVID-19 positivity. In fully adjusted models, females had higher risk than males (OR = 1.54, 95% CI 1.13-2.12), this effect was also shown in sensitivity analyses using imputed covariates. In contrast to European studies, APOE genotypes were not associated with COVID-19 positivity in this cohort, highlighting the importance of ancestrally diverse genomic research. Larger African cohorts are needed to explore sex-specific effects and ancestry-dependent modifiers of APOE in COVID-19 susceptibility.
INTRODUCTION:Serum creatinine is a commonly used biomarker for estimating glomerular filtration rate (eGFR), but is influenced by determinants unrelated to kidney function, including muscle mass, age, sex, diet, and genetic variation. Biological drivers of persistent eGFR bias in African ancestry populations remain understudied. Here, we investigated whether African-enriched genetic variation at the glycine amidinotransferase (GATM) locus, encoding the rate limiting enzyme in creatine biosynthesis, contributes to systematic bias in creatinine-based estimates of GFR. METHODS:We analyzed whole-genome sequencing and phenotypic data from population-based cohorts in Uganda (5,760 individuals) and South Africa (827 individuals). Genome-wide association studies were integrated with expression quantitative trait locus and metabolomic analyses to characterize biological mechanisms. The iohexol measured GFR was used to quantify systematic bias for creatinine-based estimates (eGFRCr). Diagnostic implications were evaluated in African American participants of the Million Veteran Program. RESULTS:We described variants associated with eGFRCr at the GATM locus that were common in African populations (allele frequency about 50%) but rare or absent in Europeans. These variants explained 1.0%-1.7% of the variance in eGFRCr and were associated with increased GATM expression and enhanced creatine synthesis, resulting in higher serum creatinine and consequently lower eGFRCr. Each allele was associated with an approximately 2 ml/min per 1.73 m2 lower eGFRCr, and the association was independent of measured GFR, age, sex, and body composition. In contrast, these variants showed no association with measured GFR or cystatin C-based eGFR, supporting a biomarker-specific effect. In African American participants, GATM variants were associated with increased odds of a CKD diagnosis using eGFRCr thresholds but not with end-stage kidney disease. CONCLUSIONS:African-enriched GATM variants are associated with measurable bias in creatinine-based GFR estimation without evidence of altered kidney function. These findings improve understanding of differences in kidney function biomarkers and may be relevant when interpreting creatinine-based estimates, particularly in populations where the variant is common.
Genetic predisposition and alcohol consumption are risk factors for increased blood pressure (BP), but their interactions influencing BP remain understudied. We conducted population-specific and cross-population meta-analyses of genome-wide gene-alcohol (GxAlc) interactions affecting BP in >1.1M individuals from multiple populations. We identified 46 GxAlc interaction loci for BP, including 21 from one-degree-of-freedom interaction tests (PGxAlc<5×10-8; or <0.05/Meff, Meff independent BP associations at P<10-5), and 25 from two-degree-of-freedom tests of main and interaction effects (PGxAlc<0.05/M2df, M2df independent 2df-associations at P2df<5×10-8), including 7 novel and 39 known BP loci. The 12q24 locus highlights the genetic effect of BRAP-rs11066001 on BP, being ~6 times larger in current drinkers than in non-drinkers. Gene prioritization with 46 GxAlc loci identified 15 genes with ≥3 lines of evidence (location, literature, druggability, functional/regulatory annotation, or pathway analyses). Several loci showed sex- and population-specific effects and revealed biological pathways of alcohol's influence on BP, suggesting mechanisms underlying alcohol-induced hypertension.
Alcohol misuse is a major global health concern with significant genetic contributions, yet studies in African populations remain limited. We conducted genome-wide association meta-analyses for alcohol drinking initiation and cessation across multiple cohorts comprising individuals from sub-Saharan Africa (N = 20,936) and those with African ancestry from UK Biobank (N = 7,726). Our analyses identified two genome-wide significant loci associated with drinking cessation ( FYB2 and FTO regions) and several suggestive associations overlapping with genes implicated in alcohol metabolism, behavioral traits, and neurodevelopment. These findings provide insights into the genetic architecture of alcohol use behaviors in diverse African populations. By focusing on alcohol drinking initiation and cessation phenotypes, this study complements prior research on alcohol dependence and consumption, highlighting genetic factors influencing vulnerability and resilience. Our results underscore the importance of including diverse ancestries in genetic studies to enhance understanding of alcohol-related behaviors and inform future investigations into prevention and intervention strategies.
The emergence of precision therapies for genetically mediated kidney diseases highlights a growing paradox — populations of African ancestry contribute essential genomic diversity but remain among the least likely to benefit from innovations that result from genomic research. Embedding equitable benefit sharing within kidney genomics is an ethical imperative and a scientific necessity.
Genomic data hold transformative potential for healthcare. Despite their unparalleled genetic diversity, African populations remain underrepresented in global genomics research, contributing to poorer health outcomes both within Africa and globally. Political, financial, and social challenges hinder progress, and therefore advocacy for public understanding of genomics, addressing social stigma, and aligning initiatives with national policies are critical steps forward. Africa hosts a large pool of highly skilled scientists and laboratories with DNA analysis capacity to harness for human genetics research and service provision, presenting opportunities for growth and innovation. This article highlights progress in African genomics, emphasizes the importance of equity, and proposes actionable steps for prioritizing, funding, and implementing national genome projects. It is a call to action. Genomic data could transform healthcare, but African populations are underrepresented, currently limiting benefits. Here the authors highlight how addressing political, financial, and social barriers and advancing equitable national genome projects is essential.
Attrition in longitudinal studies, defined as the loss of participants over time, may threaten the validity of study results. This study examined patterns, magnitude, and risk factors for attrition in the AWI-Gen cohort. The study included 12,032 participants from six African sites, with a median baseline age of 51 years (IQR: 45–57 years). Exploratory data analysis was conducted to examine attrition patterns and magnitude, and fixed-effects logistic regression models were used to identify significant factors associated with different causes of attrition (all-cause, non-death vs death, preventable vs non-preventable) while accounting for site-specific variation. Sensitivity analyses, including Bonferroni adjustments, imputation, and site-specific analyses, were used to assess model robustness. Model goodness-of-fit was evaluated, and adjusted odds ratios with 95
ABSTRACT Returning individualized microbiome results in ways that are ethical, comprehensible, and useful remains under-explored in African settings. We nested a multi-site, mixed-methods study within the AWI-Gen Wave 2 gut microbiome sub-study of 1,801 women aged 42–86 years to engage participants and provide feedback. All (1,001) participants from Agincourt and Soweto (South Africa) and Nairobi (Kenya) were invited to feedback meetings: 496 from Agincourt, 87 from Soweto, and 195 from Nairobi responded. Engagement strategies were tailored by site (small-group and home-based sessions, visual metaphors, Foldscopes, and local-language delivery). Using semi-structured discussions and structured observations analyzed thematically in MAXQDA under COREQ, five cross-cutting themes emerged: (i) understanding of microbiome reports, (ii) emotional responses to feedback, (iii) perceived health relevance, (iv) trust in research institutions, and (v) suggestions for improving engagement. Culturally grounded explanations and local-language facilitation enhanced comprehension; English-heavy sessions were associated with more confusion. Most participants expressed satisfaction and described planned or enacted dietary and lifestyle changes, while frustration centered on delays between sampling and feedback. Trust increased with transparency and individualized return of results but was often conditional on minimizing burdensome procedures such as repeat blood sampling and ensuring timely feedback. Engagement was feasible and low cost (approximately USD 29–59 per participant) with site-specific resource needs. Limitations included constrained generalizability beyond the three study sites and the study population of women aged 42–86 years. Returning individualized microbiome findings in African community settings is acceptable, feasible, and can motivate health-promoting behaviors when delivered promptly and in culturally appropriate ways. IMPORTANCE Microbiome studies rarely return individualized results in low-resource settings due to concerns about appropriate feedback and associated costs. This gap risks eroding trust and diminishing research impact. In three African communities, tailored feedback on gut microbiome profiles was provided to 778 women. By documenting a costed, multi-site engagement model and the themes influencing acceptance and actionability, this work offers a practical framework for ethically returning complex -omics results at scale in underrepresented populations—advancing scientific equity and strengthening community trust in microbiome research.
Background: The 9p21.3 locus was the first genome-wide significant signal for coronary artery disease (CAD) and replicates across multiple non-African populations, yet is absent in African ancestry cohorts. We hypothesized that ancestry-specific linkage disequilibrium (LD) and haplotype structure, rather than allele frequency or power alone, explain this discrepancy. Methods: We analyzed multi-ancestry data from European, East Asian, South Asian, Middle Eastern, African, and Admixed American groups. Within each ancestry, we performed CAD common-variant associations at 9p21.3, rare-variant tests in whole-genome-sequenced Europeans, local-ancestry inference (LAI) stratified associations, conditional and haplotype analyses, and pleiotropy assessments. Results: CAD associations at 9p21.3 were robust in Europeans, East Asians, Middle Eastern, and Admixed Americans ancestry groups, nominal in South Asians, and absent in Africans despite similar allele frequencies for lead-associated SNPs in other ancestry groups. LAI-stratified analyses showed strong signals in African heterozygotes carrying 9p21.3 European chromosomes but not in 9p21.3-homozygous Africans. The association is due to highly common variants; rare variants did not explain the locus signal. LD blocks were extended in non-Africans but fragmented into smaller blocks in Africans, with marked divergence by fixation index. Fine-mapping identified multiple independent signals, including an East Asian-specific haplotype absent in Europeans and Africans. Hamming-distance analyses revealed that risk alleles are dispersed across multiple haplotypes in homozygous African groups, consistent with weaker LD and restricting risk variation at this locus in that group. PheWAS mirrored these ancestry-specific patterns across metabolic traits. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement We gratefully acknowledge the Veterans who participated in the VA Million Veteran Program. This research is based on data from the Million Veteran Program, Office of Research and Development, and Veterans Health Administration, with support from MVP000, VA Merit Award #I01-BX003362 (Chang/Tsao) and the Department of Veterans Affairs (VA) Informatics and Computing Infrastructure (VINCI), including data analytics conducted by its Precision Medicine research team, which is funded under the research priority to Put VA Data to Work for Veterans (VA ORD 24-D4V-02). This work was also supported by the American Heart Association Predoctoral Fellowship (Award ID: 24PRE1201199) and the Lamond Fellowship. This publication does not represent the views of the Department of Veterans Affairs or the United States Government. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The Million Veteran Program study received ethical and study protocol approval from the VA Central Institutional Review Board. Ethical approval for the use of genomic and clinical data along with consent from participants for UK Biobank was obtained by the National Health Service National Research Ethics Service (ref: 264 11/NW/0382) and data use approved under application 87255. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present work are contained in the manuscript
Despite repeated calls from international governing bodies for benefit sharing in biomedical research as a means of promoting mutual trust and reciprocity, real-world examples of benefit sharing are lacking. Here, we discuss benefit-sharing initiatives instituted as part of human genomics research projects carried out in both semi-rural and urban settings in South Africa. We explore lessons learned from applying different approaches to benefit sharing, as well as recommendations for ways to enhance success for future initiatives. Ultimately, we advocate for the adoption of flexible benefit-sharing models that take into consideration community representation and engagement, benefit management mechanisms, and long-term sustainability.
Locus 9p21.3 was the first genome-wide significant locus for coronary artery disease (CAD) and replicates across multiple non-African populations, yet it is absent in African-ancestry cohorts. We analyzed multi-ancestry data from European, East Asian, South Asian, Middle Eastern, African, and admixed American groups. Ancestry was inferred using both global and local-ancestry inference (LAI) approaches. Within each ancestry, we performed CAD associations at 9p21.3 with common and rare variants, conditional and haplotype analyses, and pleiotropy assessments. CAD associations at 9p21.3 were robust in European, East Asian, Middle Eastern, and admixed American ancestry groups, nominal in South Asians, and absent in Africans despite similar allele frequencies for lead-associated SNPs. LAI-stratified analyses showed strong signals in African heterozygotes carrying 9p21.3 European chromosomes but not in homozygous Africans. The association is due to highly common variants; rare variants did not explain the locus signal. Linkage disequilibrium (LD) blocks were extended in non-Africans but fragmented into smaller blocks in Africans. Fine-mapping identified multiple independent signals, including an East Asian-specific haplotype absent in Europeans and Africans. Hamming-distance analyses revealed that risk alleles are dispersed across multiple haplotypes in homozygous-African groups. Phenome-wide association study (PheWAS) mirrored these ancestry-specific patterns across metabolic traits. The non-replication of 9p21.3 in African ancestry is not due to absent risk alleles, reduced power, or poorer imputation of risk alleles due to weaker LD; rather, it reflects greater haplotype diversity that disperses risk alleles across configurations, attenuating case-control contrast. These findings illustrate how ancestry-specific LD and multi-causal haplotype architecture modulate association detectability.
Kidney disease disproportionately affects populations of African ancestry, yet most genetic studies have focused on Europeans. Here, we present a three-stage genome-wide association study meta-analysis of estimated glomerular filtration rate in ~26,000 individuals across Eastern, Western, and Southern Africa and ~81,000 African-ancestry individuals in the diaspora. Continental African meta-analysis identifies four independent genome-wide significant loci, including two previously unreported loci. Pan-African meta-analysis identifies 19 independent loci, including three previously unreported loci. Fine-mapping reveals four loci with high causality probability, and phenome-wide analyses demonstrate pleiotropic effects on cardiometabolic and immunological traits. Notably, APOL1 high-risk variants strongly associated with kidney disease in African Americans show markedly lower frequency and attenuated effects in continental Africa, indicating potential distinct genetic architectures. Polygenic scores from genetically similar populations significantly outperformed those from distant cohorts. These findings demonstrate the necessity of conducting genomic research across diverse African populations to enable equitable health outcomes.
Type 2 diabetes is rising globally, with sub-Saharan Africa facing the steepest projected increase. In low-resource settings, limited screening contributes to high rates of undiagnosed disease, delaying intervention. Traditional risk models often rely on single-factor thresholds or population averages, which can miss combinations of demographic, anthropometric, and lifestyle factors relevant in African contexts. However, emerging evidence suggests that interactions between multiple risk factors provide more accurate, region-specific characterizations of risk. We applied a multidimensional subgroup discovery algorithm to cross-sectional data from three African populations and identified combinations of risk factor cutoffs that define high-risk subgroups. The most consistent profile, such as waist-to-hip ratio greater than 0.9, physical activity less than or equal to 2448 metabolic equivalent-minutes per week, and family history, shows elevated risk across regions and outperforms guideline-based definitions. Here we show that core anthropometric and familial factors are stable across populations, while others are region-specific, improving predictive performance for screening.
Epigenetic modifications influence gene expression levels, impact organismal traits, and play a role in the development of diseases. Therefore, variants in genes involved in epigenetic processes are likely to be important in disease susceptibility, and the frequency of variants may vary between populations with African and European ancestries. Here, we analyse an integrated dataset to define the frequencies, associated traits, and functional impact of epigenetic gene variants among individuals of African and European ancestry represented in the UK Biobank. We find that the frequencies of 88.4% of epigenetic gene variants significantly differ between these groups. Furthermore, we find that these variants map to many reported traits and diseases, and we show that allele-frequency differences can alter statistical power and the likelihood of detecting associations across ancestry groups, particularly given the substantial sample-size imbalance between the UK Biobank European-ancestry and African-ancestry subsets. Additionally, we observe that variants associated with traits are significantly enriched for quantitative trait loci that affect DNA methylation, chromatin accessibility, and gene expression. We find that methylation quantitative trait loci account for 71.2% of the variants influencing gene expression. Moreover, variants linked to biomarker traits exhibit high correlation. We therefore conclude that epigenetic gene variants associated with traits tend to differ in their allele frequencies among African and European populations and are enriched for QTLs.
AIMS:The burden of dyslipidaemia is increasing, and the association of dietary exposure, especially vegetable consumption, with dyslipidaemia among Africans is poorly characterized. This study evaluated the relationship between vegetable consumption and dyslipidaemia among Africans. METHODS AND RESULTS:The frequency of vegetable consumption (servings/week) was assessed in this study involving 13 172 participants, including 6586 pairs of dyslipidaemia cases and non-cases (matched for age within ±5 years, sex, and country), in a matched case-control design. Multivariable-adjusted conditional logistic regression models were applied to estimate odds ratios (ORs) and 95% confidence intervals (CIs) for the odds of dyslipidaemia across quartiles of frequency of vegetable consumption at a two-sided P < 0.05. The mean age was 52.18 ± 10.15 years, and 6898 (52.4%) were females. The median (IQR) vegetable consumption intake was 7.0 (2.0, 14.0) servings per week and the prevalence of dyslipidaemia by the distribution of vegetable consumption was 1776 (52.0%) for low (first quartile), 1530 (49.9%) for moderate (second quartile), 1720 (49.2%) for sufficient (third quartile), and 1560 (48.9%) for high (fourth quartile) frequency of vegetable consumption. The multivariable-adjusted OR (95%CI) of dyslipidaemia odds by the distribution of vegetable consumption were 1.00 for low, 0.89 (0.80, 0.99) for moderate, 0.84 (0.75, 0.93) for sufficient, and 0.81 (0.72, 0.92) for high; P for trend = 0.005, with a OR (95%CI) of 0.97 (0.94, 0.99) per +7 servings/week change after adjusting for age, family history of cardiovascular diseases, education, ever smoked, currently consume alcohol, physical inactivity, body mass index, diabetes mellitus status, and hypertension. A similar trend was observed for low high-density lipoprotein (<40 mg/dL): 1.00 for low, 0.90 (0.78, 1.04) for moderate, 0.93 (0.81, 1.06) for medium, and 0.80 (0.68, 0.94) for high; P for trend = 0.01, adjusting for similar covariates. CONCLUSION:Higher vegetable consumption was associated with lower odds of dyslipidaemia in this sample of Africans after accounting for multiple covariates.
Plasma proteomic scores for body mass index have been derived primarily in European populations, and it is unclear whether these scores are generalizable and useful in other ancestries. Here, we show that plasma proteomic BMI scores developed in participants of European, East Asian, South Asian and African ancestries from the UK Biobank (UKBB; N = 50,621) and the South African Middle-Aged Soweto Cohort (MASC; N = 948) are portable across ancestries, explain up to 48.7% of trait variance and significantly predict incident obesity. We identified eight protein biomarkers common across all scores, six of which showed causal associations with BMI in Mendelian randomization analyses and were associated with BMI across the Olink (UKBB, MASC) and SomaScan platforms (Qatar Biobank; N = 2410). Discrepancy analysis between the high predicted proteomic BMI and low measured BMI revealed metabolically unhealthy normal weight (MUNW) individuals, with higher visceral fat, elevated triglycerides, and low insulin sensitivity. Thus, plasma proteomic BMI scores may enhance the precision of cardiometabolic risk stratification across diverse populations.
As genomics infrastructure expands, current governance will determine whether it advances global health equity or entrenches disparities, locking biased data into future clinical tools. This Comment calls for inclusive, community-centered genomic systems, before biased data and extractive models becoming the default.