Sample size calculations can be challenging with skewed continuous outcomes in randomized controlled trials (RCTs). Standard t-test-based calculations may require data transformation, which may be difficult before data collection. Calculations based on individual and clustered Wilcoxon rank-sum tests have been proposed as alternatives, but these calculations for clustered data assume no ties in continuous outcomes, and clustered Wilcoxon rank-sum tests perform poorly with heterogeneous cluster sizes. Recent work has shown that continuous outcomes can be robustly analyzed using ordinal cumulative probability models. Analogously, sample size calculations for ordinal outcomes can be a robust design strategy for continuous outcomes. We show that Whitehead's sample size calculations for independent ordinal outcomes can naturally extend to continuous outcomes. We extend these calculations to cluster RCTs using a design effect incorporating rank intraclass correlation coefficients. Therefore, we provide a unifying, simple approach for designing individual and cluster RCTs for continuous or ordinal outcomes that makes minimal assumptions on the distribution of the still-to-be-collected outcome. We conduct simulations to evaluate our approach's performance and illustrate its application in multiple RCTs: an individual RCT with skewed continuous outcomes, a cluster RCT with skewed continuous outcomes, and a non-inferiority cluster RCT with an irregularly distributed count outcome.
The R package optimall offers a collection of functions that efficiently streamline the design process of sampling in surveys ranging from simple to complex. The package's main functions allow users to interactively define and adjust strata cut points based on values or quantiles of auxiliary covariates, adaptively calculate the optimum number of samples to allocate to each stratum using Neyman or Wright allocation, and select specific IDs to sample based on a stratified sampling design. Using real-life epidemiological study examples, we demonstrate how optimall facilitates an efficient workflow for the design and implementation of surveys in R. Although tailored towards multi-wave sampling under two- or three-phase designs, the R package optimall may be useful for any sampling survey.
Large observational datasets compiled from electronic health records are a valuable resource for medical research but are often affected by measurement error and misclassification. Valid statistical inference requires proper adjustment for these errors. Two-phase sampling with generalized raking (GR) estimation is an efficient solution to this problem that is robust to complex error structures. In this approach, error-prone variables are observed in a large phase 1 cohort, and a subset is selected in phase 2 for validation with error-free measurements. Previous research has studied optimal phase 2 sampling designs for inverse probability weighted (IPW) estimators in non-adaptive, multi-parameter settings, and for GR estimators in single-parameter settings. In this work, we extend these results by deriving optimal adaptive, multiwave sampling designs for IPW and GR estimators when multiple parameters are of interest. We propose several practical allocation strategies and evaluate their performance through extensive simulations and a data example from the Vanderbilt Comprehensive Care Clinic HIV Study. Our results show that independently optimizing allocation for each parameter improves efficiency over traditional case-control sampling. We also derive an integer-valued, A-optimal allocation method that typically outperforms independent optimization. Notably, we find that optimal designs for GR can differ substantially from those for IPW, and that this distinction can meaningfully affect estimator efficiency in the multiple-parameter setting. These findings offer practical guidance for future two-phase studies using error-prone data.
INTRODUCTION:Human papillomavirus (HPV)-associated cervical and anal cancers disproportionately affect people with HIV (PWH). This study aimed to determine the incidence trends of and risk factors for these malignancies in PWH in Latin America. METHODS:We included PWH from the Caribbean, Central and South America network for HIV epidemiology (CCASAnet) who contributed person-time between 2000 and 2019. We calculated crude and age-standardized incidence rates, examining trends over time with Poisson regression. Adjusted hazard ratios were calculated using Cox proportional hazard models with propensity score adjustment. We calculated the probability of survival after cancer diagnosis using Kaplan-Meier curves. To understand factors that influence our results, we surveyed all adult CCASAnet sites on current practices of cervical and anal cancer screening. RESULTS:Overall, 5739 females with HIV (43,417 person-years) were included in cervical cancer analyses. There were 27 incident cervical cancers: crude incidence rate of 62.2 (95% confidence interval [CI]: 34.9-89.4) per 100,000 person years. In the anal cancer analysis, 12,489 males who have sex with men (MSM), 7324 males other than MSM and 5739 females were included for a total of 25,552 PWH, contributing 157,166 person-years. Anal cancer was diagnosed in 56 individuals: crude incidence rates of 59.1 [95% CI: 33.2-85.0], 20.7 [95% CI: 11.6-29.7] and 15.2 [95% CI: 8.6-21.9] per 100,000 person-years in MSM, females and males other than MSM, respectively. Age-standardized incidence rates did not significantly change over time. Anal cancer risk decreased significantly with higher time-updated CD4 cell count. The predicted probability of 5-year survival after cancer diagnosis was 72.6% (95% CI: 48.4-86.8) for cervical cancer and 58.5% (95% CI: 44.0-70.5) for anal cancer. CONCLUSIONS:In one of the few reports outside the United States or Europe, we did not observe a decrease in age-standardized incidence rates for anal and cervical cancer between 2000 and 2019. These data support continued efforts for cancer prevention through access to gender-neutral HPV vaccination and cancer screening.
Background:Adolescents and young adults with HIV (AYAWH) represent vulnerable populations, with increased risk of virologic failure, loss to follow-up, and death. Depression and substance use in AYAWH can lead to worse outcomes, yet this overlap is not well understood. Methods:This cross-sectional study included adolescents (10-17 years) and young adults (18-24 years) with HIV in the Caribbean, Central and South America network for HIV epidemiology (CCASAnet). Participants were administered surveys to assess for depression, substance use, and antiretroviral therapy (ART) adherence. Risk factors for depression; alcohol, tobacco, and substance use; missing ART doses; viral suppression; and 1-year retention were assessed. Results:Six hundred twenty-five AYAWH were included. Depression prevelance was 16%. Males (adjusted odds ratio [aOR], 0.26; 95% CI, 0.16-0.44) and younger youth (15-year-olds vs 18-year-olds: aOR, 0.61; 95% CI, 0.40-0.95) were less likely to have depression. Fifty-eight percent reported using alcohol, 28% reported tobacco use, 17% reported cannabis use, and 4% reported cocaine use. Forty-one percent missed 1 or more doses of ART in the past week. Forty percent had detectable viral loads at the time of survey completion. Those who acquired HIV perinatally were more likely to have an unsuppressed viral load (aOR, 2.4; 95% CI, 1.24-4.62; P = .009). Only 73% of participants were retained in care following the survey; there was no statistical association between retention and age, sex, education, probable route of HIV acquisition, depression, and needing intervention for substance use. Conclusions:Substance use and depression were prevalent in AYAWH, as were missed doses of ART and detectable viral loads.
Clustered data are common in practice. Clustering arises when subjects are measured repeatedly, or subjects are nested in groups (e.g., households, schools). It is often of interest to evaluate the correlation between two variables with clustered data. There are three commonly used Pearson correlation coefficients (total, between-, and within-cluster), which together provide an enriched perspective of the correlation. However, these Pearson correlation coefficients are sensitive to extreme values and skewed distributions. They also vary with data transformation, which is arbitrary and often difficult to choose, and they are not applicable to ordered categorical data. Current nonparametric correlation measures for clustered data are only for the total correlation. Here we define population parameters for the between- and within-cluster Spearman rank correlations. The definitions are natural extensions of the Pearson between- and within-cluster correlations to the rank scale. We show that the total Spearman rank correlation approximates a linear combination of the between- and within-cluster Spearman rank correlations, where the weights are functions of rank intraclass correlations of the two random variables. We also discuss the equivalence between the within-cluster Spearman rank correlation and the covariate-adjusted partial Spearman rank correlation. Furthermore, we describe estimation and inference for the three Spearman rank correlations, conduct simulations to evaluate the performance of our estimators, and illustrate their use with data from a longitudinal biomarker study and a clustered randomized trial.
BACKGROUND:HIV treatment guidelines have evolved to recommend rapid antiretroviral therapy (ART) initiation. Data on the impact of these changes in the Americas region are scarce. METHODS:This study included data from CCASAnet sites in Brazil, Haiti, Honduras, Mexico, and Peru. ART-naïve adults who started ART from 2006 to 2022 were included. Trends in CD4 count, tuberculosis (TB), and treatment initiation were described using cumulative probability and logistic and Cox regression models. FINDINGS:Total 29,881 PLWH met inclusion criteria; 2179 (7.3%) were diagnosed with prevalent TB and 379 (1.2%) with incident TB within six months after ART initiation. For individuals without TB, enrolment CD4 count increased from 160 to 320 cells/mm3. Over the study period, TB prevalence declined from peak of 9.4% to 5.4%, and incident TB from 1.5% to 0.8%. Median time to ART initiation decreased from 476 to 1 day for PLWH without TB, and 98 to 16 days for those with prevalent TB; time to TB treatment also decreased. CONCLUSIONS:Time to ART initiation has decreased in the CCASAnet consortium, with the majority of PLWH now starting ART within a week after enrolment. There has also been a decline in the prevalence and incidence of concurrent TB disease.
Data collection procedures are often time-consuming and expensive. An alternative to collecting full information from all subjects enrolled in a study is a two-phase design: Variables that are inexpensive or easy to measure are obtained for the study population, and more specific, expensive, or hard-to-measure variables are collected only for a well-selected sample of individuals. Often, only these subjects that provided full information are used for inference, while those that were partially observed are discarded from the analysis. Recently, semiparametric approaches that use the entire dataset, resulting in fully efficient estimators, have been proposed. These estimators, however, have challenges incorporating multiple covariates, are computationally expensive, and depend on tuning parameters that affect their performance. In this paper, we propose an alternative semiparametric estimator that does not pose any distributional assumptions on the covariates or measurement error mechanism and can be applied to a wider range of settings. Although the proposed estimator is not semiparametric efficient, simulations show that the loss of efficiency to estimate the parameters associated with the partially observed covariates is minimal. We highlight the estimator's applicability to real-world problems, where data structures are complex and rich, and complicated regression models are often necessary.
Causal inference literature has extensively focused on binary treatments, with relatively fewer methods developed for multi-valued treatments. In particular, methods for multiple simultaneously assigned treatments remain understudied despite their practical importance. This paper introduces two settings: (1) estimating the effects of multiple treatments of different types (binary, categorical, and continuous) and the effects of treatment interactions, and (2) estimating the average treatment effect across categories of multi-valued regimens. To obtain robust estimates for both settings, we propose a class of methods based on the Double Machine Learning (DML) framework. Our methods are well-suited for complex settings of multiple treatments/regimens, using machine learning to model confounding relationships while overcoming regularization and overfitting biases through Neyman orthogonality and cross-fitting. To our knowledge, this work is the first to apply machine learning for robust estimation of interaction effects in the presence of multiple treatments. We further establish the asymptotic distribution of our estimators and derive variance estimators for statistical inference. Extensive simulations demonstrate the performance of our methods. Finally, we apply the methods to study the effect of three treatments on HIV-associated kidney disease in an adult HIV cohort of 2455 participants in Nigeria.
Data with measurement error in the outcome, covariates, or both are not uncommon, particularly with the increased use of routinely collected data for biomedical research. With error-prone data, often only a subsample of study data is validated; such settings are known as two-phase studies. The sieve maximum likelihood estimator (SMLE), which combines the error-prone data on all records with the validated data on a subsample, is a highly efficient and robust method to analyze such data. However, given their complexity, a computationally efficient and user-friendly tool is needed to obtain the SMLEs. The R package sleev fills this gap by making semiparametric likelihood-based inference using the SMLEs for error-prone two-phase data in settings with binary and continuous outcomes. Functions from this package can be used to analyze data with error-prone binary or continuous responses and error-prone covariates.
Inflammation is a hallmark of cancer. DNA methylation (DNAm) derived immune cell type proportions and circulating inflammation-related protein markers are precise measures of inflammatory phenotypes. We investigated the associations of DNAm-derived inflammation-related phenotypes with risks of anal cancer and precancer (high-grade squamous intraepithelial lesions, HSIL) among men with HIV. We matched 22 anal cancer cases to 37 controls with normal anal epithelium and 54 HSIL to 90 controls using propensity score methods. DNAm data were profiled by MethylationEPIC v2.0 array on the pre-diagnostic blood biospecimens and processed by R minfi and ChAMP packages. Immune cell composition, including 4 myeloid (neutrophils, eosinophils, basophils and monocytes) and 8 lymphoid cell sub-lineages (B lymphocytes naïve (Bnv), B lymphocytes memory [Bmem], T helper lymphocytes naïve [CD4nv], T helper lymphocytes memory [CD4mem], T regulatory cells [Treg], T cytotoxic lymphocytes naïve [CD8nv], T cytotoxic lymphocytes memory [CD8mem], and natural killer lymphocytes) were estimated among normalized DNA methylation β values using a reference-based deconvolution algorithm. Proportions of these 12 major immune cell types were estimated to be 100%. Ratios of CD4nv/CD4mem, CD8nv/CD8mem, Bnv/Bmem, CD8/Treg, and neutrophil/lymphocyte were computed. Additionally, 47 DNAm-predicted circulating inflammation-related protein levels were calculated using well-established prediction models. Conditional logistic regression models were built to estimate odds ratios (OR) and 95% confidence intervals (CI) on per standard deviation increase in inflammation phenotypes adjusting for cell type heterogeneities. Higher CD4nv/CD4mem (OR=2.70, 95%CI:1.09-6.70) and lower CD8/Treg (OR=0.56, 95%CI:0.31-0.99) was positively associated with anal cancer risk at nominal P < 0.05. Higher Bnv/Bmem (OR=1.47, 95%CI:1.02-2.12) was linked to increased HSIL risk at nominal P < 0.05. Elevated levels of DNAm-derived circulating protein MMP9 (OR=1.86, 95%CI: 1.20-2.88), C5 (OR=1.78, 95%CI: 1.23-2.57), B2M (OR=1.75, 95%CI:1.16-2.65), CRP (OR=2.39, 95%CI:1.32-4.33), MMP12 (OR=1.66, 95%CI:1.15-2.41), CCL11(OR=1.66, 95%CI:1.15-2.38) , TNFRSF1B (OR=1.85, 95%CI:1.20-2.87), HGF (OR=2.10, 95%CI:1.32-3.35), TGFA (OR=2.38, 95%CI:1.31-4.35), VEGFA (OR=1.82, 95%CI:1.21-2.72), CCL17 (OR=1.93, 95%CI:1.26-2.95); and decreased levels of ADAMTS13 (OR=0.47, 95%CI:0.28-0.78) were associated with increased risk of anal cancer and HSIL (combined analysis) compared to controls at false discovery rate (FDR) <5%. No significant associations were observed when comparing cancers with controls or comparing HSIL with controls, separately at FDR<5%. Our findings suggest that inflammation plays a role in anal carcinogenesis. Future research is warranted to validate our results. Shuai Xu, Bryan E. Shepherd, Tebeb Gebretsadik, Anna Junkins, Jirong Long, Qiuyin Cai, Staci L. Sudenga. Methylation-derived inflammation phenotypes and anal cancer risk among men with HIV [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 6217.
INTRODUCTION:Despite its reversal in July 2019, the World Health Organization warning issued in May 2018 of potential teratogenicity associated with dolutegravir (DTG) may have produced persistent sex disparities in access to DTG. We compared DTG uptake of people with HIV (PWH) by sex in Latin America and the Caribbean (LAC) and its potential impact on virologic outcomes. METHODS:We evaluated DTG initiation among antiretroviral therapy (ART)-naïve and -experienced cisgender PWH ≥16 years of age after DTG availability in Brazil (February/2017), Chile (August/2019), Haiti (November/2018) and Honduras (December/2018). Time was divided into pre- (before May/2018), during- (May/2018-July/2019) and post- (after July/2019) warning periods. We examined interactions of sex, age and calendar era with multivariable modified Poisson regression models and Cox proportional hazard models for the outcomes of DTG initiation among ART-naïve and ART-experienced PWH, respectively, and HIV RNA <50 copies/ml in the first year of therapy among ART-naïve PWH, adjusting for site and tuberculosis. RESULTS:Among 4622 ART-naïve PWH, 3853 (83%) initiated DTG. ART-naïve females aged 16-49 years were less likely to initiate DTG compared to males of the same age both in the pre/during-warning (adjusted prevalence ratio [aPR]: 0.75 [95% confidence interval (95% CI): 0.71-0.80]) and in the post-warning periods (aPR: 0.97 [95% CI: 0.95-1.00]). Among 16,154 ART-experienced PWH, 9236 (57%) initiated DTG. ART-experienced females 16-49 years were less likely to initiate DTG compared to males of the same age in the pre/during-warning (adjusted hazard ratio [aHR]: 0.69 [95% CI: 0.66-0.73]) and post-warning periods (aHR: 0.79 [95% CI: 0.70-0.90]). This sex difference was not observed among older ART-experienced females and males pre/during-warning (aHR: 1.06 [95% CI: 0.99-1.14]). Compared to starting ART without DTG, DTG-based ART use was associated with a higher likelihood of HIV RNA suppression in the first year (aPR = 1.10 [95% CI: 1.04-1.16]). In the post-warning period, females aged 16-49 years had a likelihood of viral suppression similar to males of the same age (aPR: 1.03 [95% CI: 0.96-1.10]), which did not change after adjusting for DTG use (aPR: 1.03 [95% CI: 0.97-1.11]). CONCLUSIONS:Despite the updated guidelines recommending DTG for all PWH, there are persistent sex disparities in the access to DTG in LAC, especially among females within the reproductive age.
Difference-in-differences (DID) approaches are widely used for estimating causal effects with observational data before and after an intervention. DID traditionally estimates the average treatment effect among the treated after making a parallel trends assumption on the means of the outcome. With skewed outcomes, a transformation is often needed; however, the transformation may be difficult to choose, results may be sensitive to the choice, and parallel trends assumptions are made on the transformed scale. Recent DID methods estimate alternative treatment effects that may be preferable with skewed outcomes. However, each alternative DID estimator requires a different parallel trends assumption. We introduce a new DID method capable of estimating average, quantile, probability, and novel Mann-Whitney treatment effects among the treated with a single unifying parallel trends assumption. The proposed method uses a semi-parametric cumulative probability model (CPM). The CPM is a linear model for a latent variable on covariates, where the latent variable results from an unspecified transformation of the outcome. Our DID approach makes a universal parallel trends assumption on the expectation of the latent variable conditional on covariates. Hence, our method avoids specifying outcome transformations and does not require separate assumptions for each estimand. We introduce the method; describe identification, estimation, and inference; conduct simulations evaluating its performance; and apply it to assess the impact of Medicaid expansion on CD4 count among people with HIV.
INTRODUCTION:Antiretroviral therapy (ART) during pregnancy and at delivery has nearly eliminated vertical transmission (VT) in some settings but previously reported VT prevalence has been as high as 15% in Latin America and the Caribbean (LAC). We evaluated VT in the Caribbean, Central and South America network for HIV epidemiology to further study the benefit of ART on VT in our region. METHODS:We retrospectively collected data on cis-gender women ≥15 years of age enrolled in HIV clinics in Brazil, Chile, Honduras and Peru from 2003 to 2018 with ≥1 pregnancy resulting in a live birth after clinic entry to examine the association of ART use at the time of delivery and VT. We used propensity-score-matched logistic regression to examine the odds of VT by ART use. Matching weights incorporated site, HIV RNA, CD4 cell count, maternal age, year and HIV diagnosis before or during pregnancy. We also examined the proportion of women who received ART during pregnancy before and after the treat-all era, as defined within each country. RESULTS:A total of 623 pregnant women with HIV contributed 727 live births. Of all births, 613 (84.3%) infants had known HIV status and there were 22 (3.6%) VT events. Four of the 22 (18%) were born to women on ART at delivery, compared to 403 of 591 (68%) infants negative for HIV. In the propensity-score-matched model, ART use at delivery was associated with 85% decreased odds of VT (odds ratio = 0.15, 95% confidence interval 0.04-0.58). In the pre-treat-all era, 37% (181/485) of women received ART within 30 days of pregnancy diagnosis, compared to 59% (75/128) during the treat-all era (p<0.001). In the pre-treat-all era, 4.3% (21/485) of infants were born HIV positive, compared to 0.8% (1/128) in the treat-all era (p = 0.055). CONCLUSIONS:We found a low prevalence of VT in our cohort, especially in the treat-all era. ART use at delivery was strongly associated with a lower odd of VT. Despite improvements, access to ART during pregnancy remained far from universal. Therefore, new strategies to ensure its effective implementation in LAC are still warranted.
Measurement error is a common challenge for causal inference studies using electronic health record (EHR) data, where clinical outcomes and treatments are frequently mismeasured. Researchers often address measurement error by conducting manual chart reviews to validate measurements in a subset of the full EHR data – a form of two-phase sampling. To improve efficiency, phase-two samples are often collected in a biased manner dependent on the patients' initial, error-prone measurements. In this work, motivated by our aim of performing causal inference with error-prone outcome and treatment measurements under two-phase sampling, we develop solutions applicable to both this specific problem and the broader problem of causal inference with two-phase samples. For our specific measurement error problem, we construct two asymptotically equivalent doubly-robust estimators of the average treatment effect and demonstrate how these estimators arise from two previously disconnected approaches to constructing efficient estimators in general two-phase sampling settings. We document various sources of instability affecting estimators from each approach and propose modifications that can considerably improve finite sample performance in any two-phase sampling context. We demonstrate the utility of our proposed methods through simulation studies and an illustrative example assessing effects of antiretroviral therapy on occurrence of AIDS-defining events in patients with HIV from the Vanderbilt Comprehensive Care Clinic.
Exposure measurement error is a ubiquitous but often overlooked challenge in causal inference with observational data. Existing methods accounting for exposure measurement error largely rely on restrictive parametric assumptions, while emerging data-adaptive estimation approaches allow for less restrictive assumptions but at the cost of flexibility, as they are typically tailored toward rigidly defined statistical quantities. There remains a critical need for assumption-lean estimation methods that are both flexible and possess desirable theoretical properties across a variety of study designs. In this paper, we introduce a general framework for estimation of causal quantities in the presence of exposure measurement error, adapted from the method of control variates. Our method can be implemented in various two-phase sampling study designs, where one obtains gold-standard exposure measurements for a small subset of the full study sample, called the validation data. The control variates framework leverages both the error-prone and error-free exposure measurements by augmenting an initial consistent estimator from the validation data with a variance reduction term formed from the full data. We show that our method inherits double-robustness properties under standard causal assumptions. Simulation studies show that our approach performs favorably compared to leading methods under various two-phase sampling schemes. We illustrate our method with observational electronic health record data on HIV outcomes from the Vanderbilt Comprehensive Care Clinic.
HIV care continuum outcome disparities by health insurance status have been noted among people with HIV (PWH). We therefore examined associations between state Medicaid expansion and HIV outcomes in the United States. Adults (≥18 years) with ≥1 visit in NA-ACCORD clinical cohorts from 2012-2017 contributed person-time annually between first and final visit or death; in each calendar year, clinical retention was ≥2 completed visits > 90 days apart, antiretroviral therapy (ART) receipt was receipt of ≥3 antiretroviral agents, and viral suppression was last measured HIV-1 RNA < 200 copies/mL. CD4 at enrollment was obtained within 6 months of enrollment in cohort. Difference-in-difference (DID) models quantified associations between Medicaid expansion changes (by state of residence) and HIV outcomes. Across 50 states, 87 290 PWH contributed 325 113 person-years of follow-up. Medicaid expansion had a substantial positive effect on CD4 at enrollment (DID = 93.5, 95% CI: 52.9, 134 cells/mm3), a small negative effect on proportions clinically retained (DID = -0.19, 95% CI: -0.037, -0.01), and no effects on ART receipt (DID = 0.001, 95% CI: -0.003, 0.005) or viral suppression (DID = -0.14, 95% CI: -0.34, 0.07). Medicaid expansion had a positive effect on CD4 at entry, suggesting more timely HIV testing and care linkage, but generally null effects on downstream HIV care continuum measures.
The probability-scale residual (PSR) is defined as E{sign(y, Y^*)}, where y is the observed outcome and Y^* is a random variable from the fitted distribution. The PSR is particularly useful for ordinal and censored outcomes for which fitted values are not available without additional assumptions. Previous work has defined the PSR for continuous, binary, ordinal, right-censored, and current status outcomes; however, development of the PSR has not yet been considered for data subject to general interval censoring. We develop extensions of the PSR, first to mixed-case interval-censored data, and then to data subject to several types of common censoring schemes. We derive the statistical properties of the PSR and show that our more general PSR encompasses several previously defined PSR for continuous and censored outcomes as special cases. The performance of the residual is illustrated in real data from the Caribbean, Central, and South American Network for HIV Epidemiology.
Background We sought to determine the prevalence of sickle cell trait (SCT) and apolipoprotein-1 ( APOL1) risk variants in people living with HIV (PLWH) in Nigeria, and to establish if SCT and APOL1 high-risk status correlate with estimated glomerular filtration rate (eGFR) and/or prevalent chronic kidney disease (CKD). Methods Baseline demographic and clinical data were obtained during three cross-sectional visits. CKD was defined as having an eGFR<60 mL/min/1.73 m2. We collected urine specimens to determine urine albumin-creatine ratio and blood samples for sickle cell genotyping, APOL1 testing, and for creatinine/cystatin C assessment. The associations between SCT, APOL1 genotype, and eGFR/CKD stages/CKD were investigated using linear/ordinal logistic/logistic regression models, respectively. Results Of 2443 participants, 599 (24.5%) had SCT, and 2291 (93.8%) had a low-risk APOL1 genotype (0 or 1 risk variant), while 152 (6.2%) had high-risk genotype (2 allele copies). In total, 108 participants (4.4%) were diagnosed with CKD. In adjusted analyses, SCT was associated with lower eGFR (adjusted mean difference [aMD]= −2.33, 95% CI -4.25, −0.42), but not with worse CKD stages, or increased odds of developing CKD. Participants with the APOL1 high risk genotype were more likely to have lower eGFR (aMD= −5.45, 95% CI -8.87, −2.03), to develop CKD (adjusted odds ratio [aOR] = 1.97, 95% CI: 1.03, 3.75), and to be in worse CKD stages (aOR = 1.60, 95% CI: 1.12, 2.29) than those with the low-risk genotype. There was no evidence of interaction between SCT and APOL1 genotype on eGFR or risk of CKD. Conclusion Our findings highlight the multifaceted interplay of genetic factors in the pathogenesis of CKD in PLWH.