Clustered data are common in practice. Clustering arises when subjects are measured repeatedly, or subjects are nested in groups (e.g., households, schools). It is often of interest to evaluate the correlation between two variables with clustered data. There are three commonly used Pearson correlation coefficients (total, between-, and within-cluster), which together provide an enriched perspective of the correlation. However, these Pearson correlation coefficients are sensitive to extreme values and skewed distributions. They also vary with data transformation, which is arbitrary and often difficult to choose, and they are not applicable to ordered categorical data. Current nonparametric correlation measures for clustered data are only for the total correlation. Here we define population parameters for the between- and within-cluster Spearman rank correlations. The definitions are natural extensions of the Pearson between- and within-cluster correlations to the rank scale. We show that the total Spearman rank correlation approximates a linear combination of the between- and within-cluster Spearman rank correlations, where the weights are functions of rank intraclass correlations of the two random variables. We also discuss the equivalence between the within-cluster Spearman rank correlation and the covariate-adjusted partial Spearman rank correlation. Furthermore, we describe estimation and inference for the three Spearman rank correlations, conduct simulations to evaluate the performance of our estimators, and illustrate their use with data from a longitudinal biomarker study and a clustered randomized trial.
Detection limits (DLs), where a variable is unable to be measured outside of a certain range, are common in research. Most approaches to handle DLs in the response variable implicitly make parametric assumptions on the distribution of data outside DLs. We propose a new approach to deal with DLs based on a widely used ordinal regression model, the cumulative probability model (CPM). The CPM is a type of semiparametric linear transformation model. CPMs are rank-based and can handle mixed distributions of continuous and discrete outcome variables. These features are key for analyzing data with DLs because while observations inside DLs are typically continuous, those outside DLs are censored and generally put into discrete categories. With a single lower DL, the CPM assigns values below the DL as having the lowest rank. When there are multiple DLs, the CPM likelihood can be modified to appropriately distribute probability mass. We demonstrate the use of CPMs with simulations and two HIV data examples. The first example models a biomarker in which 15% of observations are below a DL. The second uses multi-cohort data to model viral load, where approximately 55% of observations are outside DLs which vary across sites and over time.
Many researchers in genetics and social science incorporate information about race in their work. However, migrations (historical and forced) and social mobility have brought formerly separated populations of humans together, creating younger generations of individuals who have more complex and diverse ancestry and race profiles than older age groups. Here, we sought to better understand how temporal changes in genetic admixture influence levels of heterozygosity and impact health outcomes. We evaluated variation in genetic ancestry over 100 birth years in a cohort of 35,842 individuals with electronic health record (EHR) information in the Southeastern United States. Using the software STRUCTURE, we analyzed 2,678 ancestrally informative markers relative to three ancestral clusters (African, East Asian, and European) and observed rising levels of admixture for all clinically-defined race groups since 1990. Most race groups also exhibited increases in heterozygosity and long-range linkage disequilibrium over time, further supporting the finding of increasing admixture in young individuals in our cohort. These data are consistent with United States Census information from broader geographic areas and highlight the changing demography of the population. This increased diversity challenges classic approaches to studies of genotype-phenotype relationships which motivated us to explore the relationship between heterozygosity and disease diagnosis. Using a phenome-wide association study approach, we explored the relationship between admixture and disease risk and found that increased admixture resulted in protective associations with female reproductive disorders and increased risk for diseases with links to autoimmune dysfunction. These data suggest that tendencies in the United States population are increasing ancestral complexity over time. Further, these observations imply that, because both prevalence and severity of many diseases vary by race groups, complexity of ancestral origins influences health and disparities.
Although DNAN6-adenine methylation (6mA) is best known in prokaryotes, its presence in eukaryotes has recently generated great interest. Biochemical and genetic evidence supports that AMT1, an MT-A70 family methyltransferase (MTase), is crucial for 6mA deposition in unicellular eukaryotes. Nonetheless, the 6mA transmission mechanism remains to be elucidated. Taking advantage of single-molecule real-time circular consensus sequencing (SMRT CCS), here we provide definitive evidence for semiconservative transmission of 6mA inTetrahymena thermophila. In wild-type (WT) cells, 6mA occurs at the self-complementary ApT dinucleotide, mostly in full methylation (full-6mApT); after DNA replication, hemi-methylation (hemi-6mApT) is transiently present on the parental strand, opposite to the daughter strand readily labeled by 5-bromo-2′-deoxyuridine (BrdU). In ΔAMT1cells, 6mA predominantly occurs as hemi-6mApT. Hemi-to-full conversion in WT cells is fast, robust, and processive, whereas de novo methylation in ΔAMT1cells is slow and sporadic. InTetrahymena, regularly spaced 6mA clusters coincide with the linker DNA of nucleosomes arrayed in the gene body. Importantly, in vitro methylation of human chromatin by the reconstituted AMT1 complex recapitulates preferential targeting of hemi-6mApT sites in linker DNA, supporting AMT1's intrinsic and autonomous role in maintenance methylation. We conclude that 6mA is transmitted by a semiconservative mechanism: full-6mApT is split by DNA replication into hemi-6mApT, which is restored to full-6mApT by AMT1-dependent maintenance methylation. Our study dissects AMT1-dependent maintenance methylation and AMT1-independent de novo methylation, reveals a 6mA transmission pathway with a striking similarity to 5-methylcytosine (5mC) transmission at the CpG dinucleotide, and establishes 6mA as a bona fide eukaryotic epigenetic mark.
Cumulative probability models (CPMs) are a robust alternative to linear models for continuous outcomes. However, they are not feasible for very large datasets due to elevated running time and memory usage, which depend on the sample size, the number of predictors, and the number of distinct outcomes. We describe three approaches to address this problem. In the divide-and-combine approach, we divide the data into subsets, fit a CPM to each subset, and then aggregate the information. In the binning and rounding approaches, the outcome variable is redefined to have a greatly reduced number of distinct values. We consider rounding to a decimal place and rounding to significant digits, both with a refinement step to help achieve the desired number of distinct outcomes. We show with simulations that these approaches perform well and their parameter estimates are consistent. We investigate how running time and peak memory usage are influenced by the sample size, the number of distinct outcomes, and the number of predictors. As an illustration, we apply the approaches to a large publicly available dataset investigating matrix multiplication runtime with nearly one million observations.
Clustered data are common in biomedical research. Observations in the same cluster are often more similar to each other than to observations from other clusters. The intraclass correlation coefficient (ICC), first introduced by R. A. Fisher, is frequently used to measure this degree of similarity. However, the ICC is sensitive to extreme values and skewed distributions, and depends on the scale of the data. It is also not applicable to ordered categorical data. We define the rank ICC as a natural extension of Fisher's ICC to the rank scale, and describe its corresponding population parameter. The rank ICC is simply interpreted as the rank correlation between a random pair of observations from the same cluster. We also extend the definition when the underlying distribution has more than two hierarchies. We describe estimation and inference procedures, show the asymptotic properties of our estimator, conduct simulations to evaluate its performance, and illustrate our method in three real data examples with skewed data, count data, and three-level ordered categorical data.
The Vanderbilt-Nigeria Biostatistics Training Program (VN-BioStat) aims to establish a research and training platform for biostatisticians doing HIV-related research in Nigeria, including enhancing mid-level biostatistics capacity through annual workshops. This paper describes findings from the inaugural workshop in Kano, Nigeria. Participants were surveyed before and after the workshop to assess their self-perceived familiarity with and confidence in their abilities to use statistical software and apply specific statistical techniques, as well as to gather feedback regarding the conduct of the workshop and future topic areas. Of the 23 participants enrolled in the workshop, 22 (96%) completed both pre- and post-workshop assessments. In both pre-workshop and post-workshop surveys, participants ranked their confidence in statistical skills using Likert scales. Scores were transformed to a 0-100 scale, and averages computed. Participants also shared open-ended feedback about the workshop and suggested future topic areas. Before the training, the average participant reported having either a "beginner" (30% of participants) or "moderate" (43%) level of familiarity with R. Many participants (65%) rated themselves as having "moderate" or "expert" familiarity with SPSS. Pre-workshop averages for confidence ranged from 26 to 64, with lowest confidence in "expanding continuous covariates in regression models and interpret results" and highest confidence in "fitting and interpreting results from a linear regression model". Post-workshop averages for confidence were all above 70. The lowest post-workshop score (74) was for "fit and interpret results from a semiparametric linear transformation model". The greatest increase in confidence was observed in "expanding continuous covariates in regression models using splines and interpreting results" and the lowest increase was in "fitting and interpreting results from a linear regression model." Participants offered positive feedback on instructor effectiveness (4.9/5) and overall course quality (4.9/5). While the overall course was rated on a 0-100 scale as "moderately difficult" (mean ± SD: 40.5 ± 17.5), the participants felt the course was highly organized (87.7 ± 17.8), and the information was moderately easy to learn (81.9 ± 15.9). Suggestions for future workshops included providing supplementary resources for out-of-classroom learning and releasing codes in advance to enhance participants' preparation. Among suggestions for future workshop topics, 80% of respondents listed survival analysis. Lessons learned provide insight into how short-term training opportunities can be leveraged to build biostatistics capacity in similar settings.
Continuous response data are regularly transformed to meet regression modeling assumptions. However, approaches taken to identify the appropriate transformation can be ad hoc and can increase model uncertainty. Further, the resulting transformations often vary across studies leading to difficulties with synthesizing and interpreting results. When a continuous response variable is measured repeatedly within individuals or when continuous responses arise from clusters, analyses have the additional challenge caused by within-individual or within-cluster correlations. We extend a widely used ordinal regression model, the cumulative probability model (CPM), to fit clustered, continuous response data using generalized estimating equations for ordinal responses. With the proposed approach, estimates of marginal model parameters, cumulative distribution functions , expectations, and quantiles conditional on covariates can be obtained without pretransformation of the response data. While computational challenges arise with large numbers of distinct values of the continuous response variable, we propose feasible and computationally efficient approaches to fit CPMs under commonly used working correlation structures. We study finite sample operating characteristics of the estimators via simulation and illustrate their implementation with two data examples. One studies predictors of CD4:CD8 ratios in a cohort living with HIV, and the other investigates the association of a single nucleotide polymorphism and lung function decline in a cohort with early chronic obstructive pulmonary disease.
Regression models for continuous outcomes frequently require a transformation of the outcome, which is often specified a priori or estimated from a parametric family. Cumulative probability models (CPMs) nonparametrically estimate the transformation by treating the continuous outcome as if it is ordered categorically. They thus represent a flexible analysis approach for continuous outcomes. However, it is difficult to establish asymptotic properties for CPMs due to the potentially unbounded range of the transformation. Here we show asymptotic properties for CPMs when applied to slightly modified data where bounds, one lower and one upper, are chosen and the outcomes outside the bounds are set as two ordinal categories. We prove the uniform consistency of the estimated regression coefficients and of the estimated transformation function between the bounds. We also describe their joint asymptotic distribution, and show that the estimated regression coefficients attain the semiparametric efficiency bound. We show with simulations that results from this approach and those from using the CPM on the original data are very similar when a small fraction of the data are modified. We reanalyze a dataset of HIV-positive patients with CPMs to illustrate and compare the approaches.
PURPOSE:Age-related macular degeneration is a common form of vision loss affecting older adults. The etiology of AMD is multifactorial and is influenced by environmental and genetic risk factors. In this study, we examine how 19 common risk variants contribute to drusen progression, a hallmark of AMD pathogenesis.METHODS:Exome chip data was made available through the International AMD Genomics Consortium (IAMDGC). Drusen quantification was carried out with color fundus photographs using an automated drusen detection and quantification algorithm. A genetic risk score (GRS) was calculated per subject by summing risk allele counts at 19 common genetic risk variants weighted by their respective effect sizes. Pathway analysis of drusen progression was carried out with the software package Pathway Analysis by Randomization Incorporating Structure.RESULTS:We observed significant correlation with drusen baseline area and the GRS in the age-related eye disease study (AREDS) dataset (ρ = 0.175, P = 0.006). Measures of association were not statistically significant between drusen progression and the GRS (P = 0.54). Pathway analysis revealed the cell adhesion molecules pathway as the most highly significant pathway associated with drusen progression (corrected P = 0.02).CONCLUSIONS:In this study, we explored the potential influence of known common AMD genetic risk factors on drusen progression. Our results from the GRS analysis showed association of increasing genetic burden (from 19 AMD associated loci) to baseline drusen load but not drusen progression in the AREDS dataset while pathway analysis suggests additional genetic contributors to AMD risk.