Missing data and measurement errors in covariates are common in many longitudinal studies. In this paper, we propose and explore an approximate maximum likelihood method in the framework of the linear mixed model for longitudinal data with missing responses and covariate measurement errors. To adjust for the measurement error in a covariate, we adopt a regression calibration method. Also, to deal with nonignorable missing data, we adopt a Monte Carlo expectation-maximization method that approximates the maximum likelihood estimators of the model parameters. The finite-sample properties of the proposed estimators are studied using Monte Carlo simulations. An application is provided using actual data obtained from a health survey.
Zero-inflated Poisson (ZIP) models are typically used for analyzing count data with excess zeros. If the data are collected longitudinally, then repeated observations from a given subject are correlated by nature. The ZIP mixed model may be used to deal with excess zeros and correlations among the repeated observations. Also, it is often the case that some follow-up measurements in a longitudinal study are missing. If the missing data are informative or nonignorable, it is necessary to incorporate a missingness mechanism into the observed likelihood function for a valid inference. In this paper, we propose and explore an efficient method for analyzing count data by addressing the complex issues of excess zeros, correlations among repeated observations, and missing responses due to dropouts. The empirical properties of the proposed estimators are studied based on Monte Carlo simulations. An application is provided using some real data obtained from a health study.
In this article, we develop an innovative, robust method for jointly analyzing longitudinal count and binary responses. The method is useful for bounding the influence of potential outliers in the data when estimating the model parameters. We use a log-linear model for the count response and a logistic regression model for the binary response, where the two response processes are linked through a set of association parameters. The asymptotic properties of the robust estimators are briefly studied. The empirical properties of the estimators are studied based on simulations. The study shows that the proposed estimators are approximately unbiased and also efficient when fitting a joint model to data contaminated with outliers. We also apply the proposed method to some real longitudinal survey data obtained from a health study. L'auteur de ce travail propose une nouvelle approche robuste et innovante pour l'analyse conjointe de donn & eacute;es longitudinales discr & egrave;tes, combinant des r & eacute;ponses de d & eacute;nombrement et binaires. Cette m & eacute;thodologie vise & agrave; r & eacute;duire l'impact n & eacute;gatif potentiel des valeurs aberrantes lors de l'estimation des param & egrave;tres du mod & egrave;le conjoint. Le cadre propos & eacute; repose sur un mod & egrave;le log-lin & eacute;aire pour la composante de comptage et une r & eacute;gression logistique pour la composante binaire, les deux processsus & eacute;tant li & eacute;s par un ensemble de param & egrave;tres d'association. Les propri & eacute;t & eacute;s asymptotiques des estimateurs robustes d & eacute;velopp & eacute;s sont bri & egrave;vement examin & eacute;es. Par le biais de simulations num & eacute;riques, les auteurs & eacute;tudient le comportement empirique de leurs estimateurs, et montrent que ceux-ci sont approximativement sans biais et efficaces lors de l'ajustement d'un mod & egrave;le conjoint & aacute; des donn & eacute;es contamin & eacute;es par des valeurs aberrantes. Enfin, l'auteur applique son approche & agrave; un ensemble r & eacute;el de donn & eacute;es longitudinales issues d'une enqu & ecirc;te de sant & eacute;.
In clinical studies, often values of a covariate or biomarker are left-censored due to the limit of detection (LOD). An ordinary regression approach that fits a model by simply replacing the left-censored values of the covariate by the LOD or (1/2) LOD generally produces a biased estimator of the covariate effect. In addition, if a covariate is subject to the measurement error, then a naive approach that does not correct for the measurement error can produce an asymptotically biased estimator. In this paper, we propose and explore an innovative method for fitting a logistic regression model to binary data by correcting for both limits of detection and measurement errors in covariates. The finite-sample properties of the proposed estimators are investigated using Monte Carlo simulations. The empirical results are very encouraging, as the proposed method appears to provide unbiased and efficient estimators in the presence of covariates that are subject to the LOD and measurement error. An application is also provided using some actual cardiovascular fitness data obtained from a health survey with measurements on biomarkers and demographic variables.
We discuss the construction of optimal allocation schemes for the linear mixed model with clustered outcomes or repeated measurements often encountered in longitudinal studies. We consider both treatment and covariate effects in the mixed model, where latent pro- cesses are used to describe random cluster or subject effects. A goal of optimal design schemes is to determine proportions of sample units allocated to each treatment for a given total sample size. We develop the optimal designs in a general setting using both D- and A- optimal design criteria. Specifically, we propose a two-stage design approach to deal with unknown parameters in the linear mixed model, where the variances of the random effects across the treatment groups are considered different. We study the empirical properties of the proposed designs using Monte Carlo simulations. An application is also provided using actual clinical data from a longitudinal study. Journal of Statistical Research, Vol 56, No 2, p101-114
In many surveys and clinical trials, we obtain measurements on covariates or biomarkers that are left-censored due to the limit of detection. In such cases, it is necessary to correct for the left-censoring when studying covariate effects in regression models. The expectation-maximization (EM) algorithm is widely used for the likelihood inference in generalized linear models with censored covariates. The EM method, however, requires intensive computation involving high-dimensional integration with respect to the covariates when the dimension of the censored covariates is large. To reduce such computational difficulties, we propose and explore a Monte Carlo EM method based on the Metropolis algorithm. The finite-sample properties of the proposed estimators are studied using Monte Carlo simulations. An application is also provided using actual data obtained from a health and nutrition examination survey. Journal of Statistical Research 2021, Vol. 55, No. 2, pp. 359-375
In this paper, we propose and explore a novel semiparametric approach to analyzing longitudinal count data. We address the issue of missingness in longitudinal data and propose a weighted generalized estimations equations approach to fitting marginal mean response models for count responses with dropouts. Also, we investigate a spline regression approach to approximating the curvilinear relationship between the mean response and covariates. The asymptotic properties of the proposed estimators are studied in some detail. The empirical properties of the estimators are investigated using Monte Carlo simulations. An application is also provided using actual survey data obtained from the Health and Retirement Study (HRS).
We study robust designs for generalized linear mixed models (GLMMs) with protections against possible departures from underlying model assumptions. Among various types of model departures, an imprecision in the assumed linear predictor or the link function has a great impact on predicting the conditional mean response function in a GLMM. We develop methods for constructing adaptive sequential designs when the fitted mean response or the link function is possibly of an incorrect parametric form. We adopt the maximum likelihood method for estimating the parameters in GLMMs and investigate both I-optimal and D-optimal design criteria for the construction of robust sequential designs. To study the empirical properties of these sequential designs, we ran a series of simulations using both logistic and Poisson mixed models. As indicated in the simulation results, the I-optimal design generally outperforms the D-optimal design for all scenarios considered. Both designs are more efficient than the conventionally used uniform design and the classical D-optimal design obtained under the assumption that the fitted models are correctly specified. The proposed designs are also illustrated in an example using actual data from a dose–response experiment.
In this article, we investigate marginal models for analyzing incomplete longitudinal count data with dropouts. Specifically, we explore commonly used generalized estimating equations and weighted generalized estimating equations for fitting log-linear models to count data in the presence of monotone missing responses. A series of simulations were carried out to examine the finite-sample properties of the estimators in the presence of both correctly specified and misspecified dropout mechanisms. An application is provided using actual longitudinal survey data from the Health and Retirement Study (HRS) (HRS, 2019)
In this paper, we propose an innovative method for jointly analyzing survival data and longitudinally measured continuous and ordinal data. We use a random effects accelerated failure time model for survival outcomes, a linear mixed model for continuous longitudinal outcomes and a proportional odds mixed model for ordinal longitudinal outcomes, where these outcome processes are linked through a set of association parameters. A primary objective of this study is to examine the effects of association parameters on the estimators of joint models. The model parameters are estimated by the method of maximum likelihood. The finite-sample properties of the estimators are studied using Monte Carlo simulations. The empirical study suggests that the degree of association among the outcome processes influences the bias, efficiency, and coverage probability of the estimators. Our proposed joint model estimators are approximately unbiased and produce smaller mean squared errors as compared to the estimators obtained from separate models. This work is motivated by a large multicenter study, referred to as the Genetic and Inflammatory Markers of Sepsis (GenIMS) study. We apply our proposed method to the GenIMS data analysis.
We develop and study an innovative method for jointly modeling longitudinal response and time-to-event data with a covariate subject to a limit of detection. The joint model assumes a latent process based on random effects to describe the association between longitudinal and time-to-event data. We study the role of the association parameter on the regression parameters estimators. We model the longitudinal and survival outcomes using linear mixed-effects and Weibull frailty models, respectively. Because of the limit of detection, missing covariate (explanatory variable, x) values may lead to the non-ignorable missing, resulting in biased parameter estimates with poor coverage probabilities of the confidence interval. We define and estimate the probability of missing due to the limit of detection. Then we develop a novel joint density and hence the likelihood function that incorporates the effect of left-censored covariate. Monte Carlo simulations show that the estimators of the proposed method are approximately unbiased and provide expected coverage probabilities for both longitudinal and survival submodels parameters. We also present an application of the proposed method using a large clinical dataset of pneumonia patients obtained from the Genetic and Inflammatory Markers of Sepsis study.
Small area estimation with categorical outcomes often requires intensive computation, as the marginal likelihood does not have a closed form in general. The likelihood analysis is further complicated by deviations in distributional assumptions often arise through outliers in the data. In this paper, the author proposes a robust method for estimating the small area parameters. Finite-sample properties of the estimators are investigated using Monte Carlo simulations. The empirical study shows that the proposed robust method is very useful for bounding the influence of outliers on the small area estimators. To approximate the mean squared errors of the estimators, a parametric bootstrap method is adopted. An application is also provided using actual data from a public health survey.
The accelerated failure time model is widely used for analyzing censored survival times often observed in clinical studies. It is well-known that the ordinary maximum likelihood estimators of the parameters in the accelerated failure time model are generally sensitive to potential outliers or small deviations from the underlying distributional assumptions. In this paper, we propose and explore a robust method for fitting the accelerated failure time model to survival data by bounding the influence of outliers in both the outcome variable and associated covariates. We also develop a sandwich-type variance–covariance function for approximating the variances of the proposed robust estimators. The finite-sample properties of the estimators are investigated based on empirical results from an extensive simulation study. An application is provided using actual data from a clinical study of primary breast cancer patients.
There is growing evidence on the association between prenatal metal exposures and adverse pregnancy outcomes. Heavy metals such as arsenic (As), cadmium (Cd), mercury (Hg) and lead (Pb) are considered as endocrine disrupting chemicals and elevated exposure to these metals during pregnancy is associated with adverse effects on maternal and infant health. Nevertheless, mechanistic understanding is required to establish biological plausibility of such associations. The objective of this study was to gain insight into metal exposure-related adverse birth outcomes by understanding maternal systemic changes at the molecular level.The Maternal-Infant Research on Environmental Chemicals (MIREC) study was employed for this purpose. Third trimester maternal plasma samples were analysed for target oxidative /nitrative stress markers (e.g 8-isoprostane, 3-nitrotyrosine) by competitive enzyme immunoassay and HPLC-Coularray as well as matrix metalloproteinases, a class of enzymes and other markers of inflammation (e.g. Cytokines, cellular adhesion molecules) were measured by affinity-based multiplex array and HPLC-Fluorescence detection methods. Pearson product moment correlations, chi-squared tests and multivariate models were used to analyse the associations among maternal blood metal (Cd, Hg, Pb, As, manganese Mn) levels, plasma biomarkers, physiological changes and birth weight. Our results revealed maternal metal exposure-specific responses (p<0.05) on markers of oxidative stress pathways (e.g 8-isoprostane) and matrix metalloproteinases (MMPs), in maternal circulation. Interestingly, statistically significant (p<0.05) correlations were seen between oxidative stress pathways, MMPs and other inflammatory mediators relevant to infant birth weight changes. Our findings imply that metal exposures potentially can mediate maternal oxidative stress pathways which can alter MMP profiles and associated inflammatory processes, thus adversely impacting on infant birth weights.
In many biological experiments, certain values of a biomarker are often nondetectable due to low concentrations of an analyte or the limitations of a chemical analysis device, resulting in left‐censored values. There is an increasing demand for the analysis of data subject to detection limits in clinical and environmental studies. In this paper, we develop a novel statistical method for the maximum likelihood estimation in generalized linear models with covariates subject to detection limits. Simulations are carried out to study the relative performance of the proposed estimators, as compared to other existing estimators. The proposed method is also applied to a real dataset from the Maternal‐Infant Research on Environmental Chemicals cohort study, where we investigate how different chemical mixtures affect the health outcomes of infants and pregnant women.
Depending on the chemical and the outcome, prenatal exposures to environmental chemicals can lead to adverse effects on the pregnancy and child development, especially if exposure occurs during early gestation. Instead of focusing on prenatal exposure to individual chemicals, more studies have taken into account that humans are exposed to multiple environmental chemicals on a daily basis. The objectives of this analysis were to identify the pattern of chemical mixtures to which women are exposed and to characterize women with elevated exposures to various mixtures. Statistical techniques were applied to 28 chemicals measured simultaneously in the first trimester and socio-demographic factors of 1744 participants from the Maternal-Infant Research on Environment Chemicals (MIREC) Study. Cluster analysis was implemented to categorize participants based on their socio-demographic characteristics, while principal component analysis (PCA) was used to extract the chemicals with similar patterns and to reduce the dimension of the dataset. Next, hypothesis testing determined if the mean converted concentrations of chemical substances differed significantly among women with different socio-demographic backgrounds as well as among clusters. Cluster analysis identified six main socio-demographic clusters. Eleven components, which explained approximately 70% of the variance in the data, were retained in the PCA. Persistent organic pollutants (PCB118, PCB138, PCB153, PCB180, OXYCHLOR and TRANSNONA) and phthalates (MEOHP, MEHHP and MEHP) dominated the first and second components, respectively, and the first two components explained 25.8% of the source variation. Prenatal exposure to persistent organic pollutants (first component) were positively associated with women who have lower education or higher income, were born in Canada, have BMI ≥25, or were expecting their first child in our study population. MEOHP, MEHHP and MEHP, dominating the second component, were detected in at least 98% of 1744 participants in our cohort study; however, no particular group of pregnant women was identified to be highly exposed to phthalates. While widely recognized as important to studying potential health effects, identifying the mixture of chemicals to which various segments of the population are exposed has been problematic. We present an approach using factor analysis through principal component method and cluster analysis as an attempt to determine the pregnancy exposome. Future studies should focus on how to include these matrices in examining the health effects of prenatal exposure to chemical mixtures in pregnant women and their children.
We discuss optimal sequential designs for analyzing longitudinal data or repeated measurements in the framework of generalized linear mixed models (GLMMs). The construction of optimal designs for GLMMs typically requires intensive computation involving irreducibly high-dimensional integrals with respect to the random effects when calculating the Fisher information. In this article, we propose a simpler method for constructing optimal sequential designs based on some “pseudo” Fisher information. We investigate the performance of the proposed designs through simulation studies. The simulation results indicate that the loss of efficiency from the pseudo-optimal design construction process is minimal as compared to the computationally intensive D-optimal design. Furthermore, the proposed design appears to be as robust as the D-optimal design when possible misspecification to the assumed random effects distribution occurs. It is concluded that the proposed approach is of practical value for designing repeated-measures experiments.
In this paper, the authors investigate a robust semi-parametric mixed effects model for analyzing longitudinal data with an unspecified mean response function. The robust method, developed in the framework of the maximum likelihood, is used to bound the influence of potential outliers when estimating the model parameters. The authors also present a robust test procedure for assessing the significance of a variance component in the mixed model. An application is provided using a clinical dataset from a retinopathy of prematurity study in which longitudinal measurements were obtained from premature infants treated with supplemental oxygen. The empirical properties of the proposed estimators are also studied in simulations.
In this paper, we propose and explore a semiparametric approach to analyzing longitudinal binary data often observed in clinical studies. We applied second-order GEE approach to analyze longitudinal binary responses based on a partially linear single-index model. We use a local polynomial smoothing technique to estimate the single-index. We study the empirical properties of the proposed estimators using simulations. The empirical results demonstrate that if the true underlying model is partially linear, then our proposed method generally provides unbiased and efficient estimators. The proposed method is also applied to some real data sets obtained from longitudinal studies.