
Studies of HPV vaccine efficacy usually record infections with vaccine targeted and non-targeted strains. Contrary to blinded randomized controlled trials, confounding bias can be a threat and risk compensation may occur in observational studies. Etievant et al. (Biometrics, 2023) proposed to use cervical infections with non-targeted HPV strains to remove or reduce confounding bias in estimates of vaccine efficacy on targeted strains. However, they assumed that vaccinated women could not change their behavior after vaccination. This work investigates if the quantity estimated in practice with the method of Etievant et al. has a clear causal meaning under a more plausible setting where unmeasured sexual behavior acts as both a confounder and a mediator. Under certain assumptions, using non-targeted HPV infections can remove both confounding bias and the portion of the vaccine effect on the targeted HPV strains that is mediated through the change of behavior. In that case, the estimated quantity has a clear causal interpretation as it represents the direct immunological effect of the vaccine. Infections with non-targeted HPV strains could also be used to isolate the indirect behavioral effect of the vaccine in unblinded randomized controlled trial.
Targeted Maximum Likelihood Estimation, also referred to as Targeted Minimum Loss Estimation (TMLE), is a statistical method that enables causal inference while incorporating flexible, advanced machine learning techniques. Within the drug development space, both drug sponsors and FDA have leveraged TMLE across a range of relevant clinical settings intended for a variety of regulatory activities. The goal of this article is to provide a holistic overview of TMLE use cases as they have been applied in submissions to the Food and Drug Administration (FDA) Center for Drug Evaluation and Research (CDER) and in CDER's regulatory research activities. First, we provide an overview of regulatory submissions in which TMLE has been considered to support regulatory decision-making. Second, we summarize FDA CDER's regulatory science initiatives that advance analytic capabilities and support the use of TMLE.
Functional data analysis (FDA) provides a powerful statistical framework for analyzing complex data, such as curves or functions, over high-dimensional domains. In this paper, we focus on functional predictor regression (scalar-on-function) models applied to the prediction of birth weight using maternal dietary patterns during pregnancy. Specifically, we analyze trajectories of nine components of the Alternative Healthy Eating Index (AHEI) as multivariate functional predictors. We adopt Multivariate Functional Principal Component Analysis (MVFPCA) to obtain low-dimensional, uncorrelated representations of the multivariate functional predictors through their FPC scores. Building on MVFPCA, we develop a novel multivariate functional principal component regression (MVFPCR) model to predict birth weight effectively. Our method accommodates various regression approaches, including linear and quantile regression, depending on the distribution of birth weights. Through simulation studies and the application to a fetal growth study dataset, we demonstrate the effectiveness of our proposed model in functional predictor regression.
In right-skewed count data, the mean is disproportionately affected by a long upper tail, whereas the median remains a more representative measure of central tendency. Discrete Weibull (DW) regression links covariates to a shifted median, which in turn induces the exact integer median; however, a single DW component can fit poorly when the observed count distribution has a markedly heavier upper tail than a single-component model can accommodate. We propose a contaminated DW (cDW) regression that augments the baseline DW distribution with a more dispersed secondary component within a finite mixture while retaining a single shifted-median link. This mixture accommodates extreme counts more effectively, thereby stabilizing the median-based regression coefficients. The model accommodates general lower truncation at an arbitrary threshold c, including c = 1 for strictly positive outcomes and c = 0 for nonnegative counts, and is estimated using a straightforward Bayesian Markov chain Monte Carlo algorithm implemented in JAGS; R code accompanies the paper. Applied to hospital length-of-stay data, the cDW regression reduces the influence of outliers and achieves superior predictive performance relative to a single-component DW model, as demonstrated by leave-one-out cross-validation and a Kullback-Leibler influence diagnostic. Simulation experiments show that, under strongly heavy-tailed mixture settings, the cTDW model accurately recovers the regression coefficients and improves on the single-component TDW model. Because the added tail component can increase probability mass at both extremes, we further recommend embedding the cDW in a hurdle framework when structural zeros are present: the zero probability is modeled separately, and the heavy-tail mixture is applied only to positive counts. The cDW regression model provides a robust, median-centered alternative for analyzing skewed, possibly truncated count outcomes.
Missing data are common in real-world studies, yet the underlying missingness structure is often unknown, bringing additional uncertainty before an appropriate inference method can be applied. In this paper, we systematically examine two sources of such uncertainty: (1) the missing mechanism(s) involved and (2) the specific functional form of the missing model within a given mechanism. Focusing particularly on settings involving missing not at random (MNAR) data, we propose a two-step mixture-structure-based method, including a model filtering pre-screening step. The tasks of handling both sources of uncertainty and conducting reliable inference are then unified within a single EM-based framework. The core of our method lies in constructing a two-layer postulated mixture, which can be viewed as deliberately introducing an overfitted mixture - thereby enhancing flexibility and robustness to uncertainty. We consider two general scenarios in which the true data follow either a mixture or a non-mixture structure, and establish an identification framework for continuous finite mixtures potentially subject to MNAR. Simulation studies and a real-data application to the Medical Expenditure Panel Survey (MEPS) are utilized to demonstrate the performance of our method.
The prognostic score (PGS) is a function of observed covariates that summarizes covariates' association with potential responses. In the current study, we propose a full prognostic score (FPGS), an extension of the PGS that integrates individual prognostic scores to account for confounding adjustments in causal inference. Under effect modification, we show that FPGS and a version of FPGS using conditional expectations of the outcomes, meet the sufficiency condition for confounding adjustment to estimate the average causal effect. We present a general algorithm to implement the FPGS approach for estimation by applying linear regression, random forest regression, and XGBoost regression. When determining the average causal effect, we incorporate FPGS into semiparametric estimators including regression imputation, simple stratification, and targeted maximum likelihood estimation (TMLE). The finite-sample properties of the estimators are compared through three simulation studies. Based on the findings of the FPGS estimators, the mean squared error of the linear regression imputation estimator and the TMLE estimator comprised of linearly regressed PGS is smaller than the mean squared error of alternative estimators. In an empirical study, we analyze data from the National Health and Nutrition Examination Survey (NHANES, 2007-2008) to determine the effect of smoking on blood lead levels.
DNA methylation (DNAm) is a key epigenetic modification, and datasets capturing DNAm are typically high-dimensional. Although dimension reduction (DR) techniques are commonly applied, it remains unclear how different DR methods perform specifically on the DNAm dataset. In this study, we aim to evaluate the performance of several DR techniques to determine their performance for reducing the dimensionality of DNAm datasets. We leveraged the DNAm dataset from the STRiDE (STratification of Risk of Diabetes in Early pregnancy) prospective study, which consists of 8,62,927 CpG sites each from 258 pregnant women. Women were categorized as Normal (n = 146); and GDM (n = 112). Epigenome-wide DNA methylation profiles from peripheral blood were quantified using an Infinium Methylation EPIC array. All stated DR techniques were performed, and the retained amount of information, local neighborhood preservation criteria, and global structure-holding approaches were assessed using various statistical measures and compared. Across Shannon entropy, local-neighborhood metrics (König's measure, Spearman's ρ, trustworthiness/continuity) and global-structure metrics (Kruskal stress, Sammon's stress, residual variance), MDS and PCA consistently achieved the best performance. PLS-DA trailed closely, while ISOMAP showed moderate results and UMAP performed worst, exhibiting higher entropy, lower correlation preservation, and greater distortion of both local and global structures.
Quantifying disease-specific survival in patients with competing risks is generally done when reliable cause of death (CoD) information is available. With known CoD, cause-specific and cumulative incidence functions for competing risk data are applicable in estimating disease-specific survival. When CoD is unreliable, unknown, or subject to misspecifications, relative survival methods are used for estimating disease-specific survival. This estimator, under the independent competing risks assumption, is the ratio of all-cause survival in the disease-specific cohort group to the known expected survival from a general reference population. The disease-specific death competes with other causes of mortality, potentially creating interdependence among the CoD. The standard ratio estimate is only valid when death from disease and death from competing causes are independent. We relaxed this assumption by formulating the dependence between the times to disease-specific death and competing causes of mortality using a copula. We fit a nonparametric copula-based approach to the distribution of disease-specific death which reduces to the ratio estimator under independence. This nonparametric method is robust compared to the previously proposed copula-based parametric method. We demonstrate the utility of our method through simulation studies and an application to French breast cancer registry data.
Pooled analyses that aggregate data from multiple studies are becoming increasingly common in collaborative epidemiologic research in order to increase the size and diversity of the study population. However, biomarker measurements from different studies are subject to systematic measurement errors and directly pooling them for analyses may lead to biased estimates of the regression parameters. Therefore, study-specific calibration processes must be incorporated in the statistical analyses to address between-study/assay/laboratory variability in the biomarker measurements. We propose a likelihood-based method to evaluate biomarker-disease relationships for categorical biomarkers in matched/nested case-control studies. To account for the additional uncertainties from the calibration processes, we propose a sandwich variance estimator to obtain valid asymptotic variances of the estimated regression parameters. Extensive simulation studies with varying sample sizes and biomarker-disease associations are used to evaluate the finite sample performance of our proposed methods. As an illustration, we apply the methods to a vitamin D pooling project of colorectal cancer to evaluate the effect of categorical vitamin D levels on colorectal cancer risks.
This paper considers semiparametric estimation strategies for the nonlinear semiparametric regression model (NSRM) under the sparsity assumption by modifying the Gauss-Newton method for both low- and high-dimensional data scenarios. In the low-dimensional case, coefficients are partitioned into two parts that represent nonzero (strong signals) and sparse coefficients. In the high-dimensional case, a weighted-ridge approach is employed, and coefficients are partitioned into three parts, adding weak signals as well. Shrinkage estimators are then obtained in both cases. More importantly, in this paper, we assume that a nonlinear structure is present in the parametric component of the model, which makes the direct application of penalized least squares to the NSRM impossible. To solve this problem, we employ the iterative Gauss-Newton method to obtain the final NSRM estimators. We provide both theoretical and practical details for the suggested estimators. Asymptotic results are derived for both low- and high-dimensional cases. We conduct an extensive simulation study to evaluate the performance of the estimators in a practical setting. Moreover, we substantiate our findings with data examples from two distinct breast cancer datasets: the Breast Cancer in the United States (BCUS) and Wisconsin datasets. By demonstrating the effectiveness of our introduced estimators in these particular biostatistical contexts, our numerical study provides support for the theoretical efficacy of shrinkage estimators, suggesting their potential relevance to breast cancer research and biostatistical methodologies.
Information borrowing from historical data is gaining increasing attention in clinical trials for rare and pediatric diseases, where small sample sizes may lead to insufficient statistical power for confirming efficacy. While Bayesian information borrowing methods are well established, recent frequentist approaches, such as the test-then-pool and equivalence-based test-then-pool methods, have been proposed to determine whether historical data should be incorporated into statistical hypothesis testing. Depending on the outcome of these hypothesis tests, historical data may or may not be utilized. This paper introduces a dynamic borrowing method for leveraging historical information based on the similarity between current and historical data. Similar to Bayesian dynamic borrowing, our proposed method adjusts the degree of information borrowing dynamically, ranging from 0 to 100 %. We present two approaches to measure similarity: one using the density function of the t-distribution and the other employing a logistic function. The performance of the proposed methods is evaluated through Monte Carlo simulations. Additionally, we demonstrate the utility of dynamic information borrowing by reanalyzing data from an actual clinical trial.
The binary classification problem (BCP) aims to correctly allocate subjects in one of two possible groups. The groups are frequently defined as having or not one characteristic of interest. With this goal, we are allowed to use different types of information. There is a huge number of methods dealing with this problem; including standard binary regression models, or complex machine learning techniques such as support vector machine, boosting, or perceptron, among others. When this information is summarized in a continuous score, we have to define classification regions (or subsets) which will determine whether the subjects are classified as positive, with the characteristic under study, or as negative, otherwise. The standard (or regular) receiver-operating characteristic (ROC) curve assumes that higher values of the marker are associated with higher probabilities of being positive and considers as positive those patients with values within the intervals [c, ∞) ( c ∈ R ) , and plots the true- against the false- positive rates (sensitivity against one minus specificity) for all potential c. The so-called generalized ROC curve, gROC, allows that both higher and lower values of the score are associated with higher probabilities of being positive. The efficient ROC curve, eROC, considers the best ROC curve based on a transformation of the score. In this manuscript, we are interested in studying, comparing and approximating the transformations leading to the eROC and to the gROC curves. We will prove that, when the optimal transformation does not have relative maximum, both curves are equivalent. Besides, we investigate the use of the gROC curve on some theoretical models, explore the relationship between the gROC and the eROC curves, and propose two non-parametric procedures for approximating the transformation leading to the gROC curve. The finite-sample behavior of the proposed estimators is explored through Monte Carlo simulations. Two real-data sets illustrate the practical use of the proposed methods.
In observational studies, the treatment assignment is typically not random. Even in randomized clinical trials, the randomization may be imperfect given the limitation of sample size. In these cases, traditional statistical methods may lead to biased estimates of treatment effects, and causal inference methods are needed to obtain unbiased estimates. The doubly robust estimator (DRE) is a recent development in causal inference, but the literature on DRE for survival data is very limited, and existing methods tend to have complicated forms and may not have double robustness in the original sense. Some are constructed based on the Nelson-Aalen estimator, and to our knowledge no DRE is constructed based on the Kaplan-Meier estimator. Furthermore, in these methods, the propensity score model is often subjectively specified with a logistic model. DRE can be seriously biased if the propensity score and outcome models are slightly misspecified. Here we propose a new semiparametric robust estimator that utilizes the Kaplan-Meier estimator and Stute weighted empirical form to address these issues. Our proposed estimator is not only doubly robust in the original sense but also enhances robustness with the use of semiparametric specification. The asymptotic properties of the proposed estimator are derived, and extensive simulation studies are conducted to evaluate its finite sample performance and compare it with existing methods. Finally, we apply our proposed method to a real clinical study.
Traditional survival analysis typically assumes that all subjects will eventually experience the event of interest given a sufficiently long follow-up period. Nevertheless, due to advancements in medical technology, researchers now frequently observe that some subjects never experience the event and are considered cured. Furthermore, traditional survival analysis assumes independence between failure time and censoring time. However, practical applications often reveal dependence between them. Ignoring both the cured subgroup and this dependence structure can introduce bias in model estimates. Among the methods for handling dependent censoring data, the numerical integration process of frailty models is complex and sensitive to the assumptions about the latent variable distribution. In contrast, the copula method, by flexibly modeling the dependence between variables, avoids strong assumptions about the latent variable structure, offering greater robustness and computational feasibility. Therefore, this paper proposes a copula-based method to handle dependent current status data involving a cure fraction. In the modeling process, we establish a logistic model to describe the susceptible rate and a Cox proportional hazards model to describe the failure time and censoring time. In the estimation process, we employ a sieve maximum likelihood estimation method based on Bernstein polynomials for parameter estimation. Extensive simulation experiments show that the proposed method demonstrates consistency and asymptotic efficiency under various settings. Finally, this paper applies the method to lymph follicle cell data, verifying its effectiveness in practical data analysis.
Multi-stage models for cohort data are widely used in various fields, including disease progression, the biological development of plants and animals, and laboratory studies of life cycle development. However, the likelihood functions of these models are often intractable and complex. These complexities in the likelihood functions frequently result in significant biases and high computational costs when estimating parameters using current Bayesian methods. This paper aims to address these challenges by applying the enhanced Sequential Monte Carlo approximate Bayesian computation (ABC-SMC) method, which does not rely on explicit likelihood functions, to stage-structured development models with non-hazard rates and stage-wise constant hazard rates. Instead of using a likelihood function, the proposed method determines parameter estimates based on matching vector summary statistics. It incorporates stage-wise parameter estimations and retains accepted parameters across stages. This approach not only reduces model biases but also improves the computational efficiency of parameter estimations, despite the computational intractability of the likelihood functions. The proposed ABC-SMC method is validated through simulation studies on stage-structured development models and applied to a case study of breast development in New Zealand schoolgirls. The results demonstrate that the proposed methods effectively reduce biases in later-stage estimates for stage-structured models, enhance computational efficiency, and maintain accuracy and reliability in parameter estimations compared to the current methods.
Variable selection is challenging for high-dimensional data, in particular when sample size is low. It is widely recognized that external information in the form of complementary data on the variables, 'co-data', may improve results. Examples are known variable groups or p-values from a related study. Such co-data are ubiquitous in genomics settings due to the availability of public repositories, and is likely equally relevant for other applications. Yet, the uptake of prediction methods that structurally use such co-data is limited. We review guided adaptive shrinkage methods: a class of regression-based learners that use co-data to adapt the shrinkage parameters, crucial for the performance of those learners. We discuss technical aspects, but also the applicability in terms of types of co-data that can be handled. This class of methods is contrasted with several others. In particular, group-adaptive shrinkage is compared with the better-known sparse group-lasso by evaluating variable selection. Moreover, we demonstrate the versatility of the guided shrinkage methodology by showing how to 'do-it-yourself': we integrate implementations of a co-data learner and the spike-and-slab prior for the purpose of improving variable selection in genetics studies. We conclude with a real data example.
In this paper, a two-sample empirical likelihood method for right censored data is established. This method allows for comparisons between various functionals of survival distributions, such as mean lifetimes, survival probabilities at a fixed time, restricted mean survival times, and other parameters of interest. It is demonstrated that under some regularity conditions, the scaled empirical likelihood statistic converges to a chi-squared distributed random variable with one degree of freedom. A consistent estimator for the scaling constant is proposed, involving the jackknife estimator of the asymptotic variance of the Kaplan-Meier integral. A simulation study is carried out to investigate the coverage accuracy of confidence intervals. Finally, two real datasets are analyzed to illustrate the application of the proposed method.
The quantification of overlap between two distributions has applications in various fields of biology, medical, genetic, and ecological research. In this article, new overlap and containment indices are considered for quantifying the niche overlap between two species/populations. Some new properties of these indices are established and the problem of estimation is studied, when the two distributions are exponential with different scale parameters. We propose several estimators and compare their relative performance with respect to different loss functions. The asymptotic normality of the maximum likelihood estimators of these indices is proved under certain conditions. We also obtain confidence intervals of the indices based on three different approaches and compare their average lengths and coverage probabilities. The point and confidence interval procedures developed here are applied on a breast cancer data set to analyze the similarity between the survival times of patients undergoing two different types of surgery. Additionally, the similarity between the relapse free times of these two sets of patients is also studied.
Hyponatremia, characterized by a serum sodium concentration below 135 mEq/L, is a prevalent electrolyte imbalance associated with increased morbidity and mortality across various clinical conditions. This study employs the Holt-Winters seasonal method, a robust time series forecasting model, to predict mortality rates attributed to hyponatremia. Leveraging retrospective mortality data from a cohort of hospitals in the United States, our analysis aims to elucidate temporal patterns and trends in hyponatremia-related deaths. The findings underscore the critical role of statistical forecasting in healthcare, facilitating proactive resource allocation and targeted interventions to mitigate mortality risks associated with electrolyte imbalances. Integrating predictive analytics into clinical practice holds promise for enhancing patient care and optimizing health outcomes in populations vulnerable to hyponatremia-related complications.
Phase I trials aim to identify the maximum tolerated dose (MTD) early and proceed quickly to an expansion cohort or a Phase II trial to assess the efficacy of the treatment. We present an early completion method based on multiple dosages (adjacent dose information) to accelerate the identification of the MTD in model-assisted designs. By using not only toxicity data for the current dose but also toxicity data for the next higher and lower doses, the MTD can be identified early without compromising accuracy. The early completion method is performed based on dose-assignment probabilities for multiple dosages. These probabilities are straightforward to calculate. We evaluated the early completion method using from an actual clinical trial. In a simulation study, we evaluated the percentage of correct MTD selection and the impact of early completion on trial outcomes. The results indicate that our proposed early completion method maintains a high level of accuracy in MTD selection, with minimal reduction compared to the standard approach. In certain scenarios, the accuracy of MTD selection even improves under the early completion framework. We conclude that the use of this early completion method poses no issue when applied to model-assisted designs.