Despite the widespread use of time-to-event data in precision medicine, existing research has often neglected the presence of the cure fraction, assuming that all individuals will inevitably experience the event of interest. When a cure fraction is present, the cure rate and survival time of uncured patients should be considered in estimating the optimal individualized treatment regimes. In this study, we propose direct methods for estimating the optimal individualized treatment regimes that either maximize the cure rate or mean survival time of uncured patients. Additionally, we propose two optimal individualized treatment regimes that balance the tradeoff between the cure rate and mean survival time of uncured patients based on a constrained estimation framework for a more comprehensive assessment of individualized treatment regimes. This framework allows us to estimate the optimal individualized treatment regime that maximizes the population's cure rate without significantly compromising the mean survival time of those who remain uncured or maximizes the mean survival time of uncured patients while having the cure rate controlled at a desired level. The exterior-point algorithm is adopted to expedite the resolution of the constrained optimization problem and statistical validity is rigorously established. Furthermore, the advantages of the proposed methods are demonstrated via simulations and analysis of esophageal cancer data.
The segmented model has significant applications in scientific research when the change-point effect exists. In this article, we propose a comprehensive semiparametric framework in segmented models to test the existence and estimate the location of change points in the generalized outcome setting. The proposed framework is based on a semismooth estimating equation for the change-point estimation and an average score-type test for hypothesis testing. The root-n consistency, asymptotic normality, and asymptotic efficiency of estimators for all parameters in the segmented model are rigorously studied. The distribution of the average score-type test statistics under the null hypothesis is rigorously derived. Extensive simulation studies are conducted to assess the numerical performance of the proposed change-point estimation method and the average score-type test. We investigate change-point effects of baseline glomerular filtration rate and body mass index on bleeding after intervention using data from Blue Cross Blue Shield. This application study successfully identifies statistically significant change-point effects, with the estimated values providing clinically meaningful insights.
近年备受关注的真实世界研究以反映研究对象实际诊疗与健康状况的多源数据为基础,极大拓展了传统临床研究的适用场景。随着真实世界研究在多个国家儿科领域陆续开展,其在指导儿科临床实践、支持儿童药物研发与监管决策等方面发挥着日益重要的作用。本文系统介绍了国内外儿科真实世界研究现状,并就我国儿科真实世界研究面临的挑战进行探讨。
An optimal individualized treatment regime (ITR) is a decision rule in allocating the best treatment to each patient and, hence, maximizing overall benefits. In this paper, we propose a novel framework based on nonparametric inverse probability weighting (IPW) and augmented inverse probability weighting (AIPW) estimators of the value function when the data are subject to right censoring. In contrast to most existing approaches that are designed to maximize the expected survival time under a binary treatment framework, the proposed method targets maximizing the mean residual lifetime of patients. Specifically, the proposed IPW method searches the optimal ITR by maximizing an estimator for the overall population outcome directly, without specifying the regression model for the conditional mean residual lifetime, whereas the AIPW method integrates the model information of the mean residual lifetime to improve the robustness. Furthermore, to overcome the computational difficulty in a nonsmooth value estimator, smoothed IPW and AIPW estimators are constructed. In theory, we establish the asymptotic properties of the proposed method under suitable regularity conditions. The empirical performances of the proposed IPW and AIPW estimators are evaluated using simulation studies and are further illustrated with an application to the real-world data set from the Acquired Immunodeficiency Syndrome Clinical Trial Group Protocol 175 (ACTG175).
Estimating thresholds when a threshold effect exists has important applications in biomedical research. However, models/methods commonly used in the biomedical literature may lead to a biased estimate. For patients undergoing coronary artery bypass grafting (CABG), it is thought that exposure to low oxygen delivery (DO2) contributes to an increased risk of avoidable acute kidney injury. This research is motivated by estimating the threshold of nadir DO2 for CABG patients to help develop an evidence-based guideline for improving cardiac surgery practices. We review several models (sudden-jump model, broken-stick model, and the constrained broken-stick model) that can be adopted to estimate the threshold and discuss modeling assumptions, scientific plausibility, and implications in estimating the threshold. Under each model, various estimation methods are studied and compared. In particular, under a constrained broken-stick model, a modified two-step Newton-Raphson algorithm is introduced. Through comprehensive simulation studies and an application to data on CABG patients from the University of Michigan, we show that the constrained broken-stick model is flexible, more robust, and able to incorporate scientific knowledge to improve efficiency. The two-step Newton-Raphson algorithm has good computational performances relative to existing methods.
The linear spline model is able to accommodate nonlinear effects while allowing for an easy interpretation. It has significant applications in studying threshold effects and change-points. However, its application in practice has been limited by the lack of both rigorously studied and computationally convenient method for estimating knots. A key difficulty in estimating knots lies in the nondifferentiability. In this article, we study influence functions of regular and asymptotically linear estimators for linear spline models using the semiparametric theory. Based on the theoretical development, we propose a simple semismooth estimating equation approach to circumvent the nondifferentiability issue using modified derivatives, in contrast to the previous smoothing-based methods. Without relying on any smoothing parameters, the proposed method is computationally convenient. To further improve numerical stability, a two-step algorithm taking advantage of the analytic solution available when knots are known is developed to solve the proposed estimating equation. Consistency and asymptotic normality are rigorously derived using the empirical process theory. Simulation studies have shown that the two-step algorithm performs well in terms of both statistical and computational properties and improves over existing methods. Supplementary materials for this article are available online.
In this paper, we provide statistical guarantees for over-parameterized deep nonparametric regression in the presence of dependent data. By decomposing the error, we establish non-asymptotic error bounds for deep estimation, which is achieved by effectively balancing the approximation and generalization errors. We have derived an approximation result for Holder functions with constrained weights. Additionally, the generalization error is bounded by the weight norm, allowing for a neural network parameter number that is much larger than the training sample size. Furthermore, we address the issue of the curse of dimensionality by assuming that the samples originate from distributions with low intrinsic dimensions. Under this assumption, we are able to overcome the challenges posed by high-dimensional spaces. By incorporating an additional error propagation mechanism, we derive oracle inequalities for the over-parameterized deep fitted Q-iteration.
Under the assumption of missing at random, doubly robust (DR) estimators are consistent when either the propensity score or the outcome model is correctly specified. However, despite its appealing theoretic properties, Kang and Schafer (2007) show that the usual augmented inverse probability weighted (AIPW) DR estimator may sometimes exhibit unsatisfying behavior. We propose an alternative DR method for mean estimation. In this method, we do not directly weight outcomes by the inverse of estimated propensity scores and instead a nonparametric kernel regression is used to model residuals from an outcome regression model as a function of propensity scores. The proposed method does not suffer from the instability issue of the usual inverse propensity weighted (IPW) estimator resulted from small estimated propensities. We show that asymptotically the new estimator has the double robustness property. Moreover, we show that it is guaranteed to be more efficient than the usual AIPW DR estimator when the propensity score model is correct but the outcome model is incorrect. Our simulation studies show that it has improved finite sample performance compared to existing DR estimators. Statistica Sinica: Newly accepted Paper (accepted author-version subject to English editing)
When treatment effect heterogeneity exists, identifying the subgroup of patients who would benefit from an active treatment relative to a control is an important question. This article focuses on subgroup identification in the presence of a large dimensional set of covariates, with the number of covariates possibly greater than the sample size. We approach this problem from the perspective of optimal treatment decision rules and propose methods that can simultaneously estimate the treatment decision rule and select prescriptive variables important for treatment decision making and subgroup identification. The proposed methods are built within a robust classification framework based on doubly-robust augmented inverse probability weighted estimators (AIPWE), hence sharing the robustness property. An L1 (lasso-type) penalty is used within the classification framework to target selection of prescriptive variables. We further propose a backward elimination process for fine-tuning selection. The methods can be conveniently implemented by taking advantage of standard software for logistic regression and lasso. The methods are evaluated by extensive simulation studies, which demonstrated the superior and robust performance of the proposed methods relative to existing ones. In addition, the estimated decision rules from the proposed methods are considerably simpler than other methods. We applied various methods to identify the subgroup of patients suitable for each of the two commonly used anticoagulants in terms of bleeding risk for patients with acute myocardial infarction undergoing percutaneous coronary intervention.
Time is of the essence in evaluating potential drugs and biologics for the treatment and prevention of COVID-19. There are currently 876 randomized clinical trials (phase 2 and 3) of treatments for COVID-19 registered on clinicaltrials.gov. Covariate adjustment is a statistical analysis method with potential to improve precision and reduce the required sample size for a substantial number of these trials. Though covariate adjustment is recommended by the U.S. Food and Drug Administration and the European Medicines Agency, it is underutilized, especially for the types of outcomes (binary, ordinal, and time-to-event) that are common in COVID-19 trials. To demonstrate the potential value added by covariate adjustment in this context, we simulated two-arm, randomized trials comparing a hypothetical COVID-19 treatment versus standard of care, where the primary outcome is binary, ordinal, or time-to-event. Our simulated distributions are derived from two sources: longitudinal data on over 500 patients hospitalized at Weill Cornell Medicine New York Presbyterian Hospital and a Centers for Disease Control and Prevention preliminary description of 2449 cases. In simulated trials with sample sizes ranging from 100 to 1000 participants, we found substantial precision gains from using covariate adjustment-equivalent to 4-18% reductions in the required sample size to achieve a desired power. This was the case for a variety of estimands (targets of inference). From these simulations, we conclude that covariate adjustment is a low-risk, high-reward approach to streamlining COVID-19 treatment trials. We provide anRpackage and practical recommendations for implementation.
Identifying the optimal treatment decision rule, where the best treatment for an individual varies according to his/her characteristics, is of great importance when treatment effect heterogeneity exists. We develop methods for estimating the optimal treatment decision rule based on data with survival time as the primary endpoint. Our methods are based on a flexible semiparametric accelerated failure time model, where only the treatment contrast (ie, the difference in means between treatments) is parameterized and all other aspects are unspecified. An individual's treatment contrast is firstly estimated robustly by an augmented inverse probability weighted estimator (AIPWE). Then the optimal decision rule is estimated by minimizing the loss between the treatment contrast and the AIPWE contrast. Two loss functions with different strategies to account for censoring are proposed. The proposed loss functions distinguish from existing ones in that they are based on treatment contrasts, which completely determine the optimal treatment rule. Our methods can further incorporate a penalty term to select variables that are only important for treatment decision making, while taking advantage of all covariates predictive of outcomes to improve performance. Comprehensive simulation studies have been conducted to evaluate performances of the proposed methods relative to existing methods. The proposed methods are illustrated with an application to the ACTG 175 clinical trial on HIV-infected patients.
Most existing methods for optimal treatment regimes, with few exceptions, focus on estimation and are not designed for variable selection with the objective of optimizing treatment decisions. In clinical trials and observational studies, often numerous baseline variables are collected and variable selection is essential for deriving reliable optimal treatment regimes. Although many variable selection methods exist, they mostly focus on selecting variables that are important for prediction (predictive variables) instead of variables that have a qualitative interaction with treatment (prescriptive variables) and hence are important for making treatment decisions. We propose a variable selection method within a general classification framework to select prescriptive variables and estimate the optimal treatment regime simultaneously. In this framework, an optimal treatment regime is equivalently defined as the one that minimizes a weighted misclassification error rate and the proposed method forward sequentially select prescriptive variables by minimizing this weighted misclassification error. A main advantage of this method is that it specifically targets selection of prescriptive variables and in the meantime is able to exploit predictive variables to improve performance. The method can be applied to both single- and multiple-decision point setting. The performance of the proposed method is evaluated by simulation studies and application to a clinical trial.
A dynamic treatment regime is a sequence of decision rules, each corresponding to a decision point, that determine that next treatment based on each individual's own available characteristics and treatment history up to that point. We show that identifying the optimal dynamic treatment regime can be recast as a sequential optimization problem and propose a direct sequential optimization method to estimate the optimal treatment regimes. In particular, at each decision point, the optimization is equivalent to sequentially minimizing a weighted expected misclassification error. Based on this classification perspective, we propose a powerful and flexible C-learning algorithm to learn the optimal dynamic treatment regimes backward sequentially from the last stage until the first stage. C-learning is a direct optimization method that directly targets optimizing decision rules by exploiting powerful optimization/classification techniques and it allows incorporation of patient's characteristics and treatment history to improve performance, hence enjoying advantages of both the traditional outcome regression-based methods (Q- and A-learning) and the more recent direct optimization methods. The superior performance and flexibility of the proposed methods are illustrated through extensive simulation studies.
BACKGROUND:Genome-wide association studies (GWAS) have identified thousands of genetic variants associated with complex traits and diseases. However, most of them are located in the non-protein coding regions, and therefore it is challenging to hypothesize the functions of these non-coding GWAS variants. Recent large efforts such as the ENCODE and Roadmap Epigenomics projects have predicted a large number of regulatory elements. However, the target genes of these regulatory elements remain largely unknown. Chromatin conformation capture based technologies such as Hi-C can directly measure the chromatin interactions and have generated an increasingly comprehensive catalog of the interactome between the distal regulatory elements and their potential target genes. Leveraging such information revealed by Hi-C holds the promise of elucidating the functions of genetic variants in human diseases.RESULTS:In this work, we present HiView, the first integrative genome browser to leverage Hi-C results for the interpretation of GWAS variants. HiView is able to display Hi-C data and statistical evidence for chromatin interactions in genomic regions surrounding any given GWAS variant, enabling straightforward visualization and interpretation.CONCLUSIONS:We believe that as the first GWAS variants-centered Hi-C genome browser, HiView is a useful tool guiding post-GWAS functional genomics studies. HiView is freely accessible at: http://www.unc.edu/~yunmli/HiView .
UNLABELLED In next generation sequencing (NGS)-based genetic studies, researchers typically perform genotype calling first and then apply standard genotype-based methods for association testing. However, such a two-step approach ignores genotype calling uncertainty in the association testing step and may incur power loss and/or inflated type-I error. In the recent literature, a few robust and efficient likelihood based methods including both likelihood ratio test (LRT) and score test have been proposed to carry out association testing without intermediate genotype calling. These methods take genotype calling uncertainty into account by directly incorporating genotype likelihood function (GLF) of NGS data into association analysis. However, existing LRT methods are computationally demanding or do not allow covariate adjustment; while existing score tests are not applicable to markers with low minor allele frequency (MAF). We provide an LRT allowing flexible covariate adjustment, develop a statistically more powerful score test and propose a combination strategy (UNC combo) to leverage the advantages of both tests. We have carried out extensive simulations to evaluate the performance of our proposed LRT and score test. Simulations and real data analysis demonstrate the advantages of our proposed combination strategy: it offers a satisfactory trade-off in terms of computational efficiency, applicability (accommodating both common variants and variants with low MAF) and statistical power, particularly for the analysis of quantitative trait where the power gain can be up to ∼60% when the causal variant is of low frequency (MAF < 0.01). AVAILABILITY AND IMPLEMENTATION UNC combo and the associated R files, including documentation, examples, are available at http://www.unc.edu/∼yunmli/UNCcombo/ CONTACT yunli@med.unc.edu SUPPLEMENTARY INFORMATION Supplementary data are available at Bioinformatics online.
A dynamic treatment regime is a list of sequential decision rules for assigning treatment based on a patient's history. Q- and A-learning are two main approaches for estimating the optimal regime, i.e., that yielding the most beneficial outcome in the patient population, using data from a clinical trial or observational study. Q-learning requires postulated regression models for the outcome, while A-learning involves models for that part of the outcome regression representing treatment contrasts and for treatment assignment. We propose an alternative to Q- and A-learning that maximizes a doubly robust augmented inverse probability weighted estimator for population mean outcome over a restricted class of regimes. Simulations demonstrate the method's performance and robustness to model misspecification, which is a key concern.
Summary A treatment regime is a rule that assigns a treatment, among a set of possible treatments, to a patient as a function of his/her observed characteristics, hence “personalizing” treatment to the patient. The goal is to identify the optimal treatment regime that, if followed by the entire population of patients, would lead to the best outcome on average. Given data from a clinical trial or observational study, for a single treatment decision, the optimal regime can be found by assuming a regression model for the expected outcome conditional on treatment and covariates, where, for a given set of covariates, the optimal treatment is the one that yields the most favorable expected outcome. However, treatment assignment via such a regime is suspect if the regression model is incorrectly specified. Recognizing that, even if misspecified, such a regression model defines a class of regimes, we instead consider finding the optimal regime within such a class by finding the regime that optimizes an estimator of overall population mean outcome. To take into account possible confounding in an observational study and to increase precision, we use a doubly robust augmented inverse probability weighted estimator for this purpose. Simulations and application to data from a breast cancer clinical trial demonstrate the performance of the method.
A treatment regime maps observed patient characteristics to a recommended treatment. Recent technological advances have increased the quality, accessibility, and volume of patient-level data; consequently, there is a growing need for powerful and flexible estimators of an optimal treatment regime that can be used with either observational or randomized clinical trial data. We propose a novel and general framework that transforms the problem of estimating an optimal treatment regime into a classification problem wherein the optimal classifier corresponds to the optimal treatment regime. We show that commonly employed parametric and semi-parametric regression estimators, as well as recently proposed robust estimators of an optimal treatment regime can be represented as special cases within our framework. Furthermore, our approach allows any classification procedure that can accommodate case weights to be used without modification to estimate an optimal treatment regime. This introduces a wealth of new and powerful learning algorithms for use in estimating treatment regimes. We illustrate our approach using data from a breast cancer clinical trial.