We develop a systematic, omnibus approach to goodness-of-fit testing for parametric distributional models when the variable of interest is only partially observed due to censoring and/or truncation. In many such designs, tests based on the nonparametric maximum likelihood estimator are hindered by nonexistence, computational instability, or convergence rates too slow to support reliable calibration under composite nulls. We avoid these difficulties by constructing a regular (pathwise differentiable) Neyman-orthogonal score process indexed by test functions, and aggregating it over a reproducing kernel Hilbert space ball. This yields a maximum-mean-discrepancy-type supremum statistic with a convenient quadratic-form representation. Critical values are obtained via a multiplier bootstrap that keeps nuisance estimates fixed. We establish asymptotic validity under the null and local alternatives and provide concrete constructions for left-truncated right-censored data, current status data, and random double truncation; in particular, to the best of our knowledge, we give the first omnibus goodness-of-fit test for a parametric family under random double truncation in the composite-hypothesis case. Simulations and an empirical illustration demonstrate size control and power in practically relevant incomplete-data designs.
k-Sample versions of the Kolmogorov–Smirnov and Cramér–von Mises tests are proposed for data subject to left truncation and right censoring. Their asymptotic behaviour is studied and a bootstrap resampling plan is proposed to approximate the null distribution of the new tests. The performance of the tests with finite sample sizes is investigated in a simulation study, where the classical log-rank test is also considered. An illustration with a real dataset regarding unemployment times is discussed.
In Astronomy, Survival Analysis and Epidemiology, among many other fields, doubly truncated data often appear. Double truncation generally induces a sampling bias, so ordinary estimators may be inconsistent. In this paper, smoothing spline density estimation from doubly truncated data is investigated. For this purpose, an appropriate correction of the penalized likelihood that accounts for the sampling bias is considered. The theoretical properties of the estimator are discussed, and its practical performance is evaluated through simulations. Two real datasets are analyzed using the proposed method for illustrative purposes. Comparison to kernel density smoothing is included.
The area under the ROC curve (AUC) plays an important role in the study of the predictive capacity of regression models. It is well known that an infated AUC may result when the same data are used for training and testing the model. In this paper optimism correction of the AUC in the presence of missing data is investigated. Complete case analysis, inverse probability weighting and multiple imputation are employed to address the issue of missing data. For each of these approaches, split-sample, K-fold cross-validation and leave-one-out cross-validation are employed to correct for the optimism of the AUC. The methods are compared through intensive Monte Carlo simulations in the particular setting of binary regression. Results suggest that all estimators are consistent with the exception of complete case analysis, which may be biased when missing is not completely at random. In general, a combined application of multiple imputation and leave-one-out cross-validation is recommended.
In multiple testing several criteria to control for type I errors exist. The false discovery rate, which evaluates the expected proportion of false discoveries among the rejected null hypotheses, has become the standard approach in this setting. However, false discovery rate control may be too conservative when the effects are weak. In this paper we alternatively propose to control the number of significant effects, where 'significant' refers to a pre-specified threshold $\gamma$. This means that a $(1-\alpha)$-lower confidence bound $L$ for the number of non-true null hypothesis with p-values below $\gamma$ is provided. When one rejects the nulls corresponding to the $L$ smallest p-values, the probability that the number of false positives exceeds the number of false negatives among the significant effects is bounded by $\alpha$. Relative merits of the proposed criterion are discussed. Procedures to control for the number of significant effects in practice are introduced and investigated both theoretically and through simulations. Illustrative real data applications are given.
In survival analysis and epidemiology, among other fields, interval sampling is often employed. With interval sampling, the individuals undergoing the event of interest within a calendar time interval are recruited. This results in doubly truncated event times. Double truncation, which may appear with other sampling designs too, induces a selection bias, so ordinary statistical methods are generally inconsistent. In this paper, we introduce goodness-of-fit procedures for a regression model when the response variable is doubly truncated. With this purpose, a marked empirical process based on weighted residuals is constructed and its weak convergence is established. Kolmogorov-Smirnov- and Cramér-von Mises-type tests are consequently derived from such core process, and a bootstrap approximation for their practical implementation is given. The performance of the proposed tests is investigated through simulations. An application to model selection for AIDS incubation time as depending on age at infection is provided.
In this work logistic regression when both the response and the predictor variables may be missing is considered. Several existing approaches are reviewed, including complete case analysis, inverse probability weighting, multiple imputation and maximum likelihood. The methods are compared in a simulation study, which serves to evaluate the bias, the variance and the mean squared error of the estimators for the regression coefficients. In the simulations, the maximum likelihood methodology is the one that presents the best results, followed by multiple imputation with five imputations, which is the second best. The methods are applied to a case study on the obesity for schoolchildren in the municipality of Viana do Castelo, North Portugal, where a logistic regression model is used to predict the International Obesity Task Force (IOTF) indicator from physical examinations and the past values of the obesity status. All the variables in the case study are potentially missing, with gender as the only exception. The results provided by the several methods are in well agreement, indicating the relevance of the past values of IOTF and physical scores for the prediction of obesity. Practical recommendations are given.
In clinical and epidemiological research doubly truncated data often appear. This is the case, for instance, when the data registry is formed by interval sampling. Double truncation generally induces a sampling bias on the target variable, so proper corrections of ordinary estimation and inference procedures must be used. Unfortunately, the nonparametric maximum likelihood estimator of a doubly truncated distribution has several drawbacks, like potential nonexistence and nonuniqueness issues, or large estimation variance. Interestingly, no correction for double truncation is needed when the sampling bias is ignorable, which may occur with interval sampling and other sampling designs. In such a case the ordinary empirical distribution function is a consistent and fully efficient estimator that generally brings remarkable variance improvements compared to the nonparametric maximum likelihood estimator. Thus, identification of such situations is critical for the simple and efficient estimation of the target distribution. In this article, we introduce for the first time formal testing procedures for the null hypothesis of ignorable sampling bias with doubly truncated data. The asymptotic properties of the proposed test statistic are investigated. A bootstrap algorithm to approximate the null distribution of the test in practice is introduced. The finite sample performance of the method is studied in simulated scenarios. Finally, applications to data on onset for childhood cancer and Parkinson's disease are given. Variance improvements in estimation are discussed and illustrated.
The NPMLE of a distribution function from doubly truncated data was introduced in the seminal paper of Efron and Petrosian (J Am Stat Assoc 94:824–834, 1999). The consistency of the NPMLE depends however on the assumption of independent truncation. In this work we introduce an extension of the Efron–Petrosian NPMLE when the variable of interest and the truncation variables may be dependent. The proposed estimator is constructed on the basis of a copula function which represents the dependence structure between the variable of interest and the truncation variables. Two different iterative algorithms to compute the estimator in practice are introduced, and their performance is explored through an intensive Monte Carlo simulation study. We illustrate the use of the estimators on two real data examples.
In survival analysis, epidemiology and related fields there exists an increasing interest in statistical methods for doubly truncated data. Double truncation appears with interval sampling and other sampling schemes, and refers to situations in which the target variable is subject to two (left and right) random observation limits. Doubly truncated data require specific corrections for the observational bias, and this affects a variety of settings including the estimation of marginal and multivariate distributions, regression problems, and multi-state models. In this work multivariate Efron-Petrosian integrals for doubly truncated data are introduced. These integrals naturally arise when the goal is the estimation of the mean of a general transformation which involves the doubly truncated variable and covariates. An asymptotic representation of the Efron-Petrosian integrals as a sum of iid terms is derived and, from this, consistency and distributional convergence are established. As a by-product, uniform iid representations for the marginal nonparametric maximum likelihood estimator and its corresponding weighting process are provided. Applications to correlation analysis, regression, and competing risks models are presented. A simulation study is reported too.
In Survival Analysis, the observed lifetimes often correspond to individuals for which the event occurs within a specific calendar time interval. With such interval sampling, the lifetimes are doubly truncated at values determined by the birth dates and the sampling interval. This double truncation may induce a systematic bias in estimation, so specific corrections are needed. A relevant target in Survival Analysis is the hazard rate function, which represents the instantaneous probability for the event of interest. In this work we introduce a flexible estimation approach for the hazard rate under double truncation. Specifically, a kernel smoother is considered, in both a fully nonparametric setting and a semiparametric setting in which the incidence process fits a given parametric model. Properties of the kernel smoothers are investigated both theoretically and through simulations. In particular, an asymptotic expression of the mean integrated squared error is derived, leading to a data-driven bandwidth for the estimators. The relevance of the semiparametric approach is emphasized, in that it is generally more accurate and, importantly, it avoids the potential issues of nonexistence or nonuniqueness of the fully nonparametric estimator. Applications to the age of diagnosis of Acute Coronary Syndrome (ACS) and AIDS incubation times are included.
When analysing and presenting results of randomised clinical trials, trialists rarely report if or how underlying statistical assumptions were validated. To avoid data-driven biased trial results, it should be common practice to prospectively describe the assessments of underlying assumptions. In existing literature, there is no consensus on how trialists should assess and report underlying assumptions for the analyses of randomised clinical trials. With this study, we developed suggestions on how to test and validate underlying assumptions behind logistic regression, linear regression, and Cox regression when analysing results of randomised clinical trials. Two investigators compiled an initial draftbased on a review of the literature. Experienced statisticians and trialists from eight different research centres and trial units then participated in a anonymised consensus process, where we reached agreement on the suggestions presented in this paper. This paper provides detailed suggestions on 1) which underlying statistical assumptions behind logistic regression, multiple linear regression and Cox regression each should be assessed; 2) how these underlying assumptions may be assessed; and 3) what to do if these assumptions are violated. We believe that the validity of randomised clinical trial results will increase if our recommendations for assessing and dealing with violations of the underlying statistical assumptions are followed.
Registry data typically report incident cases within a certain calendar time interval. Such interval sampling induces double truncation on the incidence times, which may result in an observational bias. In this paper, we introduce nonparametric estimation for the cumulative incidences of competing risks when the incidence time is doubly truncated. Two different estimators are proposed depending on whether the truncation limits are independent of the competing events or not. The asymptotic properties of the estimators are established, and their finite sample performance is investigated through simulations. For illustration purposes, the estimators are applied to childhood cancer registry data, where the target population is peculiarly defined conditional on future cancer development. Then, in our application, the cumulative incidences inform on the distribution by age of the different types of cancer.
Random double truncation refers a situation in which the variable of interest is observed only when it falls within two random limits. Such phenomenon occurs in many applications of Survival Analysis and Epidemiology, among many other fields. There exist several R packages to analyze doubly truncated data which implement, for instance, estimators for the cumulative distribution, the cumulative incidences of competing risks, or the regression coefficients in the Cox model. In this paper the main features of these packages are reviewed. The relative merits of the libraries are illustrated through the analysis of simulated and real data. This includes the study of the statistical accuracy of the implemented techniques as well as of their computational speed. Practical recommendations are given.
This Special Issue on Advanced Topics in Biostatistics arose from the 38th Annual Conference of the International Society for Clinical Biostatistics (ISCB38), hosted by the city of Vigo (Spain), which took place from 9 to 13 July 2017. ISCB38 was organized by Jacobo de Uña Álvarez from Universidade de Vigo, as chair of the Local Organizing Committee. Guadalupe Gómez Melis (chair), Frank Bretz, and Jacobo de Uña Álvarez were members of the Scientific Programme Committee and formed also the board of Special Issue editors. The meeting brought together two excellent plenary talks: one by Stephen Senn (Luxembourg Institute of Health) on The Revenge of RA Fisher: Thoughts on Randomgate and its Wider Implications, and the other one by Francesca Dominici (Harvard University) on Model Uncertainty and Covariate Selection in Causal Inference. The conference included eight oral invited sessions, 50 oral and poster contributed sessions, five preconference courses, and three mini-symposia. After a peer review procedure 190 oral and 173 poster contributed presentations were selected, all of them exhibiting a high quality and relevance. In response to the call Advanced Topics in Biostatistics for a Special Issue in Biometrical Journal 37 papers were submitted. Among these, 23 papers have been accepted and are included in this volume and in the next issue of the journal (May 2019). The call for the Special Issue encouraged papers in the following areas: Survival analysis and multistate models; Statistical advances for clinical trials; High dimensional biostatistical data; Statistical methods for precision medicine, including biomarker discovery; and Bayesian methods in clinical research. The current volume covers the first four topics along 15 papers; while the May volume is dedicated to those papers in clinical research that use Bayesian techniques (six papers), plus other two papers on methods for binary data. All these papers are already accessible from the Early View Online Library of Biometrical Journal. Five out of the 23 accepted papers focus on Survival analysis and multistate models. Meira-Machado and Sestelo provide a review of recent developments in nonparametric estimation for the progressive illness-death model, including remarks on available software routines. Lauseker and Zu Eulenburg nicely explain and illustrate the potential biases of estimators for competing risks data when ignoring the presence of an intermediate state. Lambert and Bremhorst discuss identification issues in the promotion time cure model; their investigation involves simulation studies as well as illustrative clinical data analyses. Grand and coauthors introduce two different types of pseudo-values for left-truncated time-to-event data and they explore their performance through simulation studies. Alarcón-Soto and coauthors, motivated by the analysis of the time to HIV RNA viral rebound in HIV patients, present a method to fit a Cox model with mixed effects and interval censoring, investigate its properties in a simulation study and provide the R code to implement it. The topic Statistical advances for clinical trials is covered by five contributions. Jiménez and coauthors introduce and investigate a Bayesian adaptive phase I trial design for toxicity attributions in drug combinations. Heller et al. give new insights in biased model fitting when both mean and dispersion parameters depend on treatment, and they study the robustness of orthogonal parametrization with respect to misspecification of the dispersion model. Thao and Geskus discuss the relative performance of algorithms for predictive model selection in the presence of missing data that were multiply imputed; they give an illustrative example by analyzing data from three trials in patients with tuberculosis meningitis. Preussler et al. derive optimal phase II sample sizes and go/no-go decision rules in clinical trials, for settings in which at least one, two or three phase III trials need to be successful. Senn et al. investigate a random effects model for treatment in meta-analysis; they consider as example a network meta-analysis of 44 similar treatments in 10 trials. Three papers involve High dimensional biostatistical data. Miok et al. introduce an extended model and ridge estimation approach for time-course omics data; they apply their method to experimental data on HPV-induced carcinogenesis by using their own R software routine. Another contribution in omics data is reported by Csala and coauthors who propose and investigate a multiset sparse redundancy analysis framework based on partial least squares path modeling to cope with information transfer between omic sets. Goodness-of-fit tests for next-generation sequencing experiments are introduced and investigated by Jiménez-Otero et al.; they apply the proposed approach to the detection of disorders in a DNA region. The two papers devoted to Statistical methods for precision medicine deal with two issues of much interest. Moodie and coauthors introduce a Q-learning algorithm for the estimation of an optimal adaptive treatment strategy when a fraction of the population is cured. On the other hand, Chung et al. investigate selective testing strategies for haemoglobin levels in blood donors, and they give practical recommendations. Six papers are included for the May issue under the topic Bayesian methods in medical research. Martina et al. propose and simulate a Bayesian adaptive design in chronic prostatitis/chronic pelvic pain syndrome. Kopp-Schneider and coauthors employ posterior probabilities to monitor futility and efficacy in phase II trials. Rufo and coauthors use posterior probabilities and utility functions based on financial costs to assess voice disorder risks. Heyard et al. adapt Bayesian variable selection algorithms to Cox regression for discrete time-to-event data with competing risks. Verde applies the hierarchical meta-regression approach to a particular setting in which the results of randomized controlled trials are extrapolated to groups of patients enrolled in a prospective cohort study. Finally, Gray et al. describe a Bayesian regression calibration approach when a nonlinear association is modeled by transforming the exposure with a fractional polynomial model. The last two papers in the May issue approach problems with correlated binary outcomes bringing interesting discussions. Carrasco et al. compare different methods to estimate the marginal proportions and the intraclass correlations with clustered binary data. While Najera-Zuloaga et al. propose and develop a beta-binomial mixed-effects model for the analysis of correlated patient-reported outcomes and they implement the proposed approach in an R Package. Summarizing, the 23 papers included in the ISCB38 Special Issue are of great variety and interest, offering relevant advances in biostatistics, and it has been a pleasure to manage all of them along the editorial process. Last but not least, the guest editors of this Special Issue would like to express their deepest gratitude to plenty of colleagues (almost 60) who accepted to act as reviewers, performing an excellent job. This issue would not have been possible without their altruistic dedication. As usual with Biometrical Journal, the names of these colleagues will be included in the general “list of referees for 2018” to be published in one of the 2019 issues. A final thanks to the editorial team of Biometrical Journal (M.A., D.B., A.M., and M.K.) working continuously hard in the background to achieve a smooth and timely appearance of the ISCB38 Special Issue.
A recurring theme in modern statistics is dealing with high-dimensional data whose main feature is a large number, p, of variables but a small sample size. In this context our aim is to address the problem of testing the null hypothesis that the marginal distributions of p variables are the same for two groups. We propose a test statistic motivated by the simple idea of comparing, for each of the p variables, the empirical characteristic functions computed from the two samples. The asymptotic normality of the test statistic is derived under mixing conditions. In our asymptotic analysis the number of variables tends to infinity, while the size of individual samples remains fixed. In order to obtain a practical test several estimators of the variance are proposed, leading to three somewhat different versions of the test. An alternative global test based on the P-values derived from permutation tests is also proposed. A simulation study to investigate the finite sample properties of the proposed tests is carried out, and a practical illustration involving microarray data is provided.
Large scale discrete uniform and homogeneous $P$-values often arise in applications with multiple testing. For example, this occurs in genome wide association studies whenever a nonparametric one-sample (or two-sample) test is applied throughout the gene loci. In this paper we consider $q$-values for such scenarios based on several existing estimators for the proportion of true null hypothesis, $\pi_0$, which take the discreteness of the $P$-values into account. The theoretical guarantees of the several approaches with respect to the estimation of $\pi_0$ and the false discovery rate control are reviewed. The performance of the discrete $q$-values is investigated through intensive Monte Carlo simulations, including location, scale and omnibus nonparametric tests, and possibly dependent $P$-values. The methods are applied to genetic and financial data for illustration purposes too. Since the particular estimator of $\pi_0$ used to compute the $q$-values may influence the power, relative advantages and disadvantages of the reviewed procedures are discussed. Practical recommendations are given.
In this paper the R package TP.idm to compute an empirical transition probability matrix for the illness-death model is introduced. This package implements a novel nonparametric estimator which is particularly well suited for non-Markov processes observed under right censoring. Variance estimates and confidence limits are also implemented in the package.
Next-generation sequencing (NGS) experiments are often performed in biomedical research nowadays, leading to methodological challenges related to the high-dimensional and complex nature of the recorded data. In this work we review some of the issues that arise in disorder detection from NGS experiments, that is, when the focus is the detection of deletion and duplication disorders for homozygosity and heterozygosity in DNA sequencing. A statistical model to cope with guanine/cytosine bias and phasing and prephasing phenomena at base level is proposed, and a goodness-of-fit procedure for disorder detection is derived. The method combines the proper evaluation of local p-values (one for each DNA base) with suitable corrections for multiple comparisons and the discrete nature of the p-values. A global test for the detection of disorders in the whole DNA region is proposed too. The performance of the introduced procedures is investigated through simulations. A real data illustration is provided.