Crossover trials are specifically designed to evaluate treatment effects within individual participants through within-subject comparisons. In a standard AB/BA crossover trial, participants are randomly allocated to one of two treatment sequences: either the AB sequence (where patients receive treatment A first and then cross over to treatment B after a washout period) or the BA sequence (where patients receive B first and then cross over to A after a washout period). Asymptotic and approximate unconditional test procedures, based on two Wald-type statistics, the likelihood ratio statistic, and the score test statistic for the odds ratio (OR), are developed to evaluate the equality of treatment effects in this trial design. Additionally, confidence intervals for OR are constructed, accompanied by an approximate sample size calculation methodology to control the interval width at a pre-specified precision. Empirical analyses demonstrate that asymptotic test procedures exhibit robust performance in moderate to large sample sizes, though they occasionally yield unsatisfactory type I error rates when the sample size is small. In such cases, approximate unconditional test procedures emerge as a rigorous alternative. All proposed confidence intervals achieve satisfactory coverage probabilities, and the approximate sample size estimation method demonstrates high accuracy, as evidenced by empirical coverage probabilities aligning closely with pre-specified confidence levels under estimated sample sizes. To validate practical utility, two real examples are used to illustrate the proposed methodologies.
In this paper, we develop an estimation and testing procedure for comparing matched-pair ordinal outcomes in studies with confounding factors. The classification method for the categories of ordinal outcomes that is accessible for all units may be prone to mis-classification, and thus another error-free classification method that can only be affordable for a fraction of the units are used, resulting in a dataset with partial validation. The distribution of categorical variables is modelled using correlated bivariate Gaussian latent variables, and the confounding factors are adjusted as covariates. The mis-classification of ordinal outcomes is addressed by estimating the mis-classification probabilities through the partial validation structure of the dataset. The mis-classification probabilities and the other parameters are estimated by a two-stage maximum likelihood estimator, and the difference between the matched-pair ordinal outcomes are assessed by a Wald test statistic. Simulation studies were conducted to investigate the accuracy of the estimates of the model parameters, and the type I error rates and power of the proposed testing procedure. The motivating dataset from the Garki Project was analysed to demonstrate the applicability of the proposed approach.
This article investigates the confidence interval (CI) construction of proportion difference for two independent partially validated series under the double-sampling scheme in which both classifiers are fallible. Several CIs based on the variance estimates recovery method of combining confidence limits from asymptotic, bootstrap, and Bayesian methods for two independent binomial proportions are developed under two models. Simulation results show that all CIs except for the bootstrap percentile-t CI and Bayesian credible interval with uniform prior under the independence model and all CIs under the dependence model generally perform well and are recommended. Two examples are used to illustrate the methodologies.
Non inferiority (NI) trials with the possibility of multiple experimental treatments have been increasingly used to find substitutes for standard therapies (e.g., to reduce side effects). NI trials seek to determine whether an experimental treatment can provide a suitable replacement for the standard treatment by examining the clinical significance of the loss of efficacy that is associated with the new treatment. In this article, we focus on three-arm NI trials, which include a placebo to provide direct verification of the assay sensitivity. Furthermore, it is reasonable to conduct superiority tests for experimental treatments that have been confirmed to be non inferior to the standard treatment. Several methodologies have recently been developed to provide stage-wise test procedures for this purpose. However, the applicability of these methods is limited owing to their requirement of homogeneity of variance. In this article, we seek to generalize the existing methods to more practical settings that allow the treatment variance to be heterogeneous. We also discuss sample size determination when the test power is given. Clinical examples are used to illustrate our proposed procedures.
Ordinal responses are common in clinical studies. Although the proportional odds model is a popular option for analyzing ordered-categorical data, it cannot control the type I error rate when the proportional odds assumption fails to hold. The latent Weibull model was recently shown to be a superior candidate for modeling ordinal data, with remarkably better performance than the latent normal model when the data are highly skewed. In clinical trials with ordinal responses, a balanced design is common, with equal sample allocation for each treatment. However, a more ethical approach is to adopt a response-adaptive allocation scheme in which more patients receive the better treatment. In this paper, we propose the use of the doubly adaptive biased coin design to generate treatment allocations that benefit the trial participants. The proposed treatment allocation scheme not only allows more patients to receive the better treatment, it also maintains compatible test power for the comparison of treatment efficiencies. A clinical example is used to illustrate the proposed procedure.
Double sampling is usually applied to collect necessary information for situations in which an infallible classifier is available for validating a subset of the sample that has already been classified by a fallible classifier. Inference procedures have previously been developed based on the partially validated data obtained by the double-sampling process. However, it could happen in practice that such infallible classifier or gold standard does not exist. In this article, we consider the case in which both classifiers are fallible and propose asymptotic and approximate unconditional test procedures based on six test statistics for a population proportion and five approximate sample size formulas based on the recommended test procedures under two models. Our results suggest that both asymptotic and approximate unconditional procedures based on the score statistic perform satisfactorily for small to large sample sizes and are highly recommended. When sample size is moderate or large, asymptotic procedures based on the Wald statistic with the variance being estimated under the null hypothesis, likelihood rate statistic, log- and logit-transformation statistics based on both models generally perform well and are hence recommended. The approximate unconditional procedures based on the log-transformation statistic under Model I, Wald statistic with the variance being estimated under the null hypothesis, log- and logit-transformation statistics under Model II are recommended when sample size is small. In general, sample size formulae based on the Wald statistic with the variance being estimated under the null hypothesis, likelihood rate statistic and score statistic are recommended in practical applications. The applicability of the proposed methods is illustrated by a real-data example.
A stratified study is often designed for adjusting a confounding effect or effect of different centers/groups in two treatments or diagnostic tests, and the risk difference is one of the most frequently used indices in comparing efficiency between two treatments or diagnostic tests. This article presented five simultaneous confidence intervals (CIs) for risk differences in stratified bilateral designs accounting for the intraclass correlation and developed seven CIs for the common risk difference under the homogeneity assumption. The performance of the CIs is evaluated with respect to the empirical coverage probabilities, empirical coverage widths and ratios of mesial noncoverage probability and the noncoverage probability under various scenarios. Empirical results show that Wald simultaneous CI, Haldane simultaneous CI, Score simultaneous CI based on Bonferroni method and simultaneous CI based on bootstrap-resampling method perform satisfactorily and hence be recommended for applications, the CI based on the weighted-least-square (WLS) estimator, the CIs based on Mantel-Haenszel estimator, the CI based on Cochran statistic and the CI based on Score statistic for the common risk difference behave well even under small sample sizes. A real data example is used to demonstrate the proposed methodologies.
New treatments that are noninferior or equivalent to—but not necessarily superior to—the reference treatment may still be beneficial to patients because they have fewer side effects, are more convenient, take less time, or cost less. The noninferiority test is widely used in medical research to provide guidance in such situation. In addition, categorical variables are frequently encountered in medical research, such as in studies involving patient‐reported outcomes. In this paper, we develop a noninferiority testing procedure for correlated ordinal categorical variables based on a paired design with a latent normal distribution approach. Misclassification is frequently encountered in the collection of ordinal categorical data; therefore, we further extend the procedure to account for misclassification using information in the partially validated data. Simulation studies are conducted to investigate the accuracy of the estimates, the type I error rates, and the power of the proposed procedure. Finally, we analyze one substantive example to demonstrate the utility of the proposed approach.
A disease prevalence can be estimated by classifying subjects according to whether they have the disease. When gold-standard tests are too expensive to be applied to all subjects, partially validated data can be obtained by double-sampling in which all individuals are classified by a fallible classifier, and some of individuals are validated by the gold-standard classifier. However, it could happen in practice that such infallible classifier does not available. In this article, we consider two models in which both classifiers are fallible and propose four asymptotic test procedures for comparing disease prevalence in two groups. Corresponding sample size formulae and validated ratio given the total sample sizes are also derived and evaluated. Simulation results show that (i) Score test performs well and the corresponding sample size formula is also accurate in terms of the empirical power and size in two models; (ii) the Wald test based on the variance estimator with parameters estimated under the null hypothesis outperforms the others even under small sample sizes in Model II, and the sample size estimated by this test is also accurate; (iii) the estimated validated ratios based on all tests are accurate. The malarial data are used to illustrate the proposed methodologies.
In clinical studies, treatment responses are frequently measured with an ordinal scale. To compare the efficacy of these treatments, one could employ either the proportional odds model or the latent normal model. However, these two models are inadequate for comparing treatments with highly skewed ordinal responses, due to the possibility of yielding an inflated type I error rate. To overcome this problem, the latent Weibull model has been suggested for investigating the efficacy difference between two treatments. For more general applications, this model is extended to include clinical trials with more than two treatments. Two testing procedures are derived: one for multiple comparisons with a control and the other for pairwise treatment comparisons. The testing procedures are also demonstrated with two clinical examples.
An ordinal effect size measure is used to assess whether one variable is stochastically larger than the other; therefore, this measure is a useful means by which to describe the difference between two ordinal categorical distributions. In practical analysis, it is desirable to obtain data only by a gold standard test, but such tests are often limited due to high costs or ethical considerations, especially in medical research applications. However, misclassification can arise when a cheaper, faster, and/or non-invasive, but fallible test is used to collect data. The use of partially validated data obtained by double sampling has become a popular compromise between these two approaches. In this study, we develop twelve estimators of the confidence interval (CI) for an ordinal effect size measure based on partially validated data. The performance of the proposed CIs are evaluated by simulation studies in terms of the empirical coverage probability, the empirical coverage width, and the ratio of the mesial non-coverage probability and non-coverage probability. Simulation results show that the Wald CI on logit scale, the Bootstrap percentile CI and the Bias-corrected Bootstrap normal CI have outstanding performance even in small sample designs. When sample sizes are moderate, all CIs except the Wald, Bias-corrected Bootstrap percentile and logit-transformation-based Bootstrap percentile-t CIs demonstrate good coverage properties. Moreover, all CIs perform well when sample sizes are large. All methods are illustrated by analyzing a real data set from a research study of highway safety on automobile accidents.
In clinical studies, ordered categorical responses are common. To compare the efficacy of several treatments with a control for ordinal responses, the normal latent variable model has recently been proposed. This approach conceptualizes the responses as manifestations of an underlying continuous normal variable. In this article, we extend this idea to develop the multiple comparison method for use when there are two controls in the clinical trial. The proposed method is constructed such that the familywise type I error rate is controlled at a prespecified level. In addition, for a given level of test power, the procedure to evaluate the required sample size is provided. The proposed testing procedure is also illustrated by an example from a clinical study.
In clinical studies, the proportional odds model is widely used to compare treatment efficacies when the responses are categorically ordered. However, this model has been shown to be inappropriate when the proportional odds assumption is invalid, mainly because it is unable to control the type I error rate in such circumstances. To remedy this problem, the latent normal model was recently promoted and has been demonstrated to be superior to the proportional odds model. However, the application of the latent normal model is limited to compare treatments with similar underlying distributions except possibly their means and variances. When the underlying distributions are very different in skewness, both of the aforementioned procedures suffer from the undesirable inflation of the type I error rate. To solve the problem for clinical studies with ordinal responses, we provide a viable solution that relies on the use of the latent Weibull distribution, which is a member of the log‐location‐scale family. The proposed model is able to control the type I error rate regardless of the degree of skewness of the treatment responses. In addition, the power of the test also outperforms that of the latent normal model. The testing procedure draws on newly developed theoretical results related to latent distributions from the location‐scale family. The testing procedure is illustrated with two clinical examples. Copyright © 2015 John Wiley & Sons, Ltd.
Partially validated series are common when a gold-standard test is too expensive to be applied to all subjects, and hence a fallible device is used accordingly to measure the presence of a characteristic of interest. In this article, confidence interval construction for proportion difference between two independent partially validated series is studied. Ten confidence intervals based on the method of variance estimates recovery (MOVER) are proposed, with each using the confidence limits for the two independent binomial proportions obtained by the asymptotic, Logit-transformation, Agresti-Coull and Bayesian methods. The performances of the proposed confidence intervals and three likelihood-based intervals available in the literature are compared with respect to the empirical coverage probability, confidence width and ratio of mesial non-coverage to non-coverage probability. Our empirical results show that (1) all confidence intervals exhibit good performance in large samples; (2) confidence intervals based on MOVER combining the confidence limits for binomial proportions based on Wilson, Agresti-Coull, Logit-transformation, Bayesian (with three priors) methods perform satisfactorily from small to large samples, and hence can be recommended for practical applications. Two real data sets are analysed to illustrate the proposed methods.
Confidence interval construction for the Youden index of a diagnostic test based on partially validated series with dichotomous response is considered in this article. Using the Wald and Agresti-Coull, the Wilson score, logit-transformation and the method of variance estimates recovery, eight methods for constructing confidence intervals for a single Youden index and nine methods for the difference between two independent Youden indices are developed. Comparisons among the various methods with respect to their empirical coverage probabilities and confidence interval widths are conducted through simulation studies over a variety of parameter settings. Based on the simulation results, recommendations for practical use under different conditions are provided. A real aplastic anemia data set is used to illustrate the proposed methods.
Different latent variable models have been used to analyze ordinal categorical data which can be conceptualized as manifestations of an unobserved continuous variable. In this paper, we propose a unified framework based on a general latent variable model for the comparison of treatments with ordinal responses. The latent variable model is built upon the location-scale family and is rich enough to include many important existing models for analyzing ordinal categorical variables, including the proportional odds model, the ordered probit-type model, and the proportional hazards model. A flexible estimation procedure is proposed for the identification and estimation of the general latent variable model, which allows for the location and scale parameters to be freely estimated. The framework advances the existing methods by enabling many other popular models for analyzing continuous variables to be used to analyze ordinal categorical data, thus allowing for important statistical inferences such as location and/or dispersion comparisons among treatments to be conveniently drawn. Analysis on real data sets is used to illustrate the proposed methods.
In clinical studies, multiple comparisons of several treatments to a control with ordered categorical responses are often encountered. A popular statistical approach to analyzing the data is to use the logistic regression model with the proportional odds assumption. As discussed in several recent research papers, if the proportional odds assumption fails to hold, the undesirable consequence of an inflated familywise type I error rate may affect the validity of the clinical findings. To remedy the problem, a more flexible approach that uses the latent normal model with single‐step and stepwise testing procedures has been recently proposed. In this paper, we introduce a step‐up procedure that uses the correlation structure of test statistics under the latent normal model. A simulation study demonstrates the superiority of the proposed procedure to all existing testing procedures. Based on the proposed step‐up procedure, we derive an algorithm that enables the determination of the total sample size and the sample size allocation scheme with a pre‐determined level of test power before the onset of a clinical trial. A clinical example is presented to illustrate our proposed method. Copyright © 2014 John Wiley & Sons, Ltd.
Clinical trials frequently involve pairwise comparisons of different treatments to evaluate their relative efficacy. In this study, we examine methods for conducting pairwise tests of treatments with ordered categorical responses. A modified version of the Wilcoxon–Mann–Whitney test based on a logistic regression model assuming proportional odds is a popular choice for comparing two treatments. This paper discusses the extension of this test to pairwise comparisons involving more than two treatments. However, when the proportional odds assumption is not valid, the Wilcoxon–Mann–Whitney‐type test procedure cannot control the overall type I error rate at the prespecified level of significance. We therefore propose a better strategy in which a latent normal model is employed. We presented a simulated comparative study of power and the overall type I error rate to illustrate the superiority of the latent normal model. Examples are also given for illustrative purposes. Copyright © 2013 John Wiley & Sons, Ltd.
We extend generalized partially linear single-index models by incorporating a random residual effect into the nonlinear predictor so that the new models can accommodate data with overdispersion. Based on the free-knot spline techniques, we develop a fully Bayesian method to analyze the proposed models. To make the models spatially adaptive, we further treat the number and positions of spline knots as random variables. As random residual effects are introduced, many of the completely conditional posteriors become standard distributions, which greatly facilitates sampling. We illustrate the proposed models and estimation method with a simulation study and an analysis of a recreational trip data set.
It is desirable to estimate disease prevalence based on data collected by a gold standard test, but such a test is often limited due to cost and ethical considerations. Data with partial validation series thus become an alternative. The construction of confidence intervals for disease prevalence with such data is considered. A total of 12 methods, which are based on two Wald-type test statistics, score test statistic, and likelihood ratio test statistic, are developed. Both asymptotic and approximate unconditional confidence intervals are constructed. Two methods are employed to construct the unconditional confidence intervals: one involves inverting two one-sided tests and the other involves inverting one two-sided test. Moreover, the bootstrapping method is used. Two real data sets are used to illustrate the proposed methods. Empirical results suggest that the 12 methods largely produce satisfactory results, and the confidence intervals derived from the score test statistic and the Wald test statistic with nuisance parameters appropriately evaluated generally outperform the others in terms of coverage. If the interval location or the non-coverage at the two ends of the interval is also of concern, then the aforementioned interval based on the Wald test becomes the best choice.