This paper delves into the realm of ordinal classification processes for multiple raters. A Probit hierarchical model is proposed linking rater's ordinal ratings with rater diagnostic skills (bias and magnifier) and patient latent disease severity, where patient latent disease severity is assumed to follow a latent class normal mixture distribution. This model specification provides closed-form expressions for both overall and individual rater receiver operator characteristic (ROC) curves and the area under these ROC curves (AUC). We further extend the model by incorporating covariate information and adding a regression layer for rater diagnostic skill parameters and/or for patient latent disease severity. The extended covariate models also offer closed-form solutions for covariate-specific ROCs and AUCs. These analytical tools greatly facilitate traditional diagnostic accuracy analysis. We demonstrate our methods thoroughly with a practical mammography example.
Panel count data arise when recurrent events are observed periodically in a study. The response variable of interest is the number of recurrent events within different time windows instead of the exact onset times of the events. The gamma frailty Poisson process model has been proposed to accommodate the within-subject correlation and overdispersion in panel count data. Although the existing methods based on the gamma frailty Poisson process model have shown some robustness against frailty distribution misspecifications, they are also found to produce biased estimates in some other cases when the gamma frailty assumption is violated. In this paper, we generalize the gamma frailty Poisson process model to allow an unknown frailty distribution for analyzing panel count data. Specifically the frailty distribution is modeled nonparametrically by assigning a Dirichlet Process Gamma Mixture prior. An efficient Gibbs sampler is developed to facilitate the Bayesian computation. Extensive simulation results suggest that the proposed Bayesian approach has an excellent performance in estimating the regression parameters and the baseline mean function and outperforms the corresponding Bayesian method based on the gamma frailty Poisson model when the gamma frailty distribution is misspecified. The proposed method is applied to a skin cancer dataset for an illustration.
Joint modeling of longitudinal data and survival data has gained great attention in the last two decades. However, most of the existing studies have focused on right-censored survival data. In this article, we study joint analysis of longitudinal data and interval-censored survival data and conduct Bayesian variable selection in this framework. A new joint model is proposed with a shared frailty to characterize the dependence between the two types of responses, where the longitudinal response is modeled with a semiparametric linear mixed-effects submodel and the survival time is modeled by a semiparametric normal fraility probit sub-model. Several Bayesian variable selection approaches are developed by adopting Bayesian Lasso, adaptive Lasso, and spike-and-slab priors in order to simultaneously select significant covariates and estimate their effects on the two types of responses. Efficient Gibbs samplers are proposed with all unknown parameters and latent variables being sampled directly from well recognized full conditional distributions. Our simulation study shows that these methods perform well in both variable selection and parameter estimation. A real-life data application to joint analysis of blood cholesterol level and hypertension is provided as an illustration.
Arbitrarily censored data have a more general data structure than conventional right-censored data or general interval-censored data and thus are more challenging to analyze. In this article, a novel Bayesian approach is developed for analyzing arbitrarily censored data under the semiparametric proportional hazards (PH) model. The proposed method adopts M-splines and I-splines to model the baseline hazard and cumulative baseline hazard functions in the PH model, respectively. A new two-stage data augmentation involving exponential and multinomial latent variables is proposed, leading to a nice form of augmented likelihood. Based on this likelihood, an easy-to-implement Gibbs sampler is developed. Simulation studies show that the proposed method works well in estimating both regression parameters and survival functions. A numerical comparison of our method with existing Bayesian methods is also provided. Two real data sets on colorectal cancer and childhood mortality are analyzed for illustration.
Panel count data often occur in a long-term recurrent event study, where the exact occurrence time of the recurrent events is unknown, but only the occurrence count between any two adjacent observation time points is recorded. Most traditional methods only handle panel count data for a single type of event. In this paper, we propose a Bayesian semiparameteric approach to analyze panel count data for multiple types of events. For each type of recurrent event, the proportional mean model is adopted to model the mean count of the event, where its baseline mean function is approximated by monotone I-splines. The correlation between multiple types of events is modeled by common frailty terms and scale parameters. Unlike many frequentist estimating equation methods, our approach is based on the observed likelihood and makes no assumption on the relationship between the recurrent process and the observation process. Under the Poisson counting process assumption, we develop an efficient Gibbs sampler based on novel data augmentation for the Markov chain Monte Carlo sampling. Simulation studies show good estimation performance of the baseline mean functions and the regression coefficients; meanwhile, the importance of including the scale parameter to flexibly accommodate the correlation between events is also demonstrated. Finally, a skin cancer data example is fully analyzed to illustrate the proposed methods.
Both panel count data and interval-censored data arise commonly in real-life studies when subjects are examined at periodic follow-ups. Interval-censored data are studied when the exact times of the events are of interest and these exact times are not directly observed but are only known to fall within some intervals formed by the observation times. Panel count data are under investigation when the exact times of the recurrent events are not of interest but the counts of the recurrent events occurring within the time intervals are available and of interest. A novel unified Bayesian approach is developed for analyzing panel count data under the Gamma frailty Poisson process model and interval-censored data under Cox’s proportional hazards model and the proportional odds model. The baseline functions in these models share the same property of being nondecreasing positive functions and are modeled nonparametrically by assigning a Gamma process prior. Efficient Gibbs samplers are developed for the posterior computation under these three models for the two types of data. The proposed methods are evaluated in a simulation study and illustrated by three real-life data applications.
Diagnostic tests are frequently reliant upon the interpretation of images by skilled raters. In many clinical settings, however, the variability observed between experts' ratings plays a detrimental role in the degree of confidence in these interpretations, leading to uncertainty in the diagnostic process. For example, in breast cancer testing, radiologists interpret mammographic images, while breast biopsy results are examined by pathologists. Each of these procedures involves elements of subjectivity. We propose here a flexible two-stage Bayesian latent variable model to investigate how the skills of individual raters impact the diagnostic accuracy of image-related testing in large-scale medical testing studies. A strength of the proposed model is that the true disease status of a patient within a reasonable time frame may or may not be known. In these studies, many raters each contribute classifications on a large sample of patients using a defined ordinal grading scale, leading to a complex correlation structure between ratings. Our modeling approach considers the different sources of variability contributed by experts and patients while accounting for correlations present between ratings and patients, in contrast to currently available methods. We propose a novel measure of a rater's ability (magnifier) that, in contrast to conventional measures of sensitivity and specificity, is robust to the underlying prevalence of disease in the population, providing an alternative measure of diagnostic accuracy across patient populations. Extensive simulation studies demonstrate lower bias in estimation of parameters and measures of accuracy, and illustrate outperformance of the proposed model when compared with existing models. Receiver operator characteristic curves are derived to assess the diagnostic accuracy of individual experts and their overall performance. Our proposed modeling approach is applied to a large breast imaging study for known disease status and a uterine cancer dataset for unknown disease status.
The diagnostic accuracy of a test or rater has a crucial impact on clinical decision making. The assessment of diagnostic accuracy for multiple tests or raters also merits much attention. A Bayesian hierarchical conditional independence latent class model for estimating sensitivities and specificities for a large group of tests or raters is proposed, which is applicable to both with-gold-standard and without-gold-standard situations. Through the hierarchical structure, not only are the sensitivities and specificities of individual tests estimated, but also the diagnostic performance of the whole group of tests. For a small group of tests or raters, the proposed model is further extended by introducing pairwise covariances between tests to improve the fitting and to allow for more modeling flexibility. Correlation residual analysis is applied to detect any significant covariance between multiple tests. Just Another Gibbs Sampler (JAGS) implementation is efficiently adopted for both models. Three real data sets from literature are analyzed to explicitly illustrate the proposed methods.
Interval-censored data arise frequently in medical studies of diseases that require periodic examinations for symptoms of interest, such as disease-free survival (DFS). This paper provides an introduction to the R package ICBayes which implements a set of programs for analyzing case 1 interval-censored data and case 2 interval-censored data under Bayesian semiparametric framework. The main function ICBayes fits commonly-used survival regression models: proportional hazards, proportional odds, and probit. A simulation study is conducted to compare the performance of the package with two R packages that fit Bayesian proportional hazards, proportional odds, and accelerated failure time models. The use of the package is illustrated through analyzing a case 2 interval-censored breast cosmesis data and a case 1 interval-censored animal tumorigenicity data.
Panel count data commonly arise in epidemiological, social science, and medical studies, in which subjects have repeated measurements on the recurrent events of interest at different observation times. Since the subjects are not under continuous monitoring, the exact times of those recurrent events are not observed but the counts of such events within the adjacent observation times are known. A Bayesian semiparametric approach is proposed for analyzing panel count data under the proportional mean model. Specifically, a nonhomogeneous Poisson process is assumed to model the panel count response over time, and the baseline mean function is approximated by monotone I-splines of Ramsay (Stat Sci 3:425–461, 1988). Our approach allows to estimate the regression parameters and the baseline mean function jointly. The proposed Gibbs sampler is computationally efficient and easy to implement because all of the full conditional distributions either have closed form or are log-concave. Extensive simulations are conducted to evaluate the proposed method and to compare with two other bench methods. The proposed approach is also illustrated by an application to a famous bladder tumor data set (Byar, in: Pavone-Macaluso M, Smith PH, Edsmyn F (eds) Bladder tumors and other topics in urological oncology. Plenum, New York, 1980).
Many disease diagnoses involve subjective judgments by qualified raters. For example, through the inspection of a mammogram, MRI, or ultrasound image, the clinician himself becomes part of the measuring instrument. To reduce diagnostic errors and improve the quality of diagnoses, it is necessary to assess raters' diagnostic skills and to improve their skills over time. This paper focuses on a subjective binary classification process, proposing a hierarchical model linking data on rater opinions with patient true disease‐development outcomes. The model allows for the quantification of the effects of rater diagnostic skills (bias and magnifier) and patient latent disease severity on the rating results. A Bayesian Markov chain Monte Carlo (MCMC) algorithm is developed to estimate these parameters. Linking to patient true disease outcomes, the rater‐specific sensitivity and specificity can be estimated using MCMC samples. Cost theory is used to identify poor‐ and strong‐performing raters and to guide adjustment of rater bias and diagnostic magnifier to improve the rating performance. Furthermore, diagnostic magnifier is shown as a key parameter to present a rater's diagnostic ability because a rater with a larger diagnostic magnifier has a uniformly better receiver operating characteristic (ROC) curve when varying the value of diagnostic bias. A simulation study is conducted to evaluate the proposed methods, and the methods are illustrated with a mammography example.
A flexible Bayesian semiparametric accelerated failure time (AFT) model is proposed for analyzing arbitrarily censored survival data with covariates subject to measurement error. Specifically, the baseline error distribution in the AFT model is nonparametrically modeled as a Dirichlet process mixture of normals. Classical measurement error models are imposed for covariates subject to measurement error. An efficient and easy-to-implement Gibbs sampler, based on the stick-breaking formulation of the Dirichlet process combined with the techniques of retrospective and slice sampling, is developed for the posterior calculation. An extensive simulation study is conducted to illustrate the advantages of our approach.
The proportional hazards (PH) model is the most widely used semiparametric regression model for analyzing right-censored survival data based on the partial likelihood method. However, the partial likelihood does not exist for interval-censored data due to the complexity of the data structure. In this paper, we focus on general interval-censored data, which is a mixture of left-, right-, and interval-censored observations. We propose an efficient and easy-to-implement Bayesian estimation approach for analyzing such data under the PH model. The proposed approach adopts monotone splines to model the baseline cumulative hazard function and allows to estimate the regression parameters and the baseline survival function simultaneously. A novel two-stage data augmentation with Poisson latent variables is developed for the efficient computation. The developed Gibbs sampler is easy to execute as it does not require imputing any unobserved failure times or contain any complicated Metropolis-Hastings steps. Our approach is evaluated through extensive simulation studies and illustrated with two real-life data sets.
Necessary and sufficient conditions are developed for the existence of the maximum likelihood estimate (MLE) for a recognition-memory model. The propriety of posteriors is shown for a class of bounded priors. Under a constant prior, an easy-to-implement Gibbs sampler is developed and illustrated via a real data set.
Current status data commonly arise in many fields such as epidemiological studies and cross-sectional tumorigenicity studies. In this article, we propose a semiparametric Bayesian approach for analyzing current status data with the proportional odds model. The use of monotone splines for the baseline odds function and a novel data augmentation with Poisson latent variables enable simple updating all of the parameters in the posterior computation. The proposed approach shows good performance and is compared with the approach in Wang and Dunson (2010) in a simulation study. We also generalize the proposed approach to analyze clustered and multivariate current status data under the frailty proportional odds models.
The proportional hazards model is widely used to deal with time to event data in many fields. However, its popularity is limited to right-censored data, for which the partial likelihood is available and the partial likelihood method allows one to estimate the regression coefficients directly without estimating the baseline hazard function. In this paper, we focus on current status data and propose an efficient and easy-to-implement Bayesian approach under the proportional hazards model. Specifically, we model the baseline cumulative hazard function with monotone splines leading to only a finite number of parameters to estimate while maintaining great modeling flexibility. An efficient Gibbs sampler is proposed for posterior computation relying on a data augmentation through Poisson latent variables. The proposed method is evaluated and compared to a constrained maximum likelihood method and three other existing approaches in a simulation study. Uterine fibroid data from an epidemiological study are analyzed as an illustration.
Analyzing interval-censored data is difficult due to its complex data structure containing left-, interval-, and right-censored observations. An easy-to-implement Bayesian approach is proposed under the proportional odds (PO) model for analyzing such data. The nondecreasing baseline log odds function is modeled with a linear combination of monotone splines. Two efficient Gibbs samplers are developed based on two different data augmentations using the relationship between the PO model and the logistic distribution. In the first data augmentation, the logistic distribution is achieved by the scaled normal mixture with the scale parameter related to the Kolmogorov–Smirnove distribution. In the second data augmentation, the logistic distribution is approximated by a Student’s t distribution up to a scale constant. The proposed methods are evaluated by simulation studies and illustrated with an application of an HIV data set.
Interval‐censored data occur naturally in many fields and the main feature is that the failure time of interest is not observed exactly, but is known to fall within some interval. In this paper, we propose a semiparametric probit model for analyzing case 2 interval‐censored data as an alternative to the existing semiparametric models in the literature. Specifically, we propose to approximate the unknown nonparametric nondecreasing function in the probit model with a linear combination of monotone splines, leading to only a finite number of parameters to estimate. Both the maximum likelihood and the Bayesian estimation methods are proposed. For each method, regression parameters and the baseline survival function are estimated jointly. The proposed methods make no assumptions about the observation process and can be applicable to any interval‐censored data with easy implementation. The methods are evaluated by simulation studies and are illustrated by two real‐life interval‐censored data applications. Copyright © 2010 John Wiley & Sons, Ltd.
The existence of the posterior distribution for one-way random effect probit models has been investigated when the uniform prior is applied to the overall mean and a class of noninformative priors are applied to the variance parameter. The sufficient conditions to ensure the propriety of the posterior are given for the cases with replicates at some factor levels. It is shown that the posterior distribution is never proper if there is only one observation at each factor level. For this case, however, a class of proper priors for the variance parameter can provide the necessary and sufficient conditions for the propriety of the posterior.
The entire dissertation/thesis text is included in the research.pdf file; the official abstract appears in the short.pdf file (which also appears in the research.pdf); a non-technical general description, or public abstract, appears in the public.pdf file.