Abstract The overlap coefficient (OVL) is widely used to quantify the similarity between two probability distributions, with applications in agriculture, quality control, and related fields. This study evaluates four OVL estimators: two nonparametric methods, the histogram-based estimator and the kernel density estimator (KDE), and two parametric estimators under the binormal model, with and without Box–Cox transformation, in terms of bias and root mean squared error (RMSE). A simulation study examined estimator behavior under normal, approximately normalizable, and strongly non-normal distributions. Results indicate that the KDE provides more stable and accurate nonparametric estimates than the histogram method, while parametric estimators perform best when normality assumptions hold. The Box–Cox transformation provides a robust approach to transform to normality allowing for the use of the binormal model even when normality assumptions fail. The methods were also applied to a dataset on apple quality attributes to illustrate their performance in a real agricultural quality-control context. The empirical findings support the simulation results: KDE yields stable nonparametric estimates, and parametric estimators produce larger OVL values when normality is satisfied. These results offer practical guidance for selecting appropriate OVL estimators in agricultural and quality-control applications.
The overlap coefficient (OVL) quantifies the similarity between two distributions through the overlapping area of their distribution functions. It has been discussed in the literature in a variety of different contexts. One approach for testing the bioequivalence of treatments is to measure the overlap of the distributions of individual responses to therapy. In some situations, covariates can significantly influence distributional overlap. This paper develops a covariate-specific OVL estimator using linear regression with a possible Box-Cox transformation. Bootstrap-based confidence intervals for the covariate-specific OVL are proposed and evaluated through extensive simulations. The methodology is illustrated using fingerstick post-prandial blood glucose measurements as a biomarker for diabetes patients adjusted for age.
Receiver operating characteristic (ROC) curve analysis is widely used in evaluating the effectiveness of a diagnostic test/biomarker or classifier score. A parametric approach for statistical inference on ROC curves based on a Box-Cox transformation to normality has frequently been discussed in the literature. Many investigators have highlighted the difficulty of taking into account the variability of the estimated transformation parameter when carrying out such an analysis. This variability is often ignored and inferences are made by considering the estimated transformation parameter as fixed and known. In this paper, we will review the literature discussing the use of the Box-Cox transformation for ROC curves and the methodology for accounting for the estimation of the Box-Cox transformation parameter in the context of ROC analysis, and detail its application to a number of problems. We present a general framework for inference on any functional of interest, including common measures such as the AUC, the Youden index, and the sensitivity at a given specificity (and vice versa). We further developed a new R package (named 'rocbc') that carries out all discussed approaches and is available in CRAN.
Currently, there is global interest in deriving new promising cancer biomarkers that could complement or substitute the conventional ones. Clinical decisions can often be based on the cutoff that corresponds to the maximized Youden index when maximum accuracy drives decisions. When more than one classification criteria are measured within the same individuals, correlated measurements arise. In this work, we propose hypothesis tests and confidence intervals for the comparison of two correlated receiver operating characteristic (ROC) curves in terms of their corresponding maximized Youden indices. We explore delta‐based techniques under parametric assumptions, or power transformations. Nonparametric kernel‐based methods are also examined. We evaluate our approaches through simulations and illustrate them using data from a metabolomic study referring to the detection of pancreatic cancer.
The overlap coefficient (OVL) measures the similarity between two distributions through the overlapping area of their distribution functions. Given its intuitive description and ease of visual representation by the straightforward depiction of the amount of overlap between the two corresponding histograms based on samples of measurements from each one of the two distributions, the development of accurate methods for confidence interval construction can be useful for applied researchers. The overlap coefficient has received scant attention in the literature since it lacks readily available software for its implementation, while inferential procedures that can cover the whole range of distributional scenarios for the two underlying distributions are missing. Such methods, both parametric and non-parametric are developed in this article, while R-code is provided for their implementation. Parametric approaches based on the binormal model show better performance and are appropriate for use in a wide range of distributional scenarios. Methods are assessed through a large simulation study and are illustrated using a dataset from a study on human immunodeficiency virus-related cognitive function assessment.
Receiver operating characteristic (ROC) analysis is the methodological framework of choice for the assessment of diagnostic markers and classification procedures in general, in both two‐class and multiple‐class classification problems. We focus on the three‐class problem for which inference usually involves formal hypothesis testing using a proxy metric such as the volume under the ROC surface (VUS). In this article, we develop an existing approach from the two‐class ROC framework. We define a hypothesis‐testing procedure that directly compares two ROC surfaces under the assumption of the trinormal model. In the case of the assessment of a single marker, the corresponding ROC surface is compared with the chance plane, that is, to an uninformative marker. A simulation study investigating the proposed tests with existing ones on the basis of the VUS metric follows. Finally, the proposed methodology is applied to a dataset of a panel of pancreatic cancer diagnostic markers. The described testing procedures along with related graphical tools are supported in the corresponding R‐package trinROC, which we have developed for this purpose.
Evaluation of the overall accuracy of biomarkers might be based on average measures of the sensitivity for all possible specificities -and vice versa- or equivalently the area under the receiver operating characteristic (ROC) curve that is typically used in such settings. In practice clinicians are in need of a cutoff point to determine whether intervention is required after establishing the utility of a continuous biomarker. The Youden index can serve both purposes as an overall index of a biomarker's accuracy, that also corresponds to an optimal, in terms of maximizing the Youden index, cutoff point that in turn can be utilized for decision making. In this paper, we provide new methods for constructing confidence intervals for both the Youden index and its corresponding cutoff point. We explore approaches based on the delta approximation under the normality assumption, as well as power transformations to normality and nonparametric kernel- and spline-based approaches. We compare our methods to existing techniques through simulations in terms of coverage and width. We then apply the proposed methods to serum-based markers of a prospective observational study involving diagnosis of late-onset sepsis in neonates.
This article explores both existing and new methods for the construction of confidence intervals for differences of indices of diagnostic accuracy of competing pairs of biomarkers in three-class classification problems and fills the methodological gaps for both parametric and non-parametric approaches in the receiver operating characteristic surface framework. The most widely used such indices are the volume under the receiver operating characteristic surface and the generalized Youden index. We describe implementation of all methods and offer insight regarding the appropriateness of their use through a large simulation study with different distributional and sample size scenarios. Methods are illustrated using data from the Alzheimer's Disease Neuroimaging Initiative study, where assessment of cognitive function naturally results in a three-class classification setting.
The three-class approach is used for progressive disorders when clinicians and researchers want to diagnose or classify subjects as members of one of three ordered categories based on a continuous diagnostic marker. The decision thresholds or optimal cut-off points required for this classification are often chosen to maximize the generalized Youden index (Nakas et al., Stat Med 2013; 32: 995–1003). The effectiveness of these chosen cut-off points can be evaluated by estimating their corresponding true class fractions and their associated confidence regions. Recently, in the two-class case, parametric and non-parametric methods were investigated for the construction of confidence regions for the pair of the Youden-index-based optimal sensitivity and specificity fractions that can take into account the correlation introduced between sensitivity and specificity when the optimal cut-off point is estimated from the data (Bantis et al., Biomet 2014; 70: 212–223). A parametric approach based on the Box–Cox transformation to normality often works well while for markers having more complex distributions a non-parametric procedure using logspline density estimation can be used instead. The true class fractions that correspond to the optimal cut-off points estimated by the generalized Youden index are correlated similarly to the two-class case. In this article, we generalize these methods to the three- and to the general k-class case which involves the classification of subjects into three or more ordered categories, where ROC surface or ROC manifold methodology, respectively, is typically employed for the evaluation of the discriminatory capacity of a diagnostic marker. We obtain three- and multi-dimensional joint confidence regions for the optimal true class fractions. We illustrate this with an application to the Trail Making Test Part A that has been used to characterize cognitive impairment in patients with Parkinson’s disease.
Many practical problems are related to the pointwise estimation of dis- tribution functions when data contains measurement errors. Motivation for these problems comes from diverse fields such as astronomy, reliability, quality control, public health and survey data. Recently, Dattner, Goldenshluger and Juditsky (2011) showed that an estimator based on a direct inversion formula for distribution functions has nice properties when the tail of the characteristic function of the mea- surement error distribution decays polynomially. In this paper we derive theoretical properties for this estimator for the case where the error distri- bution is smoother and study its finite sample behavior for different error distributions. Our method is data-driven in the sense that we use only known information, namely, the error distribution and the data. Applica- tion of the estimator to estimating hypertension prevalence based on real data is also examined.
After establishing the utility of a continuous diagnostic marker investigators will typically address the question of determining a cut-off point which will be used for diagnostic purposes in clinical decision making. The most commonly used optimality criterion for cut-off point selection in the context of ROC curve analysis is the maximum of the Youden index. The pair of sensitivity and specificity proportions that correspond to the Youden index-based cut-off point characterize the performance of the diagnostic marker. Confidence intervals for sensitivity and specificity are routinely estimated based on the assumption that sensitivity and specificity are independent binomial proportions as they arise from the independent populations of diseased and healthy subjects, respectively. The Youden index-based cut-off point is estimated from the data and as such the resulting sensitivity and specificity proportions are in fact correlated. This correlation needs to be taken into account in order to calculate confidence intervals that result in the anticipated coverage. In this article we study parametric and non-parametric approaches for the construction of confidence intervals for the pair of sensitivity and specificity proportions that correspond to the Youden index-based optimal cut-off point. These approaches result in the anticipated coverage under different scenarios for the distributions of the healthy and diseased subjects. We find that a parametric approach based on a Box-Cox transformation to normality often works well. For biomarkers following more complex distributions a non-parametric procedure using logspline density estimation can be used.
The ROC (receiver operating characteristic) curve is frequently used for describing effectiveness of a diagnostic marker or test. Classical estimation of the ROC curve uses independent identically distributed samples taken randomly from the healthy and diseased populations. Frequently not all subjects undergo a definitive gold standard assessment of disease status (verification). Estimation of the ROC curve based on data only from subjects with verified disease status may be badly biased (verification bias). In this work we investigate the properties of the doubly robust (DR) method for estimating the ROC curve adjusted for covariates (ROC regression) under verification bias. We develop the estimator's asymptotic distribution and examine its finite sample size properties via a simulation study. We apply this procedure to fingerstick postprandial blood glucose measurement data adjusting for age.
The medical records of 3922 school children residing in the Greater Haifa Metropolitan Area in Northern Israel were analyzed. Individual exposure to ambient air pollution (SO(2) and PM(10)) for each child was estimated using Geographic Information Systems tools. Factors affecting childhood asthma risk were then investigated using logistic regression and the more recently developed Bayesian Model Averaging (BMA) tools. The analysis reveals that childhood asthma in the study area appears to be significantly associated with particulate matter of less than 10 μm in aerodynamic diameter (PM(10)) (Odds Ratio (OR) = .11; P<0.001). However, no significant association with asthma prevalence was found for SO(2) (P >0.2), when PM(10) and SO(2) were introduced into the models simultaneously. When considering a change in PM(10) between the least and the most polluted parts of the study area (9.4 μg/m(3)), the corresponding OR, calculated using the BMA analysis, is 2.58 (with 95% posterior probability limits of OR ranging from 1.52 to 4.41), controlled for gender, age, proximity to main roads, the town of a child's residence, and family's socio-economic status. Thus, it is concluded that exposure to airborne particular matter, even at relatively low concentrations (40-50 μg/m(3)), generally below international air pollution standards (55-70 μg/m(3)), appears to be a considerable risk factor for childhood asthma in urban areas. This should be a cause of concern for public health authorities and environmental decision-makers.
The accuracy of a diagnostic test is typically characterised using the receiver operating characteristic (ROC) curve. Summarising indexes such as the area under the ROC curve (AUC) are used to compare different tests as well as to measure the difference between two populations. Often additional information is available on some of the covariates which are known to influence the accuracy of such measures. We propose nonparametric methods for covariate adjustment of the AUC. Models with normal errors and non-normal errors are discussed and analysed separately. Nonparametric regression is used for estimating mean and variance functions in both scenarios. In the general noise case we propose a covariate-adjusted Mann-Whitney estimator for AUC estimation which effectively uses available data to construct working samples at any covariate value of interest and is computationally efficient for implementation. This provides a generalisation of the Mann-Whitney approach for comparing two populations by taking covariate effects into account. We derive asymptotic properties for the AUC estimators in both settings, including asymptotic normality, optimal strong uniform convergence rates and MSE consistency. The usefulness of the proposed methods is demonstrated through simulated and real data examples.
The ROC (receiver operating characteristic) curve is the most commonly used statistical tool for describing the discriminatory accuracy of a diagnostic test. Classical estimation of the ROC curve relies on data from a simple random sample from the target population. In practice, estimation is often complicated due to not all subjects undergoing a definitive assessment of disease status (verification). Estimation of the ROC curve based on data only from subjects with verified disease status may be badly biased. In this work we investigate the properties of the doubly robust (DR) method for estimating the ROC curve under verification bias originally developed by Rotnitzky, Faraggi and Schisterman (2006) for estimating the area under the ROC curve. The DR method can be applied for continuous scaled tests and allows for a non‐ignorable process of selection to verification. We develop the estimator's asymptotic distribution and examine its finite sample properties via a simulation study. We exemplify the DR procedure for estimation of ROC curves with data collected on patients undergoing electron beam computer tomography, a diagnostic test for calcification of the arteries.
The Youden Index is often used as a summary measure of the receiver operating characteristic curve. It measures the effectiveness of a diagnostic marker and permits the selection of an optimal threshold value or cutoff point for the biomarker of interest. Some markers, while basically continuous and positive, have a spike or positive mass of probability at the value zero. We provide a flexible modeling approach for estimating the Youden Index and its associated cutoff point for such spiked data and compare it with the standard empirical approach. We show how this modeling approach can be adjusted to take covariate information into account. This approach is applied to data on the Coronary Calcium Score, a marker for atherosclerosis. Published in 2007 by John Wiley & Sons, Ltd.
Andrew R. Conn合作论文数Department of Mathematical Sciences
IBM T.J. Watson Research Center;Numerical Analysis Group1