When a Medicare healthcare provider is suspected of billing abuse, a population of payments X made to that provider over a fixed timeframe is isolated. A certified medical reviewer, in a time-consuming process, can determine the overpayment Y = X - (amount justified by the evidence) associated with each payment. Typically, there are too many payments in the population to examine each with care, so a probability sample is selected. The sample overpayments are then used to calculate a 90% lower confidence bound for the total population overpayment. This bound is the amount demanded for recovery from the provider. Unfortunately, classical methods for calculating this bound sometimes fail to provide the 90% confidence level, especially when using a stratified sample. In this paper, 166 redacted samples from Medicare integrity investigations are displayed and described, along with 156 associated payment populations. The 7,588 examined (Y,X) sample pairs show (1) Medicare audits have high error rates: more than 76% of these payments were considered to have been paid in error; and (2) the patterns in these samples support an "All-or-Nothing" mixture model for (Y, X) previously defined in the literature. Model-based Monte Carlo testing procedures for Medicare sampling plans are discussed, as well as stratification methods based on anticipated model moments. In terms of viability (achieving the 90% confidence level) a new stratification method defined here is competitive with the best of the many existing methods tested and seems less sensitive to choice of operating parameters. In terms of overpayment recovery (equivalent to precision) the new method is also comparable to the best of the many existing methods tested. Unfortunately, no stratification algorithm tested was ever viable for more than about half of the 104 test populations.
Variability between raters' ordinal scores is commonly observed in imaging tests, leading to uncertainty in the diagnostic process. In breast cancer screening, a radiologist visually interprets mammograms and MRIs, while skin diseases, Alzheimer's disease, and psychiatric conditions are graded based on clinical judgment. Consequently, studies are often conducted in clinical settings to investigate whether a new training tool can improve the interpretive performance of raters. In such studies, a large group of experts each classify a set of patients' test results on two separate occasions, before and after some form of training with the goal of assessing the impact of training on experts' paired ratings. However, due to the correlated nature of the ordinal ratings, few statistical approaches are available to measure association between raters' paired scores. Existing measures are restricted to assessing association at just one time point for a single screening test. We propose here a novel paired kappa to provide a summary measure of association between many raters' paired ordinal assessments of patients' test results before versus after rater training. Intrarater association also provides valuable insight into the consistency of ratings when raters view a patient's test results on two occasions with no intervention undertaken between viewings. In contrast to existing correlated measures, the proposed kappa is a measure that provides an overall evaluation of the association among multiple raters' scores from two time points and is robust to the underlying disease prevalence. We implement our proposed approach in two recent breast-imaging studies and conduct extensive simulation studies to evaluate properties and performance of our summary measure of association.
Agreement between experts' ratings is an important prerequisite for an effective screening procedure. In clinical settings, large‐scale studies are often conducted to compare the agreement of experts' ratings between new and existing medical tests, for example, digital versus film mammography. Challenges arise in these studies where many experts rate the same sample of patients undergoing two medical tests, leading to a complex correlation structure between experts' ratings. Here, we propose a novel paired kappa measure to compare the agreement between the binary ratings of many experts across two medical tests. Existing approaches can accommodate only a small number of experts, rely heavily on Cohen's kappa and Scott's pi measures of agreement, and thus are prone to their drawbacks. The proposed kappa appropriately accounts for correlations between ratings due to patient characteristics, corrects for agreement due to chance, and is robust to disease prevalence and other flaws inherent in the use of Cohen's kappa. It can be easily calculated in the software package R. In contrast to existing approaches, the proposed measure can flexibly incorporate large numbers of experts and patients by utilizing the generalized linear mixed models framework. It is intended to be used in population‐based studies, increasing efficiency without increasing modeling complexity. Extensive simulation studies demonstrate low bias and excellent coverage probability of the proposed kappa under a broad range of conditions. Methods are applied to a recent nationwide breast cancer screening study comparing film mammography to digital mammography.
Ordinal classification scales are commonly used to define a patient's disease status in screening and diagnostic tests such as mammography. Challenges arise in agreement studies when evaluating the association between many raters' classifications of patients' disease or health status when an ordered categorical scale is used. In this paper, we describe a population-based approach and chance-corrected measure of association to evaluate the strength of relationship between multiple raters' ordinal classifications where any number of raters can be accommodated. In contrast to Shrout and Fleiss' intraclass correlation coefficient, the proposed measure of association is invariant with respect to changes in disease prevalence. We demonstrate how unique characteristics of individual raters can be explored using random effects. Simulation studies are conducted to demonstrate the properties of the proposed method under varying assumptions. The methods are applied to two large-scale agreement studies of breast cancer screening and prostate cancer severity.
Large-scale agreement studies are becoming increasingly common in medical settings to gain better insight into discrepancies often observed between experts' classifications. Ordered categorical scales are routinely used to classify subjects' disease and health conditions. Summary measures such as Cohen's weighted kappa are popular approaches for reporting levels of association for pairs of raters' ordinal classifications. However, in large-scale studies with many raters, assessing levels of association can be challenging due to dependencies between many raters each grading the same sample of subjects' results and the ordinal nature of the ratings. Further complexities arise when the focus of a study is to examine the impact of rater and subject characteristics on levels of association. In this paper, we describe a flexible approach based upon the class of generalized linear mixed models to assess the influence of rater and subject factors on association between many raters' ordinal classifications. We propose novel model-based measures for large-scale studies to provide simple summaries of association similar to Cohen's weighted kappa while avoiding prevalence and marginal distribution issues that Cohen's weighted kappa is susceptible to. The proposed summary measures can be used to compare association between subgroups of subjects or raters. We demonstrate the use of hypothesis tests to formally determine if rater and subject factors have a significant influence on association, and describe approaches for evaluating the goodness-of-fit of the proposed model. The performance of the proposed approach is explored through extensive simulation studies and is applied to a recent large-scale cancer breast cancer screening study.
Widespread inconsistencies are commonly observed between physicians' ordinal classifications in screening tests results such as mammography. These discrepancies have motivated large-scale agreement studies where many raters contribute ratings. The primary goal of these studies is to identify factors related to physicians and patients' test results, which may lead to stronger consistency between raters' classifications. While ordered categorical scales are frequently used to classify screening test results, very few statistical approaches exist to model agreement between multiple raters. Here we develop a flexible and comprehensive approach to assess the influence of rater and subject characteristics on agreement between multiple raters' ordinal classifications in large-scale agreement studies. Our approach is based upon the class of generalized linear mixed models. Novel summary model-based measures are proposed to assess agreement between all, or a subgroup of raters, such as experienced physicians. Hypothesis tests are described to formally identify factors such as physicians' level of experience that play an important role in improving consistency of ratings between raters. We demonstrate how unique characteristics of individual raters can be assessed via conditional modes generated during the modeling process. Simulation studies are presented to demonstrate the performance of the proposed methods and summary measure of agreement. The methods are applied to a large-scale mammography agreement study to investigate the effects of rater and patient characteristics on the strength of agreement between radiologists. Copyright © 2017 John Wiley & Sons, Ltd.
Many disease diagnoses involve subjective judgments by qualified raters. For example, through the inspection of a mammogram, MRI, or ultrasound image, the clinician himself becomes part of the measuring instrument. To reduce diagnostic errors and improve the quality of diagnoses, it is necessary to assess raters' diagnostic skills and to improve their skills over time. This paper focuses on a subjective binary classification process, proposing a hierarchical model linking data on rater opinions with patient true disease‐development outcomes. The model allows for the quantification of the effects of rater diagnostic skills (bias and magnifier) and patient latent disease severity on the rating results. A Bayesian Markov chain Monte Carlo (MCMC) algorithm is developed to estimate these parameters. Linking to patient true disease outcomes, the rater‐specific sensitivity and specificity can be estimated using MCMC samples. Cost theory is used to identify poor‐ and strong‐performing raters and to guide adjustment of rater bias and diagnostic magnifier to improve the rating performance. Furthermore, diagnostic magnifier is shown as a key parameter to present a rater's diagnostic ability because a rater with a larger diagnostic magnifier has a uniformly better receiver operating characteristic (ROC) curve when varying the value of diagnostic bias. A simulation study is conducted to evaluate the proposed methods, and the methods are illustrated with a mammography example.
We consider using the penny as the sampling unit in healthcare audits, where high levels of payment error—including 100 percent error—are not uncommon. The payment population is a large but finite population of pennies. Using a simple random sample of n pennies, counting the number paid in error, the hypergeometric probability distribution can be used with the principle of hypothesis test inversion to obtain a lower confidence bound for the total number of population pennies paid in error. The bound is design based, not model based: its confidence level/under-recoupment rate is mathematically guaranteed to meet or exceed the 90 percent level prescribed by the Centers for Medicare and Medicaid Services guidelines. Conservative penny sampling is robust, simple, affordable, and efficient in the frequently encountered case when payment errors are all or nothing. For example, using a sample of only 45 pennies, 95 percent of the total population payment will be recovered in a 100 percent error case. However, the efficiency of penny sampling relative to classical methods based on the central limit theorem, using claims or beneficiaries as the sampling unit, may suffer if a high percentage of partial overpayments exists in the payment population. The cost of a penny sample of size n would be less than the cost of auditing a sample of n claims or beneficiaries, however: for each sampled penny one must only investigate a single claim detail line. Detailed instructions on how to select and project a conservative penny sample are given. Tables and charts to forecast efficiency and thus aid in choosing sample size are provided. The validity of the new method is established mathematically. The relationship of conservative penny sampling to classical methods in auditing is discussed. Monte Carlo testing using real healthcare payment populations compares the efficiency of the new method relative to extrapolation methods commonly used in healthcare audits.
As the authors of the TAS article “Closed Sequential and Multistage Sampling for Binary Responses With and Without Replacement” (The American Statistician, 66, pp. 163–172), we were very excited th...
Screening and diagnostic procedures often require a physician's subjective interpretation of a patient's test result using an ordered categorical scale to define the patient's disease severity. Because of wide variability observed between physicians' ratings, many large‐scale studies have been conducted to quantify agreement between multiple experts' ordinal classifications in common diagnostic procedures such as mammography. However, very few statistical approaches are available to assess agreement in these large‐scale settings. Many existing summary measures of agreement rely on extensions of Cohen's kappa. These are prone to prevalence and marginal distribution issues, become increasingly complex for more than three experts, or are not easily implemented. Here we propose a model‐based approach to assess agreement in large‐scale studies based upon a framework of ordinal generalized linear mixed models. A summary measure of agreement is proposed for multiple experts assessing the same sample of patients' test results according to an ordered categorical scale. This measure avoids some of the key flaws associated with Cohen's kappa and its extensions. Simulation studies are conducted to demonstrate the validity of the approach with comparison with commonly used agreement measures. The proposed methods are easily implemented using the software package R and are applied to two large‐scale cancer agreement studies. Copyright © 2015 John Wiley & Sons, Ltd.
Abstract In comparing several unknowns, e.g. the means μ 1 , μ 2 , …, μ k of k > 2 populations, the classical approach would recommend starting with a test of equality, i.e. test the hypothesis H 0 : μ 1 = μ 2 = … = μ k . The process known as multiple comparisons has historically been considered as the logical next step following a rejection of this ‘omnibus’ test: if we can be confident that some differences among the μ i exist, then which pairs μ i , μ i ' are different (1 ≤ i ≠ i ′ ≤ k ) and (perhaps) how large are these differences? Modern multiple comparisons analysis retains the goal of answering these sorts of questions, though the preliminary test of equality is no longer sacrosanct.
We consider closed sequential or multistage sampling, with or without replacement, from a lot of N items, where each item can be identified as defective (in error, tainted, etc.) or not. The goal is to make inference on the proportion, pi, of defectives in the lot, or equivalently on the number of defectives in the lot D = N pi. It is shown that exact inference on pi using closed (bounded) sequential or multistage procedures with general prespecified elimination boundaries is completely tractable and not at all inconvenient using modern statistical software. We give relevant theory and demonstrate functions. for this purpose written in R (R Development Core Team 2005, available as online supplementary material). Applicability of the methodology is illustrated in three examples: (1) sharpening of Wald's (1947) sequential probability ratio test used in industrial acceptance sampling, (2) two-stage sampling for auditing Medicare or Medicaid health care providers, and (3) risk-limited sequential procedures for election audits.
Billions of dollars are lost each year to Medicare and Medicaid fraud. Using three real payment populations, we consider the operating characteristics of commonly used sampling-and-extrapolation strategies for these audits: simple random sampling using (1) the simple expansion estimator or (2) the ratio estimator; and (3) stratified sampling where the basis of stratification is the payment amount. The achieved confidence level (=rate of under-recoupment) of the lower confidence bound based on the ratio estimator fell far below the government-prescribed 90% level for all three populations in commonly encountered high denial-rate scenarios. For the expansion estimator in simple random sampling, the achieved confidence level depends on the skew of the overpayment population: if it is left skewed, the level will fall below 90%, sometimes far below; if it is right-skewed, it will exceed 90%. In the latter case, careless stratification by payment amount can destroy this conservatism. When there is strong right skew, limited stratification can sometimes preserve the 90% confidence while yielding improvements in overpayment recovery. In any population where 90% under-recoupment is not achieved by extrapolation methods based on the central limit theorem, methods based on sample counts and the hypergeometric distribution (Edwards et al., Health Serv Outcomes Res Methodol 4:241–263, 2005 ; Gilliland and Feng, Health Serv Outcomes Res Methodol 10:154–164, 2010; Edwards et al., Pennysampling. Technical Report No. 232, Dept. of Statistics, University of South Carolina, Columbia, SC, 2010 ) should be considered; these mathematically guarantee the 90% confidence level. Regardless of the sampling and extrapolation plan being considered, operating characteristics (under-recoupment rate, overpayment recovery, etc.) should be thoroughly checked in the planning stages with Monte Carlo simulation testing, which we use throughout this paper.
Consider a population of N payments by Medicare to a health care provider, each payment for $4000. Suppose that an unknown number M of the payments are not justified and the other N M payments are justified. A random sample of n payments is chosen and audited to determine if money should be recouped from the provider for the unjustified payments. Medicare guidelines state that in most situations the recoupment figure be determined by the lower end of a one-sided 90% confidence interval estimate of M based on a finding of X sample payments not to be justified. An exact lower estimate L = L(X) with the property P-M (L <= M)>= 0.90 for all M = 0, 1, ... , N - 1 is an integral component of the minimum sum method that has been used in setting recoupment figures in Medicare audits (Edwards et al. 2003).In this article, we show the simple construction of a randomized lower estimate L R that improves upon L in the sense that L-R >= L with probability 1 with P*(M) (L-R <= M) = 0.90 for all M = 0, 1, ... , N - 1. In addition, we report on coverage probability and expectation from a Monte Carlo study that compares L-R to lower estimates that are based on normal approximation.Whereas randomized confidence intervals for the parameter pi of the Binomial distribution are well-studied and discussed in the literature, this seems not to be the case with the estimation of the parameter M of the Hypergeometric distribution.
Many large-scale studies have recently been carried out to assess the reliability of diagnostic procedures, such as mammography for the detection of breast cancer. The large numbers of raters and subjects involved raise new challenges in how to measure agreement in these types of studies. An important motivator of these studies is the identification of factors that contribute to the often wide discrepancies observed between raters' classifications, such as a rater's experience, in order to improve the reliability of the diagnostic process of interest. Incorporating covariate information into the agreement model is a key component in addressing these questions. Few agreement models are currently available that jointly model larger numbers of raters and subjects and incorporate covariate information. In this paper, we extend a recently developed population-based model and measure of agreement for binary ratings to incorporate covariate information using the class of generalized linear mixed models with a probit link function. Important information on factors related to the subjects and raters can be included as fixed and/or random effects in the model. We demonstrate how agreement can be assessed between subgroups of the raters and/or subjects, for example, comparing agreement between experienced and less experienced raters. Simulation studies are carried out to test the performance of the proposed models and measures of agreement. Application to a large-scale breast cancer study is presented.
Periodicity is omnipresent in environmental time series data. For modeling estuarine water quality variables, harmonic regression analysis has long been the standard for dealing with periodicity. Generalized additive models (GAMs) allow more flexibility in the response function. They permit parametric, semiparametric, and nonparametric regression functions of the predictor variables. We compare harmonic regression, GAMs with cubic regression splines, and GAMs with cyclic regression splines in simulations and using water quality data collected from the National Estuarine Reasearch Reserve System (NERRS). While the classical harmonic regression model works well for clean, near‐sinusoidal data, the GAMs are competitive and are very promising for more complex data. The generalized additive models are also more adaptive and require less‐intervention. Copyright © 2009 John Wiley & Sons, Ltd.
Random sampling of paid medicare claims has been legally acceptable for investigating suspicious billing practices by health care providers since 1986. A population of payments made to a given provider during a given time frame is isolated and a probability sample selected for investigation. A lower confidence bound for the total amount overpaid to the provider is then used as a recoupment demand. Edwards et al. (Health Serv Outcomes Res Methodol 4:241–263, 2005) show that methods based on the Central Limit Theorem can fail badly and propose an alternative method, called the minimum sum method, for fixed sample sizes. In this paper the sampling is performed in two stages. In case of little abuse in the first stage the investigation is stopped; otherwise a second sample is examined. Based on this strategy a lower confidence bound for the total number of universe payments in error and a corresponding lower bound for the total overpayment amount are defined. Criteria for choosing the sampling parameters are considered. Relative efficiencies are studied.
The authors describe a model-based kappa statistic for binary classifications which is interpretable in the same manner as Scott's pi and Cohen's kappa, yet does not suffer from the same flaws. They compare this statistic with the data-driven and population-based forms of Scott's pi in a population-based setting where many raters and subjects are involved, and inference regarding the underlying diagnostic procedure is of interest. The authors show that Cohen's kappa and Scott's pi seriously underestimate agreement between experts classifying subjects for a rare disease; in contrast, the new statistic is robust to changes in prevalence. The performance of the three statistics is illustrated with simulations and prostate cancer data.
Quadratic response surface methodology often focuses on finding the levels of some (coded) predictor variables x = (x1, x2 ,..., xk) that optimize the expected value of a response variable y. Often the experimenter starts from some best guess or “control” combination of the predictor variables (usually coded to x = 0) and performs an experiment varying them in a region around this center point. Sa and Edwards (1993) provided simultaneous confidence intervals for the improvement in E(y) at x as compared to 0, δ(x) = E(y|x) — E(y | 0) for all x within a specified distance of 0. Here, we give a new method for inference on δ(x) in a 2nd-order rotatable design, utilizing a simulation-based critical point to give substantially sharper intervals. For large k, to achieve the same level of precision using the Sa and Edwards methods over the simulation-based method, the experimenter would need to double the experiment size. The new method allows for one-sided bounds, blocks and covariates in the design, and gives valid inference for all types of second order surfaces.