BACKGROUND AND OBJECTIVES:To improve upon existing hospital grading systems, we developed a new report card based on multivariate matching. RESEARCH DESIGN:Matched cohorts. For each focal hospital patient, we match 10 control patients treated at "well-resourced" hospitals with excellent hospital characteristics from across the nation, and 10 control patients treated at "typical" hospitals, on over 300 patient characteristics from Medicare Claims. Grades were based on outcome differences between patients at the focal hospital and their matched controls. We also create an "Analogous" match that is comprised of multiple control patients matched to each focal hospital patient with similar patient characteristics who were treated at hospitals with similar characteristics to the focal hospital, answering the question, "How would patients who looked like my patients and who were treated at hospitals like my hospital fare, compared to how my patients fared." We also report outcomes by multimorbidity status. SUBJECTS:Medicare admissions from 2017 to 2019 for heart attack, heart failure and pneumonia. To illustrate our methods, we report on 4 hospitals in the same region: a well-known "Flagship" teaching Hospital, an Affiliated Hospital within the same flagship system, a Poor-Performing Hospital that is not part of the flagship system, and a Small Hospital with unstable estimates. MEASURES:Thirty-day mortality and revisit rates. RESULTS:Report cards for each example hospital. CONCLUSIONS:Matched report cards allow users to better benchmark hospitals and see those types of patients where a specific hospital is performing poorly compared to other hospitals treating very similar patients.
In experimental design, aliasing of effects occurs in fractional factorial experiments, where certain low order factorial effects are indistinguishable from certain high order interactions: low order contrast weights may be orthogonal to one another, while their higher order interactions are aliased and not identified. In observational studies, aliasing occurs when certain combinations of covariates-for example, time period and various eligibility criteria for treatment-perfectly predict the treatment that an individual will receive, so a covariate combination is aliased with a particular treatment. In this situation, when a contrast among several groups is used to estimate a treatment effect, collections of individuals defined by contrast weights may be balanced with respect to summaries of low-order interactions between covariates and treatments, but necessarily not balanced with respect high-order interactions. We develop a theory of aliasing in observational studies, illustrate that theory in an observational study whose aliasing is more robust than conventional difference-in-differences, and develop a new form of matching to construct balanced confounded factorial designs from observational data. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
In an observational block design, there are I blocks of J individuals, typically with one treated individual and J−1 controls; however, unlike a randomized block design, individuals were not randomly assigned to treatment or control. To be convincing, an observational block design must demonstrate that an ostensible treatment effect is not actually a consequence of small or moderate unmeasured biases of treatment assignment in the absence of a treatment effect. It is known that weighting to ignore blocks with a small range of responses increases the ability to distinguish a treatment effect from a bias in treatment assignment—that is, it increases the design sensitivity. Here, it is shown that a new tactic further increases design sensitivity. The new tactic involves a conditional statistic, such that blocks with moderately large ranges are considered conditionally given that the treated individual has either the largest or smallest response in the block. The new tactic is explored: (i) in terms of an asymptotic measure, the design sensitivity, (ii) in simulation of the power of a sensitivity analysis in finite samples, and (iii) in an example. Adaptive inference is briefly discussed. An R package weightedRank implements the method, contains the data, and reproduces the empirical results.
ImportanceIn surgical patients, it is well known that higher hospital procedure volume is associated with better outcomes. To our knowledge, this volume-outcome association has not been studied in ambulatory surgery centers (ASCs) in the US. ObjectiveTo determine if low-volume ASCs have a higher rate of revisits after surgery, particularly among patients with multimorbidity. Design, Setting, and ParticipantsThis matched case-control study used Medicare claims data and analyzed surgeries performed during 2018 and 2019 at ASCs. The study examined 2328 ASCs performing common ambulatory procedures and analyzed 4751 patients with a revisit within 7 days of surgery (defined to be either 1 of 4735 revisits or 1 of 16 deaths without a revisit). These cases were each closely matched to 5 control patients without revisits (23 755 controls). Data were analyzed from January 1, 2018, through December 31, 2019. Main Outcomes and MeasuresSeven-day revisit in patients (cases) compared with the matched patients without the outcome (controls) in ASCs with low volume (less than 50 procedures over 2 years) vs higher volume (50 or more procedures). ResultsPatients at a low-volume ASC had a higher odds of a 7-day revisit vs patients who had their surgery at a higher-volume ASC (odds ratio [OR], 1.21; 95% CI, 1.09-1.36; P = .001). The odds of revisit for patients with multimorbidity were higher at low-volume ASCs when compared with higher-volume ASCs (OR, 1.57; 95% CI, 1.27-1.94; P < .001). Among patients with multimorbidity in low-volume ASCs, for those who underwent orthopedic procedures, the odds of revisit were 84% higher (OR, 1.84; 95% CI, 1.36-2.50; P < .001) vs higher-volume centers, and for those who underwent general surgery or other procedures, the odds of revisit were 36% higher (OR, 1.36; 95% CI, 1.01-1.83; P = .05) vs a higher-volume center. The findings were not statistically significant for patients without multimorbidity. Conclusions and RelevanceIn this observational study, the surgical volume of an ASC was an important indicator of patient outcomes. Older patients with multimorbidity should discuss with their surgeon the optimal location of their care.
What is the best way to split one stratum into two to maximally reduce the within-stratum imbalance in many covariates? We formulate this as an integer program and approximate the solution by randomized rounding of a linear program. A linear program may assign a fraction of a person to each refined stratum. Randomized rounding views fractional people as probabilities, assigning intact people to strata using biased coins. Randomized rounding is a well-studied theoretical technique for approximating the optimal solution of certain insoluble integer programs. When the number of people in a stratum is large relative to the number of covariates, we prove the following new results: (i) randomized rounding to split a stratum does very little randomizing, so it closely resembles the linear programming relaxation without splitting intact people; (ii) the linear relaxation and the randomly rounded solution place lower and upper bounds on the unattainable integer programming solution; and because of (i), these bounds are often close, thereby ratifying the usable randomly rounded solution. We illustrate using an observational study that balanced many covariates by forming matched pairs composed of 2016 patients selected from 5735 using a propensity score. Instead, we form 5 propensity score strata and refine them into 10 strata, obtaining excellent covariate balance while retaining all patients. An R package optrefine at CRAN implements the method. Supplementary materials are available online.
An observational block design has I blocks matched for covariates and J individuals per block, but treatments were not randomly assigned to individuals within blocks, as would have been done in an experiment. Tightening an observational block design means selecting J '<J individuals from each block, and possibly I '<= I blocks, to construct a new observational block design that, in some way, addresses unmeasured biases from nonrandom treatment assignment. Tightening must preserve covariate balance while altering the design to achieve some additional objective. An optimization algorithm is introduced that achieves this while maintaining the block structure by finely balancing covariates across blocks and through optimal subset matching. An example is considered in detail, both to motivate and illustrate the tightening of an observational block design. Two tightened designs are built from a study of light daily alcohol consumption and its possible effects on HDL cholesterol. One tightened design adjusts for an outcome tentatively presuming it was unaffected by the treatment. The second tightened design uses a differential effect to remove bias from an unobserved general disposition that promotes several treatments. An R package tightenBlock implements the method, contains the data, and in that package the help-file for the function tighten reproduces the example.
Background Observational studies of anesthetic neurotoxicity may be biased because children requiring anesthesia commonly have medical conditions associated with neurobehavioral problems. This study takes advantage of a natural experiment associated with appendicitis to determine whether anesthesia and surgery in childhood were specifically associated with subsequent neurobehavioral outcomes. Methods This study identified 134,388 healthy children with appendectomy and examined the incidence of subsequent externalizing or behavioral disorders (conduct, impulse control, oppositional defiant, attention-deficit hyperactivity disorder) or internalizing or mood or anxiety disorders (depression, anxiety, or bipolar disorder) when compared to 671,940 matched healthy controls as identified in Medicaid data between 2001 and 2018. For comparison, this study also examined 154,887 otherwise healthy children admitted to the hospital for pneumonia, cellulitis, and gastroenteritis, of which only 8% received anesthesia, and compared them to 774,435 matched healthy controls. In addition, this study examined the difference-in-differences between matched appendectomy patients and their controls and matched medical admission patients and their controls. Results Compared to controls, children with appendectomy were more likely to have subsequent behavioral disorders (hazard ratio, 1.04; 95% CI, 1.01 to 1.06; P = 0.0010) and mood or anxiety disorders (hazard ratio, 1.15; 95% CI, 1.13 to 1.17; P < 0.0001). Relative to controls, children with medical admissions were also more likely to have subsequent behavioral (hazard ratio, 1.20; 95% CI, 1.18 to 1.22; P < 0.0001) and mood or anxiety (hazard ratio, 1.25; 95% CI, 1.23 to 1.27; P < 0.0001) disorders. Comparing the difference between matched appendectomy patients and their matched controls to the difference between matched medical patients and their matched controls, medical patients had more subsequent neurobehavioral problems than appendectomy patients. Conclusions Although there is an association between neurobehavioral diagnoses and appendectomy, this association is not specific to anesthesia exposure and is stronger in medical admissions. Medical admissions, generally without anesthesia exposure, displayed significantly higher rates of these disorders than appendectomy-exposed patients. Editor’s Perspective What We Already Know About This Topic What This Article Tells Us That Is New
Do the impacts that occur when playing high school football have concussive effects that accelerate cognitive decline late in life? We examine this possibility using newly available cognitive data describing people in 2020 who graduated high school in 1957. Someone who was 18 in 1957 would be 81 in 2020. For this comparison we develop a new design for an observational study, called a triples design, and discuss its advantages and construction. A triples design consists of M blocks of size 3, where a block contains either one treated individual and two controls or two treated individuals and one control. A triples design is the simplest design that uses weights, with just two weights. The "entire number" is {1 - e(x)}/e(x), where e(x) is the propensity score at covariate x, so it is the ratio of controls-to-treated expected at x. Unlike a matched pairs design, which can remove the bias from observed covariates when the "entire number" exceeds 1, the triples design can succeed when the entire number exceeds 1/2, reflecting the possibility of matching two treated individuals to the same control. Like full matching, a triples design can match more people than can matched pairs, yet have smaller within-block covariate distances. Unlike full matching, there are no matched pairs. Like matching with multiple controls, a triples design will have a larger design sensitivity than a design which includes matched pairs, under simple models for continuous outcomes; that is, in favorable situations the design is expected to report greater insensitivity to unmeasured biases. Because there are just two weights, it is easy to construct weighted graphics for exploratory displays from triples designs. A heuristic algorithm containing network optimization constructs the design.
Objective:To compare general surgery outcomes at flagship systems, flagship hospitals, and flagship hospital affiliates versus matched controls. Summary Background Data:It is unknown whether flagship hospitals perform better than flagship hospital affiliates for surgical patients. Methods:Using Medicare claims for 2018 to 2019, we matched patients undergoing inpatient general surgery in flagship system hospitals to controls who underwent the same procedure at hospitals outside the system but within the same region. We defined a "flagship hospital" within each region as the major teaching hospital with the highest patient volume that is also part of a hospital system; its system was labeled a "flagship system." We performed 4 main comparisons: patients treated at any flagship system hospital versus hospitals outside the flagship system; flagship hospitals versus hospitals outside the flagship system; flagship hospital affiliates versus hospitals outside the flagship system; and flagship hospitals versus affiliate hospitals. Our primary outcome was 30-day mortality. Results:We formed 32,228 closely matched pairs across 35 regions. Patients at flagship system hospitals (32,228 pairs) had lower 30-day mortality than matched control patients [3.79% vs. 4.36%, difference=-0.57% (-0.86%, -0.28%), P<0.001]. Similarly, patients at flagship hospitals (15,571/32,228 pairs) had lower mortality than control patients. However, patients at flagship hospital affiliates (16,657/32,228 pairs) had similar mortality to matched controls. Flagship hospitals had lower mortality than affiliate hospitals [difference-in-differences=-1.05% (-1.62%, -0.47%), P<0.001]. Conclusions:Patients treated at flagship hospitals had significantly lower mortality rates than those treated at flagship hospital affiliates. Hence, flagship system affiliation does not alone imply better surgical outcomes.
In an observational study of the effects caused by a treatment, biases from unmeasured covariates remain a concern even after successful adjustments for measured covariates. This concern is partly addressed by demonstrating that the qualitative conclusions of the primary analysis would not be altered by small or moderate biases-that these conclusions are insen-sitive to small or moderate bias. Additionally, the concern is partly addressed by collecting additional information, such as outcomes known to be unaf-fected by the treatment, and using this information as a test of various biases. Is there a gap between these two activities? Perhaps the study is insensitive to small biases, and we can detect large biases, but the study is sensitive to mod-erate biases that cannot be detected-that is an informal description of a gap. The concept of "no gap" is defined formally in Definition 3.1, and the prob-ability of "no gap" is determined under various sampling situations. When there is no gap, ask: Are causal conclusions measurably strengthened? If so, ⠂ by how much? The answer depends upon the covering design sensitivity, ⠃, defined to be the smallest bias that can explain both the ostensible effect of the treatment on the primary outcome and the evidence of bias provided by the unaffected outcome. The covering design sensitivity is calculated in various contexts. A small observational study of the effects of light alcohol consumption on HDL cholesterol is used to illustrate ideas and methods.
In an observational study of the effects caused by a treatment, a second control group is used in an effort to detect bias from unmeasured covariates, and the investigator is content if no evidence of bias is found. This strategy is not entirely satisfactory: two control groups may differ significantly, yet the difference may be too small to invalidate inferences about the treatment, or the control groups may not differ yet nonetheless fail to provide a tangible strengthening of the evidence of a treatment effect. Is a firmer conclusion possible? Is there a way to analyze a second control group such that the data might report measurably strengthened evidence of cause and effect, that is, insensitivity to larger unmeasured biases? Evidence factor analyses are not commonly used with a second control group: most analyses compare the treated group to each control group, but analyses of that kind are partially redundant; so, they do not constitute evidence factors. An alternative analysis is proposed here, one that does yield two evidence factors, and with a carefully designed test statistic, is capable of extracting strong evidence from the second factor. The new technical work here concerns the development of a test statistic with high design sensitivity and high Bahadur efficiency in a sensitivity analysis for the second factor. A study of binge drinking as a cause of high blood pressure is used as an illustration.
We define a “flagship hospital” as the largest academic hospital within a hospital referral region and a “flagship system” as a system that contains a flagship hospital and its affiliates. It is not known if patients admitted to an affiliate hospital, and not to its main flagship hospital, have better outcomes than those admitted to a hospital outside the flagship system but within the same hospital referral region. To compare mortality at flagship hospitals and their affiliates to matched control patients not in the flagship system but within the same hospital referral region. A matched cohort study The study used hospitalizations for common medical conditions between 2018-2019 among older patients age ≥ 66 years. We analyzed 118,321 matched pairs of Medicare patients admitted with pneumonia (N=57,775), heart failure (N=42,531), or acute myocardial infarction (N=18,015) in 35 flagship hospitals, 124 affiliates, and 793 control hospitals. 30-day (primary) and 90-day (secondary) all-cause mortality. 30-day mortality was lower among patients in flagship systems versus control hospitals that are not part of the flagship system but within the same hospital referral region (difference= -0.62
The propensity score is the conditional probability of assignment to a particular treatment given a vector of observed covariates. Both large and small sample theory show that adjustment for the scalar propensity score is sufficient to remove bias due to all observed covariates. Applications include: (i) matched sampling on the univariate propensity score, which is a generalization of discriminant matching, (ii) multivariate adjustment by subclassification on the propensity score where the same subclasses are used to estimate treatment effects for all outcome variables and in all subpopulations, and (iii) visual representation of multivariate covariance adjustment by a two-dimensional plot.
To be convincing, an observational or nonrandomized study of causal effects must demonstrate that its conclusions cannot be readily explained by a small unmeasured bias in the way individuals were assigned to treatment or control. The Bahadur relative efficiency of a sensitivity analysis compares the performance of different test statistics or different research designs when sensitivity to unmeasured bias is appraised: better statistics and better designs exhibit insensitivity to larger biases. Bahadur efficiency is relevant here because, unlike Pitman efficiency, it can depict efficiency with an effect that is not trivially small: every trivially small treatment effect is sensitive to trivially small biases. The Bahadur efficiency of a sensitivity analysis has been used by various authors in the simple case of matched pairs, but the technical issues are more complex in the case of blocks larger than pairs, and they are developed here for the first time. Choosing a better test statistic for a block design, or choosing a better block size—larger than pairs—can result in enormous increases in the efficiency of a sensitivity analysis. An adaptive choice of test statistic can ensure the better Bahadur efficiency of two competing statistics. An R package weightedRank implements the methods, contains the example and reproduces its analysis. Supplementary materials for this article are available online.
Objectives Evaluate whether hospital factors, including nurse resources, explain racial differences in Medicare black and white patient surgical outcomes and whether disparities changed over time. Design Retrospective tapered-match. Setting 571 hospitals at two time points (Early Era 2003–2005; Recent Era 2013–2015). Participants 6752 black patients and three sets of 6752 white controls selected from 107 001 potential controls (Early Era). 4964 black patients and three sets of 4964 white controls selected from 74 108 potential controls (Recent Era). Interventions Black patients were matched to white controls on demographics (age, sex, state and year of procedure), procedure (demographics variables plus 136 International Classification of Diseases (ICD)-9 principal procedure codes) and presentation (demographics and procedure variables plus 34 comorbidities, a mortality risk score, a propensity score for being black, emergency admission, transfer status, predicted procedure time). Outcomes 30-day and 1-year mortality. Results Before matching, black patients had more comorbidities, higher risk of mortality despite being younger and underwent procedures at different percentages than white patients. Whites in the demographics match had lower mortality at 30 days (5.6% vs 6.7% Early Era; 5.4% vs 5.7% Recent Era) and 1-year (15.5% vs 21.5% Early Era; 12.3% vs 15.9% Recent Era). Black–white 1-year mortality differences were equivalent after matching patients with respect to presentation, procedure and demographic factors. Black–white 30-day mortality differences were equivalent after matching on procedure and demographic factors. Racial disparities in outcomes remained unchanged between the two time periods spanning 10 years. All patients in hospitals with better nurse resources had lower odds of 30-day (OR 0.60, 95% CI 0.46 to 0.78, p<0.010) and 1-year mortality (OR 0.77, 95% CI 0.65 to 0.92, p<0.010) even after accounting for other hospital factors. Conclusions Survival disparities among black and white patients are largely explained by differences in demographic, procedure and presentation factors. Better nurse resources (eg, staffing, work environment) were associated with lower mortality for all patients.
A nontechnical guide to the basic ideas of modern causal inference, with illustrations from health, the economy, and public policy. Which of two antiviral drugs does the most to save people infected with Ebola virus? Does a daily glass of wine prolong or shorten life? Does winning the lottery make you more or less likely to go bankrupt? How do you identify genes that cause disease? Do unions raise wages? Do some antibiotics have lethal side effects? Does the Earned Income Tax Credit help people enter the workforce? Causal Inference provides a brief and nontechnical introduction to randomized experiments, propensity scores, natural experiments, instrumental variables, sensitivity analysis, and quasi-experimental devices. Ideas are illustrated with examples from medicine, epidemiology, economics and business, the social sciences, and public policy.
Multivariate matching has two goals: (i) to construct treated and control groups that have similar distributions of observed covariates, and (ii) to produce matched pairs or sets that are homogeneous in a few key covariates. When there are only a few binary covariates, both goals may be achieved by matching exactly for these few covariates. Commonly, however, there are many covariates, so goals (i) and (ii) come apart, and must be achieved by different means. As is also true in a randomized experiment, similar distributions can be achieved for a high-dimensional covariate, but close pairs can be achieved for only a few covariates. We introduce a new polynomial-time method for achieving both goals that substantially generalizes several existing methods; in particular, it can minimize the earthmover distance between two marginal distributions. The method involves minimum cost flow optimization in a network built around a tripartite graph, unlike the usual network built around a bipartite graph. In the tripartite graph, treated subjects appear twice, on the far left and the far right, with controls sandwiched between them, and efforts to balance covariates are represented on the right, while efforts to find close individual pairs are represented on the left. In this way, the two efforts may be pursued simultaneously without conflict. The method is applied to our on-going study in the Medicare population of the relationship between superior nursing and sepsis mortality. The match2C package in R implements the method.