There is a dearth of safety data on maternal outcomes after perinatal medication exposure. Data-mining for unexpected adverse event occurrence in existing datasets is a potentially useful approach. One method, the Poisson tree-based scan statistic (TBSS), assumes that the expected outcome counts, based on incidence of outcomes in the control group, are estimated without error. This assumption may be difficult to satisfy with a small control group. Our simulation study evaluated the effect of imprecise incidence proportions from the control group on TBSS' ability to identify maternal outcomes in pregnancy research. We simulated base case analyses with "true" expected incidence proportions and compared these with imprecise incidence proportions derived from sparse control samples. We varied parameters that have an impact on type I error and statistical power (exposure group size, outcome's incidence proportion, and effect size). We found that imprecise incidence proportions generated by a small control group resulted in inaccurate alerting, inflation of type I error, and removal of very rare outcomes for TBSS analysis due to "zero" background counts. Ideally, the control size should be at least several times larger than the exposure size to limit the number of false positive alerts and retain statistical power for true alerts.This article is part of a Special Collection on Pharmacoepidemiology.
BACKGROUND:Concerns have been raised regarding proton pump inhibitor (PPI) use and risk of severe coronavirus disease 2019 (COVID-19). Observational studies have yielded heterogeneous results and were subject to important methodological limitations. AIMS:To examine the association between the receipt of PPIs and risk of COVID-19 hospitalizations and severe in-hospital outcomes or death. METHODS:Case-control study among Medicare fee-for-service beneficiaries 66+ years old with gastroesophageal reflux disorder (GERD). Within this population, we identified cases by an incident hospital discharge diagnosis of COVID-19 from April 1 to December 11, 2020, using the International Classification of Diseases, Tenth Revision, Clinical Modification (ICD-10-CM) U07.1, and randomly selected up to 10 controls per case, matched on date and neighborhood. We defined PPI use as a prescription providing ≥15 days of supply in the 30 days before admission, with H2-receptor antagonist (H2RA) use as the reference to account for indication. We analyzed uncomplicated hospitalizations and hospitalizations with severe outcomes (intensive/coronary care unit admission, invasive mechanical ventilation, or death), estimating odds ratios (ORs), and 95% confidence intervals (CIs) with multinomial conditional logistic regression adjusted for demographics, comorbidities, chronic medications, and health care utilization. RESULTS:We matched 25,867 uncomplicated and 12,954 severe hospitalized COVID-19 cases to 146,972 and 73,104 controls, respectively. Cases tended to be older and have more comorbidities. Relative to H2RA use, we found no association of PPI use with uncomplicated COVID-19 hospitalization (OR 0.99, 95% CI 0.93-1.06) or severe COVID-19 hospitalization (OR 1.00, 95% CI 0.91-1.10). CONCLUSIONS:Relative to H2RA use, PPI use was not associated with uncomplicated or severe COVID-19 hospitalizations among Medicare beneficiaries with GERD.
In the drug development for rare disease, the number of treated subjects in the clinical trial is often very small, whereas the number of external controls can be relatively large. There is no clear guidance on choosing an appropriate statistical method to control baseline confounding in this situation. To fill this gap, we conduct extensive simulations to evaluate the performance of commonly used matching and weighting methods as well as the more recently developed targeted maximum likelihood estimation (TMLE) and cardinality matching in small sample settings, mimicking the motivating data from a pediatric rare disease. Among the methods examined, the performance of coarsened exact matching (CEM) and TMLE are relatively robust under various model specifications. CEM is only feasible when the number of controls far exceeds the number of treated, whereas TMLE has better performance with less extreme treatment allocation ratios. Our simulations suggest bootstrap is useful for variance estimation in small samples after matching.
The International Council for Harmonization (ICH) E9(R1) addendum recommends choosing an appropriate estimand based on the study objectives in advance of trial design. One defining attribute of an estimand is the intercurrent event, specifically what is considered an intercurrent event and how it should be handled. The primary objective of a clinical study is usually to assess a product's effectiveness and safety based on the planned treatment regimen instead of the actual treatment received. The estimand using the treatment policy strategy, which collects and analyzes data regardless of the occurrence of intercurrent events, is usually utilized. In this article, we explain how missing data can be handled using the treatment policy strategy from the authors' viewpoint in connection with antihyperglycemic product development programs. The article discusses five statistical methods to impute missing data occurring after intercurrent events. All five methods are applied within the framework of the treatment policy strategy. The article compares the five methods via Markov Chain Monte Carlo simulations and showcases how three of these five methods have been applied to estimate the treatment effects published in the labels for three antihyperglycemic agents currently on the market.
Recently, retrieved-dropout-based multiple imputation has been used in some therapeutic areas to address the treatment policy estimand, mostly for continuous endpoints. In this approach, data from subjects who discontinued study treatment but remained in study were used to construct a model for multiple imputation for the missing data of subjects in the same treatment arm who discontinued study. We extend this approach to time-to-event endpoints and provide a practical guide for its implementation. We use a cardiovascular outcome trial dataset to illustrate the method and compare the results with those from Cox proportional hazard and reference-based multiple imputation methods.
Endpoints in clinical trials are often highly correlated. However, the commonly used multiple testing procedures in clinical trials either do not take into consideration the correlations among test statistics or can only exploit known correlations. Westfall and Young constructed a resampling-based stepdown method that implicitly utilizes the correlation structure of test statistics in situations with unknown correlations. However, their method requires a "subset pivotality" assumption. Romano and Wolf proposed a more general stepdown method, which does not require such an assumption. There is at present little experience with the application of such methods in analyzing clinical trial data. We advocate the application of resampling-based multiple testing procedures to clinical trials data when appropriate. We have conjectured that the resampling-based stepdown methods can be extended to a stepup procedure under appropriate assumptions and examined the performance of both stepdown and stepup methods under a variety of correlation structures and distribution types. Results from our simulation studies support the use of the resampling-based methods under various scenarios, including binary data and small samples, with strong control of Family wise type I error rate (FWER). Under positive dependence and for binary data even under independence, the resampling-based methods are more powerful than the Holm and Hochberg methods. Last, we illustrate the advantage of the resampling-based stepwise methods with two clinical trial data examples: a cardiovascular outcome trial and an oncology trial.
Measurement error arises through a variety of mechanisms. A rich literature exists on the bias introduced by covariate measurement error and on methods of analysis to address this bias. By comparison, less attention has been given to errors in outcome assessment and nonclassical covariate measurement error. We consider an extension of the regression calibration method to settings with errors in a continuous outcome, where the errors may be correlated with prognostic covariates or with covariate measurement error. This method adjusts for the measurement error in the data and can be applied with either a validation subset, on which the true data are also observed (eg, a study audit), or a reliability subset, where a second observation of error prone measurements are available. For each case, we provide conditions under which the proposed method is identifiable and leads to consistent estimates of the regression parameter. When the second measurement on the reliability subset has no error or classical unbiased measurement error, the proposed method is consistent even when the primary outcome and exposures of interest are subject to both systematic and random error. We examine the performance of the method with simulations for a variety of measurement error scenarios and sizes of the reliability subset. We illustrate the method's application using data from the Women's Health Initiative Dietary Modification Trial.
Marginal structural models are a class of causal models useful for characterizing the effect of treatment in the presence of time-varying confounding. They are more widely used than structural nested models, partly because these models are easier to understand and to implement. We extend marginal structural models to situations with clustered observations with unit- and cluster-level treatment and introduce an appropriate inferential method. We consider how to formulate models with cluster-level and unit-level treatments. For unit-level treatments, we consider cases with and without interference. We also consider the use of unit-specific inverse probability weights and certain working correlation structures to improve the efficiency of estimators in some situations. We apply our method to different scenarios including 2 or 3 units per cluster and a mixture of larger clusters. Simulation examples and data from the treatment arm of a glaucoma clinical trial were used to illustrate our method.
In many sensory organs, specialized receptors are strategically arranged to enhance detection sensitivity and acuity. It is unclear whether the olfactory system utilizes a similar organizational scheme to facilitate odor detection. Curiously, olfactory sensory neurons (OSNs) in the mouse nose are differentially stimulated depending on the cell location. We therefore asked whether OSNs in different locations evolve unique structural and/or functional features to optimize odor detection and discrimination. Using immunohistochemistry, computational fluid dynamics modeling, and patch clamp recording, we discovered that OSNs situated in highly stimulated regions have much longer cilia and are more sensitive to odorants than those in weakly stimulated regions. Surprisingly, reduction in neuronal excitability or ablation of the olfactory G protein in OSNs does not alter the cilia length pattern, indicating that neither spontaneous nor odor-evoked activity is required for its establishment. Furthermore, the pattern is evident at birth, maintained into adulthood, and restored following pharmacologically induced degeneration of the olfactory epithelium, suggesting that it is intrinsically programmed. Intriguingly, type III adenylyl cyclase (ACIII), a key protein in olfactory signal transduction and ubiquitous marker for primary cilia, exhibits location-dependent gene expression levels, and genetic ablation of ACIII dramatically alters the cilia pattern. These findings reveal an intrinsically programmed configuration in the nose to ensure high sensitivity to odors.
In assessing the efficacy of a time-varying treatment structural nested models (SNMs) are useful in dealing with confounding by variables affected by earlier treatments. These models often consider treatment allocation and repeated measures at the individual level. We extend SNMMs to clustered observations with time-varying confounding and treatments. We demonstrate how to formulate models with both cluster- and unit-level treatments and show how to derive semiparametric estimators of parameters in such models. For unit-level treatments, we consider interference, namely the effect of treatment on outcomes in other units of the same cluster. The properties of estimators are evaluated through simulations and compared with the conventional GEE regression method for clustered outcomes. To illustrate our method, we use data from the treatment arm of a glaucoma clinical trial to compare the effectiveness of two commonly used ocular hypertension medications.
347 Background: To inform the design of trials of adjuvant radiation (RT) for bladder cancer, a local failure (LF) risk grouping has been proposed and externally validated that stratifies radical cystectomy (RC) pts into 3 groups based on pathologic factors. This stratification was developed using historical surgical databases and may not reflect outcomes in the observation arm of a modern trial. The purpose of the study is to assess whether trial accrual bias or improving surgical techniques over time impact the validity of the stratification or reduce the LF risk estimates for each subgroup. Methods: The LF stratification was developed using 2 cohorts treated with RC +/- chemo: a single-institution cohort of 442 pts (1990-2008) and the multi-center SWOG 8710 cohort of 264 pts (1987-1998). To assess the impact of trial accrual bias, we excluded pts who developed LF, DM, died, or were lost to follow up <90 days from surgery or completing post-op chemo as these pts are very unlikely to be enrolled in an adjuvant RT trial. 3-yr LF rates were estimated using Gray’s test. The stratification was considered valid if all 3 risk groups had significantly different LF rates. To assess the impact of improving surgical techniques over time, a Fine-Gray regression estimated the association of LF and year of RC while controlling for risk group. Results: Using the stricter criteria, 10% of SWOG pts and 14% from the other cohort were excluded. Analysis of the remaining pts confirmed 3 subgroups with significantly different LF risk: low risk (stage ≤pT2), intermediate risk (≥pT3 with negative margins and ≥10 nodes identified at surgery), and high risk (≥pT3 with positive margins OR <10 nodes identified) with 3-yr LF rates of 7%, 17%, and 36%, respectively (p<0.01), nearly identical to the rates when not accounting for accrual bias. Year of RC was not associated with LF risk on univariate analysis or after controlling for risk group in the SWOG, single-institution, or combined cohorts. Conclusions: These results suggest that neither trial accrual bias nor improvements in surgical techniques over time invalidate the LF risk stratification or substantially affect LF estimates. The proposed stratification is promising for use in adjuvant RT trials.
Tacrolimus (TAC) remains the backbone of graft-versus-host disease (GVHD) prophylaxis in allogeneic hematopoietic stem cell transplantation (HSCT). There is significant variability in the serum concentrations of TAC attained early post-transplant due to drug interactions and genomic variation. Whether a relationship exists between early TAC concentrations and acute GVHD in reduced-intensity conditioning (RIC) HSCT remains poorly characterized. We aimed to retrospectively evaluate whether mean weekly TAC concentrations correlate with incidence of acute GVHD in 120 consecutive patients (pts) undergoing first allogeneic HSCT at the University of Pennsylvania between January 2009 and January 2014. All pts received a uniform RIC regimen of fludarabine (120 mg/m2) and busulfan (6.4 mg/kg), followed by infusion of peripheral-blood stem cells from either a related (n=53) or unrelated (n=67) donor. All pts received standard GVHD prophylaxis with oral TAC (0.06 mg/kg/day) beginning day -3 and intravenous methotrexate (15 mg/m2 day +1, 10 mg/m2 days +3, +6, +11). The dose of TAC was adjusted to a target trough of 5–15 ng/mL. The primary endpoint was the incidence of acute grade 2–4 GVHD. We analyzed mean weekly TAC concentrations as continuous variables up to 4 weeks post-HSCT and then classified pts in tertiles (e.g., week 1: ≤8.5, 8.5–12;>12 ng/mL). Univariate analyses were performed using cumulative incidence and Cox regression models. Multivariate models were constructed using the backward elimination method. The mean TAC concentrations (ranges) at 1, 2, 3, and 4 weeks after RIC HSCT were 10.2 (2.8–19.4), 10.6 (3.9–23.9), 12.7 (4.2–24.1), and 11.9 (2.9–29.2) ng/mL, respectively. The day-180 cumulative incidence of acute grade 2–4 GVHD was 42.5%. In multivariate analysis, week 1 TAC concentration was an independent predictor of acute grade 2–4 GVHD (HR 0.91; 0.84–0.97; P=.008). Interestingly, this association was driven by a lower risk of acute grade 2–4 GVHD in pts with week 1 TAC concentrations in the upper tertile (>12 ng/mL) (HR 0.47; 0.25–0.88; P=.02), as shown in Figure 1. Week 1 TAC concentrations were not predictive of relapse, chronic GVHD, overall survival or non-relapse mortality. In this cohort, 9 pts (7.5%) developed acute kidney injury (AKI) by day 14. Week 1 TAC concentrations were higher in pts who developed AKI compared to those who did not (13.0 vs. 10.0 ng/mL; P=.01). We did not observe an association between TAC concentrations at weeks 2, 3, and 4 and clinical outcomes. Higher TAC concentrations during the first week after RIC HSCT were associated with significantly reduced risk of acute grade 2–4 GVHD without increasing risk of relapse. These data highlight the importance of optimizing initial dosing of TAC in RIC HSCT recipients.
BACKGROUNDClinical trials of radiation after radical cystectomy (RC) and chemotherapy for bladder cancer are in development, but inclusion and stratification factors have not been clearly established. In this study, the authors evaluated and refined a published risk stratification for locoregional failure (LF) by applying it to a multicenter patient cohort.METHODSThe original stratification, which was developed using a single‐institution series, produced 3 subgroups with significantly different LF risk based on pathologic tumor (pT) classification and the number of lymph nodes identified. This model was then applied to patients in Southwest Oncology Group (SWOG) 8710, a randomized trial of RC with or without chemotherapy. LF was defined as any pelvic failure before or within 3 months of distant failure.RESULTSPatients in the development cohort and the SWOG cohort had significantly different baseline characteristics. The original risk model was not fully validated in the SWOG cohort, because lymph node yield was not as strongly associated with LF as in the development cohort. Regression analysis indicated that margin status could improve the model. A revised stratification using pT classification, margin status, and the number of lymph nodes identified produced 3 subgroups with significantly different LF risk in both cohorts: low risk (≤pT2), intermediate risk (≥pT3 with negative margins AND ≥10 lymph nodes identified), and high risk (≥pT3 with positive margins OR <10 lymph nodes identified) with 5‐year LF rates of 8%, 20%, and 41%, respectively, in the SWOG cohort and 8%, 19%, and 41%, respectively, in the development cohort.CONCLUSIONSA model incorporating pT classification, margin status, and the number of lymph nodes identified stratified LF risk in 2 different RC populations and may inform the design of future trials. Cancer 2014;120:1272–1280. © 2014 American Cancer Society.
A longitudinal mixture model for classifying patients into responders and non‐responders is established using both likelihood‐based and Bayesian approaches. The model takes into consideration responders in the control group. Therefore, it is especially useful in situations where the placebo response is strong, or in equivalence trials where the drug in development is compared with a standard treatment. Under our model, a treatment shows evidence of being effective if it increases the proportion of responders or increases the response rate among responders in the treated group compared with the control group. Therefore, the model has flexibility to accommodate different situations. The proposed method is illustrated using simulation and a depression clinical trial dataset for the likelihood‐based approach, and the same depression clinical trial dataset for the Bayesian approach. The likelihood‐based and Bayesian approaches generated consistent results for the depression trial data. In both the placebo group and the treated group, patients are classified into two components with distinct response rate. The proportion of responders is shown to be significantly higher in the treated group compared with the control group, suggesting the treatment paroxetine is effective. Copyright © 2014 John Wiley & Sons, Ltd.
297 Background: Trials of adjuvant radiation (RT) after radical cystectomy + pelvic node dissection (RC) are in development, but patient selection criteria have not been clearly identified. We evaluated a published model predicting the risk of local-regional failure (LF) by applying the model to a multi-center patient cohort. We also assessed the model’s ability to stratify overall survival (OS) and isolated distant metastases (DM). Methods: The original LF risk model was derived from 442 patients who had a RC +/- chemotherapy at the University of Pennsylvania (PENN). Analysis identified 3 subgroups with significantly different LF risk based on pathologic stage (pT) and number of nodes dissected (<10 vs ≥10). This risk rubric was then applied to 264 patients in the multi-center SWOG 8710 trial randomized to RC +/- neoadjuvant chemotherapy. SWOG patients differed significantly from PENN patients in mean age, pT stage, number of nodes removed, surgical margin status, and use of neoadjuvant chemotherapy. In both cohorts, LF was any pelvic failure detected before or within 3 months of DM. Secondary endpoints were OS and DM. Competing risk analysis was used. Results: The original risk stratification was not fully validated in the SWOG cohort. Regression analysis of the SWOG data found margin status was more predictive of LF than number of nodes dissected. A revised risk stratification model combining pT stage, margin status, and number of nodes removed identified 3 subgroups in both the SWOG and PENN cohorts with significantly different LF risks: low (≤pT2), intermediate (≥pT3 with negative margins AND ≥10 nodes removed), and high (≥pT3 with positive margins OR <10 nodes removed) with 5 year LF rates of 8%, 20%, and 41% in the SWOG group and 8%, 19%, and 41% in the PENN cohort. The 3 subgroups also differed significantly in OS within both cohorts with 5 year OS of 62%, 39%, and 7% for SWOG and 60%, 31%, and 10% for PENN. The model did not stratify with respect to isolated DM. Conclusions: The revised risk model combining pT stage, margin status, and number of nodes excised provides a simple rubric to stratify LF risk in two significantly different bladder cancer populations and may inform the design of adjuvant RT trials.
293 Background: Local-regional recurrences (LF) after radical cystectomy with or without chemotherapy are common in patients with locally advanced disease. Adjuvant radiation (RT) could reduce LF, but toxicity discouraged its use. Modern RT with reduced morbidity has rekindled interest but requires knowledge of pelvic failure patterns to design appropriate clinical target volumes. Methods: 5-yr LF rates after radical cystectomy plus pelvic lymph node dissection with or without chemotherapy were determined for 8 pelvic sites among 442 patients with urothelial carcinoma of the bladder. The impact on the pattern of failure of pathologic stage, margin status, nodal involvement, and extent of node dissection was assessed using competing risk statistical methods. The percentage of patients whose sites of LF would be completely encompassed within various hypothetical clinical target volumes for post-operative radiation were calculated. Results: Stage pT3-4 patients had higher 5-yr LF rates in virtually all pelvic sites compared to pT0-2 patients. Among pT3-4 patients, margin status significantly altered the pattern of failure while extent of node dissection and pathologic nodal involvement did not. Stage pT3-4 patients with negative margins failed predominantly in the iliac/obturator nodes. Failures in the cystectomy bed and presacral region were significantly higher in pT3-4 patients with positive rather than negative margins. 76% of pT3-4 patients with negative margins who failed would have had all sites of LF included within clinical target volumes encompassing the iliac/obturator nodes, but only 57% of pT3-4 patients with positive margins would have their LF sites covered by such target volumes. Including the cystectomy bed and presacral region in the clinical target volume when margins were positive increased the percentage of encompassed failures to 91%. Conclusions: In adjuvant RT protocols, the obturator and iliac regions should be targeted in pT3-4 tumors with negative margins; coverage of presacral region and cystectomy bed is advised for pT3-4 with positive margins.
Trials of adjuvant radiation (RT) after radical cystectomy + pelvic lymph node dissection (RC) are currently being developed, but inclusion and stratification criteria for such trials have not been clearly identified. We evaluated and refined a recently published risk stratification model for local-regional failure (LF) by applying the model to a multi-center patient cohort. We also evaluated the model's ability to stratify overall survival (OS) and isolated distant metastases (DM). The original stratification was derived from a single institution cohort of 442 patients treated between 1990 and 2008 with RC +/- chemotherapy at the University of Pennsylvania (PENN). Analysis identified 3 patient subgroups with significantly different LF risk based on pathologic stage (pT) and number of nodes dissected. This risk rubric was then applied to 264 patients randomized to RC +/- neoadjuvant chemotherapy in the multi-center SWOG 8710 trial. In both cohorts, LF was scored as any pelvic failure detected before or within 3 months of DM. Secondary endpoints were OS and DM. Competing risk analysis was used. SWOG patients differed significantly from PENN patients in mean age, pT stage, number of nodes removed, surgical margin status, and use of neoadjuvant chemotherapy (p < 0.01 for all comparisons). The original risk stratification was not fully validated in the SWOG cohort as the difference in LF between ≥pT3 patients who had ≥10 versus < 10 nodes removed was smaller for SWOG than PENN patients. Regression analysis of the SWOG data suggested margin status was a more important stratifying variable than number of nodes dissected. A revised risk stratification using pT stage, margin status, and number of nodes removed identified 3 subgroups with significantly different LF risk in both SWOG and PENN cohorts: low (≤pT2), intermediate (≥pT3 with negative margins AND ≥10 nodes removed), and high (≥pT3 with positive margins OR < 10 nodes removed) with 5 year LF rates of 8%, 20%, and 41% in the SWOG group and 8%, 19%, and 41% in the PENN cohort. Gray's test p < 0.01 and C-index was 0.7 for the revised model for both cohorts. The 3 subgroups were well stratified with respect to OS in both cohorts with 5 yr OS estimates of 62%, 39%, and 7% for SWOG and 60%, 31%, and 10% for PENN (log-rank p < 0.01 for all pair-wise comparisons). In contrast, the model did not stratify patients with respect to isolated DM. The revised risk grouping combining pT stage, margin status and number of nodes excised stratified LF risk in two significantly different bladder cancer populations and may inform the design of adjuvant RT trials. External validation of this revised stratification rubric is warranted.
INTRODUCTION:The obstructive sleep apnea syndrome (OSAS) is associated with increased visceral adipose tissue (VAT) in adults; however, few studies have evaluated VAT in relation to upper airway function in adolescents. We hypothesized that increased neck circumference (NC) and VAT would be associated with increased upper airway collapsibility.METHODS:Adolescents (24 obese patients with OSAS, 22 obese control patients, and 29 lean control patients) underwent abdominal magnetic resonance imaging, and measurement of upper airway pressure-flow relationships in the activated and hypotonic upper airway states.RESULTS:Patients with OSAS had a greater activated slope of the pressure-flow relationship (SPF) than control groups (P < 0.001), whereas hypotonic SPF was greater in both obese groups compared with lean control patients (P = 0.01). NC and VAT were greater in obese control patients and those with OSAS than in lean control patients (P < 0.001), but did not differ between obese patients with OSAS and obese control patients. In lean control patients and those with OSAS, increased NC was associated with increased activated SPF, whereas in obese control patients it was associated with decreased activated SPF (P = 0.03). In contrast, increased NC was associated with increased hypotonic SPF in all groups (P < 0.001). There was no significant effect of VAT on either activated or hypotonic SPF for any of the three groups.CONCLUSIONS:Increased neck circumference was associated with increased upper airway collapsibility in adolescents in the hypotonic but not activated state. These data suggest that obese adolescents without OSAS, despite a narrowed upper airway from adipose tissue, are protected from developing OSAS by upper airway neuromotor activation. Neither neck circumference nor visceral adipose tissue is useful in predicting upper airway collapsibility in obese adolescents.