The Second Generation P-Value (SGPV) measures the overlap between an estimated interval and a composite hypothesis of parameter values. We develop a sequential monitoring scheme of the SGPV (SeqSGPV) to connect study design intentions with end-of-study inference anchored on scientific relevance. We build upon Freedman's "Region of Equivalence" (ROE) in specifying scientifically meaningful hypotheses called Pre-specified Regions Indicating Scientific Merit (PRISM). We compare PRISM monitoring versus monitoring alternative ROE specifications. Error rates are controlled through the PRISM's indifference zone around the point null and monitoring frequency strategies. Because the former is fixed due to scientific relevance, the latter is a targettable means for designing studies with desirable operating characters. An affirmation step to stopping rules improves frequency properties including the error rate, the risk of reversing conclusions under delayed outcomes, and bias.
Hiring, tenure and promotion processes are powerful influences in driving innovation in research practices and in furthering adoption of open data and software. This ‘Ten simple rules’ article provides practical steps to update institutional processes to incorporate data and software, and to recognize those outputs on their own merit.
BACKGROUND: Black Americans receive a diagnosis at later stage of lung cancer more often than White Americans. We undertook a population-based study to identify factors contributing to racial disparities in lung cancer stage of diagnosis among low-income adults.RESEARCH QUESTION: Which multilevel factors contribute to racial disparities in stage of lung cancer at diagnosis?STUDY DESIGN AND METHODS: Cases of incident lung cancer from the prospective observa-tional Southern Community Cohort Study were identified by linkage with state cancer reg-istries in 12 southeastern states. Logistic regression shrinkage techniques were implemented to identify individual-level and area-level factors associated with distant stage diagnosis. A subset of participants who responded to psychosocial questions (eg, racial discrimination experiences) were evaluated to determine if model predictive power improved.RESULTS: We identified 1,572 patients with incident lung cancer with available lung cancer stage (64% self-identified as Black and 36% self-identified as White). Overall, Black participants with lung cancer showed greater unadjusted odds of distant stage diagnosis compared with White participants (OR,1.29; 95% CI, 1.05-1.59). Greater neighborhood area deprivation was associated with distant stage diagnosis (OR, 1.58; 95% CI, 1.19-2.11). After controlling for individual-and area-level factors, no significant difference were found in distant stage disease for Black vs White participants. However, participants with COPD showed lower odds of distant stage diagnosis in the primary model (OR, 0.72; 95% CI, 0.53-0.98). Interesting and complex interactions were observed. The subset analysis model with additional variables for racial discrimination experi-ences showed slightly greater predictive power than the primary model. INTERPRETATION: Reducing racial disparities in lung cancer stage at presentation will require interventions on both structural and individual-level factors.CHEST 2023; 163(5):1314-1327
BACKGROUND: Appropriate risk stratification of indeterminate pulmonary nodules (IPNs) is necessary to direct diagnostic evaluation. Currently available models were developed in populations with lower cancer prevalence than that seen in thoracic surgery and pulmonology clinics and usually do not allow for missing data. We updated and expanded the Thoracic Research Evaluation and Treatment (TREAT) model into a more generalized, robust approach for lung cancer prediction in patients referred for specialty evaluation.RESEARCH QUESTION: Can clinic-level differences in nodule evaluation be incorporated to improve lung cancer prediction accuracy in patients seeking immediate specialty evaluation compared with currently available models?STUDY DESIGN AND METHODS: Clinical and radiographic data on patients with IPNs from six sites (N = 1,401) were collected retrospectively and divided into groups by clinical setting: pulmonary nodule clinic (n = 374; cancer prevalence, 42%), outpatient thoracic surgery clinic (n = 553; cancer prevalence, 73%), or inpatient surgical resection (n = 474; cancer prevalence, 90%). A new prediction model was developed using a missing data-driven pattern submodel approach. Discrimination and calibration were estimated with cross-validation and were compared with the original TREAT, Mayo Clinic, Herder, and Brock models. Reclassification was assessed with biascorrected clinical net reclassification index and reclassification plots.RESULTS: Two-thirds of patients had missing data; nodule growth and fluorodeoxyglucose-PET scan avidity were missing most frequently. The TREAT version 2.0 mean area under the receiver operating characteristic curve across missingness patterns was 0.85 compared with that of the original TREAT (0.80), Herder (0.73), Mayo Clinic (0.72), and Brock (0.68) models with improved calibration. The bias-corrected clinical net reclassification index was 0.23. INTERPRETATION: The TREAT 2.0 model is more accurate and better calibrated for predicting lung cancer in high-risk IPNs than the Mayo, Herder, or Brock models. Nodule calculators such as TREAT 2.0 that account for varied lung cancer prevalence and that consider missing data may provide more accurate risk stratification for patients seeking evaluation at specialty nodule evaluation clinics. CHEST 2023; 164(5):1305-1314
Background Appropriate risk stratification of indeterminate pulmonary nodules (IPNs) is necessary to direct diagnostic evaluation. Currently available models were developed in populations with lower cancer prevalence than that seen in thoracic surgery and pulmonology clinics and usually do not allow for missing data. We updated and expanded the Thoracic Research Evaluation and Treatment (TREAT) model into a more generalized, robust approach for lung cancer prediction in patients referred for specialty evaluation. Research Question Can clinic-level differences in nodule evaluation be incorporated to improve lung cancer prediction accuracy in patients seeking immediate specialty evaluation compared with currently available models? Study Design and Methods Clinical and radiographic data on patients with IPNs from six sites (N = 1,401) were collected retrospectively and divided into groups by clinical setting: pulmonary nodule clinic (n = 374; cancer prevalence, 42%), outpatient thoracic surgery clinic (n = 553; cancer prevalence, 73%), or inpatient surgical resection (n = 474; cancer prevalence, 90%). A new prediction model was developed using a missing data-driven pattern submodel approach. Discrimination and calibration were estimated with cross-validation and were compared with the original TREAT, Mayo Clinic, Herder, and Brock models. Reclassification was assessed with bias-corrected clinical net reclassification index and reclassification plots. Results Two-thirds of patients had missing data; nodule growth and fluorodeoxyglucose-PET scan avidity were missing most frequently. The TREAT version 2.0 mean area under the receiver operating characteristic curve across missingness patterns was 0.85 compared with that of the original TREAT (0.80), Herder (0.73), Mayo Clinic (0.72), and Brock (0.68) models with improved calibration. The bias-corrected clinical net reclassification index was 0.23. Interpretation The TREAT 2.0 model is more accurate and better calibrated for predicting lung cancer in high-risk IPNs than the Mayo, Herder, or Brock models. Nodule calculators such as TREAT 2.0 that account for varied lung cancer prevalence and that consider missing data may provide more accurate risk stratification for patients seeking evaluation at specialty nodule evaluation clinics.
Introduction Lung cancer screening rates in the US remain low. Evidence suggests that risk-model based approaches may outperform current risk-factor based criteria in determining screening eligibility. We compared the USPSTF 2021 lung cancer screening eligibility criteria to three existing risk prediction models in a large and racially diverse population of smoking individuals. Methods We identified current and former smokers from a predominantly Black American study population recruited from community health centers across 12 southeastern US states from 2002 to 2009 and followed until 2019. Smoking behaviors were collected by in-person interviews or mailed-in questionnaires using three follow-up surveys. Incident lung cancers were ascertained from state cancer registries and National Death Index mortality records. Screening eligibility was based on USPSTF 2021 criteria or on pre-specified risk thresholds for each risk prediction model. Performance was assessed by sensitivity (the percentage of individuals who developed lung cancer that were deemed eligible for screening) and specificity (the percentage of those without lung cancer excluded from screening). Disparities by race and sex were evaluated in each subgroup. Results Of 52,911 ever smokers (64% Black, 31% White, 4% Other), a total of 1,705 developed lung cancer. More than a third of all persons with a history of smoking were screening eligible across all models: USPSTF 41%; Prostate, Lung, Colorectal and Ovarian Cancer (PLCO) m2012 39%; Lung Cancer Risk Assessment Tool (LCRAT) 43%; Lung Cancer Death Risk Assessment Tool (LCDRAT) 43%. All models exhibited nearly identical sensitivity (correctly identifying persons with lung cancer as eligible for screening) and specificity (excluding participants without lung cancer from screening): USPSTF Sensitivity: 0.61, Specificity: 0.59; PLCOm2012 Sensitivity: 0.62, Specificity: 0.60; LCDRAT Sensitivity: 0.62, Specificity: 0.58; LCDRAT Sensitivity: 0.60, Specificity: 0.58. Compared to Black persons with lung cancer, White persons with lung cancer were more likely to be eligible for screening across all models. Only 54-57% of Black persons with lung cancer were eligible for screening compared to 70-73% of White persons with lung cancer. The smallest racial disparities were observed using the USPSTF criteria (USPSTF sensitivity ratio for Whites to Blacks (SR [95% CI]): 1.27 [1.17-1.36]; PLCOm2012: 1.28 [1.17-1.39]; LCRAT: 1.29 [1.18-1.40]; LCDRAT: 1.33 [1.22-1.45]). There was no evidence of sex disparities greater than 1.1 RR in any of the models, with most estimated effects being less than 1+/- 0.05 RR. Conclusion In a racially diverse population of smoking individuals, we found similar performance of risk models vs USPFTF 2021 criteria for lung cancer screening. Racial disparities in screening eligibility persisted across all models, with the USPSTF 2021 criteria resulting in the smallest disparity. Given the poor performance of these models, there remains significant room for improvement. Citation Format: Adoma A. Manful, Megan H. Murray, Sarah F. Mercaldo, Jeffrey D. Blume, Melinda C. Aldrich. Are we there yet? Performance of risk-model based eligibility for lung cancer screening [abstract]. In: Proceedings of the 15th AACR Conference on the Science of Cancer Health Disparities in Racial/Ethnic Minorities and the Medically Underserved; 2022 Sep 16-19; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Epidemiol Biomarkers Prev 2022;31(1 Suppl):Abstract nr A106.
OBJECTIVE We present an illustrative application of methods that account for covariates in receiver operating characteristic (ROC) curve analysis, using individual patient data on D-dimer testing for excluding pulmonary embolism. STUDY DESIGN AND SETTING Bayesian nonparametric covariate-specific ROC curves were constructed to examine the performance/positivity thresholds in covariate subgroups. Standard ROC curves were constructed. Three scenarios were outlined based on comparison between subgroups and standard ROC curve conclusion: (1) identical distribution/identical performance, (2) different distribution/identical performance, and (3) different distribution/different performance. Scenarios were illustrated using clinical covariates. Covariate-adjusted ROC curves were also constructed. RESULTS Age groups had prominent differences in D-dimer concentration, paired with differences in performance (Scenario 3). Different positivity thresholds were required to achieve the same level of sensitivity. D-dimer had identical performance, but different distributions for YEARS algorithm items (Scenario 2), and similar distributions for sex (Scenario 1). For the later covariates, comparable positivity thresholds achieved the same sensitivity. All covariate-adjusted models had AUCs comparable to the standard approach. CONCLUSION Subgroup differences in performance and distribution of results can indicate that the conventional ROC curve is not a fair representation of test performance. Estimating conditional ROC curves can improve the ability to select thresholds with greater applicability.
Abstract Purpose: Black Americans experience poorer lung cancer survival than White Americans. Racial disparities in stage at diagnosis may contribute to these survival differences, but few studies have explored factors leading to racial disparities in lung cancer stage at diagnosis. We aimed to identify multilevel factors contributing to racial disparities in stage of lung cancer presentation. Methods: Using data from the Southern Community Cohort Study (SCCS), we examined factors associated with distant stage among adults diagnosed with incident lung cancer. SCCS participants were prospectively enrolled, primarily from community health centers between 2002 and 2009 across a 12-state area. Incident cancers were identified by linkage with state cancer registries through end of follow-up in 2019. Self-reported social, behavioral, and medical history information were ascertained at baseline via questionnaire. Cumulative exposure smoking histories were identified using the most recent follow-up questionnaires. Residential addresses and National Cancer Institute Comprehensive Cancer Center locations were geocoded, and residential addresses were linked to census data. Logistic and multinomial regression models were used to identify factors predictive of distant stage diagnosis. Penalized regression was used to shrink the predictor space of these models when necessary. Findings were replicated in an independent population. Results: Among 1,672 incident SCCS lung cancer cases (35% White, 61% Black, and 3% other self-reported race), a greater percentage of Black participants than White participants were diagnosed with distant stage lung cancer (56.4% vs 49.4%, respectively). Overall, Black participants had greater odds of distant vs local stage compared to White participants (odds ratio (OR) = 1.28, 95% confidence interval (CI): 1.05-1.58). Greater area deprivation was also associated with distant lung cancer stage (OR = 1.55, 95% CI: 1.17-2.04). After controlling for individual and area-level factors, there was no significant difference in the odds of distant stage disease for Black participants compared to White participants (OR = 1.03, 95% CI: 0.80-1.33). Significant interactions between race and area deprivation index were not observed. However, greater residential distance from a comprehensive cancer center was significantly associated with increased odds of distant stage disease in the final model (OR = 1.04, 95% CI: 1.00-1.08). No significant differences were observed in the odds of distant stage lung cancer among Black and White participants in the independent population. Conclusions: A greater percentage of Black participants were diagnosed with distant stage lung cancer; however, this disparity dissipated after adjusting for individual and area-level factors. Our findings suggest racial disparities in lung cancer stage at diagnosis may be ameliorated with modifiable factors, such as patient access to high quality cancer centers. Citation Format: Jennifer Richmond, Megan Hollister, Cato M. Milder, Ann G. Schwartz, Jeffrey D. Blume, Melinda C. Aldrich. Examining racial disparities in lung cancer stage of diagnosis among low-income adults living in the southeastern U.S. [abstract]. In: Proceedings of the AACR Virtual Conference: 14th AACR Conference on the Science of Cancer Health Disparities in Racial/Ethnic Minorities and the Medically Underserved; 2021 Oct 6-8. Philadelphia (PA): AACR; Cancer Epidemiol Biomarkers Prev 2022;31(1 Suppl):Abstract nr PO-236.
We introduce the ProSGPV R package, which implements a variable selection algorithm based on second-generation p-values (SGPV) instead of traditional p-values. Most variable selection algorithms shrink point estimates to arrive at a sparse solution. In contrast, the ProSGPV algorithm accounts for the estimation uncertainty – via confidence intervals – in the selection process. This additional information leads to better inference and prediction performance in finite sample sizes. ProSGPV maintains good performance even in the high dimensional case where $p>n$, or when explanatory variables are highly correlated. Moreover, ProSGPV is a unifying algorithm that works with continuous, binary, count, and time-to-event outcomes. No cross-validation or iterative processes are needed and thus ProSGPV is very fast to compute. Visualization tools are available in this package for assessing the variable selection process. Here we present simulation studies and a real-world example to demonstrate ProSGPV’s inference and prediction performance in relation to the current standards in variable selection procedures.
Many statistical methods have been proposed for variable selection in the past century, but few balance inference and prediction tasks well. Here, we report on a novel variable selection approach called penalized regression with second-generation p-values (ProSGPV). It captures the true model at the best rate achieved by current standards, is easy to implement in practice, and often yields the smallest parameter estimation error. The idea is to use an l(0) penalization scheme with second-generation p-values (SGPV), instead of traditional ones, to determine which variables remain in a model. The approach yields tangible advantages for balancing support recovery, parameter estimation, and prediction tasks. The ProSGPV algorithm can maintain its good performance even when there is strong collinearity among features or when a high-dimensional feature space with p > n is considered. We present extensive simulations and a real-world application comparing the ProSGPV approach with smoothly clipped absolute deviation (SCAD), adaptive lasso (AL), and minimax concave penalty with penalized linear unbiased selection (MC+). While the last three algorithms are among the current standards for variable selection, ProSGPV has superior inference performance and comparable prediction performance in certain scenarios.
False discovery rates (FDR) are an essential component of statistical inference, representing the propensity for an observed result to be mistaken. FDR estimates should accompany observed results to help the user contextualize the relevance and potential impact of findings. This paper introduces a new user-friendly R pack-age for estimating FDRs and computing adjusted p-values for FDR control. The roles of these two quantities are often confused in practice and some software packages even report the adjusted p-values as the estimated FDRs. A key contribution of this package is that it distinguishes between these two quantities while also offering a broad array of refined algorithms for estimating them. For example, included are newly augmented methods for estimating the null proportion of findings - an important part of the FDR estimation procedure. The package is broad, encompassing a variety of adjustment methods for FDR estimation and FDR control, and includes plotting functions for easy display of results. Through extensive illustrations, we strongly encourage wider reporting of false discovery rates for observed findings.
Variable selection has become a pivotal choice in data analyses that impacts subsequent inference and prediction. In linear models, variable selection using Second-Generation P-Values (SGPV) has been shown to be as good as any other algorithm available to researchers. Here we extend the idea of Penalized Regression with Second-Generation P-Values (ProSGPV) to the generalized linear model (GLM) and Cox regression settings. The proposed ProSGPV extension is largely free of tuning parameters, adaptable to various regularization schemes and null bound specifications, and is computationally fast. Like in the linear case, it excels in support recovery and parameter estimation while maintaining strong prediction performance. The algorithm also preforms as well as its competitors in the high dimensional setting (n>p). Slight modifications of the algorithm improve its performance when data are highly correlated or when signals are dense. This work significantly strengthens the case for the ProSGPV approach to variable selection.
Abstract Background Recent trials have suggested use of balanced crystalloids may decrease the incidence of major adverse kidney events compared to saline in critically ill adults. The effect of crystalloid composition on biomarkers of early acute kidney injury remains unknown. Methods From February 15 to July 15, 2016, we conducted an ancillary study to the Isotonic Solutions and Major Adverse Renal Events Trial (SMART) comparing the effect of balanced crystalloids versus saline on urinary levels of neutrophil gelatinase-associated lipocalin (NGAL) and kidney injury molecule-1 (KIM-1) among 261 consecutively-enrolled critically ill adults admitted from the emergency department to the medical ICU. After informed consent, we collected urine 36 ± 12 h after hospital admission and measured NGAL and KIM-1 levels using commercially available ELISAs. Levels of NGAL and KIM-1 at 36 ± 12 h were compared between patients assigned to balanced crystalloids versus saline using a Mann-Whitney U test. Results The 131 patients (50.2%) assigned to the balanced crystalloid group and the 130 patients (49.8%) assigned to the saline group were similar at baseline. Urinary NGAL levels were significantly lower in the balanced crystalloid group (median, 39.4 ng/mg [IQR 9.9 to 133.2]) compared with the saline group (median, 64.4 ng/mg [IQR 27.6 to 339.9]) (P < 0.001). Urinary KIM-1 levels did not significantly differ between the balanced crystalloid group (median, 2.7 ng/mg [IQR 1.5 to 4.9]) and the saline group (median, 2.4 ng/mg [IQR 1.3 to 5.0]) (P = 0.36). Conclusions In this ancillary analysis of a clinical trial comparing balanced crystalloids to saline among critically ill adults, balanced crystalloids were associated with lower urinary concentrations of NGAL and similar urinary concentrations of KIM-1, compared with saline. These results suggest only a modest reduction in early biomarkers of acute kidney injury with use of balanced crystalloids compared with saline. Trial registration ClinicalTrials.gov number: NCT02444988 . Date registered: May 15, 2015.
False discovery rates (FDR) are an essential component of statistical inference, representing the propensity for an observed result to be mistaken. FDR estimates should accompany observed results to help the user contextualize the relevance and potential impact of findings. This paper introduces a new user-friendly R pack-age for estimating FDRs and computing adjusted p-values for FDR control. The roles of these two quantities are often confused in practice and some software packages even report the adjusted p-values as the estimated FDRs. A key contribution of this package is that it distinguishes between these two quantities while also offering a broad array of refined algorithms for estimating them. For example, included are newly augmented methods for estimating the null proportion of findings - an important part of the FDR estimation procedure. The package is broad, encompassing a variety of adjustment methods for FDR estimation and FDR control, and includes plotting functions for easy display of results. Through extensive illustrations, we strongly encourage wider reporting of false discovery rates for observed findings.
BACKGROUND:The proliferation of genetic profiling has revealed many associations between genetic variations and disease. However, large-scale phenotyping efforts in largely healthy populations, coupled with DNA sequencing, suggest variants currently annotated as pathogenic are more common in healthy populations than previously thought. In addition, novel and rare variants are frequently observed in genes associated with disease both in healthy individuals and those under suspicion of disease. This raises the question of whether these variants can be useful predictors of disease. To answer this question, we assessed the degree to which the presence of a variant in the cardiac potassium channel gene KCNH2 was diagnostically predictive for the autosomal dominant long QT syndrome. METHODS:We estimated the probability of a long QT diagnosis given the presence of each KCNH2 variant using Bayesian methods that incorporated variant features such as changes in variant function, protein structure, and in silico predictions. We call this estimate the posttest probability of disease. Our method was applied to over 4000 individuals heterozygous for 871 missense or in-frame insertion/deletion variants in KCNH2 and validated against a separate international cohort of 933 individuals heterozygous for 266 missense or in-frame insertion/deletion variants. RESULTS:Our method was well-calibrated for the observed fraction of heterozygotes diagnosed with long QT syndrome. Heuristically, we found that the innate diagnostic information one learns about a variant from 3-dimensional variant location, in vitro functional data, and in silico predictors is equivalent to the diagnostic information one learns about that same variant by clinically phenotyping 10 heterozygotes. Most importantly, these data can be obtained in the absence of any clinical observations. CONCLUSIONS:We show how variant-specific features can inform a prior probability of disease for rare variants even in the absence of clinically phenotyped heterozygotes.
A major challenge emerging in genomic medicine is how to assess best disease risk from rare or novel variants found in disease-related genes. The expanding volume of data generated by very large phenotyping efforts coupled to DNA sequence data presents an opportunity to reinterpret genetic liability of disease risk. Here we propose a framework to estimate the probability of disease given the presence of a genetic variant conditioned on features of that variant. We refer to this as the penetrance, the fraction of all variant heterozygotes that will present with disease. We demonstrate this methodology using a well-established disease-gene pair, the cardiac sodium channel gene SCN5A and the heart arrhythmia Brugada syndrome. From a review of 756 publications, we developed a pattern mixture algorithm, based on a Bayesian Beta-Binomial model, to generate SCN5A penetrance probabilities for the Brugada syndrome conditioned on variant-specific attributes. These probabilities are determined from variant-specific features (e.g. function, structural context, and sequence conservation) and from observations of affected and unaffected heterozygotes. Variant functional perturbation and structural context prove most predictive of Brugada syndrome penetrance.
Effect size indices are useful tools in study design and reporting because they are unitless measures of association strength that do not depend on sample size. Existing effect size indices are developed for particular parametric models or population parameters. Here, we propose a robust effect size index based on M-estimators. This approach yields an index that is very generalizable because it is unitless across a wide range of models. We demonstrate that the new index is a function of Cohen’s d, $$R^2$$, and standardized log odds ratio when each of the parametric models is correctly specified. We show that existing effect size estimators are biased when the parametric models are incorrect (e.g., under unknown heteroskedasticity). We provide simple formulas to compute power and sample size and use simulations to assess the bias and standard error of the effect size estimator in finite samples. Because the new index is invariant across models, it has the potential to make communication and comprehension of effect size uniform across the behavioral sciences.
False discovery rates (FDR) are an essential component of statistical inference, representing the propensity for an observed result to be mistaken. FDR estimates should accompany observed results to help the user contextualize the relevance and potential impact of findings. This paper introduces a new user-friendly R package for computing FDRs and adjusting p-values for FDR control. These tools respect the critical difference between the adjusted p-value and the estimated FDR for a particular finding, which are sometimes numerically identical but are often confused in practice. Newly augmented methods for estimating the null proportion of findings - an important part of the FDR estimation procedure - are proposed and evaluated. The package is broad, encompassing a variety of methods for FDR estimation and FDR control, and includes plotting functions for easy display of results. Through extensive illustrations, we strongly encourage wider reporting of false discovery rates for observed findings.