Colorectal cancer (CRC) screening uptake in the Veterans Health Administration (VA) has been reported to be higher than the US general population, but CRC remains a prevalent cancer within the VA system. To examine CRC predictors and the extent to which the conventional definition of up-to-date screening applies to the population, we conducted a case-control study using VA data from 2012 to 2018. We classified patients into 5 categories: up-to-date or not up-to-date average-risk patients aged 50 to 75 (Categories 1 and 2), up-to-date or not up-to-date average-risk patients aged <50 or >75 (Categories 3 and 4), and high-risk patients (Category 5). Each CRC case was matched by age, sex, and facility with 4 controls. We performed multivariable conditional logistic regression, adjusting for race and ethnicity, diabetes, obesity, and alcohol use. Among 3714 CRC cases identified, Category 4 (odds ratio [OR] 1.40, 95% CI 1.11-1.78) and Category 5 (OR 6.23, 95% CI 5.06-7.66) patients had a higher risk of CRC compared to Category 1 patients. Compared with White patients, Black patients had a higher risk (OR 1.54, 95% CI 1.37-1.73). Diabetes (OR 1.65, 95% CI 1.51-1.81) and alcohol use disorder (OR 1.53, 95% CI 1.35-1.73) were also associated with CRC. Most CRC cases occurred in individuals aged 50 to 75, but 12.5% occurred in persons who were outside of this age range or had high-risk personal or family history. The conventional measure of CRC screening, focused on average-risk individuals aged 50 to 75, does not reflect screening status in an important minority of CRC patients.
This article considers developing false discovery rate (FDR) controlling methods for test-ing multiple hypotheses under three different classification settings of the hypotheses into groups - simultaneous multi-way classification, hierarchical classification, and a combination of these two classifications. The methods are developed in their oracle forms by considering a weighted version of the Benjamini-Hochberg (BH, 1995) method, with the weights encoding the underlying structural information about the hypotheses. They control the FDR when the p-values involved are Positively Regression Dependent on the Subset (PRDS) of null p-values, and are more powerful than the BH procedure. Data-adaptive versions of these methods are also proposed by appropriately estimating the weights under the different types of classification. The proposed data-adaptive methods control the FDR at the desired level when the p-values are independent and, as simulations show, they can be more powerful FDR controlling methods, even under certain dependency, than some existing comparable multiple testing methods. We apply the data-adaptive method under the above-mentioned combined classification setting to analyze a publicly available neuro-imaging dataset. Such data typically have complex classification structures that were not addressed in previously available multiple testing methods, as far as we know. Our proposed method, with a flexible weighting scheme, is poised to utilize more information from the data in its decision making process, than other existing multiple testing methods.& COPY; 2023 Elsevier B.V. All rights reserved.
In data collection for predictive modeling, under-representation of certain groups, based on gender, race/ethnicity, or age, may yield less-accurate predictions for these groups. Recently, this issue of fairness in predictions has attracted significant attention, as data-driven models are increasingly utilized to perform crucial decision-making tasks. Existing methods to achieve fairness in the machine learning literature typically build a single prediction model in a manner that encourages fair prediction performance for all groups. These approaches have two major limitations: i) fairness is often achieved by compromising accuracy for some groups; ii) the underlying relationship between dependent and independent variables may not be the same across groups. We propose a Joint Fairness Model (JFM) approach for logistic regression models for binary outcomes that estimates group-specific classifiers using a joint modeling objective function that incorporates fairness criteria for prediction. We introduce an Accelerated Smoothing Proximal Gradient Algorithm to solve the convex objective function, and present the key asymptotic properties of the JFM estimates. Through simulations, we demonstrate the efficacy of the JFM in achieving good prediction performance and across-group parity, in comparison with the single fairness model, group-separate model, and group-ignorant model, especially when the minority group's sample size is small. Finally, we demonstrate the utility of the JFM method in a real-world example to obtain fair risk predictions for under-represented older patients diagnosed with coronavirus disease 2019 (COVID-19).
In this article, we propose a generalized weighted version of the well-known Benjamini-Hochberg (BH) procedure. The rigorous weighting scheme used by our method enables it to encode structural information from simultaneous multi-way classification as well as hierarchical partitioning of hypotheses into groups, with provisions to accommodate overlapping groups. The method is proven to control the False Discovery Rate (FDR) when the p-values involved are Positively Regression Dependent on the Subset (PRDS) of null p-values. A data-adaptive version of the method is proposed. Simulations show that our proposed methods control FDR at desired level and are more powerful than existing comparable multiple testing procedures, when the p-values are independent or satisfy certain dependence conditions. We apply this data-adaptive method to analyze a neuro-imaging dataset and understand the impact of alcoholism on human brain. Neuro-imaging data typically have complex classification structure, which have not been fully utilized in subsequent inference by previously proposed multiple testing procedures. With a flexible weighting scheme, our method is poised to extract more information from the data and use it to perform a more informed and efficient test of the hypotheses.
There is ample research on false discovery rate (FDR) control for testing hypotheses classified according to one criterion. However, scenarios of hypotheses partitioned via two different criteria are often encountered in practice. Such two-way classification encodes more structural information in the associated multiple testing of the hypotheses than its one-way or un-classified counterparts. Unfortunately, there seems to be very little research tailored for multiple testing under that classification setting. This article proposes weighted versions of the Benjamini–Hochberg (BH) method, both in their oracle and data-adaptive forms, efficiently capturing one- or two-way classified structure of hypotheses through appropriately chosen weights. The proposed methods control FDR non-asymptotically in their oracle forms under positive regression dependence on subset (PRDS) of null p-values and in their data-adaptive forms for independent p-values. The one-way data-adaptive methods are asymptotically conservative under weak dependence. Simulations demonstrate these methods’ superior power performances over some contemporary procedures and provide evidence of their non-asymptotic conservativeness under certain dependence scenarios. The proposed two-way adaptive procedure is effectively applied to a data set from microbial abundance study.
Introduction: Digestive laboratory abnormalities related to COVID-19 have been previously described, but most reports came from single centers and findings have been conflicting. We conducted a multi-center study using data from three large urban VA centers (New York Harbor VA, New Orleans VA and Detroit VA) to examine the association between demographics and digestive laboratory values with mortality on index hospitalization among individuals diagnosed with COVID-19. Methods: We manually extracted data on individuals hospitalized for COVID-19 between December 2019 and June 2020 at the three facilities. For this analysis, data on demographics and seven digestive laboratory values (highest AST, ALT, alkaline phosphatase, total bilirubin, and INR during admission, as well as lowest hemoglobin and platelets) were analyzed in relation to index hospitalization mortality. We performed descriptive statistics and conducted a multivariable logistic regression model. Results: Out of a total of 390 individuals who were hospitalized with COVID-19, 168 (43%) died and 222 survived. The median age of patients who died was higher than those who survived (75 vs. 69 years). The vast majority (94%) of patients were male. Black patients accounted for a higher proportion of those who died than those who survived (61% vs. 55%), whereas the opposite was true for Whites (26% vs. 31%) and Hispanics (9% vs. 12%). In the multivariable model (Table), mortality was associated with older age (OR 1.07, 95% CI 1.03-1.10), higher BMI (OR 1.05, 95% CI 1.01-1.10), higher AST (OR 1.01, 95% CI 1.004-1.02), lower ALT (OR 0.99, 95% CI 0.98-0.996), higher alkaline phosphatase (OR 1.02, 95% CI 1.01-1.02), and lower hemoglobin (OR 0.83, 95% CI 0.72-0.97). Conclusion: In this multicenter VA study of patients hospitalized with COVID-19 during the first half of 2020, overall mortality was 43%. For mortality during index hospitalization, we observed a positive association with age, BMI, AST, and alkaline phosphatase, and an inverse association with ALT and hemoglobin. Every 1 unit increase in hemoglobin was associated with 17% decreased odds of death. These findings suggest that commonly used digestive laboratory tests have prognostic significance for COVID-related survival.Table 1.: Patient Retention to Second Biopsy in Practice-Integrated Sites by Study.
Introduction: Digestive laboratory abnormalities related to COVID-19 have been previously described, but most reports came from single centers and findings have been conflicting. We conducted a multi-center study using data from three large urban VA centers (New York Harbor VA, New Orleans VA and Detroit VA) to examine the association between demographics and digestive laboratory values with mortality on index hospitalization among individuals diagnosed with COVID-19. Methods: We manually extracted data on individuals hospitalized for COVID-19 between December 2019 and June 2020 at the three facilities. For this analysis, data on demographics and seven digestive laboratory values (highest AST, ALT, alkaline phosphatase, total bilirubin, and INR during admission, as well as lowest hemoglobin and platelets) were analyzed in relation to index hospitalization mortality. We performed descriptive statistics and conducted a multivariable logistic regression model. Results: Out of a total of 390 individuals who were hospitalized with COVID-19, 168 (43%) died and 222 survived. The median age of patients who died was higher than those who survived (75 vs. 69 years). The vast majority (94%) of patients were male. Black patients accounted for a higher proportion of those who died than those who survived (61% vs. 55%), whereas the opposite was true for Whites (26% vs. 31%) and Hispanics (9% vs. 12%). In the multivariable model (Table), mortality was associated with older age (OR 1.07, 95% CI 1.03-1.10), higher BMI (OR 1.05, 95% CI 1.01-1.10), higher AST (OR 1.01, 95% CI 1.004-1.02), lower ALT (OR 0.99, 95% CI 0.98-0.996), higher alkaline phosphatase (OR 1.02, 95% CI 1.01-1.02), and lower hemoglobin (OR 0.83, 95% CI 0.72-0.97). Conclusion: In this multicenter VA study of patients hospitalized with COVID-19 during the first half of 2020, overall mortality was 43%. For mortality during index hospitalization, we observed a positive association with age, BMI, AST, and alkaline phosphatase, and an inverse association with ALT and hemoglobin. Every 1 unit increase in hemoglobin was associated with 17% decreased odds of death. These findings suggest that commonly used digestive laboratory tests have prognostic significance for COVID-related survival.Table 1.: Patient Retention to Second Biopsy in Practice-Integrated Sites by Study.
Multiple testing of two-way classified hypotheses controlling false discoveries is a commonly encountered statistical problem in modern scientific research. Nevertheless, not much progress has been made yet towards improving existing multiple testing procedures by adequately adjusting them to such structural settings. This paper makes contributions to the development of local false discovery rate (Lfdr) based methodologies under these settings. More specially, it extends the two-component mixture model (Efron et al. J. Am. Statist. Assoc. 96, 1151–1160, 2001) from un-classified to two-way classified hypotheses, which captures the underlying two-way classification structure of the hypotheses and provides the foundational framework for the development of newer and potentially powerful Lfdr-based multiple testing procedures for the hypotheses.
Multiple testing literature contains ample research on controlling false discoveries for hypotheses classified according to one criterion, which we refer to as one-way classified hypotheses. Although simultaneous classification of hypotheses according to two different criteria, resulting in two-way classified hypotheses, do often occur in scientific studies, no such research has taken place yet, as far as we know, under this structure. This article produces procedures, both in their oracle and data-adaptive forms, for controlling the overall false discovery rate (FDR) across all hypotheses effectively capturing the underlying one- or two-way classification structure. They have been obtained by using results associated with weighted Benjamini-Hochberg (BH) procedure in their more general forms providing guidance on how to adapt the original BH procedure to the underlying one- or two-way classification structure through an appropriate choice of the weights. The FDR is maintained non-asymptotically by our proposed procedures in their oracle forms under positive regression dependence on subset of null $p$-values (PRDS) and in their data-adaptive forms under independence of the $p$-values. Possible control of FDR for our data-adaptive procedures in certain scenarios involving dependent $p$-values have been investigated through simulations. The fact that our suggested procedures can be superior to contemporary practices has been demonstrated through their applications in simulated scenarios and to real-life data sets. While the procedures proposed here for two-way classified hypotheses are new, the data-adaptive procedure obtained for one-way classified hypotheses is alternative to and often more powerful than those proposed in Hu et al. (2010).
A comprehensive nonparametric statistical learning framework, called LPiTrack, is introduced for large-scale eye-movement pattern discovery. The foundation of our data-compression scheme is based on a new Karhunen–Loéve-type representation of the stochastic process in Hilbert space by specially designed orthonormal polynomial expansions. We apply this novel nonlinear transformation-based statistical data-processing algorithm to extract temporal-spatial-static characteristics from eye-movement trajectory data in an automated, robust way for biometric authentication. This is a significant step towards designing a next-generation gaze-based biometric identification system. We elucidate the essential components of our algorithm through data from the second Eye Movements Verification and Identification Competition, organized as a part of the 2014 International Joint Conference on Biometrics.