
In analyses of time-to-event outcomes, we may utilize efficient sampling designs to save effort, cost, and valuable specimens. The originally proposed case-cohort design is one such design, and performs analysis on a random sample selected from the full cohort supplemented with unsampled individuals who experienced an event. In Alzheimer’s disease biomarker discovery, such cost-savings is important as samples that contain biomarker measurements can be very limited due to high burden and cost. Furthermore, quantifying differences across subgroups and ensuring that diagnostic biomarker effects can be well powered and well understood across subpopulations is an increasingly important goal. This ability to quantify subgroup differences is required in efforts to minimize health disparities and achieve precision medicine. In this paper, we emphasize the importance of such considerations and propose using a stratified case-cohort estimator to sample a subcohort with balanced subgroups across the proposed effect modifier. We provide inverse-probability-of-sampling-based estimates that consistently estimate covariate effects in the population while improving precision for assessing differences across rare subpopulations. Simulation results are presented to illustrate power gains for testing effect modification, especially as a population subgroup becomes increasingly rare, and show minimal bias for marginal covariate effect estimates. We apply this method to data from the Alzheimer’s Disease Neuroimaging Initiative (ADNI) to better assess potential heterogeneity of phosphorylated tau as a biomarker for AD progression across ethnoracial and genetic status subpopulations.
Reliable estimation of population health indicators is essential for evidence-based public health planning and policy formulation. This study applies a ranked set sampling approach with auxiliary health information to improve the estimation of population means in large-scale health surveys. Using data from the fifth round of the National Family Health Survey (NFHS-5) for the Varanasi district of Uttar Pradesh, India, two empirical applications related to maternal and child health are examined. In the first application, maternal body mass index is used as auxiliary information to estimate infant birth weight, while in the second, maternal hemoglobin level is utilized to improve estimation of child stunting prevalence. New classes of estimators incorporating logarithmic transformations of auxiliary variables are employed and evaluated in terms of mean squared error. The empirical findings demonstrate that the proposed estimators provide consistently more precise and reliable estimates than existing methods across different sample sizes. These results highlight the practical value of ranked set sampling and auxiliary information in enhancing the quality of health estimates from national surveys, with direct implications for public health monitoring and policy evaluation.
In this research, we follow the late Professor Tze Lai’s inspiration to propose a new Kiefer–Weiss–Lorden–Lai (KWLL) framework that minimizes the expected number of samples to evaluate efficacy (rejecting the null hypothesis) or futility (accepting the null hypothesis) subject to the maximum sample size and Type I error probability constraints. Compared to the traditional sample size calculation, our proposed KWLL framework structures the parameter of the alternative hypothesis as a function of design decisions, instead of a pre-specified value, while making a quick and accurate decision whether to accept or reject the null hypothesis. Moreover, we use the so-called 2-Sequential Probability Ratio Test (2-SPRT) as the foundation to develop reasonable algorithms that might provide a useful approximate solution to the KWLL framework. Finally, a case study using correlation as the test statistic is presented to illustrate the usefulness of our proposed framework and algorithm.
In this paper, we develop bivariate stochastic functional linear models (BSFLM) for gene-based association analysis of quantitative traits with high-dimensional sequencing genetic data in longitudinal studies. In longitudinal studies, the traits of different subjects are assumed to be independent for population data. The two traits of one subject, however, are correlated with each other. In addition, the multiple measurements of one trait of a subject are also correlated. To build valid models to analyze bivariate quantitative traits with sequencing genetic data in longitudinal studies, a variance-covariance structure is constructed to describe (a) variation accounting for the correlation between the two traits and (b) variation accounting for the correlation within multiple measurements of a trait on the same subject. Functional data analysis techniques are utilized to reduce high dimensionality of sequence data and draw useful genetic information. Spline models are used to approximate temporal mean functions and genetic effect functions. By intensive simulation studies, it is shown that the proposed BSFLM control type I errors well and have good power levels. We test and refine the models and related software using real datasets of multi-ethnic study of atherosclerosis.
Recent advancements in spatially resolved transcriptomics (SRT) technologies have enabled the comprehensive molecular and spatial characterization of single cells, providing valuable insights into the cellular organization of tissues. SRT techniques, such as single-molecule fluorescence in situ hybridization (FISH)-based methods (e.g., seqFISH, STARmap) and next-generation sequencing (NGS)-based methods (e.g., spatial transcriptomics, 10x Visium), allow for the measurement of gene expression across large populations of cells or tissue spots. These approaches generate high-dimensional data that integrate both molecular profiles and spatial context, which is crucial for understanding tissue structure and function in areas like development, neuroscience, and cancer biology. Identifying spatially variable (SV) genes, whose expression patterns differ across spatial locations, is a key step in analyzing these complex spatial transcriptomic maps. To enhance our understanding of the spatial profiles of SV genes, we propose a Bayesian nonparametric zero-inflated Poisson (ZIP) regression model for clustering these genes. Our model explicitly accounts for zero-inflation in the data, uses non-negative matrix factorization to uncover gene expression patterns, and incorporates Moran’s I (MI) basis functions to address potential confounding. Additionally, the model infers the number of clusters directly from the data, obviating the need for pre-specifying the number of clusters. We demonstrate the utility of this approach on two SRT datasets, showing that it provides more robust and interpretable clustering of SV genes, opening new avenues for understanding complex biological processes.
The multimodal data assembled through All of Us Research Program with enriched recruitment of populations underrepresented in biomedical research have provided a rich resource for equitable health research. The complexity of the All of Us research platform, however, is a common hurdle at the initiation of studies. This tutorial aims to demonstrate to researchers a practical, transparent and reproducible workflow for streamlining data curation and analysis. This tutorial covers modules: (1) design of All of Us studies; (2) selecting study cohort; (3) curating multimodal data (electronic health records, surveys, genomics, and wearable devices); (4) harmonizing disease onset outcomes in risk models. We showcased the proposed workflow using a case study for genetic risk prediction of rheumatoid arthritis for minority race/ethnicity groups. Step-by-step guides are provided in the supplementary materials and the online repository with links to SQL/R/Python example code and All of Us user interface operations.
Modern biomedical data are increasingly collected across multiple institutions and time periods, creating opportunities for improved statistical inference through knowledge transfer, but also posing challenges for privacy, scalability, and distributional heterogeneity. We propose a Renewable Federated Incremental Transfer framework, termed ReFIT, for sequentially integrating information from streaming source datasets to improve model estimation and prediction in a target population with limited samples. ReFIT builds upon a density ratio model to account for covariate shift between the source and target populations and employs a renewable updating strategy that allows model parameters to be incrementally refined as new source data become available, using only summary-level information from prior sources. This framework ensures privacy preservation and computational efficiency while adapting to evolving data environments. Beyond improving predictive performance, ReFIT also quantifies predictive uncertainty within a conformal prediction framework, yielding valid prediction intervals that adapt as new information accumulates. Extensive simulation studies demonstrate that ReFIT achieves higher predictive accuracy and better uncertainty quantification than models trained on target or source data alone. The method remains robust under nonlinear model misspecification and varying degrees of source-target shift. Moreover, as ReFIT incrementally integrates additional source data, the conformal prediction intervals become progressively narrower without sacrificing coverage, evidencing improved statistical efficiency with growing information. In an electronic health record application for breast cancer prediction, ReFIT substantially improves prediction for the Hispanic population by sequentially leveraging information from non-Hispanic White patients collected over multiple time periods. These results highlight the potential of ReFIT as a general and practical framework for privacy-preserving, adaptive, and scalable learning from distributed and periodically updated biomedical data.
Functional Principal Component Analysis (FPCA) is a popular tool for representing functional data in lower-dimensions, along highly-interpretable components depicting directions of variation in the data across time. We extend FPCA to contrastive settings involving multiple groups with a focus on time-dynamic components which are shared across samples versus those that are unique to a particular sample. The proposed Contrastive Latent Functional Model (cLFM) integrates FPCA with contrastive learning principles to capture common directions of variation across datasets while isolating distinct features specific to each sample. The model employs a computationally efficient Expectation-Maximization (EM) algorithm for parameter estimation, ensuring orthogonality between shared and unique functional spaces. Simulation studies demonstrate the efficacy of the proposed method across different scenarios of shared and unique variation, varying sample size and error variance. Applied to neurodevelopmental and clinical datasets, including EEG studies and longitudinal studies of kidney function, cLFM reveals novel insights into group differences between autism vs neurotypical development and among mild to severe albuminuria subgroups of chronic kidney disease patients. Bridging contrastive analysis with functional data, this framework advances the capacity to disentangle complex variation patterns in multi-group functional studies in biomedical research.
This paper discusses simultaneous estimation and variable selection for a class of Cox–Aalen transformation models, one of the most general types of models commonly used in regression analysis of failure time data. For the problem, we propose a penalized maximum likelihood estimation procedure with the use of the minimum information criterion (MIC). Unlike most of penalized methods, the proposed approach does not need the selection of the tuning parameter, which makes it both stable and efficient. For the implementation of the proposed method, an EM algorithm is developed and an extensive simulation study is conducted and suggests that it works well in practical situations. Finally, it is applied to a motivating breast cancer study.
Little research has been conducted on housekeeping (HK) genes’ methylation patterns in breast cancer (BRCA) and uterine corpus endometrial carcinoma (UCEC). Using publicly available data, we conducted comprehensive analyses of HK genes’ methylation patterns at the cytosine guanine (CG) site level. Findings are summarized below. First, HK sites exhibited lower differential methylation (DM) rates than overall DM rates. UCEC demonstrated a higher DM rate than BRCA, and they shared a small percentage of DM sites and 18 common DM HK genes that are cancer-related. Second, DM sites’ locations were different between BRCA and UCEC, with notable disparities in CpG islands and shores. Third, UCEC had more highly co-methylated CG pairs than BRCA. Normal and tumor samples had different co-methylation distributions in both BRCA and UCEC. Fourth, some CG sites’ co-methylation patterns did change between normal and tumor, with some minor shifts regarding overall co-methylation differences between the Alive and Dead samples in both BRCA and UCEC. Fifth, only a small number of CG sites kept strong and stable co-methylation patterns in BRCA and UCEC, respectively, and these two cancers only shared a small percentage of these co-methylated CG sites. Despite the small common CG list, 11 HK genes and 27 non-HK genes were identified as common in BRCA and UCEC. 19 of these 38 genes were from the protocadherin gamma cluster (PCDHG@). Finally, 1 HK gene (CUX1) and 4 non-HK genes (TP73, FOXP1, BANP, and MIR199A1) may play dual roles as both oncogenes and tumor suppressor genes.
The accelerated failure time (AFT) model offers a directly interpretable, time-ratio alternative to Cox regression and is attractive for clustered right-censored data. We develop an embedded-likelihood AFT framework that (i) transforms observed times so that, at the truth, the transformed failure times are i.i.d. and independent of covariates, and (ii) uses these transformed data to derive score-type estimating equations under two common semiparametric embeddings: proportional hazards (PH) and proportional odds (PO). Estimation proceeds via a two-level algorithm that separates outer updates for the regression vector from inner profiling of the shared frailty, scale, and frailty-variance parameters. We provide sandwich standard errors and optional cluster bootstrap inference including the consistency and asymptotic normality for the PH/PO embeddings. Comprehensive simulations spanning multiple cluster sizes, censoring levels, and error/frailty distributions indicate that the embedded-likelihood estimators match, and occasionally improve upon, penalized partial likelihood (PPL) for regression effects, with near-nominal coverage using either sandwich or bootstrap standard errors. Heavy-tailed errors mainly inflate variance, while frailty-law misspecification primarily impacts the frailty variance but leaves regression estimates well centered with robust inference. In a Diabetic Retinopathy Study reanalysis, PH and PO-embedded AFT fits yield effect estimates consistent with PPL, slightly better AIC/BIC for the Cox-AFT embedding, and a clinically meaningful between-patient heterogeneity captured by the frailty. The embedded-likelihood AFT formulation provides a unified, interpretable, and computationally practical route to likelihood-based inference for clustered censored data, delivering performance competitive with established approaches while retaining the advantages of the AFT perspective.
This paper addresses three-category and multi-category diagnostic problems, emphasizing the importance of identifying pseudo-categories before multi-dimensional ROC analysis. When the distributions of diagnostic variable for adjacent categories are undiagnosable, they should be combined into one category, removing nuisance parameters. First, this work introduces a test method called Nonparametric Test by Bootstrap AUC (NTBA), rooted in bootstrap technique and nonparametric AUC estimation theory. This method helps determine whether two adjacent categories can be fused and is applicable in multi-category problems. For three-category medical diagnostics, the hypothesis test is formulated by H_0 :P( X < Y) = 0.5VSH_1 :P( X < Y) 0.5. If H_0 is not rejected, merging them to two categories is justified. Second, three existing tests have been introduced and compared for the above H_0 to assess which can achieve optimal size and power similar to NTBA. Simulation experiments compare the performance of NTBA with three popular tests under various distribution conditions. Finally, the Wilcoxon–Mann–Whitney test, despite lacking the theoretical foundation of AUC, can achieve results comparable to those of NTBA. This discovery opens new applications for WMW in the fusion of categories for multi-way ROC analysis. Additionally, two medical datasets are utilized to evaluate the feasibility of fusing two categories using four different hypothesis testing methods.
Artificial intelligence and machine learning are transforming drug development by streamlining clinical trials. This review examines their role in enhancing trial design through parameter optimization and endpoint development, as well as improving trial conduct via optimized site selection, patient recruitment, and safety monitoring. It also details AI/ML applications in data analytics—including adaptive designs (dose finding, randomization), data cleaning, synthetic control generation, and multimodal data integration—as well as predicting trial success and enrollment. The paper concludes by emphasizing the necessity of robust technical validation, ethical oversight, and regulatory alignment for the safe and effective integration of these technologies.
A new model for joint analysis of recurrent and terminal events is proposed. Conditional on a random effect, the intensity for the recurrent event process has a multiplicative form and a Cox model is formulated for the terminal event process. The two processes are linked via a copula function which defines a joint model for the random effect and the terminal event. When these processes are subject to right-censoring, the observed data likelihood can be maximized directly or via an EM algorithm; the latter facilitates joint semiparametric analysis. Variance estimates are derived based on Louis’ method. A computationally more convenient two-stage estimation procedure is also investigated. Simulation studies demonstrate good finite sample performance under simultaneous and two-stage estimation, and an application to a study of effect of pamidronate on reducing skeletal complications in patient with skeletal metastases illustrates the use of this model.
Certain rare cancers such as ovarian or pancreatic cancer would benefit if detected early at a stage when they are resectable. Unfortunately, approved biomarkers for these cancers are not adequate for screening the general population, and it is unlikely that a single marker will meet the performance criteria for screening. Determining a combination of biomarkers for early detection of rare cancers is a challenge. Often model selection suffers from overfitting in the discovery phase, which leads to poor performance upon validation. Since ovarian cancer has a poor prognosis, we aim to identify biomarkers that perform robustly in early cancer detection discovery and validation phases. Stability selection methods have been used to prevent overfitting and to reliably select truly expressed biomarkers. Ensemble learning methods provide robust prediction results in the face of model misspecification. We present a novel framework with a biomarker selection stage with stability selection and prediction stage using an ensemble of machine learning (ML) methods, namely the stability selection ensemble learning (STABEL). The ensemble consists of random forest (RF), logistic regression (LR), linear discriminant analysis (LDA), and support vector machine (SVM). Simulated data, along with ovarian cancer data from Prostate, Lung, Colorectal, and Ovarian (PLCO) Cancer Screening Trials are used to test the methodology. Our results demonstrate the power of integrating ensemble learning with stability selection to select cancer biomarkers and achieve strong predictive performance. The STABEL technique outperforms or performs similarly to the existing methods in terms of sensitivity, specificity, prediction accuracy, and area under the curve.
It is often not feasible to conduct randomized clinical trials (RCTs) with a sufficiently large sample size in a pediatric study. A common strategy to improve feasibility is to borrow information from a comparable external adult study. However, adult trial subjects may not be directly comparable to pediatric patients due to differences in baseline characteristics, prognostic factors, comorbidities, and previous treatments. In this article, we propose a propensity score integrated borrowing-by-parts power prior to estimate the effect of treatment in a pediatric study using a Bayesian approach. Specifically, we first estimate the propensity scores to assess the comparability between the pediatric and adult study. We then form strata based on these propensity scores and apply a stratum-specific borrowing-by-parts power prior, allowing for greater flexibility in borrowing information from different components of the data. In addition, we introduce data-driven discounting parameters to determine the level of borrowing. An extensive simulation study is carried out to demonstrate the proposed approach. To illustrate the practical implementation of our approach, we apply it to a case study of systemic lupus erythematosus (SLE) disease.
Hospital medications can significantly influence patient recovery and length of stay (LoS) in COVID-19, highlighting the need to understand drug effects in diverse patient subgroups. Therefore, identifying drug interactions and the types of medications administered during hospitalization is crucial, as it can be associated with LoS, and reducing the length of hospitalization can lead to improved hospital service delivery. We analyzed a prospective cohort of 1998 COVID-19 patients hospitalized at Khorshid Hospital, Iran (February 2020–February 2021), using an innovative hybrid modeling approach. Our framework combines model-based recursive partitioning (MOB trees) with (1) semiparametric Fine–Gray (FG) models for interpretable covariate effect estimation and (2) non-parametric Generalized Additive Models for competing risks (GAM-CR) to model complex nonlinear medication effects. This dual modeling strategy, building upon rigorous feature selection, successfully identified clinically meaningful patient subgroups with differential treatment responses, enabling personalized assessment of LoS based on individual medication regimens and clinical characteristics. The results indicate that the hospital LoS among severe patients is about 3 days, significantly higher than in non-severe patients (p-value < 0.05). Moreover, the effect of average daily medication dosage on LoS varies notably across four clinical subgroups. MOB tree analysis using FG model (FG-MOB) indicated that older age consistently prolonged hospitalization across all patient subgroups, reducing the probability of discharge by approximately 1–1.3
Instrumental variable (IV) analysis is widely used in fields such as economics and epidemiology to address unobserved confounding and measurement error when estimating the causal effects of intermediate covariates on outcomes. However, extending the commonly used two-stage least squares (TSLS) approach to survival settings is nontrivial due to censoring. This paper introduces a novel extension of TSLS to the semiparametric accelerated failure time (AFT) model with right-censored data, supported by rigorous theoretical justification. Specifically, we propose a generalized estimating equation (GEE) approach that combines Leurgans’ synthetic variable method with adaptive weighting, establish the asymptotic properties of the resulting estimator, and derive a consistent variance estimator, enabling valid causal inference. Simulation studies are conducted to evaluate the finite-sample performance of the proposed method across different scenarios. The results show that it outperforms the naïve unweighted GEE method, a parametric IV approach, and a one-stage estimator without IV. The proposed method is also highly scalable to large datasets, achieving a 300- to 1500-fold speedup relative to a Bayesian parametric IV approach. We further illustrate its utility through a real-data application using the UK Biobank data.
The kidney paired donation (KPD) program provides an innovative solution to overcome incompatibility challenges in kidney transplants by matching incompatible donor-patient pairs and facilitating kidney exchanges. To address unequal access to transplant opportunities, there are two widely used fairness criteria: group fairness and individual fairness. However, these criteria do not consider protected patient features, which refer to characteristics legally or ethically recognized as needing protection from discrimination, such as race and gender. Motivated by the calibration principle in machine learning, we introduce a new fairness criterion: the matching outcome should be conditionally independent of the protected feature, given the sensitization level. We integrate this fairness criterion as a constraint within the KPD optimization framework and propose a computationally efficient solution using linearization strategies and column-generation methods. Theoretically, we analyze the associated price of fairness using random graph models. Empirically, we compare our fairness criterion with group fairness and individual fairness through both simulations and a real-data example.