Objectives:Diabetes affects over 500 million people globally and glycemia is inadequately managed. Metformin is the most frequently prescribed initial treatment for type 2 diabetes globally, yet glycemic response trajectories to metformin in routine real-world care and predictors of treatment response have not been well described. We aimed to identify glycemic response trajectories in adults prescribed metformin monotherapy as initial type 2 diabetes treatment and predictors of poor glycemic response to metformin. Design:Observational cohort study using latent class mixed models to identify hemoglobin A1c (HbA1c) trajectory classes, followed by random forests machine learning to predict trajectory class membership. Setting:US Veterans Affairs Healthcare System. Participants:Adults treated with metformin alone for >30 days after diabetes diagnosis with a minimum of two HbA1c measurements from 90 days prior to two years after the first metformin prescription (N=140,413). Exposures:Demographic, laboratory, vital sign, and comorbidity data were included as predictors of metformin response trajectory. Main Outcomes and Measures:We included all HbA1c measurements (487,604 total) for two years after metformin initiation to define metformin glycemic response trajectories. Results:We identified three HbA1c trajectories: stably low (89.7% of sample, mean HbA1c decrease from 7.2% to 6.6%), brisk response (7.1% of sample, mean HbA1c decrease from 11.4% to 7.0%), and non-response (3.1% of sample, mean HbA1c increase from 8.9% to 10.8%). Of those in the stably low and brisk response classes at 2 years, 91% maintained HbA1c at approximately 7% on metformin alone for 5 years after drug initiation. Prediction models could accurately predict brisk response (91% accuracy) but not metformin non-response (59% accuracy). Conclusions:Most individuals treated initially with metformin monotherapy have a beneficial and durable glycemic response. Predicting individuals who will not respond to metformin may be challenging but is evident within six months with recommended glycemic surveillance. The findings support current guidelines for HbA1c surveillance when initiating diabetes treatment.
High-throughput spatial omics technologies enable molecular profiling within intact tissue architecture, yet identifying concise, predictive, and biologically interpretable marker panels for cell types, tissue domains, and disease-associated tissue classes remains challenging. This limitation hinders the development of actionable panels for targeted validation and downstream translation. Existing pipelines rely largely on univariate differential-expression analyses, which ignore joint molecular structure and provide limited predictive insight. Multivariate machine-learning methods, including random forest, XGBoost, elastic net, and specialized single-cell panel-selection approaches, can capture predictive patterns but typically lack explicit spatial modeling and probabilistic feature selection, relying instead on model-specific importance scores or user-specified panel sizes. We develop rapid kernel machine regression (RKMR), a scalable framework for spatial-omics marker discovery that integrates nonlinear kernel modeling, spike-and-slab variable selection, and spatial dependence. RKMR uses automatic relevance determination (ARD) kernels and sparsity-inducing priors to capture nonlinear marker-outcome relationships and implicit feature interactions while producing approximate posterior inclusion probabilities (PIPs) that quantify model-based uncertainty in feature inclusion. To scale inference to large spatial datasets, RKMR combines low-rank kernel approximations with stochastic variational optimization. In simulations, RKMR consistently achieves higher AUPRC than competing methods across a range of molecular-signal and spatial-effect settings. Across spatial transcriptomics and scRNA-seq datasets, RKMR identifies parsimonious marker sets that recover reported cell-type signatures and reproducible tissue-layer markers. These results establish RKMR as a scalable and uncertainty-aware framework for translating high-dimensional spatial omics data into robust, experimentally actionable marker panels.
A bstract Identifying spatially variable genes (SVGs), genes whose expression varies coherently across tissue space, is a central analytic goal in spatially resolved transcriptomics. Current methods rank spatially variable genes using either significance probabilities from parametric models or effect sizes such as the proportion of spatial variance, but parametric approaches impose distributional assumptions, such as Gaussian processes or negative binomial models, that may be violated for sparse or zero-inflated data. Furthermore, most detection tools cannot jointly model multiple biological replicates, and no existing framework provides both nonparametric significance probabilities and stabilized effect-size estimates with formal uncertainty quantification. Here, we introduce CytoKspace, a nonparametric framework that combines a sparse exponential kernel constructed from nearest-neighbor graphs with a quadratic-form test statistic and adaptive permutation testing. CytoKspace employs a multi-stage adaptive permutation schedule that yields substantial computational savings over fixed-permutation baselines, an adaptive shrinkage layer built on empirical Bayes estimation that stabilizes raw spatial effect sizes and provides posterior estimates with local false sign rates, and a scalable multi-sample extension via Fisher combination of significance probabilities and inverse-variance-weighted meta-analysis that accommodates studies with multiple biological replicates. In extensive simulations across a broad range of sample sizes, gene counts, spatially variable gene fractions, and effect sizes, as well as in applications to two real datasets from the Visium and seqFISH platforms, CytoKspace demonstrates competitive sensitivity, well-calibrated false positive rates, and practical computational requirements compared to existing methods. A software implementation of our method is freely available at https://github.com/Ghoshlab/CytoKspace .
Causal effect estimation from observational data requires careful adjustment for confounding. Classical estimators such as inverse probability weighting and augmented inverse probability weighting are effective under favorable model specification, but may become unstable when treatment assignment and outcome mechanisms are complex, non-linear, and high-dimensional. Machine learning and representation learning approaches improve flexibility, yet joint training can allow outcome-related information to influence treatment-side representations, which is undesirable from a causal perspective. We propose MOCA (Modular One-way Causal Attention), a transformer-based framework that separates treatment and outcome modeling through a modular design, and performs confounder adjustment using a one-way attention mechanism. A cutting-feedback strategy, implemented via gradient detachment, prevents the outcome loss from updating the treatment module. This design preserves directional information flow while retaining the representational power of transformer architectures for causal inference. Across multiple simulated scenarios, including linear, nonlinear, heavy-tailed, hidden confounding, and high-dimensional settings, MOCA shows competitive or improved performance relative to IPW, AIPW, X-learner, TARNet, and DragonNet. We further illustrate the method on the Infant Health and Development Program dataset and the Dehejia-Wahba dataset as real-world benchmarks. These results suggest that modular attention with one-way information flow provides a promising and interpretable direction for causal inference with modern deep learning models.
The introduction of gene therapies to treat and potentially cure various diseases has led to new considerations in the design, analysis and conduct of clinical trials involving these treatments. In this article, we consider the use of causal inference techniques to improve our understanding of aspects of these trials, focusing primarily on single-arm studies. While the ICH E9 R1 guidance provide a strong connection to causal inference ideas, we show how other aspects of gene therapy trials are amenable to causal thinking. We provide a review of the potential outcomes model and its assumptions. Next, we summarize the ICH E9 R1 guidelines on estimands. We then describe the implications of various aspects of causal inference for gene therapy clinical trials. Finally, we discuss how current and future clinical trials involving gene therapy in cystic fibrosis can benefit from causal inferential thinking.
Objectives: HIV is associated with reduced cardiorespiratory fitness (CRF), a strong predictor of future cardiovascular events and mortality. Supervised exercise training programs improve CRF, but the ideal exercise training intensity is unknown among PWH. We hypothesized that high-intensity interval training (HIIT) would result in greater improvements in CRF than continuous moderate-intensity exercise (CME) without increased risk of adverse events among older sedentary adults with HIV. Methods: We conducted a multicenter randomized clinical trial of people ages 50 years and older with treated, virally-suppressed HIV, self-reported fatigue, and sedentary lifestyle. Participants were randomized 1:1 to 16 weeks of supervised HIIT or CME. Cardiopulmonary exercise testing (CPET) using a treadmill graded exercise protocol was performed prior to randomization and at 16 weeks to assess CRF, measured as peak oxygen consumption (VO 2 ). CPETs were interpreted blinded to treatment assignment. The trial was preregistered at ClinicalTrials.gov NCT04550676. Results: Among 212 PWH screened, 142 consented, 127 completed baseline CPET, 118 were randomized (110 with adequate quality CPET): 54 to HIIT and 56 to CME. Median age was 56.5 years, 12% female, and 44% were obese (BMI ≥30). At baseline, absolute peak VO 2 was 2.3 L/min and relative to body weight was 25 ml/kg/min (91% predicted peak VO 2 ). In the intention-to-treat analysis adjusted for age, sex, and site, assignment to HIIT was associated with a 0.14 L/min greater increase in absolute peak VO 2 from baseline to 16 weeks compared to CME (95% CI 0.01 to 0.27, p =0.03) and a 1.60 ml/kg/min greater increase in relative peak VO 2 (95% CI 0.23 to 2.97; p =0.023). Results were consistent in unadjusted and per-protocol analyses. Six serious adverse events (SAEs) occurred during the 16-week intervention: 4 unrelated (including 1 death), 2 related (both syncope) in the CME arm; participants were evaluated by cardiology and resumed exercise without further SAEs. There were more nonserious AEs in the HIIT group (40 versus 22 in CME group), mostly mild-moderate musculoskeletal injuries. Conclusions: Among older sedentary people with HIV, HIIT is associated with a greater improvement in CRF assessed with CPET compared to CME, without a greater risk of serious adverse events. Exercise recommendations for aging people with HIV should consider the greater cardiopulmonary benefit of HIIT compared to CME.
The E-value has been proposed as a tool for sensitivity analysis for unmeasured confounding in observational studies (VanderWeele and Ding). As it was developed in the setting of risk ratio comparisons, using E-values for proportional hazards models requires approximating a risk ratio from a hazard ratio. We consider the impact of the bias induced when using the suggested approximation method, as well as how doing so may exacerbate recognized E-value interpretation issues. A running example helps illustrate how challenges of biased E-values may arise in time-to-event studies and also ways in which alternative approaches can avoid the bias.
Clinical decisions for determining optimal patient-specific interventions are complicated prediction tasks that rely on health care professionals' understanding of physiological mechanisms and their dynamics. These decisions are challenged by (a) observational data sparsity and (b) patient heterogeneity. Here, we focus on estimating and forecasting specific physiological properties-that are not explicitly present in clinical observations-to provide additional features using only data available bedside at the time of decision-making. Mechanistic models of physiological system(s), e.g., physiological ordinary differential equation (ODE) models, provide pathways to compensate for data sparsity by synchronizing the model with observations of an individual patient using data assimilation (DA). However, DA used in a standard computational workflow to estimate constant model parameters from presently-known data is less effective at optimizing state forecasts of the model governed by physiological processes that evolve before new observations are available. Stated simply, we cannot forecast the future evolution of the model because we cannot forecast model parameters. To support next-generation clinical decision support, we develop a new DA and machine learning (ML) hybrid pipeline to estimate and forecast individual future physiological processes by forecasting ODE model parameters. This pipeline overcomes model and DA workflow limitations by stacking a DA-estimated posterior empirical distribution of physiological parameters with longitudinal ML forecasting models. We work within the context of glycemic management in an ICU using EHR data to construct and test a use case. We use synthetic data and real-world clinical data to validate the integrated pipeline and quantify uncertainties.
Association between intensity adjusted O&G well site activity within 16-kilometer radius of maternal residence and ALL for cases and controls born between 1992 and 2019 in Colorado for mother’s residential address confirmed in commercial database.
In this article, we develop a weighted approach to estimation for right-censored time to event data in the presence of external predictions available from a prediction model. There are several advantages to the proposed approach. First, the method allows for arbitrary forms for the external prediction model. Second, the methodology can be fit easily using standard software packages that allow for subject-specific weights. Third, all that is needed from the external models are access to predictions and not the actually prediction equation. A complication is that inference becomes challenging, so we develop new theoretical results along with a perturbation-based method for inference. The methodology is applied to three publicly available datasets.
Study population exposure distributions between IA-IDW exposure groups within 16-km buffer by case-control status and exposure window (451 cases and 2,706 controls)
An accurate and early brain tumor identification and classification system is essential for making a proper and on-time treatment decision. Manually, brain tumors can be diagnosed and classified by analysis of histopathological reports of biopsy samples, but it is time-consuming and has a high chance of wrong identification and classification, which may put the patient in danger. Recent research and developments in Artificial Intelligence, Machine Learning, and Deep Learning have started to play important roles in medical diagnosis and support doctors in making smart treatment decisions. Convolutional Neural Networks (CNNs) are a powerful deep learning-based technique that has recently been applied to the automatic detection and classification of brain tumors. In this literature review, we have reviewed the research papers related to automatic brain tumor identification and classification approaches based on deep learning. A lot of research based on DL has been done in the last few years to implement brain tumor classification and segmentation models with significant results. Some of the drawbacks may be found in various previously proposed methods. So, more improved deep learning models may be implemented in further research by comparing, changing, and combining the various previously proposed models.
Emerging technologies, such as artificial intelligence, emphasize the importance of quantitative methods and their applications when conducting increasingly data-intensive research. The scientific discipline of statistical practice is critical for achieving rigor in research addressing important domain-level questions. The misconception that statistical practice is not a science but rather a service threatens scientific rigor and compromises the quality of its contribution to clinical and translational research. The authors call on academic and research leadership to recognize statistical practice as a key scientific discipline and to further ensure that the field and its scientists are nurtured. To that end, academic homes for faculty of statistical practice and its scholarship need to be fostered with clear pathways to hire, retain, promote, and tenure faculty who are largely team scientists. This goal can be achieved inclusively by avoiding the creation of separate faculty lines that might imply differing levels of value. In contrast, we must recognize and appreciate equally significant intellectual contributions made by various types of faculty members. Additionally, engagement of statistical practitioners as peer scientists and coleaders in collaborative research will ensure higher quality and rigor of scientific endeavors. Finally, leaders of statistical practice must be included in critical discussions around strategic development for academic and other research organizations for those institutions to achieve their missions in this modern era of data-intensive research.
High-throughput bulk and single-cell omics technologies enable comprehensive molecular profiling, yet identifying compact, biologically interpretable marker sets that distinguish cell types, conditions, or disease states remains challenging. Standard pipelines rely on univariate differential expression tests, which ignore gene-gene dependencies and nonlinear effects, while multivariate machine-learning (ML) methods often lack principled feature selection and uncertainty quantification. The Bayesian kernel machine regression (BKMR) framework offers an appealing alternative because it (a) captures nonlinear gene-outcome relationships and higher-order interactions, and (b) enables automatic relevance determination (ARD) through sparsity-inducing priors. However, we show that the traditional latent Gaussian process (GP) formulation of BKMR is inadequate for discrete outcomes (e.g., cell-type labels), leading to biased inference and unstable variable selection. We propose a copula-based Bayesian kernel machine regression (CBKMR) model that uses outcome-appropriate discrete marginals while a Gaussian copula captures kernel-induced dependence across observations. To ensure scalability to modern single-cell datasets, we further introduce a nearest-neighbor GP-based variant, NNCBKMR, which reduces computational complexity from 𝒪 N 3 to nearly linear in N . Simulation studies show that CBKMR more accurately captures nonlinear effects and yields stronger marker-selection performance than BKMR and top ensemble ML methods (e.g., random forests, XGBoost). Applications to multiple scRNA-seq datasets demonstrate that CBKMR identifies concise marker panels that align closely with expert-annotated gene signatures while providingposterior uncertainty for principled decision-making.
Background:Mendelian randomization (MR) uses genetic instruments (GI) to infer causality between exposures, like dietary intake, and health outcomes. Almost all MR of dietary intake use the full set of genome-wide significant (GWS) variants in the GI, and therefore, causal estimates are likely biased by variants that act indirectly on diet. Objective:First, we performed an assessment of the diet MR literature to evaluate the applications and approaches common in the field. Second, using conventional two-sample MR techniques with GWS variants, we evaluated whether MR could detect expected associations between six diet-health relationships supported by existing nutrition science literature. Third, we developed and tested methods for refining the GI using filtering and mediation-based approaches. Methods:Studies that performed MR of foods or beverages on any health outcome were identified in PubMed. We recorded how the GI was created, what dietary intake traits were studied, how the exclusion restriction assumption was evaluated, and what sensitivity tests were performed. We tested if conventional MR methods could detect established diet-health relationships by selecting a biomarker and disease outcome for each dietary trait (six positive controls total). This included oily fish intake on triglycerides (TG) and cardiovascular disease (CVD), alcohol intake on alanine aminotransferase (ALT) and liver cirrhosis, and white vs whole grain or brown bread on LDL cholesterol (LDL-C) and CVD. To refine the GI to better estimate the direct effect of diet by removing or accounting for the indirect effects of confounders, we tested two phenome-wide association study (PheWAS) based GI filtering approaches and a mediation approach via multivariable MR (MVMR). Causal inferences were estimated by the inverse variance weighted (IVW) and weighted median (WM) estimators and by MR-CAUSE. Results:There is a strong and rapidly expanding interest in applying MR to dietary intake exposures (178 studies identified with 76 published in 2024). Existing studies showed a wide range of methodological rigor, especially with respect to GI specificity, which raised concerns whether MR using GWS GIs can adequately evaluate diet-health relationships. In empirical testing, conventional two-sample MR methods on GWS GIs only identified the relationships between oily fish on TG and white vs whole grain or brown bread on LDL-C using the WM estimator, whereas no relationships were identified by the IVW estimator. Filtering the GI improved the ability to detect the expectation for diet-biomarker pairs (IVW, oily fish on TG: β=-0.12 [95% CI -0.18 to -0.054]; IVW, white vs whole grain or brown bread on LDL-C: β = 0.11 [95% CI 0.058 to 0.16]) but not diet-disease pairs. MR-CAUSE identified the only diet-disease association - white vs. whole grain or brown bread on CVD (γ=0.17 [95% credible interval, 0.09 to 0.25]). Furthermore, MR-CAUSE found that many diet-health relationships were impacted by confounding. We evaluated which traits contributed to confounding via the PheWAS results and found that body composition traits were the most prevalent confounders. The PheWAS output was used to prioritize traits for MVMR and rescued the expected direct effect of alcohol on ALT (β= 0.028 [95% CI 0.017 to 0.039]). Conclusion:MR studies of diet's causal role in health have flooded the literature; however, our inconsistent associations with positive and negative controls using multiple tests and filtering methods signal a need for caution. More thoughtful curation of the GI is critical to reduce confounding due to health and environmental factors when evaluating the causal effect of diet on health.
Older adults have decreased vaccine efficacy, but the adjuvanted recombinant VZV-gE zoster vaccine (RZV) is highly efficacious. We investigated memory-like innate immune responses after RZV and after the zoster vaccine live (ZVL), which is much less efficacious. RZV increased NK, monocyte, and DC activation in response to in vitro VZV-gE stimulation for up to 5 years post-vaccination, while ZVL increased only DC responses to VZV for up to 90 days. In purified monocyte and NK cell cocultures, RZV recipients showed increased responses to VZV-gE, HCMV and HSV antigenic stimulation post-vaccination. ATAC-seq analysis of purified monocytes revealed decreased accessibility in areas of the TGFβ1 gene. scRNA-seq and immunoproteomics confirmed decreased TGFβ1 transcription and translation, respectively. Exogenous supplementation and inhibition of TGFβ1 modulated in vitro monocyte responses to VZV-gE. In conclusion, RZV generated homologous (VZV-gE) and heterologous (HCMV, HSV) trained immunity in monocytes through genomic repression of the regulatory cytokine TGFβ-1. Cytokine modulation may represent a novel mechanism of generating trained immunity in myeloid cells.
Supplemental Figure 3 illustrates associations between IA-IDW summed from conception to latency in 16.1 km (10 mile) buffer and childhood acute lymphocytic leukemia for cases and controls born in Colorado between 1992 and 2019 for rural birth address, urban birth address, males, females, Hispanic mother and Non-Hispanic mother.