We propose a training-free conditional sampling method for flow matching models based on importance sampling. Because a naïve application of importance sampling suffers from weight degeneracy in high-dimensional settings, we modify and incorporate a resampling technique in sequential Monte Carlo (SMC) during intermediate stages of the generation process. To encourage generated samples to diverge along distinct trajectories, we derive a stochastic flow with adjustable noise strength to replace the deterministic flow at the intermediate stage. Our framework requires no additional training, while providing theoretical guarantees of asymptotic accuracy. Experimentally, our method significantly outperforms existing approaches on conditional sampling tasks for MNIST and CIFAR-10. We further demonstrate the applicability of our approach in higher-dimensional, multimodal settings through text-to-image generation experiments on CelebA-HQ.
Missense variants (MVs) influence clinical phenotypes, but our understanding of their phenotypic consequences remains constrained. Existing computational approaches to interpret MVs predominantly assess their pathogenicity, without considering phenotypic heterogeneity. We present a machine-learning-based method, PheMART, to predict the clinical phenotypic consequences of MVs. PheMART integrates comprehensive variant and phenotype characterizations by leveraging a robust combination of multiple resources involving protein language models, protein-protein interactions, protein domains, medical knowledge graphs and electronic health records. Exploiting contrastive learning, PheMART establishes connections between MVs and 4,179 phenotypes by jointly projecting them into a cohesive low-dimensional metric space where proximity signifies relevance. Besides substantially outperforming existing models, PheMART aids in diagnosing individuals with rare diseases by effectively pinpointing clinical diagnoses and causative MVs. As a resource to the community, we provide a database of phenotypic predictions for 5.1 million putative pathogenic amino acid alterations.
A primary advantage of neural networks lies in their feature learning characteristics, which is challenging to theoretically analyze due to the complexity of their training dynamics. We examine feature learning and its potential benefits for generalization from a statistical perspective. After reviewing the neural tangent kernel (NTK) theory and recent results in kernel regression, which address the generalization issue of sufficiently wide neural networks, we examine limitations and implications of the fixed kernel theory (as the NTK theory) and review recent theoretical advancements in feature learning. Moving beyond theories with fixed features, we consider neural networks as adaptive feature models. Finally, we propose an over-parameterized Gaussian sequence model as a prototype for the adaptive feature model to study feature learning characteristics and motivate their future analysis for neural networks.
In the FDR-controlling literature, mirror statistics offer a flexible alternative to p-value based procedures. When prior information is available, however, it is unclear how to incorporate mirror statistics in a principled way, and the standard equal split used by data-splitting methods can be inefficient. In this paper, we characterize a broader class of mirror statistics for any fixed splitting scheme and establish asymptotic FDR control under mild weak-dependence conditions using a two-stage procedure inspired by . Within this class, we derive a Bayes-optimal mirror statistic. Theoretically, we demonstrate its power advantage through analyses in the Rare/Weak signal model. Building upon this Bayes-optimal mirror statistic, we propose PRADAS (PRior-Assisted DAta Splitting) that treats split ratio as a stopping time and recasts the data-splitting as an optional stopping over a natural filtration; the optimal stopping rule is characterized by the Snell envelope and computed efficiently via a Longstaff–Schwartz regression approximation. Both simulations and real data examples demonstrate the effectiveness of our proposed framework.
Modeling the time-varying covariance structures of high-dimensional variables is critical across diverse scientific and industrial applications; however, existing approaches exhibit notable limitations in either modeling flexibility or inferential efficiency. For instance, change-point modeling fails to account for the continuous time-varying nature of covariance structures, while GARCH and stochastic volatility models suffer from over-parameterization and the risk of overfitting. To address these challenges, we propose a Bayesian factor modeling framework designed to enable simultaneous inference of both the covariance structure of a high-dimensional time series and its time-varying dynamics. The associated Expectation-Maximization (EM) algorithm not only features an exact, closed-form update for the M-step but also is easily generalizable to more complex settings, such as spatiotemporal multivariate factor analysis. We validate our method through simulation studies and real-data experiments using climate and financial datasets.
Accurate identification of affected tissues of human diseases is important for the derivation of disease etiology and the development of new treatment strategies. In this study, we develop a logistic regression-based method named DEDUCE (disease tissue detection using logistic regression) that combines genomics big data and machine learning to address this important problem. The central hypothesis is that most disease-associated genes are expressed specifically in affected tissues. DEDUCE takes advantage of newly emerged data on disease-related genes as well as tissue-specific gene expression data. The unique feature of DEDUCE is that it takes into account the strength of gene-disease associations. When we applied DEDUCE to a total of 3261, 324 gene-disease associations collected from DisGeNET covering 30,170 diseases and 21,666 genes, we identified 216 significant tissue-disease pairs composed of 120 unique diseases and 37 unique tissues. Many of them shed light on potential explanations for disease pathogenesis. The results showed great consistency with previous findings and were proven effective by empirical plots and gene set enrichment analysis. Overall, DEDUCE has shown great potential in uncovering novel pathogenesis mechanisms of complex diseases. In-depth analysis and experimental validation were required to fully understand these discovered tissue-trait associations and their enriched genes.
We introduce a class of generic spike-and-slab priors for high-dimensional linear regression with grouped variables and present a Coordinate-ascent Variational Inference (CAVI) algorithm for obtaining an optimal variational Bayes approximation. Using parameter expansion for a specific, yet comprehensive, family of slab distributions, we obtain a further gain in computational efficiency. The method can be easily extended to fitting additive models. Theoretically, we present general conditions on the generic spike-and-slab priors that enable us to derive the contraction rates for both the true posterior and the VB posterior for linear regression and additive models, of which some previous theoretical results can be viewed as special cases. Our simulation studies and real data application demonstrate that the proposed method is superior to existing methods in both variable selection and parameter estimation. Our algorithm is implemented in the R package GVSSB.
Identifying DNA binding sites remains a critical task in bioinformatics, with applications ranging from gene regulation studies to drug design. Although progress has been made in computational techniques, we still face challenges such as data complexity and prediction accuracy. In this paper, we introduce OptimDase, a new algorithm. It integrates feature encoding with optimum decision-making frameworks to improve DNA binding site prediction. OptimDase integrates multi-scale scanning and feature selection strategies, making it highly effective for both classification and regression tasks. Our experiments demonstrate that OptimDase achieves superior performance with an accuracy of 0.8943 in classification tasks and an RMSE of 0.0054 in regression tasks, outperforming existing algorithms in key evaluation metrics. These results highlight OptimDase's portability and robustness, making it an effective solution for identifying DNA binding sites and advancing the applications of drug design.
As spatially resolved transcriptomics (SRT) datasets increasingly span multiple adjacent or replicated slices, effective joint analysis across slices is needed to reconstruct tissue structures and identify consistent spatial gene expression patterns. This requires resolving spatial correspondences between slices while capturing shared transcriptomic features, two tasks that are typically addressed in isolation. Multi-slice analysis remains challenging due to physical distortions, technical variability, and batch effects. To address these challenges, we introduce Joint Alignment and Deep Embedding for multi-slice SRT (JADE), a unified computational framework that simultaneously learns spatial location-wise alignments and shared low-dimensional embeddings across tissue slices. Unlike existing methods, JADE adopts a roundtrip framework in which each iteration alternates between alignment and embedding refinement. To infer alignment, we employ attention mechanisms that dynamically assess and weight the importance of different embedding dimensions, allowing the model to focus on the most alignment-relevant features while suppressing noise. To the best of our knowledge, JADE is the first method that jointly optimizes alignment and representation learning in a shared latent space, enabling robust multi-slice integration. We demonstrate that JADE outperforms existing alignment and embedding methods across multiple evaluation metrics in the 10x Visium human dorsolateral prefrontal cortex (DLPFC) and Stereo-seq axolotl brain datasets. By bridging spatial alignment and feature integration, JADE provides a scalable and accurate solution for cross-slice analysis of SRT data.
We present a method for fitting monotone curves using cubic B-splines, which is equivalent to putting a monotonicity constraint on the coefficients. We explore different ways of enforcing this constraint and analyze their theoretical and empirical properties. We propose two algorithms for solving the spline fitting problem: one that uses standard optimization techniques and one that trains a Multi-Layer Perceptrons (MLP) generator to approximate the solutions under various settings and perturbations. The generator approach can speed up the fitting process when we need to solve the problem repeatedly, such as when constructing confidence bands using bootstrap. We evaluate our method against several existing methods, some of which do not use the monotonicity constraint, on some monotone curves with varying noise levels. We demonstrate that our method outperforms the other methods, especially in high-noise scenarios. We also apply our method to analyze the polarization-hole phenomenon during star formation in astrophysics. The source code is accessible at \texttt{\url{https://github.com/szcf-weiya/MonotoneSplines.jl}}.
OBJECTIVES:To address the limitations of existing models for research and population health applications in older adults with type 2 diabetes, we developed and validated cardiovascular disease (CVD) and heart failure risk models using linked Medicare claims and electronic health records (EHRs; 2013-2020). STUDY DESIGN AND SETTING:The study included adults aged >65 years with type 2 diabetes and ≥1 HbA1c measurement before cohort entry (defined as the date of a physician/outpatient visit). Using least absolute shrinkage and selection operator and extreme gradient boosted machine learning algorithms, we predicted 1-year risks of a composite cardiovascular event (myocardial infarction, stroke, coronary artery revascularization, or hospitalization for heart failure). Separate models were developed for patients with and without baseline CVD using claims-only and claims-EHR predictors. Models were trained on 70% of the data and validated on 30%. Model performance was evaluated using c-statistics for discrimination, scaled Brier scores, and calibration curves. We externally validated the models in Clinformatics commercial and Medicare Advantage claims data. RESULTS:There were 14,776 patients with baseline CVD (mean [SD] age: 77 [8] years) and 10,679 without baseline CVD (mean [SD] age: 74 [7] years). Claims-only models achieved a c-statistic of 0.75 and a Brier score of 0.09 in patients with baseline CVD, while in those without baseline CVD, the c-statistic was 0.73, and the Brier score was 0.01. For both subgroups, calibration intercepts were ∼0, with slopes ∼1. Claims-EHR models provided similar performance. CONCLUSION:In older adults with diabetes, our models predicted 1-year cardiovascular outcomes with good discrimination and accuracy, independently of CVD history. PLAIN LANGUAGE SUMMARY:Older adults with type 2 diabetes have a high risk of heart disease, heart failure, and death, yet it is difficult to predict who is most at risk. Most existing prediction tools are designed for use during a single clinic visit, not for large health care databases that researchers use to study treatment safety and effectiveness. In this study, we developed computer-based models using Medicare claims data and, for some models, additional information from EHRs. These models predicted the chance of having a major heart event or dying within 1 year. We created separate models for people with and without existing heart disease because their risk factors differ. Our models accurately predicted risk in both groups. Adding EHR data did not improve performance compared to using claims data alone. This means that claims-only models can still be useful for researchers studying treatments in large health care databases. These models can help identify people at higher risk, guide research on diabetes medications, and support better planning for health care resources.
OBJECTIVES:To examine the comparative risk of malignancy, venous thromboembolism (VTE), and heart failure (HF) associated with biologic/targeted synthetic disease-modifying antirheumatic drugs (b/ts DMARDs) in patients with rheumatoid arthritis (RA). METHODS:We conducted an observational cohort study using 3 US insurance claims databases: Medicare (2009-2019), MarketScan (2009-2020), and Optum's de-identified Clinformatics Data Mart Database (CDM, 2009-2022). We included adults with RA initiating abatacept (reference), tumor necrosis factor inhibitors (TNFi), rituximab, interleukin-6 inhibitors (IL-6i), or Janus kinase inhibitors (JAKi). We used an as-treated approach as the primary analysis to estimate outcome incidence. Inverse probability of treatment weighting was applied to adjust for confounding. Database-specific hazard ratios (HR) with 95% confidence intervals (CI) were estimated using Cox-proportional hazard models, then combined through random-effects meta-analysis. RESULTS:We identified 26 908 abatacept, 11 176 IL-6i, 115 437 TNFi, 14 045 JAKi, and 12 097 rituximab initiators. Weighted HR (95% CI) of malignancy was 0.73 (0.60-0.88) for IL-6i, 0.85 (0.60-1.19) for JAKi, 1.30 (1.14-1.49) for rituximab, and 0.93 (0.85-1.02) for TNFi, compared to abatacept. Weighted HR (95% CI) for VTE was 0.92 (0.62-1.35), 1.17 (0.73-1.86), 1.43 (1.50-1.95), and 1.16 (0.93-1.46), respectively. Weighted HR (95% CI) for HF was 1.00 (0.69-1.46), 1.24 (0.62-2.51), 1.52 (1.04-2.22), and 1.54 (1.22-1.94), respectively. CONCLUSION:We observed increased risks of malignancy, VTE, and HF among rituximab initiators; an increased risk of HF among TNFi initiators; and a lower risk of malignancy among IL-6i initiators, all compared to abatacept initiators. These findings should be interpreted with caution due to the potential influence of residual confounding.
It is increasingly recognized that participation bias can pose problems for genetic studies. Recently, to overcome the challenge that genetic information of nonparticipants is unavailable, it is shown that by comparing the IBD (identity by descent) shared and not-shared segments between participating relative pairs, one can estimate the genetic component underlying participation. That, however, does not directly address how to adjust estimates of heritability and genetic correlation for phenotypes correlated with participation. Here, we demonstrate a way to do so by adopting a statistical framework that separates the genetic and nongenetic correlations between participation and these phenotypes. Crucially, our method avoids making the assumption that the effect of the genetic component underlying participation is manifested entirely through these other phenotypes. Applying the method to 12 UK Biobank phenotypes, we found eight that have significant genetic correlations with participation, including body mass index, educational attainment, and smoking status. For most of these phenotypes, without adjustments, estimates of heritability and the absolute value of genetic correlation would have underestimation biases.
The knockoff filter is a recent false discovery rate (FDR) control method for high-dimensional linear models. We point out that knockoff has three key components: ranking algorithm, augmented design, and symmetric statistic, and each component admits multiple choices. By considering various combinations of the three components, we obtain a collection of variants of knockoff. All these variants guarantee finite-sample FDR control, and our goal is to compare their power. We assume a Rare and Weak signal model on regression coefficients and compare the power of different variants of knockoff by deriving explicit formulas of false positive rate and false negative rate. Our results provide new insights on how to improve power when controlling FDR at a targeted level. We also compare the power of knockoff with its propotype - a method that uses the same ranking algorithm but has access to an ideal threshold. The comparison reveals the additional price one pays by finding a data-driven threshold to control FDR.
The varying coefficient model has received broad attention from researchers as it is a powerful dimension reduction tool for non-parametric modeling. Most existing varying coefficient models fitted with polynomial spline assume equidistant knots and take the number of knots as the hyperparameter. However, imposing equidistant knots appears to be too rigid, and determining the optimal number of knots systematically is also a challenge. In this article, we deal with this challenge by utilizing polynomial splines with adaptively selected and predictor-specific knots to fit the coefficients in varying coefficient models. An efficient dynamic programming algorithm is proposed to find the optimal solution. Numerical results show that the new method can achieve significantly smaller mean squared errors for coefficients compared with the equidistant spline fitting method.
Although China has seen strong reductions in air pollution levels in the last decade, PM _2.5 concentrations still exceed the WHO Guideline several times, causing a substantial burden of mortality and morbidity. With many ‘low hanging fruits’ in terms of abatement measures already taken, further improvements will be more difficult and likely require different strategies than pursued so far. This study looks into the trends expected under current energy policies and air pollution control legislation and analyses the source contributions to ambient PM _2.5 in China, with a special focus on the megacity of Beijing. Although reductions are foreseen, China appears not yet on track to meet its long-term targets for greenhouse gas emissions nor the future national air quality standards. Going beyond current policies, we analyze effects of measures which tackle both issues and quantify health co-benefits from further decarbonization policies required to meet the national target of reaching carbon neutrality by 2060, as well as the potential for further air pollution mitigation.
A central topic in functional data analysis is how to design an optimaldecision rule, based on training samples, to classify a data function. We exploit the optimal classification problem when data functions are Gaussian processes. Sharp nonasymptotic convergence rates for minimax excess mis-classification risk are derived in both settings that data functions are fully observed and discretely observed. We explore two easily implementable classifiers based on discriminant analysis and deep neural network, respectively, which are both proven to achieve optimality in Gaussian setting. Our deepneural network classifier is new in literature which demonstrates outstanding performance even when data functions are non-Gaussian. In case of discretely observed data, we discover a novel critical sampling frequency thatgoverns the sharp convergence rates. The proposed classifiers perform favorably in finite-sample applications, as we demonstrate through comparisonswith other functional classifiers in simulations and one real data application.
Metamaterial design, encompassing both microstructure topology selection and geometric parameter optimization, constitutes a high-dimensional optimization problem, with computationally expensive and time-consuming design evaluations. Bayesian optimization (BO) offers a promising approach for black-box optimization involved in various material designs, and this work presents several advanced techniques to adapt BO to address the challenges associated with metamaterial design. First, variational autoencoders (VAEs) are employed for efficient dimensionality reduction, mapping complex, high-dimensional metamaterial microstructures into a compact latent space. Second, mutual information maximization is incorporated into the VAE to enhance the quality of the learned latent space, ensuring that the most relevant features for optimization are retained. Third, trust region-based Bayesian optimization (TuRBO) dynamically adjusts local search regions, ensuring stability and convergence in high-dimensional spaces. The proposed techniques are well incorporated with conventional Gaussian processes (GP)-based BO framework. We applied the proposed method for the design of electromagnetic metamaterial microstructures. Experimental results show that we achieve a significantly high probability of finding the ground-truth topology types and their geometric parameters, leading to high accuracy in matching the design target. Moreover, our approach demonstrates significant time efficiency compared with traditional design methods.
Genomes contain conserved non-coding sequences that perform important biological functions, such as gene regulation. We present a phylogenetic method, PhyloAcc-C, that associates nucleotide substitution rates with changes in a continuous trait of interest. The method takes as input a multiple sequence alignment of conserved elements, continuous trait data observed in extant species, and a background phylogeny and substitution process. Gibbs sampling is used to assign rate categories (background, conserved, accelerated) to lineages and explore whether the assigned rate categories are associated with increases or decreases in the rate of trait evolution. We test our method using simulations and then illustrate its application using mammalian body size and lifespan data previously analyzed with respect to protein coding genes. Like other studies, we find processes such as tumor suppression, telomere maintenance, and p53 regulation to be related to changes in longevity and body size. In addition, we also find that skeletal genes, and developmental processes, such as sprouting angiogenesis, are relevant.
Rong Chen (陈嵘)合作论文数Rutgers University16