Motivated by investigating how daily temperature affects soybean yield, this article proposes a simultaneous functional quantile regression (FQR) model with a unique twist-a locally sparse bivariate slope function that is indexed by both quantile and time, linked to a functional predictor. The slope function's local sparsity means it holds non-zero values only in certain segments of its domain, remaining zero elsewhere. These zero-slope regions, which vary by quantile, indicate times when the functional predictor has no discernible impact on the response variable. This feature boosts the model's interpretability. Unlike traditional FQR models, which fit one quantile at a time and have several limitations, our proposed method can handle a spectrum of quantiles simultaneously. We tested the new approach through simulation studies, demonstrating its clear advantages over standard techniques. To validate its practical use, we applied the method to soybean yield data, pinpointing the time periods when daily temperature doesn't affect yield. This insight could be crucial for agricultural planning and crop management. Supplementary materials for this article are available online.
We propose a model for directed, real-valued networks in which the dependence between the two directed edges within each dyad is captured by a one-parameter Frank copula selected on the basis of empirical dependence structure and predictive likelihood, while the marginal mean of each edge is decomposed into additive sender and receiver effects. The model is closed-form, tractable under No-U-Turn Sampling (NUTS) via NumPyro and JAX, and designed for replicated networks (multiple networks observed on a common node set, each belonging to one of G comparison groups). We derive several theoretical properties that clarify the interpretation and estimabil-ity of the dyadic reciprocity parameter. The copula parameter ρ admits a closed-form, margin-invariant interpretation as a monotone function of Kendall’s τ via the Debye function. Under a standard normalisation of the sender effects, the full parameter vector is identifiable whenever N ≥ 3. With fixed N and growing number of replicated networks S, the MLE of ρ is √S-consistent and asymptotically normal under a simplified model serving as a parametric benchmark. The model is applied to effective connectivity networks in the default mode network from resting-state fMRI for 112 participants across three diagnostic groups: Alzheimer’s disease, mild cognitive impairment, and cognitively normal. Posterior estimates of ρ are strongly negative for all three groups (ρ ( g ) ≈ −3.7 to −3.8, 95% credible intervals well below zero), indicating that within each dyad, the two directed edges are negatively associated: participants with a stronger effective connection Aij tend to have a weaker reverse connectionAji, consistent with a hierarchical regulation structure in the DMN. The three groups have nearly identical posterior distributions for ρ; the posterior probability of the strict ordering ρ(NL) < ρ(MCI) < ρ(AD) is 0.20, close to the uniform chance level of 1/6.
Recent advances in scientific machine learning (SciML) have enabled neural operators (NOs) to serve as powerful surrogates for modeling the dynamic evolution of physical systems governed by partial differential equations (PDEs). While existing approaches focus primarily on learning simulations from the target PDE, they often overlook more fundamental physical principles underlying these equations. Inspired by how numerical solvers are compatible with simulations of different settings of PDEs, we propose a multiphysics training framework that jointly learns from both the original PDEs and their simplified basic forms. Our framework enhances data efficiency, reduces predictive errors, and improves out-of-distribution (OOD) generalization, particularly in scenarios involving shifts of physical parameters and synthetic-to-real transfer. Our method is architecture-agnostic and demonstrates consistent improvements in normalized root mean square error (nRMSE) across a wide range of 1D/2D/3D PDE problems. Through extensive experiments, we show that explicit incorporation of fundamental physics knowledge significantly strengthens the generalization ability of neural operators. We promise to release models and data upon acceptance.
Accurate survival prediction in kidney transplantation is critical yet challenging due to the complex interplay between functional biomarkers and patient characteristics under censoring. To address this, we propose a functional censored quantile neural network (FunCQNet), a novel framework that integrates deep neural networks with a censoring-adjusted sequential quantile loss to approximate interaction-dependent coefficient functions. We further introduce a conformal inference approach to rigorously assess the significance of scalar-functional interactions, ensuring interpretability alongside predictive power. Extensive simulations demonstrate that FunCQNet robustly recovers functional effects under varying noise and censoring levels. When applied to kidney transplant data, the model yields precise multi-quantile predictions and reveals clinically significant, age-dependent interaction patterns between donor type and recipient survival.
The relationship between temperature and electricity consumption may vary with household wealth, highlighting a critical issue of inequality in energy usage. To address this, we propose a Generalized Functional Additive Nonlinear Model with Multimodal Interaction effects (FANMI), which leverages Functional Principal Component Analysis (FPCA) to capture the complex interplay between functional and scalar covariates—referred to here as multimodal interaction. Our estimation procedure combines quasi-likelihood methods with B-spline approximation to efficiently fit the FANMI model. We establish the optimal convergence rate and derive the asymptotic normality of the resulting estimators. Furthermore, we develop two hypothesis testing procedures: one to evaluate the overall goodness-of-fit of the FANMI model, and another to test whether certain bivariate nonparametric interaction functions can be simplified to univariate forms. The asymptotic distributions of the proposed test statistics are also derived. Extensive simulation studies are conducted to assess the finite-sample performance of both the estimation and testing methods. Finally, we illustrate the practical utility of FANMI by analyzing the joint effects of temperature and GDP on electricity consumption and the electricity Gini coefficient.
In the real estate market, property prices are recorded only at the time of sale, leading to highly sparse transaction data. Recovering house price trajectories from such sparse data is therefore essential for accurate property valuation. This study models the latent price trajectory as functional data and applies functional principal component analysis (FPCA) to recover the underlying price trajectory. Real estate transaction data from the Kitsilano neighborhood in Vancouver is collected to validate the FPCA model. The first three eigenfunctions are employed to reconstruct the price trajectories, along with their corresponding confidence intervals. A 10-fold cross-validation procedure is conducted to assess the accuracy of the recovered trajectories. The results demonstrate that FPCA can reconstruct house price trajectories with high precision.
Reliable prognoses for progressive illnesses often require analysis of a series of medical images rather than a single snapshot. Classical joint models and many recent deep-learning approaches either discard censored observations, ignore temporal dependencies between adjacent scans or offer limited insight into the image regions that drive risk. We introduce SurLonMamba, a joint survival model that couples a computationally efficient selective state-space model for sequential image processing and survival prediction. Its vision encoder extracts spatial features from each image scan; the sequence encoder propagates these representations through time, capturing long-range progression patterns at linear computational complexity; and the survival module converts the resulting sequence summary into a hazard score trained by minimizing the negative log partial likelihood, thereby leveraging all censored data. The architecture naturally supports participant-level interpretation through occlusion sensitivity and survival risk prediction. Extensive simulations and experiments on the Alzheimer’s Disease Neuroimaging Initiative (ADNI) cohort show that SurLonMamba exhibits strong predictive performance, while highlighting brain regions consistently linked to conversion to dementia.
Motivated by an application to study the impact of temperature, precipitation and irrigation on soybean yield, this article proposes a sparse semiparametric functional quantile model. The model is called “sparse” because the functional coefficients are only nonzero in the local time region where the functional covariates have significant effects on the response under different quantile levels. To tackle the computational and theoretical challenges in optimizing the quantile loss function added with a concave penalty, we develop a novel convolution-smoothing-based locally sparse estimation (CLoSE) method, to do three tasks in one step, including selecting significant functional covariates, identifying the nonzero region of functional coefficients to enhance the interpretability of the model and estimating the functional coefficients. We establish the functional oracle properties and simultaneous confidence bands for the estimated functional coefficients, along with the asymptotic normality for the estimated parameters. In addition, because it is difficult to estimate the conditional density function given the scalar and functional covariates, we propose the split wild bootstrap method to construct the confidence interval of the estimated parameters and simultaneous confidence band for the functional coefficients. We also establish the consistency of the split wild bootstrap method. The finite sample performance of the proposed CLoSE method is assessed with simulation studies. The proposed model and estimation procedure are also illustrated by identifying the active time regions when the daily temperature influences the soybean yield. Supplementary materials accompanying this paper appear online.
Abstract Integrating longitudinal data with survival models is a prevalent strategy for dynamic survival risk prediction while accounting for subjects' longitudinally observed variables. However, existing methods primarily focus on scalar longitudinal data and seldom tackle the complexities associated with high‐dimensional longitudinal imaging data. This article introduces a new approach that effectively incorporates longitudinal medical images as features for dynamic risk prediction. This approach ensures interpretability, enhances computational efficiency, and performs robustly even with small datasets. Our method achieves high prediction accuracy, as validated through extensive simulation studies and a real‐world application to Alzheimer's disease data. To the best of our knowledge, this is the first attempt to use longitudinal medical images for predicting dynamic survival risk.
Sparse functional data arise when measurements are observed infrequently and at irregular time points for each subject, often in the presence of measurement error. These characteristics introduce additional challenges for functional principal component analysis. In this paper, we propose a new approach for extracting functional principal components from such data by combining basis expansion with maximum likelihood estimation. Orthogonality of the estimated eigenfunctions is preserved throughout the optimization using modified Gram-Schmidt orthonormalization. An information criterion is proposed to select both the optimal number of basis functions and the rank of the covariance structure. Principal component scores are subsequently estimated via conditional expectation, enabling accurate reconstruction of the underlying functional trajectories across the full domain despite sparse observations. Simulation studies demonstrate the effectiveness of the proposed method and show that it performs favorably compared with existing approaches. Its practical utility is illustrated through applications to CD4 cell count data from the Multicenter AIDS Cohort Study and somatic cell count data from Irish research dairy cattle. Supplementary materials, including technical details, additional simulation results, and the R package mGSFPCA, are available online.
Sensor devices have been increasingly used in engineering and health studies recently, and the captured multi-dimensional activity and vital sign signals can be studied in association with health outcomes to inform public health. The common approach is the scalar-on-function regression model, in which health outcomes are the scalar responses while high-dimensional sensor signals are the functional covariates, but how to effectively interpret results becomes difficult. In this study, we propose a new Functional Adaptive Double-Sparsity (FadDoS) estimator based on functional regularization of sparse group lasso with multiple functional predictors, which can achieve global sparsity via functional variable selection and local sparsity via zero-subinterval identification within coefficient functions. We prove that the FadDoS estimator converges at a bounded rate and satisfies the oracle property under mild conditions. Extensive simulation studies confirm the theoretical properties and exhibit excellent performances compared to existing approaches. Application to a Kinect sensor study that utilized an advanced motion sensing device tracking human multiple joint movements and conducted among community-dwelling elderly demonstrates how the FadDoS estimator can effectively characterize the detailed association between joint movements and physical health assessments. The proposed method is not only effective in Kinect sensor analysis but also applicable to broader fields, where multi-dimensional sensor signals are collected simultaneously, to expand the use of sensor devices in health studies and facilitate sensor data analysis.
To address the challenge of utilizing patient data from other organ transplant centers (source cohorts) to improve survival time estimation and inference for a target center (target cohort) with limited samples and strict data-sharing privacy constraints, we propose the Similarity-Informed Transfer Learning (SITL) method. This approach estimates multivariate functional censored quantile regression by flexibly leveraging information from each source cohort based on its similarity to the target cohort. Furthermore, the method is adaptable to continuously updated real-time data. We establish the asymptotic properties of the estimators obtained using the SITL method, demonstrating improved convergence rates. Additionally, we develop an enhanced approach that combines the SITL method with a resampling technique to construct more accurate confidence intervals for functional coefficients, backed by theoretical guarantees. Extensive simulation studies and an application to kidney transplant data illustrate the significant advantages of the SITL method. Compared to methods that rely solely on the target cohort or indiscriminately pool data across source and target cohorts, the SITL method substantially improves both estimation and inference performance.
Alzheimer's disease (AD) is a progressive neurodegenerative disorder that leads to memory loss, cognitive decline, and behavioral changes, without a known cure. Neuroimages are often collected alongside the covariates at baseline to forecast the prognosis of the patients. Identifying regions of interest within the neuroimages associated with disease progression is thus of significant clinical importance. One major complication in such analysis is that the domain of the brain area in neuroimages is irregular. Another complication is that the time to AD is interval-censored, as the event can only be observed between two revisit time points. To address these complications, we propose to model the imaging predictors via bivariate splines over triangulation and incorporate the imaging predictors in a flexible class of semiparametric transformation models. The regions of interest can then be identified by maximizing a penalized likelihood. A computationally efficient expectation-maximization algorithm is devised for parameter estimation. An extensive simulation study is conducted to evaluate the finite-sample performance of the proposed method. An illustration with the AD Neuroimaging Initiative dataset is provided.
Structural magnetic resonance imaging (MRI) is one of the primary predictors of Alzheimer's disease risk, enabling the identification of patients with similar risk profiles for precision medicine treatment. Motivated by the need for flexible modeling in AD research, we propose a latent-class model that addresses the heterogeneity within study populations. This model allows for varying relationships between covariates and survival outcomes, accommodating the dynamics of AD progression. The imaging predictors are characterized by bivariate splines over triangulation to accommodate the irregular domain of the brain images. We develop a generalized expectation-maximization (EM) algorithm that combines the computational methods for logistic regression and penalized proportional hazards models to implement the proposed approach. We demonstrate the advantages of the proposed method through extensive simulation studies and provide an application to the Alzheimer's Disease Neuroimaging Initiative (ADNI) study, which helps to reveal different subtypes or stages of the disease process in Alzheimer's Disease.
The generalized linear model (GLM) is a popular modeling choice for pricing non-life insurance policies. However, high-cardinality categorical insurance data presents significant challenges for these GLM rate-making models. Additionally, insurance regulators often require rating territories, which are clusters of insurance policies' geographic locations for setting insurance rates, to meet certain standards. For instance, (1) the credibility standard ensures that the number of policies in a territory is large enough to be credible and representative, (2) the contiguity standard requires the locations in each territory to be geographically adjacent to promote a logical and practical spatial grouping, and (3) the cardinality standard specifies an acceptable range for the number of territories in a geographic area. To address these challenges, this article proposes a nested GLM framework for non-life insurance rate-making applications. In this framework, neural network models with categorical embedding layers are constructed to model the residual deviance from simple GLMs, using high-cardinality categorical variables as input. Low-dimensionly features extracted from the neural network model effectively translate categorical variables into meaningful numerical representations, capturing their effects on the initial model's residuals. The features corresponding to the location-related variable are further converted into a contiguous territory rating variable via spatially constrained clustering models. By incorporating outcomes from these models, the nested GLM not only satisfies regulatory requirements but also enhances the model's predictive power, while maintaining the interpretability from the (generalized) linear form. The construction of a nested Poisson GLM is presented in this article. Its performance is demonstrated using a real-life Brazil auto insurance data to model claim frequency.
We develop a robust Bayesian functional principal component analysis (RB-FPCA) method that utilizes the skew elliptical class of distributions to model functional data, which are observed over a continuous domain. This approach effectively captures the primary sources of variation among curves, even in the presence of outliers, and provides a more robust and accurate estimation of the covariance function and principal components. The proposed method can also handle sparse functional data, where only a few observations per curve are available. We employ annealed sequential Monte Carlo for posterior inference, which offers several advantages over conventional Markov chain Monte Carlo algorithms. To evaluate the performance of our proposed model, we conduct simulation studies, comparing it with well-known frequentist and conventional Bayesian methods. The results show that our method outperforms existing approaches in the presence of outliers and performs competitively in outlier-free datasets. Finally, we demonstrate the effectiveness of our method by applying it to environmental and biological data to identify outlying functional observations. The implementation of our proposed method and applications are available at https://github.com/SFU-Stat-ML/RBFPCA .
Sample size determination in Bayesian randomized phase II trial design often relies on computationally intensive search methods, presenting challenges in terms of feasibility and efficiency. We propose a novel approach that greatly reduces the computing time of sample size calculations for Bayesian trial designs. Our approach innovatively connects group sequential design with Bayesian trial design and leverages the proportional relationship between sample size and the squared drift parameter. This results in a faster algorithm. By employing regression analysis, our method can accurately pinpoint the required sample size with significantly reduced computational burden. Through theoretical justification and extensive numerical evaluations, we validate our approach and illustrate its efficiency across a wide range of common trial scenarios, including binary endpoint with Beta-Binomial model, normal endpoint, binary/ordinal endpoint under Bayesian generalized linear model, and survival endpoints under Bayesian piecewise exponential models. To facilitate the use of our methods, we create an R package named "BayesSize" on GitHub.
Statistical models are an essential tool to model, forecast, and understand the hydrological processes in watersheds. In particular, the understanding of time lags associated with the delay between rainfall occurrence and subsequent changes in streamflow is of high practical importance. Since water can take a variety of flow paths to generate streamflow, a series of distinct runoff pulses may combine to create the observed streamflow time series. Current state-of-the-art models are not able to sufficiently confront the problem complexity with interpretable parametrization, thus preventing novel insights about the dynamics of distinct flow paths from being formed. The proposed Gaussian Sliding Windows Regression Model targets this problem by combining the concept of multiple windows sliding along the time axis with multiple linear regression. The window kernels, which indicate the weights applied to different time lags, are implemented via Gaussian-shaped kernels. As a result, straightforward process inference can be achieved since each window can represent one flow path. Experiments on simulated and real-world scenarios underline that the proposed model achieves accurate parameter estimates and competitive predictive performance, while fostering explainable and interpretable hydrological modelling.