
ABSTRACT Advancements in 3D scanning and cloud infrastructure enable the acquisition and analysis of high‐density body surface datasets. In this work, we propose a novel methodology that extends linear discriminant analysis (LDA) to Kendall's shape space for the classification of 3D objects, specifically human body shapes. Our approach adapts LDA to the non‐Euclidean geometry of shape space, generalizing Euclidean distributional assumptions and incorporating parallel transport to improve the estimation of inter‐cluster variability. Through several simulation studies, we demonstrate the effectiveness of the proposed methodology before applying it to classify female body shapes across various age ranges.
Biological age (BA) is a quantity of interest in insurance, and this paper explores the statistical foundations for estimating an individual's BA, proposing a semi-parametric Bayesian model for its estimation using multiple quantitative and qualitative phenotypes. BA is modeled explicitly as a latent biological state that influences observable phenotypes. The proposed approach employs hierarchical Bayesian models based on a deep learning (DL) model to accommodate potential non-linear relationships between phenotypes and BA. To explore the statistical features, we start by investigating the identifiability issues and deriving the Fisher information. Then, models are becoming increasingly complicated compared to the Bayesian DL model. A numerical illustration for clarity in the exposition, along with applications to both simulated and real datasets, is used to demonstrate the main insights from this study: explicit estimation of BA under a Bayesian approach avoids the identifiability problem and leads to more precise and interpretable estimation with respect to implicit approaches where BA is the implicit response of the individual's observed phenotypes.
ABSTRACT Transfer learning has attracted considerable attention in various fields, as it effectively alleviates the problem of insufficient data in individual prediction tasks. In this paper, we propose a transfer learning method for linearized maximum rank correlation estimation under the single‐index model framework (denoted as T‐lmrc). The core idea of the proposed method is to improve the fitting accuracy and estimation reliability of the target data by screening and fusing informative auxiliary datasets. To address the problem that informative auxiliary sources are difficult to pre‐determine in practical applications, we specially design a transferable source detection process to accurately identify auxiliary sources that are helpful for the target task, eliminate invalid auxiliary sources, and avoid the interference of invalid information on the estimation results. On this basis, we further strictly prove the consistency of the proposed transferable source detection procedure under mild theoretical conditions, providing a solid theoretical guarantee for the effectiveness of the method. Extensive numerical experiments demonstrate that, regardless of whether the target model is correctly specified, the proposed T‐lmrc method outperforms the conventional linearized maximum rank correlation estimator using only target data (S‐lmrc), the method that directly merges target and auxiliary source data without screening (A‐lmrc), and the AIC weighting‐based method (Aic‐lmrc). The practical application value of the proposed method is further verified by its application to housing rental data in Shanghai and Beijing, respectively.
This paper introduces two novel, communication-efficient Bayesian frameworks, FLamb and OSCLamb, for clustering high-dimensional data in federated learning (FL) settings. Traditional clustering methods often struggle with the non-IID data and privacy constraints inherent in distributed environments. Our proposed methods extend the Latent Mixture for Bayesian (Lamb) model to address these challenges, enabling robust dimension reduction and variable selection without sharing raw data. FLamb is an iterative FL algorithm where a central server aggregates sufficient statistics from distributed sites to build a global consensus model. While generally more accurate, its performance can be sensitive to the number of participating sites and communication overhead. In contrast, OSCLamb is a communication-efficient, single-round decentralized framework that uses peer-to-peer consensus averaging, significantly reducing latency and proving more robust in settings with extreme data heterogeneity. Our simulation studies demonstrate the trade-offs between the methods, with FLamb achieving higher accuracy in less heterogeneous environments and OSCLamb offering superior speed and stability under challenging conditions. We validate our approaches on two real-world high-dimensional datasets, a single-cell RNA sequencing dataset and an EEG dataset, where both methods demonstrate compelling clustering performance. A key advantage of our Bayesian approach, particularly FLamb, is the ability to provide comprehensive posterior uncertainty quantification for the cluster structure, offering more interpretable and reliable results in decentralized analyses.
Transfer learning is a powerful tool for improving estimation and prediction accuracy in high-dimensional problems, particularly in scenarios where the target data is limited. Current works, however, remain limited in their ability to robustly accommodate heavy-tailed data and outliers while retaining statistical efficiency. In this paper, we propose a robust and efficient transfer learning framework for a high-dimensional linear model based on a penalized convoluted rank estimation approach. We first develop a two-step transfer learning algorithm when the set of transferable source datasets is known, and establish non-asymptotic - and -error bounds for the resulting estimator. Our theoretical analysis shows that, when the target and sources are sufficiently close, these bounds could be improved over those of the classical penalized estimators that rely solely on target data. To address the more realistic setting in which transferable sources are unknown, we further devise a transferable source detection algorithm and prove its consistency, along with the corresponding estimation error bounds. Furthermore, we analyze the asymptotic relative efficiency of the proposed de-sparsifying estimator. Extensive simulation studies and a real-data analysis demonstrate that the proposed method is robust to heavy-tailed data and outliers while remaining efficient under light-tailed noise.
Accurate prediction of disease transmission is challenged by its dynamic nature, influenced by factors such as population density. This study introduces a novel approach that enhances predictions of disease surveillance data by integrating likelihood weighting into the integrated nested Laplace approximation (INLA) framework, specifically tailored to account for population density within a spatiotemporal Bayesian methodology. Our method prioritizes recent information for non-stationary outbreak time series online prediction by employing calibrated discounting on historical data through weight adjustments on their likelihoods. Empirical analysis of real COVID-19 daily case count data from Massachusetts counties demonstrates the effectiveness of this approach, revealing improved prediction accuracy compared to existing methods. The INLA-based method with weighted smoothing offers a promising avenue for enhancing infectious disease forecasting models, with significant potential applications in public health decision-making and resource allocation.
In studies with right-censored data involving a cure fraction, it is conventional to model the cured and non-cured groups separately. However, this modeling approach is unable to employ a single model to capture the latent relationship between covariates and failure time. Furthermore, when covariates possess an interdependent network structure, traditional cure models cannot capture the relationship between the network structure and the failure time. Therefore, this paper constructs a semi-parametric cure model incorporating a network structure by integrating the exponential family graphical model and the non-mixture cure model under right-censored data. A Sieve maximum likelihood estimation method based on Bernstein polynomials is proposed for estimating unknown parameters and identifying the network structure. The robustness of the proposed method is validated through numerical simulations conducted under various settings. Finally, the proposed method is applied to the study of head and neck squamous cell carcinoma data, revealing the potential influence of the inter-gene network structure on HNSCC progression.
We develop a distribution-valued framework for modeling, forecasting, and monitoring traffic flow counts by treating each day as a probability distribution summarized by jittered empirical quantile signatures. Inference is conducted under the 2-Wasserstein geometry, which in one dimension is isometric to the metric on quantile functions. This representation preserves the empirical distribution of within-day traffic intensities beyond mean aggregation while deliberately abstracting away from the chronological ordering of the intraday curve. We introduce Wasserstein-based distributional regression, one-step-ahead forecasting, and a Wasserstein CUSUM statistic for change-point detection and localization. Our theory provides finite-sample and asymptotic guarantees under the two-stage sampling structure of traffic data, with error bounds that separate the roles of the number of days and the within-day resolution . Simulations show competitive performance under location shifts and substantial gains under dispersion or shape changes. An analysis of publicly available interstate traffic volumes illustrates quantile-dependent covariate effects and interpretable regime changes via quantile-shift diagnostics.
We address the task of identifying anomalous observations by analyzing digits under the lens of Benford's law. Motivated by the crucial objective of providing reliable statistical analysis of customs declarations, we answer one major and still open question: How can we detect the behavior of operators who are aware of the prevalence of the Benford's pattern in the digits of regular observations and try to manipulate their data in such a way that the same pattern also holds after data fabrication? This challenge arises from the ability of highly skilled and strategically minded manipulators in key organizational positions or criminal networks to exploit statistical knowledge and evade detection. For this purpose, we write a specific contamination model for digits, obtain new relevant distributional results and derive appropriate goodness-of-fit statistics for the considered adversarial testing problem. Along our path, we also unveil the peculiar relationship between two simple conformance tests based on the distribution of the first digit. We show the empirical properties of the proposed tests through a simulation exercise and application to data from international trade transactions. Although we cannot claim that our results are able to anticipate data fabrication with certainty, they surely point to situations where more substantial controls are needed. Furthermore, our work can reinforce trust in data integrity in many critical domains where mathematically informed misconduct is suspected.
Estimating the hypervolume occupied by multivariate data is a fundamental problem in statistics and data science, with applications ranging from ecology and machine learning to multi-objective optimization and Bayesian inference. Traditional approaches rely on geometric approximations, kernel density estimation, or convex-hull constructions, which often suffer from restrictive assumptions or do not scale well in higher dimensions. We introduce a novel methodology for hypervolume estimation based on finite Gaussian mixture models. The proposed approach defines the hypervolume as a high-probability region of the fitted mixture density and estimates its volume using efficient Monte Carlo techniques, such as Latin hypercube sampling and importance sampling. An automatic, data-driven procedure selects the density threshold that determines the region over which the hypervolume is computed. Across simulations, the proposed mixture-based estimator proves broadly applicable and achieves accuracy, flexibility, and computational efficiency equal to or superior to those of existing methods. Applications to anomaly detection and ecological niche estimation illustrate the method's practical utility and interpretability in complex multivariate settings.
In this article, we introduce a matrix-variate skew normal distribution and its extended version for modeling asymmetric matrix-valued data. We investigate the main theoretical properties of these models and develop an EM-type algorithm for maximum likelihood estimation. The proposed methodology is assessed through simulation studies and illustrated with an application to historical Dow Jones dividend data, highlighting its practical relevance for modeling asymmetric dependence structures.
We evaluate four large language models (gpt-4o, gpt-4.1, o3, o4-mini) on six biomedically inspired tasks motivated by domains that often exhibit heavy-tailed or strongly skewed behavior. We compare three prompt styles-Natural (N), Mixed (M), and Constrained (C)-that differ in distributional pressure; temperature and top- are varied where supported. The study is designed to quantify how prompt-conditioned distributional pressure modulates the tail-risk profile of generated numeric outputs, rather than to assess fidelity to an external biomedical ground-truth distribution. Outliers are flagged using a stabilized MAD- rule (), and threshold robustness is summarized by the normalized area under outlier-threshold curves (nAURC). Tail behavior is assessed via complementary diagnostics (exceedances, high quantiles, Moors' kurtosis, and a secondary Pareto-equivalent exceedance index), with uncertainty quantified using pooled binomial aggregation (Wilson/Newcombe intervals), logit-based CIs for nAURC, dominance areas over threshold grids, and random-effects meta-analysis across problems. Three results emerge. (1) Prompt style is the primary driver of extremes: Under our evaluation criteria, Natural prompts yield the most conservative tail-risk profile: they consistently minimize outlier rate and nAURC and produce the lightest tails. The ordering between Mixed and Constrained varies by model, while Constrained often produces heavier tails, as expected under explicit tail instructions. (2) These patterns persist across thresholds and problems: Natural robustly dominates both Mixed and Constrained in robustness-curve dominance area, and meta-analytic gaps in outlier share are consistently positive for C-N and M-N across models. (3) Sampling (temperature/top-) modulates risk within a fixed (model, prompt) cell: effects are consistently positive for outlier incidence in our temperature contrasts and can be substantial in some settings (notably gpt-4.1 under Natural prompting), but they do not overturn the qualitative prompt ordering in pooled summaries. Overall, prompt design is the most reliable lever for controlling outliers and heavy-tail risk in LLM-generated numeric data, with sampling parameters providing additional combination-specific control. These conclusions apply most directly to the heavy-tailed biomedical-inspired settings studied here.
This paper proposes a new class of M-estimators based on an innovative objective function that provides highly robust and efficient estimates. The resulting estimator, referred to as the robust AKY estimator, is introduced as an alternative to the random coefficient regression (RCR) estimator in the presence of outliers. The results show that, for normal and clean data, the proposed robust AKY estimator performs almost as well as the RCR estimator. However, it demonstrates significantly greater resistance to outliers when applied to contaminated datasets within the random coefficients panel data (RCPD) model. To evaluate its performance, a Monte Carlo simulation study was conducted under various data-generation scenarios with different levels of outlier contamination. The results were compared with those of the non-robust RCR estimator and several existing robust M-estimators, including Huber, Hampel, Andrew, and Bisquare. In addition, the proposed robust AKY estimator was evaluated using a real insurance dataset. The findings from both the simulation study and the empirical application indicate that the robust estimators (Huber, Hampel, Andrew, Bisquare, and AKY) outperform the RCR estimator in the presence of outliers in the RCPD model. Moreover, the proposed robust AKY estimator is more efficient than the other robust M-estimators.
This study presents a framework to perform unsupervised time-event probabilistic classification using time series data of large cross-sectional dimension. These datasets often exhibit complexities such as non-linearities, structural breaks, asynchronicity, missing data, and outliers; which hampers their analysis and modeling. To address these challenges, the proposed approach integrates symbolic analysis, compositional data analysis, and Markov-switching time series modeling into a unified methodology. A Monte Carlo simulation study demonstrates the robustness of the method in various challenging scenarios. The practical applicability of the framework is illustrated through two economic case studies: (i) identifying recurrent recession and expansion regimes in the US economy using state-level data, and (ii) detecting breakpoints in high-volatility episodes in the US stock market using data from all assets in the S&P 500 index.
Causal mediation analysis is an effective method for understanding the mechanism between the exposure and the outcome, often assuming that the mediation model is consistent for each individual in the target population. In practice, however, the natural indirect effect (NIE) may vary across individuals due to their distinct characteristics. As a result, the population can be partitioned into subgroups according to the varying sizes of the NIEs. Distinguishing subgroups within the study population enables the development of more precise and targeted treatment strategies. In this paper, we propose an identifiable mixture mediation model with latent subgroups for the survival data, where the outcome follows an accelerated failure time model and the mediator is Gaussian distributed. We further employ three information criteria including the AIC, BIC, and singular BIC (sBIC) to select the number of subgroups, followed by the expectation-maximization (EM) algorithm to estimate the model parameters and NIEs. Simulation study shows that the sBIC is the most robust and efficient criterion for selecting the number of subgroups; therefore, we recommend the sBIC-EM algorithm for practical use. Lastly, we apply our algorithm to the lung cancer data and discover two latent groups with opposing NIEs.
Because of the complexity of data sets in practice, there has been much interest in developing statistical analysis tools for problems involving high-dimensional covariates. Examples of these models include partial linear additive models (PLAMs) and single-index models (SIMs). A common feature of these models is that they achieve dimension reduction to circumvent the "curse of dimensionality" while retaining the flexibility of the nonparametric regression. In the statistical and machine learning literature, fitting the additive parts in PLAM models and the link function in SIM models by nonparametric methods usually requires smooth additive components and regular link functions, and it is usually achieved using kernel methods or spline smoothing. In this work, we present a novel intrinsically interpretable combination of these two models with competitive predictive performance. We relax the smoothness assumptions and develop a nonparametric estimation procedure of the additive components and the link function that uses wavelet bases expansions adapted to non-equispaced designs. Simulation studies and real data analyses are employed to demonstrate the usefulness of the approach. Computer codes are provided as .
We propose an approach for fitting multi-cluster data using copula-based Dirichlet process mixture models (DPM). Unlike conventional finite mixture models, our framework uses Sklar's theorem to accommodate heterogeneous marginal distributions and complex inter-variable dependencies. We adopt a slice-sampling MCMC scheme to enable full Bayesian inference, which makes the posterior distribution on the number of clusters and the cluster-specific copula parameters simultaneously available. Simulation studies show that this DPM-copula approach can accurately capture and recover heavy-tailed or skewed clusters, while the Gaussian mixture model cannot. We apply our method to real NBA player-level advanced statistics from the 2021-2024 seasons, demonstrating how the model discovers distinct subgroups that exhibit different correlation structures and marginal shapes. These insights show the advantages of a flexible, copula-based approach for multi-cluster data analysis in sports analytics and beyond.
Community discovery is a central problem in the analysis of dynamic social networks. Traditional community discovery methods mainly focus on the formation and dissolution of links between nodes, and therefore often fail to capture the richer spatial structure and temporal dependency underlying network evolution. To address this limitation, we propose STEC-Net, a spatiotemporal graph neural framework for community discovery in dynamic social networks. STEC-Net integrates spatial structure and temporal dynamics within a unified embedding architecture. First, Graph Convolutional Networks (GCNs) are used to learn snapshot-level node representations from network topology. To adapt the spatial encoder to structural evolution, a GRU-based weight evolution mechanism is introduced to update the GCN parameters over time. Then, a second Gated Recurrent Unit (GRU) is employed to model temporal dependencies across snapshot embeddings and to learn spatiotemporal node representations. Finally, a Self-Organizing Map (SOM) is applied to the learned embeddings to cluster nodes and infer their community affiliations. Experiments on four types of dynamic networks show that STEC-Net consistently outperforms traditional community discovery methods in terms of purity, normalized mutual information, homogeneity, and completeness. These results demonstrate that STEC-Net can effectively uncover evolving community structures in dynamic social networks.
Image features from digital kidney biopsies may serve as novel biomarkers of kidney function in glomerular disease. Every subject's biopsy contains a different number of histologic objects, and for every object, a common set of image features is measured across all subjects. These data can be represented by a matrix for each subject with the row dimension representing the objects and the column dimension representing the image features which are computed per object. However, no existing regression method can select informative features of outcomes when feature matrices are unbalanced across subjects and features are correlated. Therefore, we developed the Random CLUstering Structured lasSO (Random CLUSSO), a bootstrap-based and L1-penalized scalar-on-matrix approach. We demonstrated through simulations that Random CLUSSO has lower bias and a higher true positive rate for identifying informative image features relative to existing approaches when features are correlated. Finally, we applied Random CLUSSO to predict kidney function using image features from kidney biopsies of subjects with glomerular disease from the Nephrotic Syndrome Study Network (NEPTUNE) and Cure Glomerulonephropathy (CureGN) studies.
Source identification is an inferential problem that evaluates the likelihood of opposing propositions regarding the origin of items. The specific source problem refers to a situation where the researcher aims to assess if a particular source originated the items or if they originated from an alternative, unknown source. Score-based likelihood ratios offer an alternative method to assess the relative likelihood of both propositions when formulating a probabilistic model is challenging or infeasible, as in the case of pattern evidence in forensic science. However, the lack of available data and the dependence structure created by the current procedure for generating learning instances can lead to reduced performance of score likelihood ratio systems. To address these issues, we propose a resampling plan that creates synthetic items to generate learning instances under the specific source problem. Simulation results show that our approach achieves a high level of agreement with an ideal scenario where data is not a limitation and learning instances are independent. We also present two applications in forensic sciences-handwriting and glass analysis-illustrating our approach with both a score-based and a machine learning-based score likelihood ratio system. These applications show that our method may outperform current alternatives in the literature.