
We introduce a criterion, named the mean L1 -distance (ML1D) criterion, to construct uniform designs in experiments with mixtures. This criterion allows for a flexible number of design points and produces a more uniform pattern within the experimental region, in terms of representative points of the uniform distribution across that region. We further explore the optimal Scheff & eacute;-type simplex-lattice designs under the ML1D criterion and show that a connection exists between uniform mixture designs and optimal Scheff & eacute;-type simplex-lattice designs. An efficient algorithm is proposed to generate uniform designs under the ML1D criterion. Simulations and applications highlight the advantages of the proposed designs, supporting their use for modelling and prediction in mixture experiments. Our method combines model-based and uniform design principles to enable flexible and efficient mixture experiments.
A fundamental concept in information theory is information entropy, which is used to quantify the degree of uncertainty associated with random variables. Building upon this, relative entropy serves as a measure of the discrepancy between two probability distributions and has been widely studied in statistics. The function of relative entropy in the field of parameter estimation has been extensively investigated. This paper extends existing research by deriving minimum mean relative entropy estimators for the parameters of the Gamma, Laplace, and Rayleigh distributions. Furthermore, we introduce the residual mean relative entropy, a novel measure based on mean relative entropy, and apply it to model comparison. To estimate this measure, we employ kernel density estimation and bootstrap methods. Simulation experiments and empirical analysis demonstrate that the proposed residual mean relative entropy provides an effective new criterion for model comparison.
The comparative analysis of single-cell RNA sequencing (scRNA-seq) datasets across various biological conditions, technology platforms and tissue types reveals crucial insights into cellular heterogeneity and tissue architecture. Recent advances in high-resolution technologies have enabled the profiling of gene expression at the single-cell level, yet inherent heterogeneities between platforms and differences in cell type composition make data integration challenging. We present the FIRM R package, which offers a streamlined workflow for the flexible integration of scRNA-seq data using a re-scaling algorithm that accounts for the effects of cell type composition. FIRM achieves accurate mixing of shared cell type identities and superior preservation of the original structure without overcorrection, generating robust integrated datasets for downstream exploration and analysis.
This paper investigates the asymptotic properties of the Kaplan–Meier and hazard estimators for censored survival time data. We conduct this analysis under the assumption of m-widely acceptable (m-WA) dependence, a generalized form of weak correlation. Using the Fuk–Nagaev inequality, we establish strong consistency and strong representation results for these estimators. Our findings show that the rate of strong consistency is near [Formula: see text] and the remainder term in the strong representation is of the same order. These results generalize and extend existing work for other types of dependent data, such as linearly extended negative quadrant-dependent (LENQD) and extended negative dependent (END) sequences, thereby broadening the theoretical foundation for these widely used statistical tools.
In the era of big data, ensuring data privacy has emerged as a significant challenge in large-scale data applications. Currently, differential privacy is one of the most promising privacy preserving algorithms, as it provides an explicit measure of the degree of privacy protection. Although the development of differential privacy is still in its early stages within the field of statistics, it is expected to play an integral role in future research. Motivated by this, this paper first provides a review of the development of privacy models, including the detailed introduction and interpretation of the differential privacy framework. In addition, we present the applications of several commonly used noise mechanisms and elaborate on the parallel and sequential composition theorems in differential privacy. Finally, this paper also discusses potential future research on differential privacy for online data analysis and statistical inference.
Anomaly detection in sequence data is widely applicable across various domains and has significant commercial value to the financial industry. This paper studies its utility as a means of controlling credit card delinquency risk. Transactions that deviate from the regular data sequence are a common precursor of payment difficulty. Current detection methods, however, do not effectively identify abnormal transactions from such data, making it difficult to control the overdue payment risk. Therefore, in this paper, we propose a Multiple Instance Learning-based Anomaly Detection (MILAD) method with well designed learning networks to address this problem. Comparing the performance of the MILAD and Deep Autoencoding Gaussian Mixture Model (DAGMM) method, which is currently the most commonly used unsupervised deep learning algorithm for credit card risk control, we observe that the proposed MILAD is able to effectively control the overdue risk by leveraging both transaction and payment information.
Diagnostic testing typically involves two types of classification. Binary tests separate individuals into diseased or non-diseased groups, while multi-class methods, like tree or umbrella ordering, compare one class's biomarker levels to those of other classes. Kullback-Leibler divergence (KL), which measures the difference between two distributions, has been considered a valuable index for assessing the diagnostic performance of biomarkers. In this work, we derive and propose the total rule-in and rule-out Kullback-Leibler divergence (TTKL(c)) as a measure of accuracy, obtained by dichotomizing a continuous biomarker, and as an optimization criterion for cut-off point selection under tree or umbrella ordering. We have established a connection between the proposed TTKL(c) measure and the extended Youden index, which is the most used criterion for cut-off point selection. Additionally, we present both theoretical and numerical derivations for scenarios involving a single cut-off point under extended tree ordering. Graphically, KL divergence is represented through the information graph. Using simulation methods, we conducted a power study to compare the performance of our proposed methods with the extended Youden index under tree ordering, as well as the accuracy of optimal cut-off selection. This analysis provides insights into the effectiveness of TTKL(c) as a robust criterion for selecting cut-off points in multi-class diagnostic settings. A comprehensive data analysis of lung cancer data illustrates the proposed applications.
Burglary, as a prevalent and detrimental crime type, poses a major threat to public safety and property security. Accurate prediction of burglary occurrence is therefore critical. Although deep learning has achieved notable progress in crime prediction, the influence of varying urban spatial characteristics on predictive performance and resource efficiency remains underexplored. This study analyzes burglary prediction in eight representative cities from China, the United States, and Canada, uniformly employing the model based on convolutional neural networks and long short-term memory networks (CNN-LSTM) under a consistent spatio-temporal scale. A comprehensive evaluation framework based on precision, hit rate, prediction accuracy index (PAI), and prediction effectiveness index (PEI) was established, focussing on the validity of PEI and variations in optimal prediction thresholds across cities. Furthermore, standard deviation ellipses, Moran index, and kernel density analysis were applied to quantify spatial characteristics and explore their associations with PEI and resource allocation. The results indicate that cities with more concentrated spatial distributions and stronger spatial autocorrelation exhibit superior predictive efficiency and resource utilization. This study enriches the analytical scope of crime prediction efficiency and supports a shift from 'accuracy-oriented' to 'efficiency-oriented' modelling for intelligent allocation of urban public safety resources.
Products often operate in dynamic environments, and field failure data is frequently heavily censored, posing significant challenges in the assessment of product reliability. To enhance the accuracy of field reliability predictions, we introduce a novel joint modelling approach that combines accelerated life tests (ALT) and field failure data. We capture the stochastic influence of dynamic environmental factors on product aging using an exponential dispersion process and present a methodology for jointly modelling ALT and field failure data. Our approach is grounded in the cumulative exposure principle, providing a clear and intuitive physical interpretation. We offer point and interval estimates for model parameters and reliability using maximum likelihood and Bayesian methods, validating their effectiveness through comprehensive simulation studies. Finally, we demonstrate the performance and practical application of our proposed joint model through the analysis of a real dataset.
This article introduces a novel approach to integrating correlation matrix information from training samples to construct a classification rule for testing samples. Traditional discriminant analysis methods that rely solely on mean vectors tend to perform poorly when the mean of the training samples is not indicative of the testing samples. To address this limitation, we propose a new discriminant analysis method called Correlation-matrix driven Discriminant Analysis (CorrDA). By considering the correlation matrices of different classes in the training samples, we can capture the unique patterns among the classes. CorrDA utilizes the Bayes classifier and mixture models to effectively incorporate the correlation matrix information derived from the training samples, thereby improving the discriminant analysis performance on the testing data. Through the analysis of COVID-19 datasets and extensive simulation studies, we provide empirical evidence demonstrating the superior performance of CorrDA.
This study quantifies risk spillover effects from multi-dimensional energy markets to China's Guangdong carbon market by constructing an EGARCH-CQR-based CoVaR model, which integrates the Exponential Generalized Autoregressive Conditional Heteroskedasticity (EGARCH) framework to capture volatility leverage effects and clustering, combined with quantile regression for precise characterization of cross-market tail dependencies. Empirical analysis reveals significant structural heterogeneity in energy-to-carbon risk spillovers: traditional energy markets, such as oil and coke, exhibit ‘high-intensity, high-volatility’ shock patterns that transmit abrupt short-term risks during global crises like the COVID-19 outbreak, whereas new energy markets, including new energy vehicles and wind power, demonstrate ‘low-intensity, persistent’ spillover dynamics reflecting stronger market resilience. Additionally, China's ‘Dual Carbon’ policy reinforcement is identified as a critical policy transmission channel that significantly intensifies risk linkages between high-carbon energy sectors and the carbon market, with model validation confirming the robustness and coverage capability of the proposed GARCH-CQR-CoVaR framework.
Instrumental variable (IV) methods are widely used to address unmeasured confoundings in structural equation models. In this paper, we focus on the settings where a possibly large number of instruments and a weak correlation between the instruments and the endogenous variable exist. Specifically, we propose a novel two-stage least squares (2SLS) model averaging approach to estimate the coefficient of an endogenous variable. Differing from existing literature, our model averaging estimation allows multiple exogenous variables to be included in both stages simultaneously. Theoretically, we study the consistency and asymptotic distributions of the estimated weights and the proposed model averaging estimator. Importantly, we discover that the proposed model averaging estimator produces an asymptotic bias when the endogenous variable and exogenous variables are correlated. Then, we construct a debiased estimator and establish its consistency and asymptotic normality to make statistical inference. Furthermore, we present an equivalent interpretation of the debiased estimator from another construction. Finally, numerical simulations and a real data analysis are conducted to illustrate our proposal.
Accurately assessing reliability and predicting the life of products is an important part of Prediction and Prognostics and Health Management. According to historical degradation data of products, the degradation model of two-stage nonlinear phase-dependent Wiener process is proposed to evaluate the dependability and forecast the lifetime of individual products. Firstly, a two-stage phase-dependent degradation model based on a nonlinear Wiener process was developed to characterize nonlinear nature of the degradation of product performance by simultaneously taking the correlation of degradation stages and the degradation heterogeneity of products into account; Secondly, the model parameters are estimated using EM algorithm based on historical degradation data of products; Finally, simulation studies are used to demonstrate the superiority of the approach for assessing offline reliability, and the validity of the method proposed in this paper is verified through an example of high-voltage pulse capacitors. The results show that compared with the degradation model based on phase-independent nonlinear Wiener processes, the two-stage nonlinear phase-dependent Wiener process degradation model can better capture actual degradation paths of products and provide a more reasonable assessment of lifetime reliability of products.
Accurate prediction of the remaining useful life (RUL) of lithium batteries is essential for ensuring efficient equipment maintenance and energy management, particularly as these batteries serve as a core driver in the new energy technology revolution. While deep learning models such as Convolutional Neural Networks (CNNs), Long Short-Term Memory Networks (LSTMs), and their variants have demonstrated significant success in RUL prediction, they often face challenges related to inadequate modelling of long-term dependencies in complex degradation data. To overcome these limitations, this paper proposes a novel hybrid architecture that integrates the Kolmogorov-Arnold Network (KAN) with an Extended Long Short-Term Memory Network (xLSTM). The KAN component enhances high-dimensional function approximation and improves parameter efficiency by substituting traditional linear weights with B-spline-parameterized univariate functions. Meanwhile, the xLSTM introduces exponential gating mechanisms and covariance update rules to more effectively capture high-order long-term dependencies. Experimental results on the NASA lithium battery aging dataset demonstrate that the proposed KAN-xLSTM model significantly outperforms CNN, LSTM, and xLSTM models in prediction accuracy, particularly for battery groups with large capacity fluctuations.
Antimicrobial resistance (AMR) has already been identified as an urgent issue affecting global health. Many research interests lie in monitoring the change in AMR in both human and animal populations. Moreover, it is important to study the correlation in AMR between the two populations. In this study, we develop a hierarchical latent class mixture model for the detection of linear changes and correlation analysis across populations with antimicrobial resistance. We propose Bayesian methods to estimate the unknown parameters in the proposed model. The simulation study is conducted to evaluate the empirical performance of the proposed method. Finally, we employ the proposed model and methodology to analyze the datasets obtained from the National Antimicrobial Resistance Monitoring System (NARMS).
Quantile regression is essential for analyzing the relationship between conditional quantiles of independent and dependent variables, widely applied in economics, education, social science, and beyond. However, traditional methods often struggle with nonlinearity, high dimensionality, and other complexities in data. To address these challenges, this paper innovatively integrates the cascade architecture into quantile regression forests, proposing the deep quantile forest method. Unlike existing deep-quantile regression estimators, our approach requires fewer hyperparameters and less training data while offering better performance and enhanced interpretability. Extensive numerical simulations and real-world data experiments demonstrate that our proposed method outperforms competing methods, showcasing its effectiveness and robustness in handling the complex structures of high-dimensional data.
Space-filling designs with superior low-dimensional properties are highly required in computer experiments. Strong orthogonal arrays (SOAs) represent a class of such designs that outperform ordinary orthogonal arrays in their stratification properties within low dimensions. Nevertheless, current methods for constructing high-strength SOAs are rare, and they typically rely on regular designs, thereby limiting the number of runs in the final arrays to prime powers. This study presents new construction methods for three types of SOAs: SOAs of strength three, column-orthogonal SOAs (OSOAs) of strength three and three minus. The resulting designs have run sizes of twice an odd prime power without replications, filling the gaps in run sizes left by existing constructions. The projection properties of Addelman-Kempthorne orthogonal arrays are instrumental in the development of these construction methods.
Tukey's boxplot is widely used for outlier detection; however, its classic fixed-fence rule tends to flag an excessive number of outliers as the sample size grows. To address this, we introduce two new R packages, ChauBoxplot and AdaptiveBoxplot, which implement more robust and statistically principled outlier detection methods. We illustrate their advantages and practical implications through comprehensive simulation studies and a real-world analysis of provincial university admission rates from China's National College Entrance Examination. Based on these findings, we provide practical guidance to help practitioners select appropriate boxplot methods, achieving a balance between interpretability and statistical reliability.
Traditional multivariate parametric control charts often perform inadequately in detecting shifts in the covariance matrix when the data deviate from normality. In this paper, we propose a multivariate nonparametric exponentially weighted moving average (SGLGEWMA) control chart, incorporating a Sparse Group Lasso penalty, which is capable of detecting shifts in the covariance matrix across a wide range of data types, including discrete, continuous, and mixed distributions. The proposed approach projects multivariate data into a Euclidean space and then computes an approximate Alt's likelihood ratio, regularized via the Sparse Group Lasso. The resulting EWMA statistic monitors process shifts. Monte Carlo simulations demonstrate that SGLGEWMA outperforms both the Lasso-based LGShewhart and the Ridge-based RGEWMA control charts under various distributions, with enhanced efficacy in high-dimensional scenarios. Sensitivity analyses are performed on the tuning parameters ( $ \lambda _1 $ lambda 1, $ \lambda _2 $ lambda 2) and smoothing parameter rho, to evaluate their impact on monitoring performance. Additionally, a simulation study and an illustrative example involving covariance monitoring in wafer semiconductor manufacturing are presented to demonstrate the practical application of the proposed chart. Empirical results confirm that the proposed control chart promptly identifies abnormal fluctuations and issues timely alerts, highlighting both its theoretical significance and practical utility.
The integrated nested Laplace approximation (INLA) algorithm provides a computationally efficient approach for approximate Bayesian inference, overcoming the limitations of traditional Markov chain Monte Carlo (MCMC) methods. This paper reviews INLA algorithm and provides a systematic review of six key books that explore the theoretical foundations, practical implementations, and diverse applications of INLA. These six books cover spatial and spatio-temporal modelling, general Bayesian inference, SPDE-based spatial analysis, geospatial health data, regression modelling, and dynamic time series. In addition, these books highlight the versatility of INLA method in handling complex models while maintaining high computational efficiency. This paper begins with an introduction to the INLA method and algorithm, followed by a systematic review of six key publications in the field.