
The non-identifiability of the dependent competing risks model has received a lot of attention as it often renders empirical analysis uninformative. This paper contributes a novel route to establish identifiability under a Gumbel copula model, where the main risk of interest follows the semiparametric Cox proportional hazards model and another dependent risk or dependent censoring is left unspecified. It is shown that the Gumbel copula parameter is identifiable, provided that there exist two covariates that only affects the main risk.
Normal approximations justified by the central limit theorem constitute a cornerstone of statistical inference, particularly in settings where tail probabilities govern decision thresholds and error calibration. Although asymptotically valid under minimal conditions, such approximations exhibit non-uniform convergence, with tail regions displaying markedly slower rates of convergence in the presence of mild distributional asymmetry. This short communication investigates the finite-sample behavior of standardized sample means and establishes that Gaussian tail approximations systematically underestimate upper-tail exceedance probabilities when skewness is present. The effect persists at sample sizes routinely encountered in applied work and admits a precise asymptotic characterization via Edgeworth expansions. A simple monotone corrective approximation, constructed to preserve closed-form evaluation, invertibility, and probabilistic admissibility, is proposed as a practical finite-sample refinement.
Functional repeated-measures data arise when curves are observed under multiple conditions for the same subjects. Existing inference procedures typically use L^2 -norm statistics that emphasize between-condition variability. We propose an F-type test statistic constructed as the ratio of the integrated squared difference between the sample mean functions to a scalar measure of within-condition variability. The asymptotic null and alternative distributions of the proposed statistic are established, and the null distribution is approximated by an F-distribution with data-driven degrees of freedom via moment matching. Simulations demonstrate that the test achieves good Type I error control and offers greater computational efficiency than bootstrap- or permutation-based methods. An application to mortality data illustrates its practical utility.
This paper examines the asymptotic properties of the Quasi Maximum Likelihood Estimator (QMLE) in high-dimensional heteroskedastic and approximate factor models with random common factors. We demonstrate that the asymptotic behaviour of the QMLE for both the covariance of common factors and their predictors differs substantially from the existing literature, where common factors are typically treated as fixed. A unified Central Limit Theorem (CLT) is established for the predictors of common factors, which subsumes the results of Bai and Li (2012) and (2016) as special cases. The theoretical findings are supported by Monte Carlo simulations and an empirical application using the Fama-French 10 × 10 portfolio dataset.
Unilateral or bilateral data from paired organs of each subject are commonly encountered in clinical trials. Compared to using unilateral or bilateral data alone, combined data can provide more information. To analyze the effects of explanatory variables, including stratum and group variables (e.g., age, gender) on the response probability in the stratified combined data, this paper develops a Bayesian framework for multinomial logistic regression based on Dallal’s model. Within this framework, we employ the NUTS and HMC algorithms for posterior sampling and compute Bayes factors to conduct hypothesis testing via the Savage-Dickey density ratio. Subsequently, simulation studies compare the Bayesian methods with the Fisher scoring algorithm under various sample sizes and conduct a prior sensitivity analysis using three priors. Results show that the Bayesian approach (especially NUTS) outperforms traditional frequentist methods and yields reliable inference for stratified unilateral and bilateral paired data, even at small sample sizes. Applications to two real cases are used to demonstrate our Bayesian approach.
Estimating basis covariance and precision matrices is crucial to compositional data analysis. However, missing data in practical sampling often distorts the correlation structure among key variables. By integrating thresholding and the CLIME algorithm, this paper develops novel estimators for both basis covariance and basis precision matrices in the presence of missing values. We establish convergence rates of estimators under both exponential and polynomial moment conditions over a larger parameter space. Simulation studies and a real-data application validate that the proposed estimators exhibit high efficiency and robustness under missing completely at random mechanism.
This paper studies a unified first-order mixed integer-valued autoregressive (INAR(1)) model that integrates binomial and negative binomial thinning operators, enabling flexible modeling of count time series exhibiting heterogeneous dispersion and regime-like dynamics. The proposed model encompasses several existing INAR specifications as special cases within a single coherent structure. We establish some fundamental probabilistic properties of the model, including strict stationarity, ergodicity and conditional moment structures. Parameter estimation is carried out via an empirical likelihood (EL) approach to avoid full distributional specification, and the asymptotic properties of the resulting estimators are rigorously derived. Furthermore, we propose an EL-based testing procedure to assess the equivalence of the two thinning parameters, which provides a formal tool for model simplification and structural interpretation. Monte Carlo simulations demonstrate that the proposed method achieves reliable finite-sample performance. An application to a real dataset highlights the empirical relevance of the model and its advantages over existing alternatives.
Distribution-free control charts play a vital role in statistical process control, particularly when the underlying distribution of process data is unknown or only partially characterized. Most existing distribution-free control charts are designed to monitor univariate or multivariate processes separately. In this paper, we combine conformal inference theory with the exponentially weighted moving average (EWMA) control scheme to develop a novel distribution-free control chart that is applicable to both univariate and multivariate processes. The key idea is to construct a series of distribution-free test statistics based on conformal p-value, whose distributions are independent of the underlying data distribution. The proposed method imposes no assumptions on the distribution or dimensionality of the data, consequently, it is applicable to data of any dimension and any distribution. Furthermore, it is computationally efficient, straightforward to implement, and effective in detecting shifts in location as well as changes in the overall distribution. Simulation studies show that the proposed chart achieves competitive performance for relatively large mean shifts and strong performance for the variance and covariance shifts considered in this study, although its sensitivity to small mean shifts is comparatively limited. A semiconductor manufacturing example further demonstrates the practical applicability of the proposed method.
This paper considers the problem of testing for the overall significance of a large number of remaining predictors in high-dimensional linear models, given that a few predictors are known to have significant effect on the response. We propose a novel test based on Bayes factor, assuming that the errors are normally distributed. The proposed test is also applicable to testing the global significance of the linear models. The asymptotic normality of the Bayes factor-based test statistic under the null hypothesis and the local alternatives is established. Furthermore, we derive the asymptotic local power function of the proposed test and compare it with that of existing methods. Monte Carlo simulation results show that the limiting approximation to the null distribution is highly accurate in finite samples, and the proposed test outperforms several well-known tests in the literature in terms of Type I error rate and the empirical power. The practical effectiveness of the proposed test is illustrated through two real-data applications.
In this paper, we propose a computationally efficient and theoretically justified group least absolute shrinkage and selection operator (Group LASSO; GLASSO) method for estimating multiple change-points in a piecewise stationary generalized integer-valued autoregressive process. The proposed method is particularly suitable for finite samples with many closely spaced change-points. We further develop an efficient implementation that combines least angle regression and optimal partitioning (OP). The overall computational complexity is O(Kn+K^2) when OP is used and O(Kn+K^3) when the backward elimination algorithm is used. In addition, we propose an iterative procedure for selecting a data-driven order p̃ , which achieves satisfactory performance with relatively low computational cost. Simulation studies and a real data analysis demonstrate that the proposed method and iterative procedure perform well in practice and support the theoretical results.
In this paper, we propose a new efficient estimator for the weighted exponential family and its underlying components. These components constitute a flexible class of distributions within the standard exponential family, characterized by positive support and a generator function. For such models, maximum likelihood estimators (MLEs) are often unavailable in closed form and must be derived through numerical optimization. To address this limitation, an asymptotically efficient closed-form estimator was developed for these distributions. Monte Carlo simulations demonstrate that the proposed estimator achieves performance nearly identical to the numerically computed MLE while consistently outperforming previously proposed closed-form estimators.
High-dimensional data with complex dependence structures are routinely collected in clinical and social science studies, where leveraging such structure can reveal latent pathways or regulatory networks and improve variable selection performance. In this work, we develop a generalized-distribution-based Bayesian approach for consistent network-guided variable selection in high-dimensional linear regression. Our novel approach simultaneously incorporates hierarchical spike-and-slab priors on the regression coefficients and the elements of the inverse covariance matrix. We further establish strong selection consistency of the proposed methodology and propose a highly scalable Gibbs sampler for posterior computation. Through simulation studies, we demonstrate that our method achieves superior performance relative to state-of-the-art alternatives. We analyze amplitude of low-frequency fluctuation (ALFF) neuroimaging data for individuals with autism spectrum disorder (ASD) to identify key brain regions associated with cognitive performance, providing insight into neural mechanisms related to ASD and benchmarking against existing methods.
We study the spiked tensor model whose noise entries follow general distributions with zero mean, unit variance, and finite fourth moment. We consider the asymmetric (independent-entry) model. Let s_1=⟨ u^(1),v_0⟩ be the signal projection at the first iteration of the tensor power iteration method. We prove that 𝔼[s_1]=β m_0^k-1 and Var(s_1)=1 for all admissible distributions of the noise entries. Furthermore, under a mild delocalization condition on the initialization of the power method, the centered projection s_1-β m_0^k-1 is asymptotically Gaussian: s_1-β m_0^k-1d→𝒩(0,1) . Thus, the signal projection at the first iteration of the tensor power iteration method is asymptotically universal. We also establish concentration bounds for the projection s_1 and provide a heuristic approximation formula for the normalized first iterate of the tensor power method. Finally, simulations are presented to support the theoretical results.
In this paper, we propose a robust goodness-of-fit test based on the empirical characteristic function of the sample median. The test is specifically designed to address the challenges of statistical inference in small to moderate sample sizes, where traditional methods may be affected by a few extreme observations or by endpoint effects (e.g., bounded support, heavy tails). By leveraging the inherent robustness of the median and the descriptive power of the characteristic function, the proposed test exhibits stable and reliable performance across a wide range of settings. Our main contribution is the development of a goodness-of-fit procedure that combines a median-based subsampling scheme with the empirical characteristic function, resulting in a test that is both robust to outliers and effective in small-sample regimes. Although our theoretical results are derived under a two-dimensional asymptotic regime with both the number of subsamples n and the within-subsample size N increasing, the test is calibrated at fixed (n, N) using Monte Carlo simulation. The method is designed for scenarios with small within-subsample sizes N (single digits to low tens) and a small-to-moderate number of subsamples n (about 10–50). In simulations and applications from reliability and quality control, this regime yields accurate size and competitive power.
We propose a nonparametric β -model for modelling the evolution of node degrees in dynamic networks. The model has n unknown parameter functions; therefore, a statistical analysis is challenging. We develop an adaptive weighted approach for estimating n parameter functions considering the network similarity between nearby time points. The proposed estimator performs well, even when the coefficient functions are piecewise smooth. We establish the consistency and asymptotic normality of the proposed estimator, which has smaller variance than the point-wise maximum likelihood estimator. In dynamic network data, it is also important to detect time points where the behavior of some nodes exhibits a sudden change in structure. We further develop inference methods for the change-point detection problem based on the proposed estimator. It dramatically improves statistical power by pooling information from nearby time points and can precisely identify the locations of change-points with a probability tending toward one. We evaluate the finite sample performance of the proposed method using extensive simulation studies and illustrate its application using an ant social organization dataset.
This paper introduces a novel divergence measure between two probability distributions, parameterized by a constant α∈ [0,1] . The proposed divergence generalizes well-known measures such as the Kullback–Leibler (KL) and Jensen–Shannon (JS) divergences, providing a flexible and unified framework for distribution comparison. We analyze its key mathematical properties, including convexity, differentiability, and symmetry, and explore its relationships with other divergence measures and invariance characteristics. Furthermore, we demonstrate its practical effectiveness through applications in clustering, anomaly detection, and machine learning, supported by experiments on the Iris and MNIST datasets. Detailed results and visualizations highlight the advantages of the proposed divergence, particularly in adaptive and dynamic scenarios.
In this paper, we consider the phase transition of Wilks’ phenomenon for the likelihood ratio statistic on testing the high-dimensional block compound symmetry covariance structure. When both the sample size and the dimension tend to infinity, we establish the necessary and sufficient condition for Wilks’ phenomenon, and also obtain the asymptotic bias of the chi-squared approximation and the precise expansion of the asymptotic distribution under a local asymptotic regime. Some numerical simulations are shown to illustrate the theoretical results. The results provide practical guidance for testing on high-dimensional block compound symmetry covariance structure.
Regression extremiles are of great practical importance in risk management as they satisfy the coherency axiom and take the severity of tail losses into account. Yet the existing work mainly focuses on the univariate extremile regression in the low-dimensional framework. High-dimensional data subject to heavy-tailed phenomena are commonly encountered in various scientific fields and pose new challenges for extremile regression. In this article, we propose a (penalized) robust linear extremile regression model (remire) in the multivariate setting, incorporating the Huber loss function in place of the squared loss to enhance robustness for high-dimensional heavy-tailed data. In the regularized framework, we adopt the folded concave penalty for variable selection, which is implemented via a local adaptive majorize-minimization algorithm. Theoretically, we establish the convergence properties of the penalized remire estimator. The proposed method exhibits desirable properties and performs well in finite samples in terms of coefficient estimation and model selection, as demonstrated through comprehensive numerical studies. We further illustrate its practical utility through an application to childhood malnutrition data.
Under the design-based approach, we carry out the analysis of the Theil index decomposition by focusing on the joint inference for the within-group and the between-group components. First, we express the two population components as statistical functionals in terms of a discrete measure. Subsequently, we derive the corresponding influence functions and obtain their properties. On the basis of such findings, we introduce estimators for the Theil index components and their variance-covariance matrix, as well as a pivotal quantity for implementing confidence ellipses. By means of a Monte Carlo simulation study, we show the suitable performance of the component estimators and the variance-covariance matrix estimator and assess the adequate coverage of the confidence ellipse. Finally, we apply our methodology to the data collected during the 2021 Household Budget Survey in Italy. More precisely, considering energy consumption at the household level as the target variable and adopting the NUTS 2 grouping, we estimate the components of the Theil index and the corresponding variability for such survey data.
Mixed hitting-time models allow the analysis of competing risks through optimal stopping decisions interpreted as crossing times of latent Lévy processes with heterogeneous thresholds. In this paper, we consider a bivariate time model with dependent default, where observation times are subject to censoring and share a common latent process given by a Lévy subordinator. We establish the identifiability of the model and propose different estimators for the marginal distributions and the joint survival function. We establish their asymptotic properties and evaluate the finite-sample performance of our results through a simulation study on synthetic data, followed by an application using real data.