We propose a new test for the heterogeneity of the spatial autoregressive parameter in semiparametric varying-coefficient spatial autoregressive models. Our specification test is built on the difference of parametric and nonparametric estimates of the spatial autoregressive coefficient, where the two estimates are obtained by the sieve GMM estimation method. Under mild conditions, we derive the limiting null distribution, the local power property and consistency of the test statistic. Numerical simulations show promising performance of the proposed test for finite samples in the considered cases, and the crime data of Tokyo is analyzed to illustrate the usefulness of the test.
Transfer learning is a powerful tool for improving estimation and prediction accuracy in high-dimensional problems, particularly in scenarios where the target data is limited. Current works, however, remain limited in their ability to robustly accommodate heavy-tailed data and outliers while retaining statistical efficiency. In this paper, we propose a robust and efficient transfer learning framework for a high-dimensional linear model based on a penalized convoluted rank estimation approach. We first develop a two-step transfer learning algorithm when the set of transferable source datasets is known, and establish non-asymptotic - and -error bounds for the resulting estimator. Our theoretical analysis shows that, when the target and sources are sufficiently close, these bounds could be improved over those of the classical penalized estimators that rely solely on target data. To address the more realistic setting in which transferable sources are unknown, we further devise a transferable source detection algorithm and prove its consistency, along with the corresponding estimation error bounds. Furthermore, we analyze the asymptotic relative efficiency of the proposed de-sparsifying estimator. Extensive simulation studies and a real-data analysis demonstrate that the proposed method is robust to heavy-tailed data and outliers while remaining efficient under light-tailed noise.
Testing the equality of two multivariate distributions is critical to ensuring valid and reasonable statistical conclusions. Since the true alternative is unknown in practice, developing a powerful testing strategy remains an open question. We propose a composite kernel-based test that flexibly assigns weights to candidate tests, thereby effectively enhancing power. Our test is robust and adept at handling various alternative hypotheses. Furthermore, our test is data-adaptive and sensitive to detecting distributional differences across both local and global features. The asymptotic distributions are investigated under both the null and the alternative hypotheses. In addition, we present the asymptotic behavior of the test statistic under a range of Pitman local alternative hypotheses. To improve the finite sample properties of the proposed test, we employ a permutation method to obtain critical values or p-values. We also show the theoretical validity of the permutation method. Compared with current state-of-the-art test methods, the proposed test exhibits satisfactory performance, as evidenced by extensive simulations and real-world data.
This paper proposes a novel approach to enhance the versatility of the max-sum test in high-dimensional data analysis by combining two distinct rank correlation coefficients: Spearman's rho and Chatterjee's xi. We uncovered the independence between the max-type test and the sum-type test by deriving their joint distribution. This insight enables the development of a comprehensive max-sum test that tackles both sparse and dense alternative correlation structures in an adaptive manner. Leveraging the asymptotic independence between the two coefficients and the intrinsic highlights of two single-coefficient tests, we have strategically implemented Cauchy combination principles to devise a multifunctional testing methodology. This approach can accommodate monotonic and nonmonotonic data types and thus offers a versatile solution to a broad spectrum of analytical requirements. This versatility of our proposed method has been impressively demonstrated through a diverse range of simulation data studies and two real-world data analyses, underscoring its effectiveness and practical utility.
In this study, we first propose the functional quadratic spatial autoregressive model, which can effectively capture spatial dependencies, linear relationships and nonlinear interactions in functional data. Then, we develop a generalized statistical inference framework for this model, where the functional component is treated by principal component analysis, followed by parameter estimation employing the generalized method of moments. Subsequently, we construct two residual-based test statistics to assess the model’s goodness of fit. Under some regularity conditions, we derive the asymptotic properties, encompassing the asymptotic normality of estimators for the finite parametric vector, the optimal convergence rate for nonparametric functions, and the asymptotic distributions of the proposed test statistics under null hypothesis and local or global alternative hypothesis. To determine the critical values for model checking statistics, we introduce a wild bootstrap procedure, and the asymptotic validity of this bootstrap-based testing approach is also discussed. The finite-sample performance of our methodology is evaluated through Monte Carlo simulations, and its practical usefulness is illustrated through two real-data applications to Spanish meteorological data and economic growth data. The results substantiate the efficacy of our approach in the real-world scenarios.
Existing research consistently demonstrates a nonlinear relationship between crash frequency and annual average daily traffic (AADT). However, conventional logarithmic models with power-law assumptions prove inadequate in characterising rate variations across wide AADT ranges and fail to address prevalent zero-inflation in crash data. This study develops a partially linear zero-inflated negative binomial model incorporating non-decreasing constraints, employing Gibbs sampling with P & oacute;lya-Gamma augmentation for efficient Bayesian inference. The proposed model offers a flexible framework for detecting multi-stage relationships under the premise that monotonicity holds, thereby facilitating potential AADT control strategies. Through empirical analysis of three Texas crash datasets, we first compare the proposed model against five baseline alternatives, demonstrating superior performance across multiple goodness-of-fit metrics. Subsequently, we evaluate competing hypotheses using maintenance districts as spatial proxies, providing potential evidence that missing crash coordinates linked to varying administrative reporting practices account for a portion of the excess zeros. The analysis reveals a novel finding that differs from previous studies: crash counts follow three stages with increasing AADT-an initial increase, a middle plateau, and a slow increase at the end. The observed pattern aligns with phenomena reported in traffic safety literature, including congestion-related risk moderation and behavioural adaptation, which are not readily represented in conventional parametric frameworks.
Predicting stock prices remains challenging under conditions of market volatility, structural breaks, and data scarcity, where deterministic models often fail to provide reliable predictions or meaningful uncertainty estimates. To address this, we propose a novel time series forecasting framework that integrates Multi-Pass Bayesian Estimation (MPBE) into Bayesian Long Short-Term Memory (LSTM) networks. MPBE enables explicit quantification of predictive uncertainty, mitigates overfitting, and enhances robustness in noisy financial data capabilities absent in standard point-estimate models. To the best of our knowledge, this is the first application of MPBE in Bayesian LSTMs. We evaluate two variants, Markov Chain Monte Carlo (MCMC) and Variational Inference (VI), on four real-world datasets: Apple, CRISPR Therapeutics, KOSPI, and the S&P 500. The MCMC model achieves strong accuracy on KOSPI (RMSE: 0.0158, MSE: 0.00025), while the VI variant delivers superior overall performance, with higher R2 (0.897 for Apple, 0.934 for KOSPI) and lower MAPE (3.59% for Apple, 2.31% for S&P 500), alongside improved computational efficiency. Our framework provides well-calibrated, interpretable uncertainty, critical for risk-aware finance, and outperforms state-of-the-art baselines, confirming its effectiveness and practical viability.
Traditional causal inference methods face challenges when dealing with causal effect estimation in the presence of missing data. This paper proposes the Missing Imputation Generative Adversarial Network Average Treatment Effect (MIGANATE) model which provides a possible solution for this problem. MIGANATE estimates average treatment effects (ATE) with missing data imputation using Generative Adversarial Networks. This model consists of two sub-models. The first sub-model, Missing Imputation Generative Adversarial Network (MIGAN), uses adversarial learning between a generator and a discriminator to generate imputed values that approximate the distribution of the true data, thereby addressing the missing data issue and providing complete data for subsequent ATE estimation. The second sub-model, GANATE, comprises a counterfactual module and a treatment effect module. The counterfactual generator produces counterfactual outcomes, which are then passed to the treatment effect module, where the Generative Adversarial Networks (GAN) is trained by minimizing an empirical loss function to estimate the ATE. A numerical study compares the effectiveness of the MIGANATE model with the Inverse Probability of Treatment Weighting (IPTW) and Covariate Balancing Propensity Score (CBPS) methods in estimating ATE. The results show that MIGANATE can provide more stable and accurate causal effect estimates in the presence of missing data. In addition, the effectiveness of the proposed method is illustrated by a real data example.
Zero-inflated models are commonly used for longitudinal count data with excess zeros. When the data exhibit overdispersion, conventional zero-inflated Poisson mixed models may fail to provide an adequate fit. Moreover, missing outcomes induced by subject dropout represent a common challenge in longitudinal studies, necessitating the incorporation of informative dropout into modeling frameworks to mitigate bias. To address these issues, we propose a novel Bayesian model for analyzing longitudinal count data characterized by excess zeros, overdispersion, and informative dropout. Unlike frequentist methods, our approach treats unobserved dropout outcomes as latent variables through data augmentation, thereby transforming high-dimensional integration problems into posterior sampling tasks. By integrating Pólya-Gamma data augmentation, we develop an efficient Gibbs sampling algorithm. The simulation results demonstrate that ignoring missing data can lead to severe bias, even reversing time trend estimates, whereas the proposed method maintains accurate parameter estimation. In a health research application, our model identified a positive temporal trend in hospitalization duration that was not captured by models ignoring informative dropout. It suggests a possible population-wide deterioration in health habits or systemic and behavioral shifts, pointing to the need for targeted policy interventions.
The two-sample test problem is a fundamental problem in statistical inference that attempts to test whether two probability measures are different based on corresponding samples. Consequently, many statistical methods have been proposed when random vectors are multivariate or even high-dimensional. For this problem, we introduce a randomly projected maximum mean discrepancy (MMD) in a reproducing kernel Hilbert space to characterize the distance between the distributions of two random vectors. The multivariate random vectors are projected onto univariate random variables and projected MMD statistics are constructed. The collection of projected MMD indexed by the unit sphere, and hence we treat it as the U-process. Theorems include the asymptotic theory of test statistics under the null hypothesis and the alternative hypothesis. Combining continuous mapping theory with projected MMD statistics, a class of test statistics is proposed, which includes the Cramér-von Mises and Kolmogorov-Smirnov methods as special cases. Since the limit null distribution of the test statistic depends on the data generation process, we apply the permutation test procedure to determine a critical value. Furthermore, the empirical size and power of the test statistics are evaluated by numerical simulations. Finally, we illustrate our method by applying it to real data sets with two-sample problem.
"Small sample size, high dimension" data bring tremendous challenges to epilepsy Electroencephalography (EEG) data analysis and seizure onset prediction. Commonly, sparsity technique is introduced to tackle the problem. In this paper, we construct a indicator matrix acting as prior knowledge to assist logistic regression model with group lasso penalty to implement seizure prediction. The proposed method selects the feature at the group level, and it achieves the seizure prediction based on the important feature groups, recognizes the unknown clusters properly and performs well for both synthetic data following Bernoulli distribution and dataset CHB-MIT.
There are lots of methods for complete independence test for high-dimensional random vector. However, it is difficult for practitioners to choose a powerful test because the true alternative hypothesis is unknown. Combining the L-2-type with the L-infinity-type test statistic, we propose a rank-based test method. From a technical point of view, the proposed test is distribution-free and consequently the corresponding critical values or p-values can be obtained by Monte Carlo methods. Compared with permutation or bootstrap test methods, the proposed test method saves calculation cost. Simulation results show that the resulting method has excellent performance with finite sample size. We also provide a real data application to demonstrate the practicality and effectiveness of the proposed test method.
In this paper, we introduce a new class of heterogeneous spatial autoregressive models (heterogeneous SAR models) where the variance parameters are modeled in terms of covariates. In order to estimate the model parameters, as well as their corresponding standard error estimates, we proposed a computational efficient MCMC method which combines the Gibbs sampler with Metropolis-Hastings algorithm. The proposed estimate method is illustrated through numerous simulations and is applied to the Boston housing data.
Interval-valued functional data, a new type of data in symbolic data analysis, depicts the characteristics of a variety of big data and has drawn the attention of many researchers. Mean regression is one of the important methods for analyzing interval-valued functional data. However, this method is sensitive to outliers and may lead to unreliable results. As an important complement to mean regression, this paper proposes an interval-valued scalar-on-function linear quantile regression model. Specifically, we constructed two linear quantile regression models for the interval-valued response and interval-valued functional regressors based on the bivariate center and radius method. The proposed model is more robust and efficient than mean regression methods when the data contain outliers as well as the error does not follow the normal distribution. Numerical simulations and real data analysis of a climate dataset demonstrate the effectiveness and superiority of the proposed method over the existing methods.
The partially linear varying coefficient spatial autoregressive model is a semi-parametric spatial autoregressive model in which the coefficients of some explanatory variables are variable, while the coefficients of the remaining explanatory variables are constant. For the nonparametric part, a local linear smoothing method is used to estimate the vector of coefficient functions in the model, and, to investigate its variable selection problem, this paper proposes a penalized robust regression estimation based on exponential squared loss, which can estimate the parameters while selecting important explanatory variables. A unique solution algorithm is composed using the block coordinate descent (BCD) algorithm and the concave-convex process (CCCP). Robustness of the proposed variable selection method is demonstrated by numerical simulations and illustrated by some housing data from Airbnb.
Recent research and substantive studies have shown growing interest in expectile regression (ER) procedures. Similar to quantile regression, ER with respect to different expectile levels can provide a comprehensive picture of the conditional distribution of a response variable given predictors. This study proposes three composite-type ER estimators to improve estimation accuracy. The proposed ER estimators include the composite estimator, which minimizes the composite expectile objective function across expectiles; the weighted expectile average estimator, which takes the weighted average of expectile-specific estimators; and the weighted composite estimator, which minimizes the weighted composite expectile objective function across expectiles. Under certain regularity conditions, we derive the convergence rate of the slope function, obtain the mean squared prediction error, and establish the asymptotic normality of the slope vector. Simulations are conducted to assess the empirical performances of various estimators. An application to the analysis of capital bike share data is presented. The numerical evidence endorses our theoretical results and confirm the superiority of the composite-type ER estimators to the conventional least squares and single ER estimators.
Chatterjee's new coefficient proposed by Chatterjee [2021, 'A New Coefficient of Correlation', Journal of the American Statistical Association, 116(536), 2009-2022.] is used to measure the degree of dependence between two scalars by rank correlation. However, the independence test based on Chatterjee's rank correlation may lose power since it only considers the distance between the nearest neighbours and ignores the other neighbours. In this paper, we propose an improvement to Chatterjee's new coefficient by incorporating the inverse distance-weighting, and further obtain the asymptotic normality of the improved coefficient and the Berry-Esseen bound under the null hypothesis. The proposed method is evaluated on the simulated as well as the real data on Yeast Gene Expression. The results show that the proposed method is superior to Chatterjee's new coefficient under various alternative hypotheses.
Testing high-dimensional data independence is an essential task of multivariate data analysis in many fields. Typically, the quadratic and extreme value type statistics based on the Pearson correlation coefficient are designed to test dense and sparse alternatives for evaluating high-dimensional independence. However, the two existing popular test methods are sensitive to outliers and are invalid for heavy-tailed error distributions. To overcome these problems, two test statistics, a Spearman's footrule rank-based quadratic scheme and an extreme value type test for dense and sparse alternatives, are proposed, respectively. Under mild conditions, the large sample properties of the resulting test methods are established. Furthermore, the proposed two test statistics are proved to be asymptotically independent. The max-sum test based on Spearman's footrule statistic is developed by combining the proposed quadratic with extreme value statistics, and the asymptotic distribution of the resulting statistical test is established. The simulation results demonstrate that the proposed max-sum test performs well in empirical power and robustness, regardless of whether the data is sparse dependence or not. Finally, to illustrate the use of the proposed test method, two empirical examples of Leaf and Parkinson's disease datasets are provided.
In this article, we consider the complete independence test of high-dimensional data. Based on Chatterjee coefficient, we pioneer the development of quadratic test and extreme value test which possess good testing performance for oscillatory data, and establish the corresponding large sample properties under both null hypotheses and alternative hypotheses. In order to overcome the shortcomings of quadratic statistic and extreme value statistic, we propose a testing method termed as power enhancement test by adding a screening statistic to the quadratic statistic. The proposed method do not reduce the testing power under dense alternative hypotheses, but can enhance the power significantly under sparse alternative hypotheses. Three synthetic data examples and two real data examples are further used to illustrate the performance of our proposed methods.