
In the era of big data, matrix data appear frequently in modern scientific applications, such as electroencephalography (EEG) data, excitation-emission matrix (EEM) data, and colorimetric sensor array data. Some application data, such as stocks data, may not appear to be in matrix form, but we can reformulate it into matrix data. To the end of finding the group structure in these datasets, we propose a matrix sparse group Lasso variable selection and estimation method for data with high-dimensional predictors. The method is carried out through a penalized matrix regression model with an arbitrary group structure for the regression coefficient matrix. For the new proposed model, we establish the oracle bounds for the prediction error, the estimation error, and the order of sparsity. In order to get the estimator of the new model, we design a mixed coordinate descent algorithm. The numerical studies show that our model can select the group structure and sparsity accurately and effectively. Especially, we used the new model to predict the movements of some stocks selected from NASDAQ Stock Market.
The paper deals with models in which the dependent variable, some explanatory variables, or both represent unobserved sensitive data. We introduce a novel discretization method that reveals sufficient information from the sensitive variable to approximate the parameter(s) of interest. Multiple discretization schemes are employed, and we show convergence in distribution for the unobserved variable. The asymptotic properties of the OLS estimator for linear models are derived and discussed. Monte Carlo simulations support our theoretical findings and demonstrate finite-sample properties. Finally, we contrast our method with other alternative methods for estimating the Australian gender wage gap.
Categorical variables collected through surveys often share a high degree of information, leading to substantial redundancy. If not properly accounted for, redundant information may result in misleading outcomes when distance-based techniques are applied, such as data visualization and profiling via Multidimensional Scaling (MDS). This issue typically arises when additive dissimilarity coefficients are employed in datasets characterized by moderate to strong associations among variables. In this work, we propose new dissimilarities for categorical data that explicitly account for the association structure of the data. These dissimilarities are subsequently combined with a robust distance for numerical variables, yielding a flexible and robust metric for multivariate heterogeneous data. The performance of the proposed metrics is evaluated under adverse scenarios involving underlying correlation structures and outlier contamination, and is compared with that of the classical Gower distance using MDS representations and a Nearest Neighbor classifier. In addition, the proposed methodology is illustrated through three real-data applications, addressing both profiling and classification tasks. The results show that the proposed distances are effective in isolating outlying units and in improving classification accuracy.
Missing values in multivariate dependent variables may occur during data collection, requiring imputation methods capable of handling complex inter-variable relationships. We propose a nonparametric copula-based method for imputing dependent multivariate missing data, called NPCoImp. By leveraging the empirical beta copula to calculate the conditional cumulative probability function of the missing variables given the observed ones, NPCoImp imputes data while respecting the underlying distributional shape - particularly radial symmetry - and accordingly adjusts the multivariate values used for imputation. NPCoImp is highly flexible and can handle multivariate missing data with any type of missingness pattern. The performance of NPCoImp has been evaluated through an extensive Monte Carlo study and compared with classical imputation methods, as well as with its direct competitor, the CoImp algorithm, and a machine learning-based technique, the missForest. Our findings indicate that NPCoImp is particularly effective in preserving complex dependence structure irrespective of the sample size, the proportion of missing data and the strength of the dependence. The strong performance of the proposed method is further supported by empirical case studies in the agricultural sector. Finally, the NPCoImp algorithm has been implemented in the R package CoImp, which is available on CRAN.
The R-learner is a widely used method for estimating heterogeneous treatment effects (HTE) and offers considerable flexibility in causal inference applications. However, its effectiveness decreases when treatment assignment is endogenous, as it relies exclusively on observed covariates. To address this limitation, the RLASSO-IV method is proposed, which extends the R-learner by incorporating instrumental variables (IV) to mitigate endogeneity. The least absolute shrinkage and selection operator (LASSO) regularization approach is employed to facilitate variable selection and enhance the accuracy of treatment effect estimation. Extensive simulation studies indicate that RLASSO-IV achieves lower estimation error than existing machine learning-based methods for causal inference. Our proposed method is further evaluated on real-world application examining the effect of education on wages, the results show that it captures treatment effect heterogeneity more accurately and robustly than benchmark approaches. These results highlight the potential of RLASSO-IV as a reliable approach for individualized causal effect estimation in the presence of endogeneity.
Dyadic data analysis is increasingly used to model intra- and inter-personal mechanisms across disciplines such as psychology, education, healthcare, and the social sciences. Examples of dyads include husband–wife, doctor–patient, parent–child, and coach–athlete relationships. Such data typically violate the standard independence assumption due to the inherent interdependence between dyad members, requiring more sophisticated modeling approaches. In this paper, we introduce a novel statistical framework—the Dyadic Multivariate Mixture Model (DMMM)—designed to address these complexities in the context of ordinal multivariate responses. The DMMM integrates techniques from dyadic analysis, finite mixtures, copula-based dependence modeling, and ordinal regression, offering a flexible and parsimonious way to capture latent interpersonal traits and heterogeneous patterns of interaction. While the motivating application concerns coach–athlete relationships in competitive swimming—where both members of the dyad responded to a series of Likert-scale items—the model itself is general and applicable across a wide range of domains where dyadic ordinal data are observed. These include clinical psychology (e.g., therapist–client), organizational settings (e.g., supervisor–employee), and educational contexts (e.g., teacher–student). The proposed methodology allows for the identification of latent dyadic profiles, the modeling of interdependence between responses, and the inclusion of covariates at multiple levels. A case study involving elite swimmers demonstrates the model’s ability to address three key research questions: measuring the strength of relational ties, identifying group differences in communication dynamics, and evaluating personality-related covariates that shape latent relational patterns.
Realized kernels (RK) of Barndorff-Nielsen et al. (Econometrica 76:1481–1536, 2008) are consistent estimators of daily integrated volatility (IV) under dependent exogenous and endogenous market microstructure noise (MMN), if the bandwidth is properly selected. We propose to calculate RK with an iterative plug-in (IPI) bandwidth selector by adapting the idea of Barndorff-Nielsen et al. (Econometr J 12:C1–C32, 2009). Under independent MMN, the IPI bandwidth converges (relatively) at the rate n^-1/5 . The selected bandwidth under dependent MMN also converges up to a bias factor. In both cases the resulting RK achieves its best rate of convergence of the order O_p(n^-1/5) in the current context. The proposal is applied to a huge number of data examples. Its nice practical performance is further confirmed in a simulation with the approach of Barndorff-Nielsen et al. (Econometr J 12:C1–C32, 2009) and its resulting RK as comparisons. It shows that the IPI-algorithm and the resulting RK perform better than the comparisons, respectively.
Motivated by the analysis of well-being equitable and sustainable indicators across Italian NUTS3 areas from 2004 to 2019, we introduce a family of parsimonious hidden Markov models for clustering multivariate longitudinal data. This approach offers an alternative to existing model-based clustering methods by capturing both temporal dynamics and latent heterogeneity among territorial units. Given the multivariate dimensionality of the data, we adopt a factor model representation to reparameterize the covariance structure, reducing the number of parameters and enhancing interpretability. Parameter estimation is carried out using an alternating expectation conditional maximization (AECM) algorithm. Our results reveal a substantial degree of heterogeneity, identifying 15 clusters across NUTS3 areas. Despite this, Italian well-being patterns show a remarkable degree of temporal stability. To further simplify interpretation and support policy analysis, we apply cluster merging techniques (discussed in the Appendix C), which group the original clusters into four broader macro-clusters, each characterized by distinct well-being profiles across subsets of indicators.
This study reflects on the inconsistency of the fixed-design residual bootstrap procedure for GARCH models under dependent innovations. We introduce a recursive- design residual block bootstrap procedure to accurately quantify the uncertainty around parameter estimates and the next period’s volatility. A simulation study provides evidence for the validity of the recursive-design residual block bootstrap under dependent innovations. The resulting confidence intervals are not only valid but also potentially narrower than those obtained by the inconsistent fixed-design bootstrap. In an empirical illustration, we demonstrate the practical relevance by showing residual dependence and differences in confidence intervals between the two bootstrap procedures.
There are practical situations in which continuous data on the positive real line contain an excess of zeros. In this context, building on the flexibility of the beta prime (BP) distribution, we introduce the zero-adjusted BP distribution that accommodate the presence of zeros and propose zero-adjusted BP regression to deal with the issue of regression estimation when there are data with zeros in the dependent variable. The maximum likelihood method is used to estimate the model parameters. Also, we consider residual analysis. Simulation studies are conducted to evaluate its finite sample performance. Finally, a real dataset from a population-based study of fumonisin production by Fusarium verticillioides in corn grains, conducted by the Institute of Biomedical Sciences at the University of São Paulo, is discussed in detail.
We extend the generalized information criteria framework for model selection to high-dimensional sparse statistical jump models, a recent class of statistically robust and computationally efficient alternatives to hidden Markov models. Specifically, we derive expressions for the model fit and complexity to construct suitable information criteria for hyperparameter selection. In extensive simulation studies, we demonstrate that our approach selects the correct hyperparameters with high probability. Finally, providing an empirical application, we infer the key features that drive the return dynamics of the world equity market. We find that a three-state model best describes the dynamics of MSCI developed and emerging markets indexes.
This paper concentrates on statistical models that feature both main effects and interaction terms, where the main effects exhibit a non-sparse structure while the interactions are sparse. This type of data structure is prevalent in various disciplines, including biology and genetics, and it often poses challenges due to the intricate interplay of interactions. We introduce a two-stage methodological method to address these modeling complexities, ensuring both accurate modeling and prediction. Our proposed method adeptly handles the complex correlations among interaction terms. Both numerical and theoretical evidence support the efficacy of our method across diverse problem domains. In extensive numerical comparisons with a range of established techniques, our method demonstrates superior statistical properties. We further apply our algorithm to a protein dataset related to Alzheimer’s disease, providing a practical tool for biological diagnosis.
Recent studies have suggested incorporating individual heterogeneity in the ordered choice models. This study proposes a new Bayesian model with individual heterogeneity in the dynamic panel-ordered probit model, which is the first attempt to capture individual heterogeneity for panel ordinal responses. The validity and accuracy of the newly proposed model are evaluated using simulated data. We further compare standard and individual heterogeneity models in analyzing subjective well-being in China, using panel data from the China Family Panel Studies. The empirical results demonstrate that the newly proposed model can measure the individual heterogeneity for panel ordinal responses fairly well.
Load-sharing systems arise in many different reliability applications, for instance, when modeling tensile strength of fibrous composites in textile industry or lifetimes of redundant technical systems in engineering. Sequential order statistics serve as a flexible model for the ordered component failure times of such systems and allow the residual lifetime distribution of the components to change after each component failure. In a proportional hazard rate setting, the model consists of some baseline distribution function and several model parameters describing successive adjustments of the hazard rates of the operating components. This work provides nonparametric confidence bands for the baseline distribution function, where the model parameters may be known or unknown. In case of known model parameters, we show how to construct exact confidence bands based on Kolmogorov-Smirnov type statistics, which are distribution-free with respect to the baseline distribution. If the model parameters are unknown, finite sample inference turns out to be infeasible, and asymptotic confidence bands for the baseline distribution function are derived. As a technical tool, we extend the existing asymptotic theory of semiparametric estimators based on the profile-likelihood approach.
Univariate and multivariate general linear regression models, subject to linear inequality constraints, arise in many scientific applications. The linear inequality restrictions on model parameters are often available from phenomenological knowledge and motivated by machine learning applications of high-consequence engineering systems (Agrell, 2019; Veiga and Marrel, 2012). Some studies on the multiple linear models consider known linear combinations of the regression coefficient parameters restricted between upper and lower bounds. In the present paper, we consider both univariate and multivariate general linear models subjected to this kind of linear restrictions. So far, research on univariate cases based on Bayesian methods is all under the condition that the coefficient matrix of the linear restrictions is a square matrix of full rank. This condition is not, however, always feasible. Another difficulty arises at the estimation step by implementing the Gibbs algorithm, which exhibits, in most cases, slow convergence. This paper presents a Bayesian method to estimate the regression parameters when the matrix of the constraints providing the set of linear inequality restrictions undergoes no condition. For the multivariate case, our Bayesian method estimates the regression parameters when the number of the constrains is less than the number of the regression coefficients in each multiple linear models. We examine the efficiency of our Bayesian method through simulation studies for both univariate and multivariate regressions. After that, we illustrate that the convergence of our algorithm is relatively faster than the previous methods. Finally, we use our approach to analyze two real datasets.
This paper presents an area-level temporal bivariate linear mixed model, incorporating correlated time effects for estimating socioeconomic indicators in small areas. The model is applied through the residual maximum likelihood method, leading to the derivation of empirical best linear unbiased predictors for these indicators. Additionally, an approximation of the mean square error matrix (MSE) is provided and four MSE estimators are proposed. The first estimator involves a plug-in approach to the MSE approximation, while the remaining estimators are based on parametric bootstrap procedures. To assess the performance of the fitting algorithm, predictors, and MSE estimators, three simulation experiments are carried out. An application to real data from the 2016 to 2022 Spanish Living Conditions Survey is conducted. The focus is on estimating poverty proportions and gaps for the year 2022, categorized by provinces and sex.
Respondent-driven sampling (RDS) is a network-based sampling technique widely used to survey hidden or hard-to-reach populations. Existing estimators rely on strong assumptions intended to ensure the unbiasedness of RDS estimators. Even though survey practitioners employ various practical tools to account for them, these assumptions are often unrealistic and challenging to satisfy in practice. In particular, when working with real-world survey data obtained through this methodology, RDS estimators may demonstrate poor performance, exhibiting notable bias. In this context, several strategies to adjust for selection biases inherent in non-probability samples, such as those that use propensity score adjustment, statistical matching and double robust estimators, are available and can outperform conventional RDS estimators. This study investigates the performance of such estimators with a simulated population designed to deviate from certain assumptions of respondent-driven sampling, and with an empirically collected network dataset. We examine how the number of seeds affects the performance of the RDS methodology by using different number of seeds to the first simulated population. We also examine network structures featuring individuals with exceptionally high centrality, along with structures where connections are either concentrated within small groups or more evenly distributed across groups. Our analysis reveals that these alternative estimators can outperform conventional RDS estimators, particularly in scenarios where the underlying assumptions of RDS are violated. This study represents the first application of these methods and techniques within an RDS framework.
Clustering categorical data presents unique challenges that traditional techniques do not adequately address. This paper proposes an extension of the fuzzy C-modes algorithm. By incorporating a noise cluster and integrating spatial contiguity relationships among units, the algorithm’s robustness is significantly enhanced. Performance evaluations using synthetic data demonstrate the efficacy of the proposed algorithm in handling both global and local outliers. Furthermore, the paper discusses the application of the algorithm to real-world data on sustainable urban mobility in the Italian provincial capitals during 2021, highlighting its practical relevance and potential impact in real-world scenarios.