Parametric factor copula models typically work well in modeling multivariate dependencies due to their flexibility and ability to capture complex dependency structures. However, accurately estimating the linking copulas within these mod- els remains challenging, especially when working with high-dimensional data. This paper proposes a novel approach for estimating linking copulas based on a non-parametric kernel estimator. Unlike conventional parametric methods, our approach utilizes the flexibility of kernel density estimation to capture the un- derlying dependencies more accurately, particularly in scenarios where the un- derlying copula structure is complex or unknown. We show that the proposed estimator is consistent under mild conditions and demonstrate its effectiveness through extensive simulation studies. Our findings suggest that the proposed approach offers a promising avenue for modeling multivariate dependencies, par- ticularly in applications requiring robust and efficient estimation of copula-based models.
In this paper, we address the problem of robust direction -of-arrival (DOA) estimation in unknown spatially correlated noise fields using sensor arrays composed of subarrays in sparse configurations. In such arrays, the noise covariance matrix has a block-diagonal structure. The proposed robust DOA estimation method is derived from a parametric distribution divergence, cr-divergence. Our approach can be viewed as an extension of existing Maximum Likelihood (ML) iterative procedures. The degree of robustness is controlled by the parameter a: as a 1, the proposed method converges to the traditional ML approach, while for a < 1, our method effectively mitigates the impact of potential outliers. Moreover, simulation studies show that our robust DOA estimation not only handles two types of outliers better than the ML method, but also exhibits high breakdown point properties.
In this article, a copula-based method for mixed regression models is proposed, where the conditional distribution of the response variable, given covariates, is modelled by a parametric family of continuous or discrete distributions, and the effect of a common latent variable pertaining to a cluster is modelled with a factor copula. We show how to estimate the parameters of the copula and the parameters of the margins, and we find the asymptotic behaviour of the estimation errors. Numerical experiments are performed to assess the precision of the estimators for finite samples. An example of an application is given using COVID-19 vaccination hesitancy from several countries. Computations are based on R package CopulaGAMM.
Canonical correlation analysis (CCA) is a widely used mutivariate statistical technique for exploring the relationship between two multivariable datasets. It extracts existing relationship information by finding pairs of linear combinations from the two sets of variables with maximum correlation. In some applications however, the observed datasets may be contaminated by outliers and the standard CCA methods are sensitive to the presence of outliers in the datasets. In this paper, a robust CCA (RCCA) algorithm is presented. It is obtained using the interpretation of CCA as a latent variable model with two Gaussian random vectors and a robust loss function derived from the alpha-divergence as an alternative to maximum likelihood. Compared to existing robust CCA approaches, the proposed loss has the advantage of belonging to class of redescending M-estimators, guaranteeing inference stability for large deviation from the Gaussian nominal noise model. Experimental results on simulated and real datasets show that the proposed RCCA outperforms some existing robust and standard CCA methods.
In this paper, we introduce a new class of models for spatial data obtained from max-convolution processes based on indicator kernels with random shape. We show that this class of models have appealing dependence properties including tail dependence at short distances and independence at long distances. We further consider max-convolutions between such processes and processes with tail independence, in order to separately control the bulk and tail dependence behaviors, and to increase flexibility of the model at longer distances, in particular, to capture intermediate tail dependence. We show how parameters can be estimated using a weighted pairwise likelihood approach, and we conduct an extensive simulation study to show that the proposed inference approach is feasible in high dimensions and it yields accurate parameter estimates in most cases. We apply the proposed methodology to analyse daily temperature maxima measured at 100 monitoring stations in the state of Oklahoma, US. Our results indicate that our proposed model provides a good fit to the data, and that it captures both the bulk and the tail dependence structures accurately.
We introduce a new copula model for non-stationary replicated spatial data. It is based on the assumption that a common factor exists that controls the joint dependence of all the observations from the spatial process. As a result, our proposal can model tail dependence and tail asymmetry, unlike the Gaussian copula model. Moreover, we show that the new model can cover a full range of dependence between tail quadrant independence and tail dependence. Although the log-likelihood of the model can be obtained in a simple form, we discuss its numerical computational issues and ways to approximate it for drawing inference. Using the estimated copula model, the spatial process can be interpolated at locations where it is not observed. We apply the proposed model to temperature data over the western part of Switzerland, and we compare its performance with that of its stationary version and with the Gaussian copula model.
We study the class of dependence models for spatial data ob-tained from Cauchy convolution processes based on different types of ker-nel functions. We show that the resulting spatial processes have appealing tail dependence properties, such as tail dependence at short distances and independence at long distances with suitable kernel functions. We derive the extreme-value limits of these processes, study their smoothness prop-erties, and detail some interesting special cases. To get higher flexibility at sub-asymptotic levels and separately control the bulk and the tail depen-dence properties, we further propose spatial models constructed by mixing a Cauchy convolution process with a Gaussian process. We demonstrate that this framework indeed provides a rich class of models for the joint modeling of the bulk and the tail behaviors. Our proposed inference ap-proach relies on matching model-based and empirical summary statistics, and an extensive simulation study shows that it yields accurate estimates. We demonstrate our new methodology by application to a temperature dataset measured at 97 monitoring stations in the state of Oklahoma, US. Our results indicate that our proposed model provides a very good fit to the data, and that it captures both the bulk and the tail dependence structures accurately.
Gaussian factor models allow the statistician to capture multivariate dependence between variables. However, they are computationally cumbersome in high dimensions and are not able to capture multivariate skewness in the data. We propose a copula model that allows for arbitrary margins, and multivariate skewness in the data by including a non-Gaussian factor whose dependence structure is the result of a one-factor copula model. Estimation is carried out using a two-step procedure: margins are modelled separately and transformed to the normal scale, after which the dependence structure is estimated. We develop an estimation procedure that allows for fast estimation of the model parameters in a high-dimensional setting. We first prove the theoretical results of the model with up to three Gaussian factors. Then, simulation results confirm that the model works as the sample size and dimensionality grow larger. Finally, we apply the model to a selection of stocks of the S&P500, which demonstrates that our model is able to capture cross-sectional skewness in the stock market data.
We propose a new class of extreme-value copulas which are extreme-value limits of conditional normal models. Conditional normal models are generalizations of conditional independence models, where the dependence among observed variables is modeled using one unobserved factor. Conditional on this factor, the distribution of these variables is given by the Gaussian copula. This structure allows one to build flexible and parsimonious models for data with complex dependence structures, such as data with spatial dependence or factor structure. We study the extreme-value limits of these models and show some interesting special cases of the proposed class of copulas. We develop estimation methods for the proposed models and conduct a simulation study to assess the performance of these algorithms. Finally, we apply these copula models to analyze data on monthly wind maxima and stock return minima.
Factor copula models involve latent variables that explain much of the dependence in the observed variables. Their log-likelihoods can involve one-dimensional or multi-dimensional integration. For the one-factor copula with weak residual dependence and for the oblique factor copula model, we show that, under some mild assumptions, proxy variables that are unweighted averages computed from the observed variables can be used for the latent variables when the dimension is large. Then alternative log-likelihoods without integrals can be used for parameter estimation. The proxy variables can help to select appropriate linking copulas in some factor copula models and to perform numerically faster maximum likelihood estimation of parameters. Simulation studies show that parameter estimates obtained using the proxy variable approach are close to those obtained using the maximum likelihood approach. The proxy variable approach is used to analyze a financial data set of stock returns in a single sector.
Multivariate statistical monitoring charts are efficient tools for assessing the quality of a process by identifying abnormalities. Most commonly used multivariate monitoring charts, such as the Hotelling T2 rule, however, assume the availability of uncorrelated Gaussian observations. Unfortunately, very often, real data do not satisfy these assumptions, and thus limit the usefulness of these techniques in practice. Furthermore, in many real applications, changes can occur in the shape of the multivariate distribution of the process while its mean or variance remains the same. Conventional process monitoring charts, such as the T2 chart, fail to detect such changes in the distribution. In this article, we develop new copula-based multivariate monitoring techniques for possibly autocorrelated, non-Gaussian data that can detect changes in the shape of a multivariate distribution that are usually overlooked by conventional monitoring charts. Using synthetic data, we demonstrate the effectiveness of the developed charts over a conventional monitoring chart. Results indicate that the proposed charts are very promising because copula-based charts are, in practice, designed to monitor the entire distribution of a process instead of the individual components. The developed monitoring charts are validated through practical application on data from a decentralized wastewater treatment plant in Golden, CO, United States.
A new class of copula models with dynamic dependence is introduced; it can be used when one can assume that there exist a common latent factor that affects all of the observed variables. Conditional on this factor, the distribution of these variables is given by the Gaussian copula with a time-varying correlation matrix, and some observed driving variables can be used to model dynamic correlations. This structure allows one to build flexible and parsimonious models for multivariate data with non-Gaussian dependence that changes over time. The model is computationally tractable in high dimensions and the numerical maximum likelihood estimation is feasible. The proposed class of models is applied to analyze three financial data sets of bond yields, CDS spreads and stock returns. The estimated model is used to construct projected distributions and, for the bond yield and CDS spread datasets, compute the expected maximum number of investments in distress under different scenarios.
We propose three methods for estimating the joint tail probabilities based on a d-variate copula with dimension d ≥ 2. For the first two methods, we use two different tail expansions of the copula which are valid under mild regularity conditions. We estimate the coefficients of these expansions using the maximum likelihood approach with appropriate data beyond a threshold in the tail. For the third method, we propose a family of tail-weighted measures of multivariate dependence and use these measures to estimate the coefficients of the second tail expansion using regression. This expansion is then used to estimate the joint tail probabilities when the empirical probabilities cannot be used because of lack of data in the tail. The three proposed methods can also be used to estimate tail dependence coefficients of a multivariate copula. Simulation studies are used to indicate when the methods give more accurate estimates of the tail probabilities and tail dependence coefficients. We apply the proposed methods to analyze tail properties of a data set of financial returns.
We consider a special case of factor copula models with additive common factors and independent components. These models are flexible and parsimonious with O ( d ) parameters where d is the dimension. The linear structure allows one to obtain closed form expressions for some copulas and their extreme‐value limits. These copulas can be used to model data with strong tail dependencies, such as extreme data. We study the dependence properties of these linear factor copula models and derive the corresponding limiting extreme‐value copulas with a factor structure. We show how parameter estimates can be obtained for these copulas and apply one of these copulas to analyse a financial data set.
We propose a new copula model that can be used with replicated spatial data. Unlike the multivariate normal copula, the proposed copula is based on the assumption that a common factor exists and affects the joint dependence of all measurements of the process. Moreover, the proposed copula can model tail dependence and tail asymmetry. The model is parameterized in terms of a covariance function that may be chosen from the many models proposed in the literature, such as the Matern model. For some choice of common factors, the joint copula density is given in closed form and therefore likelihood estimation is very fast. In the general case, one-dimensional numerical integration is needed to calculate the likelihood, but estimation is still reasonably fast even with large data sets. We use simulation studies to show the wide range of dependence structures that can be generated by the proposed model with different choices of common factors. We apply the proposed model to spatial temperature data and compare its performance with some popular geostatistics models.
The multivariate Husler-Ress copula is obtained as a direct extreme-value limit from the convolution of a multivariate normal random vector and an exponential random variable multiplied by a vector of constants. It is shown how the set of Husler-Reiss parameters can be mapped to the parameters of this convolution model. Assuming there are no singular components in the Husler-Reiss copula, the convolution model leads to exact and approximate simulation methods. An application of simulation is to check if the Husler-Reiss copula with different parsimonious dependence structures provides adequate fit to some data consisting of multivariate extremes. (C) 2017 Elsevier Inc. All rights reserved.
We propose a new copula model for replicated multivariate spatial data. Unlike classical models that assume multivariate normality of the data, the proposed copula is based on the assumption that some factors exist that affect the joint spatial dependence of all measurements of each variable as well as the joint dependence among these variables. The model is parameterized in terms of a cross-covariance function that may be chosen from the many models proposed in the literature. In addition, there are additive factors in the model that allow tail dependence and reflection asymmetry of each variable measured at different locations and of different variables to be modeled. The proposed approach can therefore be seen as an extension of the linear model of coregionalization widely used for modeling multivariate spatial data. The likelihood of the model can be obtained in a simple form and therefore the likelihood estimation is quite fast. The model is not restricted to the set of data locations, and using the estimated copula, spatial data can be interpolated at locations where values of variables are unknown. We apply the proposed model to temperature and pressure data and compare its performance with the performance of a popular model from multivariate geostatistics.
We propose a new copula model that can be used with replicated spatial data. Unlike the multivariate normal copula, the proposed copula is based on the assumption that a common factor exists and affects the joint dependence of all measurements of the process. Moreover, the proposed copula can model tail dependence and tail asymmetry. The model is parameterized in terms of a covariance function that may be chosen from the many models proposed in the literature, such as the Matérn model. For some choice of common factors, the joint copula density is given in closed form and therefore likelihood estimation is very fast. In the general case, one-dimensional numerical integration is needed to calculate the likelihood, but estimation is still reasonably fast even with large data sets. We use simulation studies to show the wide range of dependence structures that can be generated by the proposed model with different choices of common factors. We apply the proposed model to spatial temperature data and compare its performance with some popular geostatistics models. Some key words: copula; heavy tails; non-Gaussian random field; spatial statistics; tail asymmetry. Short title: Factor Copula Models for Replicated Spatial Data CEMSE Division, King Abdullah University of Science and Technology, Thuwal 23955-6900, Saudi Arabia. E-mail: pavel.krupskiy@kaust.edu.sa, raphael.huser@kaust.edu.sa, marc.genton@kaust.edu.sa This research was supported by the King Abdullah University of Science and Technology (KAUST).
We propose measures of copula reflection and permutation asymmetry for data with positive quadrant dependence. We first define the measures of reflection asymmetry using a weighting function and then extend this approach to construct measures of permutation asymmetry for bivariate data. We define the corresponding statistical tests based on these measures and find that the proposed tests have higher statistical power comparing to some other tests for permutation and reflection symmetry studied in the literature. In addition, the measures can be used to summarize dependence structure of a multivariate data set in a few numbers and to select a more appropriate copula in the model.