We consider the problem of recovering a low-rank matrix in a distributed setting, based on a convex loss function and non-convex matrix factorization. We use a linearized and decentralized alternating direction method of multipliers (ADMM) algorithm to compute the consensus solution. We establish local linear convergence (up to the approximation error when the unconstrained solution is not exactly low-rank) of the method despite the optimization problem being non-convex due to the factorization. Numerical examples are presented to illustrate the performance.
In this article, we delve into the quantile regression and homogeneity detection of a varying index coefficient panel data model, which incorporates fixed individual effects and exhibits nonlinear time trends. Using spline approximation, we obtain estimators for the trend functions, link functions, and index parameters, and subsequently establish the corresponding convergence rates and asymptotic normality. Observing that subjects within a group may share identical trend functions, we are motivated to further explore potential homogeneity in these trends. To this end, we propose a homogeneity identification algorithm based on binary segmentation. For the determination of the thresholding parameter in homogeneity identification, we propose a generalized Bayesian information criterion. Furthermore, we introduce a penalized method to discern the constant and linear structures within the nonparametric functions of our model. By leveraging grouped observations, we achieve more efficient estimation and improve the asymptotic properties of the estimators. To demonstrate the finite sample performance of our proposed approach, we conduct simulation studies and apply our methodology to a real-world dataset comprising Air Pollution Data and Integrated Surface Data (APD&ISD). Supplementary materials for this article are available online.
Download This Paper Open PDF in Browser Add Paper to My Library Share: Permalink Using these links will ensure access to this page indefinitely Copy URL Copy DOI
Longitudinal data arise frequently in many economic studies and epidemiological research. In this paper, we investigate the partially linear additive model for longitudinal data in the framework of quantile regression. To incorporate the within-subject correlation, we develop an estimation procedure using quadratic inference function (QIF) and polynomial spline approximation for unknown nonparametric functions. The theoretical properties of the resulting estimators are established, where the nonparametric functions achieve the optimal convergence rate and the parametric components are asymptotically normal even when the number of parameters in the linear part is diverging. We also propose a variable selection procedure based on penalization. Since the objective function is discontinuous, a practical estimation procedure is proposed using induced smoothing and we prove that the smoothed estimator is asymptotically equivalent to the original estimator. The proposed methods are evaluated via simulation studies and a real data application.
In this paper, we consider high-dimensional quantile tensor regression using a general convex decomposable regularizer and analyze the statistical performances of the estimator. The rates are stated in terms of the intrinsic dimension of the estimation problem, which is, roughly speaking, the dimension of the smallest subspace that contains the true coefficient. Previously, convex regularized tensor regression has been studied with a least squares loss, Gaussian tensorial predictors and Gaussian errors, with rates that depend on the Gaussian width of a convex set. Our results extend the previous work to nonsmooth quantile loss. To deal with the non-Gaussian setting, we use the concept of Rademacher complexity with appropriate concentration inequalities instead of the Gaussian width. For the multi-linear nuclear norm penalty, our Orlicz norm bound for the operator norm of a random matrix may be of independent interest. We validate the theoretical guarantees in numerical experiments. We also demonstrate advantage of quantile regression over mean regression, and compare the performance of convex regularization method and nonconvex decomposition method in solving quantile tensor regression problem in simulation studies.
In kernel-based learning, the random projection method, also called random sketching, has been successfully used in kernel ridge regression to reduce the computational burden in the big data setting, and at the same time retain the minimax convergence rate. In this work, we consider its use in sparse multiple kernel learning problems where a closed-form optimizer is not available, which poses significant technical challenges, for which the existing results do not carry over directly. Even when random projection is not used, our risk bound improves on the existing results in several aspects. We also illustrate the use of random projection via some numerical examples.
This paper constructs an evaluation index system of urban financial centrality from five dimensions: service level, anti risk ability, openness, collection scale and development environment, and uses Entropy TOPSIS method and obstacle model to measure the level of urban financial centrality in Guangdong-Hong Kong-Macao Greater Bay Area in 2020 and explore its obstacle factors. The study finds that from the comprehensive level, the level of urban financial centrality in Guangdong-Hong Kong-Macao Greater Bay Area is generally low and the difference is obvious, and the spatial distribution shows the characteristics of strong in the east and weak in the west, strong inside and weak outside. From the sub dimension, except Hong Kong and Shenzhen, which perform well and are relatively balanced in the five dimensions, other cities have development weaknesses. From the perspective of obstacle factors, the level of financial services and the degree of financial openness are the main obstacle factors for the improvement of financial centrality in most cities in the Guangdong-Hong Kong-Macao Greater Bay Area. In view of this, it is necessary to continue to improve the urban financial network, strengthen the quality and efficiency of financial services, and innovate the financial opening system, so as to improve the level of urban financial centrality in the region.
The Anti-Monopoly Law serves as a competition policy to correct market order, providing a favorable institutional environment for corporate governance. We use China’s listed companies and a difference-in-difference strategy to examine the impact of the Anti-Monopoly Law on executives’ excess perks. The results indicate that after implementing the Anti-Monopoly Law, executives of companies with high monopoly power decrease their excess perks. The impact channel analysis suggests that alleviating information asymmetry is an important pathway. Heterogeneity analysis results reveal that the effect of the Anti-Monopoly Law is more pronounced in state-owned companies and companies in the region with low institutional quality.
This paper examines how asymmetric information affects peer-to-peer lending in China. We find that default rates rise significantly with interest rates. Specifically, borrowers who select interest rates above the legal maximum private lending rate of 15.4% are more likely to default. A lack of interest rate caps to prevent adverse selection is responsible for the market failure of Chinese peer-to-peer lending. However, there is no strong evidence of the moral hazard effect in relation to interest rate and loan size. In addition, the credit scoring based on applicant characteristics can mitigate the asymmetric information problem, but not eradicate it completely.
Functional linear regression is at the centre of research attention involving curves as units of observation. In this article, we consider distributed computation in fitting functional linear regression with functional responses. We show that the aggregated estimator by simple averaging has the same convergence rate as the estimator using the entire data. Some simulation results are reported for illustration.
In many data analytic problems, repeated measurements with a large number of covariates are collected and conditional quantile modeling for such correlated data are often of significant interest, especially in medical applications. We propose a quadratic inference functions based approach to take into account the correlations within clusters and use smoothing to make the objective function amenable to computation. We show that the asymptotic properties of the estimators are the same whether or not smoothing is applied, established in the “diverging p, large n” setting. The cluster sizes are also allowed to diverge with sample size n. Simulation results are presented to demonstrate the effectiveness of the proposed estimator by taking into account the within-cluster correlations and we use a longitudinal data set to illustrate the method.
We attempt to examine the impact of the social security contributions on firms' market performance in the context of the 2010 Social Security Act. Using a sample of Chinese A-share listed firms, we use a difference-in-difference strategy to identify a causal relationship between the impact of social security contributions on firm performance. The empirical results show that imposing social security burdens significantly reduces firms’ market performance, especially for non-SOE firms, firms in competitive industries and firms with strong financing constraints.
The paper examines the impact and mechanism of statutory pension insurance contribution rates on firms' ESG performance. Using comprehensive data on statutory pension insurance contribution rates and Chinese listed companies from 2012 to 2019, we find that high contribution rates significantly reduce ESG performance, and the finding remains robust after robustness tests. We explain how pension insurance contributions undermine ESG performance through social responsibility and layoffs. Regarding heterogeneity, labor-intensive and small firms are more susceptible to increased pension insurance.
This paper examine the impact of noise trading on stock liquidity in China. We construct a theoretical model including noise traders, rational traders, and insiders, and test the model using transaction data on individual stocks in the CSI-300 (China-Shanghai-Shenzhen-300-Stock Index) in 2020. We find that noise trading has a negative effect on stock liquidity. As evidenced by the transmission mechanism test, noise trading lowers stock liquidity by increasing stock price volatility. Furthermore, when there is a higher level of insider trading, noise trading has a greater negative impact on stock liquidity. Dividends, however, reduce noise trading, thereby boosting stock liquidity.
In modern scientific applications, more and more data sets contain natural matrix predictors and traditional regression methods are not directly applicable. Matrix regression has been adapted to such data structure and received increasing attention in recent years. In this paper, we consider estimation of the conditional quantile in high-dimensional regularized matrix regression with a nuclear norm penalty and establish the convergence rate of the estimator. In order to construct a quantile matrix regression estimator in the distributed setting or for massive data sets, we propose a regularized communication-efficient surrogate loss (CSL) function. The proposed CSL method only needs the worker machines to compute the gradient based on local data and the central machine solves a regularized estimation problem. We prove that the estimation error based on the proposed CSL method matches the estimation error bound of the centralized method that analyzes the entire data set. An alternating direction method of multipliers algorithm is developed to efficiently obtain the distributed CSL estimator. The finite-sample performance of the proposed estimator is studied through simulations and an application to Beijing Air Quality data set.
Understanding why extreme events occur is often of major scientific interest in many fields. The occurrence of these events naturally depends on explanatory variables, but there is a severe lack of flexible models with asymptotic theory for understanding this dependence, especially when variables can affect the outcome nonlinearly. This article proposes a novel semiparametric tail index regression model to fill the gap for this purpose. We construct consistent estimators for both parametric and nonparametric components of the model, establish the corresponding asymptotic normality properties for these components that can be applied for further inference, and illustrate the usefulness of the model via extensive Monte Carlo simulation and the analysis of return on equity data and Alps meteorology data.
We consider partially linear quantile regression with a high-dimensional linear part, with the nonparametric function assumed to be in a reproducing kernel Hilbert space. We establish the overall learning rate in this setting, as well as the rate of the linear part separately. Our proof relies heavily on the empirical processes and the Rademacher complexity in the semi-nonparametric setting as analytic tools. Some simulation studies and a real data analysis are presented for illustration.
Although semiparametric models, in particular varying-coefficient models, alleviate the curse of dimensionality by avoiding estimation of fully nonparametric multivariate functions, there would typically still be a large number of functions to estimate. We propose a dimension reduction approach to estimating a large number of nonparametric univariate functions in varying-coefficient models, in which these functions are constrained to lie in a finite-dimensional subspace consisting of the linear span of a small number of smooth functions. The proposed methodology is put in the context of quantile regression, which provides more information on the response variable than the more conventional mean regression. Finally, we present some numerical illustrations to demonstrate the performances.
Varying-coefficient regression is a popular statistical tool that models the way a certain variable modulates the effect of other predictors nonlinearly. However, a majority of the VC regression models consider univariate responses; the case of multivariate responses have received relatively lesser attention. In this paper, we propose a robust multivariate varying-coefficient model based on rank loss that models the relationships among different responses via reduced-rank regression and penalized variable selection. Some asymptotic results are also established for the proposed methods. Using synthetic data, we investigate the finite sample performance and robustness properties of the estimator. We also illustrate our methodology by application to a real dataset on periodontal disease.
This article is concerned with statistical inference of partially linear additive regression models where the covariates in parametric component are measured with errors. Using polynomial spline approximations, we propose bias-corrected least squares estimators for parameters and establish the asymptotic normality, and show that the estimators of unknown functions achieve optimal nonparametric convergence rate. Moreover, we propose a variable selection procedure to identify significant regressors and derive the oracle property of penalized estimators. Finally, we propose two-stage local polynomial estimation for additive functions and show the corresponding asymptotical normality. Monte carlo studies and real data analysis illustrate the performance of our approaches.