In the presence of covariates affected by measurement errors, we first propose the corrected least product relative error score (CLAPRES) function to mitigate the effects of measurement errors on parameter estimation in multiplicative regression models. This method is invariant under scale transformations of the positive response and the covariates. To address the challenge of massive datasets with measurement errors, we explore an optimal subsampling algorithm based on the CLAPRES method and derive the optimal subsampling probabilities under the A- and L-optimality criteria. The consistency and asymptotic normality of the subsampling CLAPRES estimators are established. Numerical studies demonstrate the effectiveness of the CLAPRES method.
Regression coefficients clustering has gained increasing importance in various applications. However, most of the existing studies fail to accommodate skewed or asymmetrically distributed data with heteroscedasticity. In the context of longitudinal data, this paper investigates regression coefficients clustering by combining the asymmetric least squares loss with the multi-direction separation penalty method. The proposed method enables the identification of subgroups in which individuals share similar covariate effects, while allowing different important variables to be selected for different individuals. Meanwhile, the proposed method effectively addresses heteroscedasticity issues and captures more comprehensive distributional characteristics compared to ordinary least squares regression. The paper establishes theoretical properties, including the consistency of the estimator and its oracle property. To efficiently compute the proposed estimator, we develop an algorithm that combines cyclic coordinate descent with the alternating direction method of multipliers. Simulation studies and a practical example are provided to demonstrate the superiority of the proposed approach.
Causality is widely concerned in real life. Given empirical evidence for the dependence of an effect variable on a cause variable, we can typically provide bounds for the "probability of causation". This article is to further deal with the problem of estimating the probability that one event was a cause of another on the basis of partial mediation. For the multi-valued variables, we show how to optimally bound these quantities from data and make minimal assumptions concerning the data-generating process. We use simulation studies to evaluate the performance of the proposed results.
Subsampling techniques have been promoted in massive data and can substantially reduce the computing time. However, existing subsampling techniques do not consider the case of dirty data, especially the inaccuracy of covariates due to measurement errors, which will lead to the inconsistent estimators of regression coefficients. Therefore, to eliminate the influence of measurement errors on parameter estimators for massive data, this paper combined the corrected score function with the subsampling technique. The consistency and asymptotic normality of the estimators in the general subsampling are also derived. In addition, optimal subsampling probabilities are obtained based on the general subsampling algorithm using the A-optimality and L-optimality criteria and the truncation method, and then an adaptive two-step algorithm is developed. The effectiveness of the proposed method is demonstrated through numerical simulations and two real data analyses.
Timely and accurate prediction of epileptic seizures allows healthcare professionals to implement interventions promptly prior to a seizure, preventing secondary injuries. Many previous deep learning network models have achieved varying degrees of success in seizure prediction. However, most of the existing prediction models do not fully extract and integrate the temporal and spatial feature information in multichannel electroencephalogram (EEG). To address this issue, this study proposes a multiscale convolutional attention network for epileptic seizure prediction. Specifically, the network consists of a weighted attention mechanism and a multiscale convolutional module, which are designed respectively for highlighting the spatial information in important EEG channels and capturing both the local and the global temporal features. Tested on the CHB-MIT scalp EEG dataset, the seizure prediction method based on the multiscale convolutional attention network achieved an average sensitivity of 97.5%, an average false prediction rate of 0.30/h, and an average prediction time of 26.04 minutes. The results of multiple experiments further validate the feature extraction capability to multichannel EEG and the seizure prediction performance of the multiscale convolutional attention network.
. The local linear approximation algorithm is an effective algorithm for computing a global solution of the folded concave penalization problem. However, the effectiveness of this method is highly dependent on a reasonably good initial estimator. It will lose efficacy when the correlation among predictors is high. In this paper, we propose a new local linear approximation ridge algorithm designed to deal with highly correlated predictors. The ridge estimator is chosen as an initial estimator, the local linear approximation ridge algorithm is stable and effective. Simulation studies and a real data analysis show that the proposed algorithm has better performance than the local linear approximation algorithm in the presence of highly correlated predictors.
In this paper, a model averaging method is proposed for varying-coefficient models with response missing at random by establishing a weight selection criterion based on cross-validation. Under certain regularity conditions, it is proved that the proposed method is asymptotically optimal in the sense of achieving the minimum squared error.
An improved weighted expectile average estimator for the regression coefficient has been obtained based on the covariate balancing propensity score (CBPS), when the responses of linear models are missing at random. The asymptotic normality of the proposed method has been proved, and the estimation effect of the method is further illustrated by numerical simulation.
In this article, we proposed a weighted expectile average estimator for linear models with missing covariates. The asymptotic normality of the proposed estimator was established in theory. Further, we derived the explicit optimal weight and thus obtained the resulted optimal weighted expectile estimator. In order to examine the finite-sample performance of the proposed estimator, simulation studies and a real data analysis were conducted and the results were compared with other existing competitors in the literature.
In this paper, we consider variable selection for a class of semiparametric spatial autoregressive models based on exponential squared loss (ESL). Using the orthogonal projection technique, we propose a novel orthogonality-based variable selection procedure that enables simultaneous model selection and parameter estimation, and identifies the significance of spatial effects. Under appropriate conditions, we show that the proposed procedure is consistent and the resulting estimator has oracle properties. Furthermore, some simulation studies and an analysis of the Boston housing price data are also carried out to examine the finite-sample performance of the proposed method.
In this paper, three smoothed empirical log-likelihood ratio functions for the parameters of nonlinear models with missing response are suggested. Under some regular conditions, the corresponding Wilks phenomena are obtained and the confidence regions for the parameter can be constructed easily.
Objective: To find potential diagnostic biomarkers for ovarian cancer (OC), a prospective analysis of the expression of five biomarkers in patients with intermediate-risk and their correlation with the occurrence of OC was conducted.Method: A prospective observational study was carried out, patients who underwent surgical treatment with benign or malignant ovarian tumors in our hospital from January 2020 to February 2021 were included in this study, and a total of 263 patients were enrolled. Based on the postoperative pathological results, enrolled patients were divided into ovarian cancer group and benign tumor group (n = 135). The ovarian cancer group was further divided into a mid-stage group (n = 46) and an advanced-stage group (n = 82). The basic information of the three groups of patients was collected, the preoperative imaging data of the patients were collected to assess the lymph node metastasis, the preoperative blood samples were collected to examine cancer antigen 125 (CA125), carbohydrate antigen 19–9 (CA19–9), Neutrophil to lymphocyte ratio (NLR), platelet to lymphocyte ratio (PLR), and the postoperative pathological data were sorted and summarized.Result: The average during of disease in the advanced ovarian cancer group was 0.55 ± 0.18 years higher than the benign tumor group (0.43 ± 0.14 years), p < 0.001. In the advanced ovarian cancer group, the ratio of patients with the tumor, node, metastasis (TNM) stage IV (64.63%), with tumor Grade stage II and III (93.90%), and without lymph node metastasis (64.63%) was respectively more than that in the mid-stage group (accordingly 0.00, 36.96, 23.91%) (p < 0.001); The ratio of patients with TNM grade III in the mid-stage group (73.91%) was more than that in the advanced group (35.37%) (p < 0.001). The levels of the five biomarkers: CA19-9, CA125, NLR, PLR, and BDNF were different among the three groups (p < 0.001).Conclusion: CA19-9, CA125, NLR, PLR, BDNF are five biomarkers related to the occurrence of ovarian cancer and are risk factors for it. These five biomarkers and their Combined-Value may be suitable to apply in the diagnosis and the identification of ovarian cancer in patients with intermediate-risk.
In this article, two types of weighted quantile estimators were proposed for nonlinear models with missing covariates. The asymptotic normality of the proposed weighted quantile average estimators was established. We further calculated the optimal weights and derived the asymptotic distributions of the correspondingly resulted optimal weighted quantile estimators. Numerical simulations and a real data analysis were conducted to examine the finite sample performance of the proposed estimators compared with other competitors.
In this paper, we consider the multiple robust estimation of the parameters in the varying-coefficient partially linear model with response missing at random. The multiple robust estimation method is proposed, and the multiple robustness of the proposed method is proved. Numerical simulations are conducted to investigate the finite sample performance of the proposed estimators compared with other competitors.
In this paper, we consider the statistical inferences for varying coefficient partially nonlinear model with missing responses. Firstly, we employ the profile nonlinear least squares estimation based on the weighted imputation method to estimate the unknown parameter and the nonparametric function, meanwhile the asymptotic normality of the resulting estimators is proved. Secondly, we consider empirical likelihood inferences based on the weighted imputation method for the unknown parameter and nonparametric function, and propose an empirical log-likelihood ratio function for the unknown parameter vector in the nonlinear function and a residual-adjusted empirical log-likelihood ratio function for the nonparametric component, meanwhile construct relevant confidence regions. Thirdly, the response mean estimation is also studied. In addition, simulation studies are conducted to examine the finite sample performance of our methods, and the empirical likelihood approach based on the weighted imputation method (IEL) is further applied to a real data example.
In this paper, we study the weighted quantile average estimation technique for the parameter in additive partially linear models with missing covariates, which is proved to be an efficient method. The proposed method is based on optimally combining information over different quantiles via multiple quantile regression. We establish asymptotic normality of the weighted quantile average estimators when the selection probability is known, estimated using the non-parametrical method and parametrical method, respectively. Moreover, we compute optimal weights by minimizing asymptotic variance and then obtain the corresponding optimal weighted quantile average estimators. To examine the finite performance of our proposed method, we use the numerical simulations and apply to model time sober for the patients from a rehabilitation center. Simulation results and data analysis further verify that the proposed method is an efficient and safe alternative to both the WCQR method and WLS method.
In applications, predictors are naturally grouped with some variables in nonzero groups being irrelevant, simultaneously variable selection at both the group and within-group levels is more desirable. In addition, to achieve a robust estimation against outliers in both covariates and responses, combining the excellent properties of weighted least absolute deviation (WLAD) and least squares, we propose an adjusted WLAD (AWLAD) regression estimator with the adaptive group bridge penalty. Importantly, we demonstrate that the AWLAD estimator enjoys the oracle property when the number of parameters grows with the sample size. Simulation studies and a real data analysis indicate that the AWLAD has superior performance in the finite sample cases.
Nonresponse is a very common phenomenon in survey sampling. Nonignorable nonresponse - that is, a response mechanism that depends on the values of the variable having nonresponse - is the most difficult type of nonresponse to handle. This article develops a robust estimation approach to estimating equations (EEs) by incorporating the modelling of nonignorably missing data, the generalized method of moments (GMM) method and the imputation of EEs via the observed data rather than the imputed missing values when some responses are subject to nonignorably missingness. Based on a particular semiparametric logistic model for nonignorable missing response, this paper proposes the modified EEs to calculate the conditional expectation under nonignorably missing data. We can apply the GMM to infer the parameters. The advantage of our method is that it replaces the non-parametric kernel-smoothing with a parametric sampling importance resampling (SIR) procedure to avoid nonparametric kernel-smoothing problems with high dimensional covariates. The proposed method is shown to be more robust than some current approaches by the simulations.
Nonparametric models are popular owing to their flexibility in model building and optimality in estimation. However nonparametric models have the curse of dimensionality and do not use any of the prior information. How to sufficiently mine structure information hidden in the data is still a challenging issue in model building. In this paper, we propose a parametric family of estimators which allows for penalizing deviation from linear structure. The new estimator can automatically capture the linear information underlying regressions function to avoid the curse of dimensionality and offers a smooth choice between the full non-parametric models and parametric models. Besides, the new estimator is the linear estimator when the model has linear structure, and it is the local linear estimator when the model has no linear structure. Compared with the complete nonparametric models, our estimator has smaller bias due to using linear structure information of the data. The new estimator is useful in higher dimensions; the usual nonparametric methods have the curse of dimensionality. Based on the projection framework, the theoretical results give the structure of the new estimator and simulation studies demonstrate the advantages of the new approach.
Varying coefficient partially linear models are usually used for longitudinal data analysis, and an interest is mainly to improve efficiency of regression coefficients. By the orthogonality estimation technology and the empirical likelihood inference method, we propose a new orthogonality-based empirical likelihood inference method to estimate parameter and nonparametric components in a class of varying coefficient partially linear instrumental variable models with longitudinal data. The proposed procedure can separately estimate the parametric and nonparametric components, and the resulting estimators do not affect each other. Under some mild conditions, we establish some asymptotic properties of the resulting estimators. Furthermore, the finite sample performance of the proposed procedure is assessed by some simulation experiments and a real data analysis.