Differences between near-infrared (NIR) spectroscopy instruments make it difficult to apply calibration models universally across multiple instruments; hence, calibration transfer (CT) is crucial. To ensure that a model developed on one instrument is also applicable to a new instrument, this study establishes a CT method based on nonparametric techniques, referred to as nonparametric varying-coefficient regression calibration transfer (NVT). This method uses a varying-coefficient model (VCM) to build a functional relationship between the master and slave spectra using a set of standard sample spectra and employs B-splines as basis functions for function fitting. This functional relationship helps transfer the slave spectra of other samples into the master spectra, reducing the spectral differences caused by instrument variations. The performance of NVT was tested on the determination of moisture, oil, protein, and starch in corn, and the content of total plant alkaloids, reducing sugars, total sugars, and total nitrogen in tobacco using NIR spectroscopy. NVT was compared with two common CT methods: spectral space transformation (SST) and piecewise direct standardization (PDS). The results show that NVT can effectively eliminate some spectral differences caused by different instruments and significantly improve analytical accuracy. Compared with that of PDS, the CT effect of NVT is significantly improved, whereas compared with that of SST, it is slightly improved. This method is insensitive to parameters, making it easy to select parameters and providing a new idea for CT method design.
For clinical trials with continuous outcomes, researchers may opt to report the whole or part of the five-number summary rather than the sample mean and standard deviation, especially when the outcome data are skewed. To include such studies in meta-analysis, several popular methods have been proposed in the literature that convert the five-number summary back to the sample mean and standard deviation. Nevertheless, most existing methods are based on the normality assumption, which may not hold for the clinical studies with the five-number summary being reported. Recently, Shi et al. (2023, Stat Methods Med Res, 32, 1338-1360) and Balakrishnan et al. (2023, Math Methods Stat, 32, 260-273) proposed methods for detecting the skewness of data that utilize the whole or part of the five-number summary, together with the sample size. In this article, we show that the max-type test of Shi et al. is not only structurally complex but also conservative in controlling the type I error rate, whereas the test of Balakrishnan et al. has difficulty controlling the type I error rate with small sample sizes. Inspired by these findings, we develop a novel test statistic that leverages the ratio of two tails, a measure known for its heightened sensitivity to data skewness. Simulation results demonstrate that our ratio-based test is both less conservative and more powerful compared to the existing methods, especially when the alternative distribution exhibits mild skewness. Additionally, simulated meta-analyses and real-world data examples are presented to demonstrate the utility of our new method in the context of meta-analysis.
In this article, we propose a bias-corrected double penalized quadratic inference functions method to simultaneously identify model structure, estimate parameters, and perform variable selection for varying coefficient errors-in-variables (EV) models with longitudinal data. Unlike the linear models or the partial linear varying coefficient models, the proposed method does not assume in advance whether each regression coefficient is constant or varying. Instead, it represents each coefficient as a nonparametric function and identifies whether it is constant or varying using the proposed method. By employing a B-spline basis to approximate the unknown coefficient functions, the proposed method integrates a bias-corrected quadratic inference function with two penalized terms to achieve structure identification, estimation, and variable selection. Under certain regularized conditions, the consistency and sparsity properties of the estimator are established. Moreover, a three-step iterative algorithm is developed to implement the proposed method in practice. Simulation studies and a real data analysis demonstrate the superior finite-sample performance of the method.
In this paper, we introduce a novel partial functional quadratic regression model and address the problem of robust statistical inference within this framework. We propose an M-estimation approach based on functional principal component analysis, incorporating three robust loss functions: Huber, Tukey’s bisquare, and the exponential squared loss. A fast iterative algorithm is developed to compute the corresponding estimators. The proposed method provides a highly efficient and robust alternative to the ordinary least squares method, and can be conveniently implemented using existing R software packages. Furthermore, we construct a robust testing procedure to assess the significance of the quadratic term in the model and employ the bootstrap procedure to evaluate the null distribution of test statistic and compute its corresponding p-value. Finally, we demonstrate the finite sample performance of our methodology through simulation studies and an empirical application to a real-world spectroscopy dataset.
In this paper, we proposed a bias-corrected double penalised least squares function method to investigate model identification and selection for varying coefficient errors-in-variables (EV) models. Without making assumptions about whether the regression coefficients in the model are constant or varying coefficients, the proposed method first approximates the nonparametric regression coefficients using a B-spline basis function, then does bias-correction for the unobserved covariates and establishes a double penalised least-squares function to identify, estimate and select the varying and nonzero constant coefficients simultaneously. Under some regularity conditions, the proposed method is consistent in both identification and selection of nonzero constant and varying coefficients. Further, the resulting estimators of varying coefficients possess the optimal convergence rate of nonparametric function estimation, and the estimators of nonzero constant coefficients are consistent and asymptotically normal. Finally, the finite sample performance of the proposed method is evaluated by simulation studies and a real data analysis.
In this paper, we consider model estimation and variable selection for partial linear errors-in-variables models with longitudinal data through empirical likelihood and quadratic inference function methods. We propose a bias-corrected penalized empirical likelihood method that addresses both measurement errors in covariates and unknown within-subject correlations while performing model estimation and variable selection. Under some regularity conditions, the resulting estimators possess the oracle property, and the nonparametric function estimator achieves optimal convergence rates. Numerical results including simulation studies and real example analysis demonstrate that the proposed method makes sense in finite samples, validating its theoretical properties and practical applicability.
In this paper, we propose a model identification and selection method for varying coefficient errors-in-variables (EV) models with missing responses, termed the imputation-based bias-corrected double-penalized estimating equation (ibbcDPEE) method. The proposed method does not need to assume in advance whether the regression coefficients in models are constants or varying coefficients. First, it utilizes B-spline basis functions to approximate the nonparametric regression coefficients. Subsequently, the bias-corrected double-penalized estimating equation (bcDPEE) is constructed based on the observed responses, while accounting for the bias in the unobserved covariates. The missing responses are then imputed via the kernel estimation technique. Lastly, the ibbcDPEE is constructed to do model identification and selection simultaneously. Under some regularity conditions, the proposed method can consistently identify and select varying coefficients and nonzero constant coefficients. Moreover, the estimators of the varying coefficients achieve the optimal convergence rate of nonparametric function estimation. The finite sample performance of the proposed method is evaluated through simulation studies and a real data analysis.
Ultrahigh-dimensional data analysis has received great achievement in recent years. When the data are stored in multiple clients and the clients can be connected only with each other through a network structure, the implementation of ultrahigh-dimensional analysis can be numerically challenging or even infeasible. In this work, we study decentralised federated learning for ultrahigh-dimensional data analysis, where the parameters of interest are estimated via a large amount of devices without data sharing by a network structure. In the local machines, each parallel runs gradient ascent to obtain estimators via the sparsity-restricted constrained methods. Also, we obtain a global model by aggregating each machine's information via an alternating direction method of multipliers (ADMM) using a concave pairwise fusion penalty between different machines through a network structure. The proposed method can mitigate privacy risks from traditional machine learning, recover the sparsity and provide estimates of all regression coefficients simultaneously. Under mild conditions, we show the convergence and estimation consistency of our method. The promising performance of the method is supported by both simulated and real data examples.
Functional regression has been a hot topic in statistical research. However, not much work has been done when response variables are cross-sectionally dependent variables and explanatory variables contain a real-valued scalar variable and a functional-valued random variable. In this paper, we consider a new functional partially linear spatial autoregressive model. Based on the functional principal components analysis and basis function approximation, we obtain the estimators of the unknown parameter and functions through the instrumental variables estimation method. The asymptotic normality and convergence rates of estimators are proved under some mild conditions. In addition, we illustrate the finite sample performance of the proposed estimation method through simulation study and a real data analysis.
Functional regression allows for a scalar response to be dependent on a functional predictor; however, not much work has been done when response variables are dependence spatial variables. In this paper, we introduce a new partial functional linear spatial autoregressive model which explores the relationship between a scalar dependence spatial response variable and explanatory variables containing both multiple real-valued scalar variables and a function-valued random variable. By means of functional principal components analysis and the instrumental variable estimation method, we obtain the estimators of the parametric component and slope function of the model. Under some regularity conditions, we establish the asymptotic normality for the parametric component and the convergence rate for slope function. At last, we illustrate the finite sample performance of our proposed methods with some simulation studies.
This paper studies the estimation and inference of a partially linear varying coefficient spatial autoregressive panel data model with fixed effects. By means of the basis function approximations and the instrumental variable methods, we propose a two-stage least squares estimation procedure to estimate the unknown parametric and nonparametric components, and meanwhile study the asymptotic properties of the proposed estimators. Together with an empirical log-likelihood ratio function for the regression parameters, which follows an asymptotic chi-square distribution under some regularity conditions, we can further construct accurate confidence regions for the unknown parameters. Simulation studies show that the finite sample performance of the proposed methods are satisfactory in a wide range of settings. Lastly, when applied to the public capital data, our proposed model can also better reflect the changing characteristics of the US economy compared to the parametric panel data models.
在教育信息化和全球化的时代背景下,以及我国"双一流"高校建设的浪潮下,从新的人才培养目标和培养途径开展大学数学多元化教学新模式的探索势在必行.本文通过分析现有教学模式存在的不足,提出了充分利用网络教育资源和在线学习平台,探索线下课堂教学和线上学生自主学习相结合,并有机融合数学实验、数学建模、思政教育的大学数学多元化教学新模式.
In this article, minimum average variance estimation (MAVE) based on local modal regression is proposed for partial linear single-index models, which can be robust to different error distributions or outliers. Asymptotic distributions of the proposed estimators are derived, which have the same convergence rate as the original MAVE based on least squares. A modal EM algorithm is provided to implement our robust estimation. Both simulation studies and a real data example are used to evaluate the finite sample performance of the proposed estimation procedure.
In this paper, we consider the variable selection problem in functional linear regression with interactions. Our goal is to identify relevant main effects and corresponding interactions associated with the response variable. Heredity is a natural assumption in many statistical models involving two-way or higher-order interactions. Inspired by this, we propose an adaptive group Lasso method for the multiple functional linear model that adaptively selects important single functional predictors and pairwise interactions while obeying the strong heredity constraint. The proposed method is based on the functional principal components analysis with two adaptive group penalties, one for main effects and one for interaction effects. With appropriate selection of the tuning parameters, the rates of convergence of the proposed estimators and the consistency of the variable selection procedure are established. Simulation studies demonstrate the performance of the proposed procedure and a real example is analyzed to illustrate its practical usage.
The heterogeneous treatment effect (HTE) is estimated by using the semiparametric regression method. Firstly, a flexible semiparametric single-index model is considered by assuming the nonparametric link function and the interaction between treatment and covariates, and the index parameter vector and the unknown link function are estimated by using the rMAVE method. Then a HTE estimator can be obtained based on the estimators of index parameter vector and the link function. The consistency and asymptotic normality of the HTE estimator are established under some regularity conditions. Secondly, a hypothesis test is developed for the existence of HTE, and the bootstrap procedure is utilized to evaluate the null distribution of test statistic. Finally, simulation studies and a real data analysis are conducted to assess the performance of our proposed method.
In this paper, we develop and study a novel testing procedure that has more a powerful ability to detect mean difference for functional data. In general, it includes two stages: first, splitting the sample into two parts and selecting principle components adaptively based on the first half-sample; then, constructing a test statistic based on another half-sample. An extensive simulation study is presented, which shows that the proposed test works very well in comparison with several other methods in a variety of alternative settings.
部分线性模型是一类非常重要的半参数回归模型,由于它既含有参数部分又含有非参数部分,与常规的线性模型相比具有更强的适应性和解释能力.文章研究带有局部平稳协变量的固定效应部分线性面板数据模型的统计推断.首先提出一个两阶段估计方法得到模型中未知参数和非参数函数的估计,并证明估计量的渐近性质,然后运用不变原理构造出非参数函数的一致置信带,最后通过数值模拟研究和实例分析验证了该方法的有效性.
本文研究协变量随机缺失下异方差半参数变系数模型约束估计问题.首先在完全数据情形下,利用profile最小二乘方法构造模型参数和非参数分量的约束估计量;其次利用非参数核估计方法构造方差函数的约束估计量;随后基于逆概率加权法和加权profile最小二乘法构造模型参数和非参数分量的自适应逆概率加权profile最小二乘约束估计量;最后在一定正则条件下证明自适应逆概率加权profile最小二乘约束估计量的渐近性质,并通过蒙特卡洛数值模拟验证有限样本表现.
This paper considers the estimation for a partial index additive regression model, when the response variable and covariates in the index part are observed with additive distortion measurement errors. For the index parameter, the dimension-reduction based estimators with or without additive distortion measurement errors are proposed. This new estimation method is further adopted to the partial linear models for parameter estimation. We study the asymptotic properties of the proposed estimators. Simulation studies are conducted to compare the proposed estimation methods.
In this paper, we study the estimation for the partial linear single-index varying-coefficient model, which is a natural extension of the partially linear varying-coefficient model. A stepwise estimation procedure is developed to obtain asymptotic normality estimators of the index parameter vector, the coefficient parameter vector, and the coefficient function vector. The asymptotic properties of the resulting estimators are established under some conditions. A simulation study is conducted to assess the performance of the stepwise estimation procedure and the results show that our proposed procedure performs well in finite samples. Furthermore, a real data example is also used to illustrate our proposed method.