Quasi-likelihood nonlinear models with random effects (QLNMWRE) include generalized linear models with random effects and quasi-likelihood nonlinear models as special cases. In this paper, some regularity conditions analogous to those given by Breslow and Clatyton (1993) are proposed. On the basis of the proposed regularity conditions and Laplace approximation, the existence, the strong consistency and asymptotic normality of the approximate maximum quasi-likelihood estimation of the fixed effects are proved in QLNMWRE.
Oxidoreductase is an enzyme that widely exists in organisms. It plays an important role in cellular energy metabolism and biotransformation processes. Oxidoreductases have many subclasses with different functions, creating an important classification task in bioinformatics. In this paper, a dataset of 2640 oxidoreductase sequences was used to perform an analysis and comparison. The idea of dipeptides was introduced to process the Position Specific Score Matrix (PSSM), since each dipeptide consists of two amino acids and each column of PSSM corresponds to the information of one amino acid. Two kinds of dipeptide scores were proposed, the Standardization Normal Distribution PSSM (SND-PSSM) and the Correlation Coefficient PSSM (CC-PSSM). Recursive Feature Elimination (RFE) is used to extract features from the SND-PSSM and CC-PSSM, and the two sets of extracted features are combined to form a new feature matrix, the RFE-SND-CC-PSSM. The results show that, with the proposed method and a kernel-based nonlinear SVM classifier, the accuracy can reach 95.56% by the Jackknife test. Our method greatly improves the accuracy of oxidoreductase subclass prediction. Using this method to predict the categories of the 6 major types of enzymes effectively improves its prediction accuracy to 94.54%, indicating that this method has general applicability to other protein problems. The results show that our method is effective and universally applicable, and might be complementary to the existing methods.
Membrane proteins occupy an important position in the life activities of humans and other species. The elucidation of membrane protein types provides clues for understanding the structure and function of proteins. With the fusion of various protein information including amino acid classification, physicochemical property, and evolutionary information, this paper proposes a system for predicting membrane protein types. In this system, a new feature selection method called MIC-GA is proposed to deal with the curse of high-dimensional features. The findings show that this approach is effective in reducing feature dimensions and improves prediction accuracy. Ensemble method based on stacked generalization is also used to solve the problem of feature heterogeneity. The performance of the present method is evaluated on two benchmark datasets. The overall prediction accuracies of eight types are 89.23% and 93.49% using jackknife test and independent test, respectively. The final experimental results show that our method is more effective than the existing methods for prediction of membrane protein types.
Quasi-likelihood nonlinear models (QLNMs) are an extension of generalized linear model and include a widen class of models as special cases. This article investigates some diagnostic methods in QLNMs. An equivalency between a case-deletion model and a mean-shift outlier model in QLNM is established. Two simulation study and a real dataset are used to illustrate the proposed diagnostic methods.
In this paper, we establish the asymptotic properties of maximum quasi-likelihood estimator (MQLE) in quasi-likelihood non linear models (QLNMs) with stochastic regression under some mild regular conditions. We also investigate the existence, strong consistency, and asymptotic normality of MQLE in QLNMs with stochastic regression.
Count data with excess zeros are often encountered in many medical, biomedical and public health applications. In this paper, an extension of zero-inflated Poisson mixed regression models is presented for dealing with multilevel data set, referred as hierarchical mixture zero-inflated Poisson mixed regression models. A stochastic EM algorithm is developed for obtaining the ML estimates of interested parameters and a model comparison is also considered for comparing models with different latent classes through BIC criterion. An application to the analysis of count data from a Shanghai Adolescence Fitness Survey and a simulation study illustrate the usefulness and effectiveness of our methodologies.
在带自适应设计的拟似然非线性模型中,在响应变量的矩条件尽可能弱和其它正则条件下,证明了以概率为1,当n充分大时,拟似然方程有一个解βn,它收敛于参数真值β0.
This paper proposes some mild regularity conditions analogous to those given by Wu (1981) and Chang (1999). On the basis of the proposed regularity conditions, the strong consistency as well as convergence rate for maximum quasi-likelihood estimator (MQLE) is obtained in quasi-likelihood nonlinear models (QLNMs) with stochastic regression.
The quasi-likelihood function proposed by Wedderburn [Quasi-likelihood functions, generalized linear models, and the Gauss-Newton method. Biometrika. 1974; 61: 439-447] broadened the application scope of generalized linear models (GLM) by specifying the mean and variance function instead of the entire distribution. However, in many situations, complete specification of variance function in the quasi-likelihood approach may not be realistic. Following Fahrmeir's [Maximum likelihood estimation in misspecified generalized linear models. Statistics. 1990; 21: 487-502] treating with misspecified GLM, we define a quasi-likelihood nonlinear models (QLNM) with misspecified variance function by replacing the unknown variance function with a known function. In this paper, we propose some mild regularity conditions, under which the existence and the asymptotic normality of the maximum quasi-likelihood estimator (MQLE) are obtained in QLNM with misspecified variance function. We suggest computing MQLE of unknown parameter in QLNM with misspecified variance function by the Gauss-Newton iteration procedure and show it to work well in a simulation study.
In many practical applications, count data often exhibit greater or less variability than allowed by the equality of mean and variance, referred to as overdispersion/underdispersion, and there are several reasons that may lead to the overdispersion/underdispersion such as zero inflation and mixture. Moreover, if the count data are distributed as a generalized Poisson or a negative binomial distribution that accommodates extra variation not explained by a simple Poisson or a binomial model, then the dispersion occurs too. In this paper, we deal with a class of two‐component zero‐inflated generalized Poisson mixture regression models to fit such data and propose a local influence measure procedure for model comparison and statistical diagnostics. At first, we formally develop a general model framework that unifies zero inflation, mixture as well as overdispersion/underdispersion simultaneously, and then we mainly investigate two types of perturbation schemes, the global and individual perturbation schemes, for perturbing various model assumptions and detecting influential observations. Also, we obtain the corresponding local influence measures. Our method is novel for count data analysis and can be used to explore these essential issues such as zero inflation, mixture, and dispersion related to zero‐inflated generalized Poisson mixture models. On the basis of the results of model comparison, we could further conduct the sensitivity analysis of perturbation as well as hypothesis test with more accuracy. Finally, we employ here a simulation study and a real example to illustrate the proposed local influence measures. Copyright © 2012 John Wiley & Sons, Ltd.
This paper investigates K independent 2 × 2 contingency tables with structural zero in clinical and epidemiological studies. By using the measurement of risk ratio, we present three test statistics: the Wald-based, the logarithmic-transformation-based and the Score-based statistics to test the homogeneity hypothesis across strata. Power-controlled sample size formulae and corresponding numerical experiment results are obtained based on the above three tests. The simulation studies show that generally the sample sizes associated with the Score-based and the logarithmic-transformation-based tests are more accurate to guarantee the pre-specified power than the Wald-based one, which develop the hypothesis testing methods based on the contingency tables.
Structural equation models are widely used in social, educational, medical, marketing, and behavioral sciences. In these fields, missing data are commonly encountered. To deal with this problem, structural equation models with missing data have been proposed. In the application of this kind of models, one of the most important issues is model selection. In this paper, we proposed an alternative Bayesian criterion-based method called the Lv measure for model selection of structural equation models with missing data. A simulation study and a real example are presented to demonstrate the efficiency and the application of the Lv measure, and results based on Bayes factor in the real example are also presented to illustrate the performance of the Lv measure for model selection.
Incomplete correlated 2 × 2 tables are common in some infectious disease studies and two‐step treatment studies in which one of the comparative measures of interest is the risk ratio (RR). This paper investigates the two‐stage tests of whether K RRs are homogeneous and whether the common RR equals a freewill constant. On the assumption that K RRs are equal, this paper proposes four asymptotic test statistics: the Wald‐type, the logarithmic‐transformation‐based, the score‐type and the likelihood ratio statistics to test whether the common RR equals a prespecified value. Sample size formulae based on hypothesis testing method and confidence interval method are proposed in the second stage of test. Simulation results show that sample sizes based on the score‐type test and the logarithmic‐transformation‐based test are more accurate to achieve the predesigned power than those based on the Wald‐type test. The score‐type test performs best of the four tests in terms of type I error rate. A real example is used to illustrate the proposed methods.
This paper studies a compound interval hypothesis about risk ratio in an incomplete correlated \(2\times 2\) table. Asymptotic test statistics of the Wald-type and the logarithmic transformation are proposed, with methods of the sample estimation and the constrained maximum likelihood estimation (CMLE) being considered. Score test statistic is also considered for the interval hypothesis. The approximate sample size formulae required for a specific power for these tests are presented. Simulation results suggest that the logarithmic transformation test based on CMLE method outperforms the other tests in terms of true type I error rate. A real example is used to illustrate the proposed methods.
Quasi-likelihood nonlinear models(QLNM) include generalized linear models as a special case.This paper proposes some sufficient conditions of weak consistency of maximum quasi- likelihood estimator(MQLE) in QLNM,in which the condition of the moment is weaker than that of strong consistency of MQLE in the existing literature.
In this paper, the DEA models with undesirable outputs are considered and an effective bootstrap approach is developed for statistical inference. The most attractive point for our method is that it not only can allow statistical noise and inherent dependency to be investigated simultaneously, but also can further enhance the statistical foundation of DEA analysis with the case of undesirable outputs. Finally, an empirical example is used to illustrate our methodology proposed above.
The present paper proposes a semiparametric reproductive dispersion nonlinear model (SRDNM) which is an extension of the nonlinear reproductive dispersion models and the semiparameter regression models. Maximum penalized likelihood estimates (MPLEs) of unknown parameters and nonparametric functions in SRDNM are presented. Assessment of local influence for various perturbation schemes are investigated. Some local influence diagnostics are given. A simulation study and a real example are used to illustrate the proposed methodologies.
文章着重研究了带有有序分类变量的结构方程模型的模型选择问题,并将一个基于贝叶斯准则的统计量称为测度,应用到此类模型中进行模型选择。通过实例分析说明了上述方法的应用,并给出了根据贝叶斯因子进行模型选择的结果。
In this article, we consider a varying-coefficient reproductive dispersion linear model (VCRDLM). By the local likelihood approach, the estimates of parameters of interest are given, and the determination of the local weight and the smoothing parameter as well as some statistical inferences are investigated in VCRDLM. Our results can be regarded as a further generalization of the results in [11].