This paper presents a comprehensive investigation of statistical inference and predictive analysis for two Poisson-exponential distributions using a joint type-II censored sample. The Poisson-exponential distribution is widely used in reliability and survival analysis where failure times follow an exponential distribution within each groups. Parameter estimation is performed using the expectation-maximization algorithm, with approximate confidence intervals derived from the observed Fisher information matrix. Bayesian estimators are obtained via importance sampling under squared error, linear-exponential, and generalized entropy loss functions, along with corresponding credible intervals. A shrinkage pretest estimator combining Bayesian and maximum likelihood approaches is also proposed. The paper provides the best unbiased and Bayesian predictors for future failure times based on the observed joint type-II censored data, including point and interval predictions. Extensive simulations across various sample sizes evaluate the proposed methods, and their practical application is illustrated using two real datasets.
Univariate and multivariate general linear regression models, subject to linear inequality constraints, arise in many scientific applications. The linear inequality restrictions on model parameters are often available from phenomenological knowledge and motivated by machine learning applications of high-consequence engineering systems (Agrell, 2019; Veiga and Marrel, 2012). Some studies on the multiple linear models consider known linear combinations of the regression coefficient parameters restricted between upper and lower bounds. In the present paper, we consider both univariate and multivariate general linear models subjected to this kind of linear restrictions. So far, research on univariate cases based on Bayesian methods is all under the condition that the coefficient matrix of the linear restrictions is a square matrix of full rank. This condition is not, however, always feasible. Another difficulty arises at the estimation step by implementing the Gibbs algorithm, which exhibits, in most cases, slow convergence. This paper presents a Bayesian method to estimate the regression parameters when the matrix of the constraints providing the set of linear inequality restrictions undergoes no condition. For the multivariate case, our Bayesian method estimates the regression parameters when the number of the constrains is less than the number of the regression coefficients in each multiple linear models. We examine the efficiency of our Bayesian method through simulation studies for both univariate and multivariate regressions. After that, we illustrate that the convergence of our algorithm is relatively faster than the previous methods. Finally, we use our approach to analyze two real datasets.
In this paper, a novel hybrid framework, called the Sparse and Pruned Approach to Random Forest (SPARF), is proposed that enhances prediction accuracy by automatically pruning the ensemble of trees generated by the RF algorithm. Unlike traditional RF, the proposed framework applies non-concave penalties, namely SCAD and GSCAD, to identify and eliminate redundant trees. The core innovation lies in integrating SCAD-based techniques with RF and aggregating remaining trees. This model leverages the sparsity of SCAD and GSCAD to provide an interpretable model with high predictive accuracy. The performance of this method is evaluated on two real datasets and Monte Carlo simulation, where a combined RF, SCAD, and GSCAD model is created to reduce and automatically select RF trees. In real-world datasets, the SPARF-SC model achieves an approximately 9.68% reduction in RMSE compared to RF, while SPARF-GSC achieves a reduction of approximately 8.46% in RMSE.
Recent studies in machine learning are based on models in which parameters or state variables are restricted by a restricted boundedness. These restrictions are based on prior information to ensure the validity of scientific theories or structural consistency based on physical phenomena. The valuable information contained in the restrictions must be considered during the estimation process to improve the accuracy of the estimation. Many researchers have focused on linear regression models subject to linear inequality restrictions, but generalized linear models have received little attention. In this paper, the parameters of beta Bayesian regression models subjected to linear inequality restrictions are estimated. The proposed Bayesian restricted estimator, which is demonstrated by simulated studies, outperforms ordinary estimators. Even in the presence of multicollinearity, it outperforms the ridge estimator in terms of the standard deviation and the mean squared error. The results confirm that the proposed Bayesian restricted estimator makes sparsity in parameter estimating without using the regularization penalty. Finally, a real data set is analyzed by the new proposed Bayesian estimation method.
The Bell regression model (BRM) is a statistical model that is often used in the analysis of count data that exhibits overdispersion. In this study, we propose a Bayesian analysis of the BRM and offer a new perspective on its application. Specifically, we introduce a G-prior distribution for Bayesian inference in BRM, in addition to a flat-normal prior distribution. To compare the performance of the proposed prior distributions, we conduct a simulation study and demonstrate that the G-prior distribution provides superior estimation results for the BRM. Furthermore, we apply the methodology to real data and compare the BRM to the Poisson and negative binomial regression model using various model selection criteria. Our results provide valuable insights into the use of Bayesian methods for estimation and inference of the BRM and highlight the importance of considering the choice of prior distribution in the analysis of count data.
This paper investigates the use of shrinkage estimators in the generalized Poisson hurdle (GPH) model for count data analysis. The GPH model effectively handles data with both excess zeros and over- or underdispersion. We propose shrinkage estimators to improve parameter estimation in this model and analyze their asymptotic properties, including biases and risks. An extensive comparison through Monte Carlo simulations evaluates the efficacy of the suggested estimators against the maximum likelihood estimator, employing a simulated relative efficiency criterion. In addition, we apply the estimators to two real-world datasets. Our findings illustrate that the suggested shrinkage estimators yield superior results compared to the traditional maximum likelihood estimator.
Ensembling is a powerful technique to obtain the most accurate results. In some cases, the large number of learners in ensemble learning mostly increases both computational load during the test phase and error rate. To solve this problem, in this paper we propose an Ensemble of Reduced Deep Regression (ERDeR) model, which is a combination of Deep Regressions (DRs), shrinkage methods, and ensemble approaches. The framework of the proposed model contains three phases. The first phase includes base regressions in which parallel DRs are used as learners. The role of these DRs is to extract features of input data and make prediction. In the second phase, to automatically reduce and select the most suitable DRs, shrinkage methods such as Least Absolute Shrinkage and Selection Operator (LASSO) and Elastic Net (EN) are employed. These models are compared with the non-shrinkage model. The last phase is ensemble phase, which consists of three different ensemble methods namely Multi-Layer Perceptron (MLP), Weighted Average (WA), and Simple Average (SA). These ensemble methods are used to aggregate the remaining learners from previous steps. Finally, the proposed model is applied to Monte Carlo simulation data and three real datasets including Boston House Price, Real Estate Valuation and Gold Price per Ounce. The results show that after applying the shrinkage methods the error rate is significantly reduced and the model accuracy is increased. Accordingly, the results of combining shrinkage methods and ensemble approaches not only decreased the computational load during test phase, but also increased the model accuracy.
The prevalence of high-dimensional datasets has driven increased utilization of the penalized likelihood methods. However, when the number of observations is relatively few compared to the number of covariates, each observation can tremendously influence model selection and inference. Therefore, identifying and assessing influential observations is vital in penalized methods. This article reviews measures of influence for detecting influential observations in high-dimensional lasso regression and has recently been introduced. Then, these measures under the elastic net method, which combines removing from lasso and reducing the ridge coefficients to improve the model predictions, are investigated. Through simulation and real datasets, illustrate that introduced influence measures effectively identify influential observations and can help reveal otherwise hidden relationships in the data.
In modeling count data with overdispersion and extra zeros, zero-inflated negative binomial (ZINB) regression model is useful. In a regression model, the multicollinearity problem arises when there are some high correlations between predictor variables. This problem leads to the maximum likelihood method will not be an efficient estimator. The ridge and Liu-type estimators have been proposed to combat the multicollinearity problem so that the Liu-type estimator is better. In this paper, we proposed the Liu-type shrinkage estimators, namely linear shrinkage, preliminary test, shrinkage preliminary test, Stein-type, and positive Stein-type Liu estimators to estimate the count parameters in the ZINB model, when some of the predictor variables have not a significant effect to predict the response variable so that a sub-model may be sufficient. The asymptotic distributional biases and variances of the proposed estimators are nicely demonstrated. We also compared the performance of the Liu-type shrinkage estimators along with the Liu-type unrestricted estimator by using an extensive Monte Carlo simulation study. The results show that the performances of the proposed estimators are superior to those based on Liu-type unrestricted estimators. We also applied the proposed estimation methods to Expenditure and Default Data.
In this paper, we propose the application of shrinkage strategies to estimate coefficients in the Bell regression models when prior information about the coefficients is available. The Bell regression models are well-suited for modelling count data with multiple covariates. Furthermore, we provide a detailed explanation of the asymptotic properties of the proposed estimators, including asymptotic biases and mean squared errors. To assess the performance of the estimators, we conduct numerical studies using Monte Carlo simulations and evaluate their simulated relative efficiency. The results demonstrate that the suggested estimators outperform the unrestricted estimator when prior information is taken into account. Additionally, we present an empirical application to demonstrate the practical utility of the suggested estimators.
We suggest Stein-type estimators for zero-inflated Bell regression models by incorporating information on model parameters. These estimators combine the advantages of unrestricted and restricted estimators. We derive the asymptotic distributional properties, including bias and mean squared error, for the suggested shrinkage estimators. Monte Carlo simulations demonstrate the superior performance of our shrinkage estimators across various scenarios. Furthermore, we apply the suggested estimators to analyze a real dataset, showcasing their practical utility.
In the field of chemical data modeling, it is common to encounter response variables that are constrained to the interval (0, 1). In such cases, the beta regression model is often a more suitable choice for modeling. However, like any regression model, collinearity can present a significant challenge. To address this issue, the Liu-type estimator has been used as an alternative to the maximum likelihood estimator, but it suffers from bias. In this paper, we introduce the Jackknifed Liu-type estimator and its modified version, which demonstrate improved bias reduction compared to the original Liu-type estimator. We assess the theoretical and numerical performance of these estimators through Monte Carlo simulations and real-data examples from the field of chemistry. Our findings highlight the significant improvements offered by the proposed estimators in terms of accuracy and reliability.
This paper, using the signature technique and a generalized Farlie-GumbelMorgenstern (FGM) copula function, presents a generic mean residual lifetime (MRL) model for the reliability analysis of a load-sharing coherent system. The present approach differs from earlier models in that in addition to load-sharing phenomenon it simultaneously considers the effect of operating conditions on the system. Further, using the developed model and the renewal-reward argument, an age replacement policy is investigated. The proposed MRL model and the behavior of the optimal solution as the model parameters change are illustrated through numerical examples.
Objectives In 2015, the Iranian Ministry of Health and Medical Education (MoHME) developed and introduced an operational plan (OP) mandatory for implementation for all medical sciences universities. Through this program, a set of indicators were defined for the annual assessment of the universities' performance. This study examined the effect of OP implementation on universities' performance and the healthcare system.Methods We compared seven key performance indicators before and after OP implementation related to public health, education, research, nursing, treatment, food and drug and student affairs. Descriptive statistics, paired t-tests and Wilcoxon tests were used for data analyses using SPSS 26.Results Of 32 studied indicators, ten indicators were not affected by the OP. Of 22 indicators, including public health (7 indicators), education (2 indicators), research (2 indicators), nursing (5 indicators), treatment (1 indicator), food and drugs (4 indicators), and student (1 indicator) that changed, 17 have improved.Conclusion OP implementation reached some of its intended goals. The existing OP program process should be replaced with evidence-informed health policies and programs to achieve significant changes in health system functioning and overall public health.
In this paper, we consider the multicollinearity problem in the gamma regression model when model parameters are linearly restricted. The linear restrictions are available from prior information to ensure the validity of scientific theories or structural consistency based on physical phenomena. In order to make relevant statistical inference for a model any available knowledge and prior information on the model parameters should be taken into account. This paper proposes therefore an algorithm to acquire Bayesian estimator for the parameters of a gamma regression model subjected to some linear inequality restrictions. We then show that the proposed estimator outperforms the ordinary estimators such as the maximum likelihood and ridge estimators in term of pertinence and accuracy through Monte Carlo simulations and application to a real dataset.
This article is improved the random forest algorithm by selecting the most appropriate penalized regression methods, and it is tried to improve the post-selection boosting random forest (PBRF) algorithm using elastic net regression. The proposed method with the highest efficiency is called Reducing and Aggregating Random Forest Trees by Elastic Net (RARTEN). The introduced method consists of three steps. In the first step, the random forest algorithm is used as a predictor. In the second step, Elastic Net, as a penalized regression method, is applied to reduce the number of trees and improve the random forest and PBRF. In the last step, selected trees are aggregated. The obtained results of the real data and Monte Carlo simulation are evaluated using various statistical performance criteria. The simulation study shows that the RARTEN with 7%, 5%, and 8.5% reduction in the linear, nonlinear, and noise model, respectively improve the accuracy of the traditional random forest and the proposed method by Wang. In addition, this method has a significant reduction compared to other penalized regression methods. Moreover, the real data results show that the proposed method in our study with a reduction of almost 16% confirms the validity of the proposed model.
To assess the potential of wind energy in a specific area, statistical distribution functions are commonly used to characterize wind speed distributions. The selection of an appropriate wind speed model is crucial in minimizing wind power estimation errors. In this paper, we propose a novel method that utilizes the T-X family of continuous distributions to generate two new wind speed distribution functions, which have not been previously explored in the wind energy literature. These two statistical distributions, namely the Weibull-three parameters-log-logistic (WE3-LL3) and log-logistic-three parameters-Weibull (LL3-WE3) are compared with four other probability density functions (PDFs) to analyze wind speed data collected in Tabriz, Iran. The parameters of the considered distributions are estimated using maximum likelihood estimators with the Nelder-Mead numerical method. The suitability of the proposed distributions for the actual wind speed data is evaluated based on criteria such as root mean square errors, coefficient of determination, Kolmogorov-Smirnov test, and chi-square test. The analysis results indicate that the LL3-WE3 distribution demonstrates generally superior performance in capturing seasonal and annual wind speed data, except for summer, while the WE3-LL3 distribution exhibits the best fit for summer. It is also observed that both the LL3-WE3 and WE3-LL3 distributions effectively describe wind speed data in terms of the wind power density error criterion. Overall, the LL3-WE3 and WE3-LL3 models offer a highly accurate fit compared to other PDFs for estimating wind energy potential.
The main objective of this paper is to apply linear and pretest shrinkage estimation techniques to estimating the parameters of two 2-parameter Burr-XII distributions. Further more, predictions for future observations are made using both classical and Bayesian methods within a joint type-II censoring scheme. The efficiency of shrinkage estimates is compared to maximum likelihood and Bayesian estimates obtained through the expectation-maximization algorithm and importance sampling method, as developed by Akbari Bargoshadi et al. (2023) in "Statistical inference under joint type-II censoring data from two Burr-XII populations" published in Communications in Statistics-Simulation and Computation". For Bayesian estimations, both informative and non-informative prior distributions are considered. Additionally, various loss functions including squared error, linear-exponential, and generalized entropy are taken into account. Approximate confidence, credible, and highest probability density intervals are calculated. To evaluate the performance of the estimation methods, a Monte Carlo simulation study is conducted. Additionally, two real datasets are utilized to illustrate the proposed methods.
In this article, we improve parameter estimation in the zero-inflated Poisson regression model using shrinkage strategies when it is suspected that the regression parameter vector may be restricted to a linear subspace. We consider a situation where the response variable is subject to right-censoring. We develop the asymptotic distributional biases and risks of the shrinkage estimators. We conduct an extensive Monte Carlo simulation for various combinations of the inactive predictors and censoring constants to compare the performance of the proposed estimators in terms of their simulated relative efficiencies. The results demonstrate that the shrinkage estimators outperform the classical estimator in certain parts of the parameter space. When there are many inactive predictors in the model, as well as when the censoring percentage is low, the proposed estimators perform better. The performance of the positive Stein-type estimator is superior to the Stein-type estimator in certain parts of the parameter space. We evaluated the estimators' performance using wildlife fish data.