
In this paper, we investigate a new approach of fuzzy regression analysis based on support vectors when the available data and error variable are fuzzy quantities. In this approach, based on the concept of the distance between two parallel hyperplanes, we obtain the marginal hyperplanes and then, based on some constraints on the fuzzy data, we present an optimization problem to estimate the parameters of fuzzy regression model. The proposed method is investigated in two cases: with fuzzy fixed error and with fuzzy variable errors. To evaluate the proposed support vector fuzzy regression (SVFR) models, we present two indices of goodness of fit. Based on these indices, the presented SVFR models are compared with some other approaches on the numerical and simulated examples.
The SIRD (Susceptible-Infected-Recovered-Deceased) model is a standard framework for analyzing infectious disease dynamics. Classical continuous-time Markov formulations assume constant transition rates and memoryless (exponential) sojourn-times, which may oversimplify empirical epidemic processes. In this study, the Markov SIRD model is employed as a baseline, with transition parameters estimated analytically via maximum likelihood. To assess the validity of the exponential sojourn-time assumption, a semi-Markov framework is introduced exclusively for duration analysis of the infected state. Specifically, the sojourn-times associated with recovery (I-R) and death (I-D) transitions are modeled using exponential and Weibull distributions and compared using likelihood-based criteria. Using COVID-19 data from the Special Region of Yogyakarta, Indonesia, the results show that Weibull distributions provide a substantially better fit than the exponential assumption for both recovery and mortality durations. These findings indicate significant deviations from the memoryless assumption underlying the Markov model. This study does not construct a full dynamic semi-Markov epidemic simulator; instead, the semi-Markov framework is used to statistically characterize and evaluate the temporal structure of infected-state durations. The results highlight the importance of realistic sojourn-time modeling for understanding epidemic progression and for assessing the limitations of classical Markov-based epidemic models.
The accurate diagnosis of infectious diseases such as COVID-19 requires statistically reliable classification methods capable of handling complex, heterogeneous, and imbalanced data. In this study, several statistical and machine learning algorithms-Logistic Regression, Linear Discriminant Analysis, K-Nearest Neighbors, Decision Tree, and Random Forest-were comparatively evaluated using clinical and laboratory data from 506 hospitalized patients in Rafsanjan, Iran. The dataset included 27 categorical and 11 quantitative variables. To address class imbalance, the Synthetic Minority Over-sampling Technique (SMOTE) was employed. Model performance was assessed using a comprehensive set of criteria, including accuracy, sensitivity, specificity, positive and negative predictive values (NPV), and the area under the ROC curve. The comparative analysis showed that Random Forest and Logistic Regression achieved the best overall performance, while SMOTE improved sensitivity and NPV at the expense of specificity. The findings emphasize the importance of appropriate imbalance correction and multi-metric evaluation in developing statistically robust diagnostic models for medical data.
The generalized linear models (GLMs) use Gamma-Pareto regression Model (G-PRM) to address the sensitivity of influential observations. Difference of Fits (DF-FITS) is a popular technique for identifying influential observations. We apply DFFITS to the G-PRM with various residuals. We present illustrative real and simulated data. One class of adjusted Pearson residuals is more effective in detecting influential observations, considering small or large dispersion parameters. We calculate detection percentages to evaluate the proposed procedure's performance, replicating the process 10,000 times.
The prediction of credit risk is of great economic importance for banks and financial institutions, leading to the utilization of various methods in developing predictive models. This study introduces a credit risk prediction model that combines the support vector machine (SVM) with a genetic algorithm (GA) to aid credit decision-making by managers. While SVM is a reliable classification method, its performance can be influenced by factors such as model shape, parameter setting, and feature selection. To address these challenges, a novel approach is proposed that employs GA to optimize feature selection and parameter settings within the SVM framework. The proposed model is compared against alternative models including neural network, logistic regression, random forest, and decision tree. The study utilizes data from Bank of Yazd Province, with a sample size of 1876 customers divided into two groups: those who defaulted on their credit obligations and those who fulfilled them. The results demonstrate that the GA-SVM model serves as a suitable alternative for credit risk prediction, outperforming other models in terms of predictive power. Furthermore, the proposed model offers the benefit of feature selection, enabling financial institutions to identify potential risks and implement preventive measures. The use of GA in conjunction with SVM also facilitates the identification of optimal SVM parameter values, thereby enhancing the overall performance of the model. In conclusion, the proposed GA-SVM model emerges as a valuable tool for credit decision-making and risk management within banks and financial institutions. Further optimization can be achieved by exploring other meta-heuristic optimization algorithms.
In this paper, we investigate the stochastic properties of spacings among order statistics derived from a sample of independent, non-negative random variables that are divided into two groups with different distributions. Previous studies have shown that when these distribution functions are exponential distributions with specified hazard rates, the likelihood ratio ordering holds among the spacings under specific conditions. The present work extends these results by considering more general continuous distribution functions. We identify the necessary conditions on the parent distribution functions for preserving the likelihood ratio ordering among spacings in general settings. The comparison results enhance our understanding of stochastic ordering theory and provide valuable insights for applications in reliability, survival analysis, and related fields, aiding in the development of more flexible and accurate statistical models.
In this paper, we propose a generalized cross entropy between the the ith order statistic and the parent random variable X, defined using the quantile function. This method is more flexible than traditional PDF-based measures, particularly in situations where estimating the underlying density is difficult or unreliable. We investigate the properties of this measure and present examples to illustrate these concepts. Furthermore, we introduce a residual version of the quantile based generalized cross entropy between the the ith order statistic and the parent random variable X, along with some characterization results. Comparative analyses using simulation and real data indicate that the proposed measure provides improved interpretability and robustness relative to the quantile-based Kerridge inaccuracy measure. This study effectively connects theoretical development with practical application, contributing to the field of statistical analysis.
In design of experiments, optimal design is an important approach that maximizes the chances of experimental success. A- and D-optimality are well-known criteria for identifying optimal designs. In nonlinear models, these criteria depend on unknown parameters, complicating the design derivation. This paper uses the Bayesian method to address this, deriving A- and D-optimal designs for EMAX, log-linear, and LINEXP models with three or four parameters, using uniform priors. Optimal designs with minimum support points are obtained, with varying weights. These designs serve as benchmarks for evaluating practical alternative designs. Two alternatives were assessed, showing over 80% efficiency in most models compared to A- and D-optimal designs. The computations in this study were performed using a numerical nonlinear approach, specifically the NLPSolve method, which is included in the Optimization package in Maple software.
The term "functional data" refers to data where the units of observation are functions defined over a time interval. The fundamental philosophy behind functional data is that the repeated measurements for each individual are considered as a stochastic process over time. One of the commonly used analyses for such data is functional principal component analysis. In this study, since the intracranial pressure was measured over time in patients with aneurysmal subarachnoid hemorrhage, functional principal component analysis was employed to identify the main factors contributing to increased intracranial pressure. The first four functional principal components account for 87.8 percent of the total variation in the intracranial pressure curve. The first, second, third, and fourth principal components explain approximately 52.3, 21.9, 8, and 5.6 percent of the overall variation, respectively. These four components are linked to the total Glasgow Coma Scale score, diastolic blood pressure, age, and systolic blood pressure, respectively.
Recently Alizadeh Noughabi and Shafaei Noughabi (2024) introduced some estimators for the varextropy of an absolutely continuous random variable. In this paper, we propose other nonparametric estimators for the varextropy function. Additionally, we prove asymptotic properties of two estimators given in Alizadeh Noughabi and Shafaei Noughabi (2024). Asymptotic properties of the proposed estimators are established under suitable regularity conditions. Moreover, a simulation study is performed to compare the performance of the proposed estimators based on mean squared error (MSE) and bias. Furthermore, by using the proposed estimators some tests are constructed for uniformity. It is shown that the varextropy-based test proposed in this paper performs well in terms of power when compared to other uniformity hypothesis tests. Real datasets are utilized to evaluate the performance of the varextropy estimators.
This paper investigates the optimal risk management strategies in a general compound Poisson risk model consisting of safety loading of insurer and reinsurance to minimize the infinite-time ruin probability. Its price process is perturbed by a geometric Brownian motion with the drift and volatility of risky asset. In addition, we allow this company to buy proportional reinsurance to reduce the underlying risk and invest its surplus in a risky asset whose price is driven by correlated Brownian motions. We focus on the possibility of an insurance company utilizing optimal controls and study the optimization problem of minimizing the infinite-time ruin probability in a financial market. For the diffusion approximation of risk model, we obtain an analytic expression for the minimum nfinite-time ruin probability and the corresponding optimal controls by using the martingale approach. Since it is not easy to derive the explicit expression for the infinite-time ruin probability of perturbed risk model, we obtain the identically distributed and have an exponentially decaying tail. Moreover, we study the effect of investment on the ruin probability in both perturbed risk models. Finally, some numerical examples are conducted to illustrate the effects of model parameters on the optimal risk management strategies and on the financial market.
Heavy-tailed distributions have recently gained prominence in science, economics, and industry as robust alternatives to the Gaussian distribution, particularly for modeling data with extreme variability or outliers. Several studies in the literature have introduced and examined the Pakes generalized Linnik distribution and its related distributions due to their flexibility in capturing heavy-tailed behavior. However, most existing works focus exclusively on the special case of symmetric random variables-a significant limitation, given that real-world data often exhibit skewness. To address this gap, this paper proposes a new class of generalized skewed Linnik distributions, extending previous symmetric models to accommodate asymmetric data structures. We investigate their theoretical properties, including moments, tail behavior, and stability under linear transformations. Furthermore, we develop an autoregressive (AR) model based on this framework, enabling time-series analysis with skewed, heavy-tailed innovations. Additionally, we introduce a novel class of geometric skewed Linnik distributions, which arise as the limit of random sums and exhibit unique dependence structures. The practical utility of these models is demonstrated through theoretical derivations and potential applications in finance, risk assessment, and signal processing. Our results broaden the scope of Linnik-based models, offering more accurate tools for skewed, heavy-tailed data analysis.
The escalating public health costs are a significant concern for governments globally. The efficient management of those costs is critical, with health insurance systems playing a pivotal role. However, the insurance industry faces challenges due to the heterogeneous data, leading to inconsistent outputs for identical inputs. Traditional predictive methods such as Artificial Neural Networks and Adaptive Neuro-Fuzzy Inference Systems (ANFIS) often fail to address these inconsistencies. This study proposes a novel two-stage model to determine insurance premiums, incorporating equity considerations and advanced computational techniques. We advocate for an expenditure-based premium calculation as a superior alternative to the traditional salary-based approach. This method aligns premiums more closely with household expenses, promoting fairness and efficiency. Our results demonstrate that the expenditure-based strategy outperforms the salary-based one in controlling costs for both the insured and the insurer. Specifically, the error metrics, including Mean Absolute Error and Root Mean Square Error, show significant improvement in our model compared to the ANFIS method. To enhance the model's accuracy, we integrate sampling techniques to mitigate the data heterogeneity and employ genetic algorithms to optimize the weights of the neural network. The genetic algorithm iteratively evolves the network parameters, ensuring robust performance even in diverse data. Our findings indicate that this integrated approach significantly reduces prediction errors and enhances the overall reliability of the premium calculation process. In conclusion, the proposed model offers a robust framework for premium determination, addressing the inherent data heterogeneity in the insurance industry. This study provides a valuable contribution to the field by demonstrating a practical and effective solution for improving the accuracy and fairness of insurance premium calculations.
Errors in factor levels often occur in response surface modeling. A design in which these errors have minimal effect is the desired design. This study evaluates the prediction capability of second-order Orthogonal Array Composite Design (OACD) and Orthogonal Uniform Composite Design (OUCD) with and without errors in factor levels for 3 <= k <= 5 factors using 2, 3, and 5 center points in the cuboidal region. Design optimality criteria (in terms of G-and IV-optimality values) and quantile dispersion plots are used to examine the prediction capability of these designs. The results show that OUCD is the preferred design in terms of G-optimality, while IV-optimality and quantile plots indicate that OACD is the preferred design in both the presence and absence of errors in factor levels.
The problem of small area estimation is how to produce reliable estimates of characteristics of interest such as means, counts and quantiles. It is usually assumed that the observed values and the auxiliary values follow the linear regression model and the sampling errors are dependent and follow the autoregressive model. However, in practice, there are many situations in the observed values and the auxiliary values follow the non-linear regression model. We assume that the true model is unknown and consider some non-nested, non-linear or linear regression models as rival models and select an optimal model based on extensions of the model selection tests such as and model selection based on latent variables and proposes a global model selection test for small-area estimation. A numerical example and real data analysis were carried out to illustrate the procedures obtained theoretically.
This work extends an existing multivariate homogeneously weighted moving average (MHWMA)-control chart to a multivariate double homogeneously weighted moving average (MDHWMA)-control chart aimed at a more efficient monitoring of the process mean vector. Like the MHWMA-control chart, the MDHWMA-control chart statistic assigns a specific weight to the current observation, and the remaining weight is evenly assigned among the previous observations but unlike the MHWMA-control chart, the MDHWMA-control chart statistic utilizes the information contained in the observations twice. We present the design structure of the MDHWMA-control chart and on the basis of the average run length, standard deviation of the average run length and the median run lengths (ARL, SDRL & MRL) compare the performance with the MHWMA-control chart in relations to Hotelling's chi(2)-chart, multivariate cumulative sum (MCUSUM)-chart and the multivariate exponentially weighted moving average (MEWMA)-chart. The comparison showed that the proposed (MDHWMA)control chart has a better performance than the competing charts especially for small shifts. .
This paper investigates the reliability and parametric inference for the inverse power Maxwell distribution under progressive Type-II censored sample. Under the frequentist approach, the maximum likelihood estimate, least square, and weighted least square methods are considered for estimating the model parameters and any parametric function involved in this model. Approximate confidence intervals for parameters and any of their functions are created via a variance-covariance matrix. Bayes estimates are obtained using Lindley's approximation and Markov chain Monte Carlo (MCMC) technique under squared error loss function. Additionally, the highest posterior density (HPD) credible intervals are constructed using MCMC approximation techniques. A comprehensive Monte Carlo simulation study is conducted to assess the efficiency of the proposed methodologies. Furthermore, three optimality criteria are presented to choose the most suitable progressive scheme from various sampling plans. The practical utility of the proposed methods is demonstrated using two real-world datasets: the failure times of mechanical components and the strength of glass fiber.
In this paper, we first consider the propagation characteristics of a spatiotemporal Gaussian pulse using both simplified analytical and numerical approaches. Second, we focus on the counterintuitive presence of a possible dimple property associated with these spatio-temporal Gaussian pulses and the corresponding intensity functions. The analytical results are supported with numerical simulations of the exact pulsed-beam solution and various plots.
In this paper, we propose a class of bivariate distributions as a general solution to a functional equation. This general class of distributions proposed includes many well studied bivariate distributions. It also enjoys a proportional reversed hazards model for the distribution of the component-wise maxima. Characterizations of this general class based on a functional equation, conditional mean and conditional variance are studied. The simulation algorithm to generate bivariate pairs from the members of this general class is provided. It is shown that these properties find applications in developing simple univariate procedures in lieu of complicated bivariate goodness of fit procedures for members of the proposed class. The univariate goodness of fit procedure for the American Football dataset of the National Football League has been illustrated.
The restricted mean survival time (RMST) in the context of length-biased data is an important addition to clinical studies. The RMST is a widely used measurement for evaluating survival over a specific period, and the area under the survival function is a key component of this metric. However, when the data under study are length-biased, traditional parametric and classical methods for examining the RMST are not applicable. Nonparametric and semi-parametric methods are used to address this issue. We utilize the empirical likelihood (EL) method to investigate RMST. Our proposed EL procedure provides a reliable approach for inferential analysis of RMST in the presence of length-biased data. We have shown that the limiting distribution of the empirical log-likelihood ratio is a chi-square distribution with one degree of freedom. We also demonstrated that the likelihood ratio exhibits weak convergence to a mean-zero Gaussian process, which we used to construct a confidence band. In our simulation section, we compared the confidence intervals obtained from the normal approximation (NA) and EL methods. We showed that the EL method has a better coverage probability than the NA method. Additionally, we provided a real data application using bank customers' monthly taxes to illustrate further the effectiveness of our proposed method.