
This study compares parameter estimation methods for the unit inverse Weibull distribution under ranked set sampling (RSS) and simple random sampling (SRS) techniques. We examine Maximum Product Spacing Estimation, Ordinary Least Squares Estimation, Maximum Likelihood Estimation, Weighted Least Squares Estimation, Anderson-Darling Estimation, Left-Tail Anderson-Darling Estimation, Right-Tail Anderson-Darling Estimation, Cramer-von Mises Estimation, Minimum Spacing Absolute Distance Estimation, Minimum Spacing Square Distance Estimation, Minimum Spacing Absolute-Log Distance Estimation, and Minimum Spacing Square Log Distance Estimation. Monte Carlo simulations evaluate estimator performance using mean squared error (MSE), bias (AB), and mean absolute relative error (MARE). A COVID-19 dataset validates the practical applicability of the methods. Results show RSS-based estimators consistently outperform SRS counterparts across all metrics and estimation techniques. RSS demonstrates superior accuracy, reduced AB and lower MSE, particularly in small sample scenarios. These findings establish RSS as the preferred approach for unit inverse Weibull parameter estimation, providing significant improvements in statistical efficiency and reliability for practical applications.
Public health represents nowadays one of the major global challenges, especially with regards to the predictive analysis of biometric data. In this framework, Body Mass Index (referred to as BMI ) represents one of the most recognized indicators for monitoring, understanding and forecasting the general health of the population. An analysis of BMI could have important social and economic impacts, allowing policy-makers to foster more and more sustainable economic development and to develop preventive strategies (and effective health policies) aimed at promoting social well-being. This contribution proposes the use of computational approaches to model BMI using both traditional statistical methods and AI techniques, namely linear regression, random forest, decision trees and neural networks. In particular, a comparative analysis of the predictive performance of the models mentioned above is proposed by discussing the significance of different health-related indicators on BMI. The computational analysis is conducted on a dataset consisting of health parameters of a sample of women in Belgium collected between 2000 and 2001. The results demonstrate the effectiveness of linear and AI-based in BMI and valuable information for makers interested in quantitative assessment of health parameters and disease prediction to promote strategic choices that can contribute to improving collective well-being and reducing health and economic disparities, generating benefits at both the individual and community levels.
MultiClass Classification (MCC) is a foundational task in machine learning, especially within high-dimensional domains like text classification. While most studies focus on predictive accuracy, the growing demands of large-scale models have raised urgent questions about their environmental cost. In line with the Green AI paradigm, this work examines not only the performance but also the carbon and energy efficiency of various classifier-strategy combinations for MCC. Using two real-world textual datasets we systematically evaluate strategies such as One-Vs-Rest (OVA), One-Vs-One (OVO), Best-of-Best (BOB), and Error-Correcting Output Codes (ECOC), across classifiers ranging from simple Na & imath;ve Bayes to complex Artificial Neural Networks. We introduce emissions-per-accuracy metrics to measure the environmental efficiency of each configuration. Our findings show that while models like Random Forest incur high computational and ecological costs, simpler classifiers such as Logistic Regression and Na & imath;ve Bayes achieve comparable performance with drastically lower emissions. OVA consistently offers the best trade-off between speed and accuracy, while OVO and BOB prove more robust to class imbalance. Notably, Threshold-based Na & imath;ve Bayes paired with OVO demonstrates strong performance and sustainability. By integrating environmental considerations into MCC evaluation, this study highlights the importance of choosing classifier-strategy pairs that are not only effective but also computationally and ecologically responsible.
In this article, we introduce a new three-parameter Frechet distribution via the modified Lehmann Type-II class of distributions and investigate its properties, inferential methods, and real-world application. Fundamental distributional properties such as the quantile function, moments, moment generating function, entropy, and order statistics have been discussed. Inferential results have been obtained within the classical and Bayesian frameworks, utilizing a progressive censoring scheme. In the classical estimation framework, maximum likelihood estimation and maximum product spacing estimation are considered to obtain the estimates using the Newton-Raphson method. In addition, approximate confidence intervals have been derived using maximum likelihood estimation estimates. Meanwhile, we have used informative and non-informative prior via likelihood and product spacing functions to find the Bayes estimates in the Bayesian framework. Since the posterior distributions cannot be expressed in closed form, we employ a combination of Gibbs sampling and the Metropolis-Hastings algorithm to obtain the Bayes estimates. Furthermore, credible intervals are constructed using the Bayes estimates under informative and non-informative Prior. A comprehensive simulation study is carried out to assess the performance of the proposed estimation techniques. To demonstrate the practical utility of the proposed model, a real-life dataset is analyzed, showcasing the effectiveness of the proposed methodologies.
This study aimed to classify the cancer types across different regions in Jordan using a machine learning-algorithm and based on the Decision Tree Analysis (DTA) technique. The model employed patients' demographic information-specifically gender, age, and region of residence-as independent variables (IVs) to assess their interaction with cancer types as the dependent variable (DV). The objective was to determine the predictive relationship between these demographic factors and cancer types to support regional cancer profiling and inform targeted public health planning. The research utilized secondary data from the Ministry of Health, the Directorate of Non-Communicable Diseases, and the Jordan Cancer Registry for the years 2020 to 2021. A total of 9,547 cancer cases were analyzed using the DTA model, which effectively identified significant incidence patterns, with the central region of Jordan accounting for the highest number of cases (n = 6,815; 71.4%). The model classified cancer into 25 distinct types based on demographic attributes, with breast cancer being the most prevalent, particularly among middle-aged females residing in the central region. The DTA model demonstrated high efficacy in handling and stratifying large-scale medical data, predicting cancer type interactions, categorizing and labeling datasets, and suggesting potential category mergers. These findings have important implications for the development of focused cancer prevention strategies and the efficient allocation of healthcare resources. However, a key limitation of the study is the incomplete characterization of cancer patient attributes across all Jordanian regions.
This study investigates tourist engagement with garden visitation during health crises, aiming to understand the evolving role of public green spaces in fostering well-being and mitigating disruption. Utilising a self-completion questionnaire, data was collected from 536 randomly selected respondents across seven public gardens in Irbid, Jordan, during the COVID-19 pandemic. Analysis included descriptive statistics, reliability and regression analysis, and various comparative tests. The typical visitor was identified as a young, highly educated woman, often married with children, who had visited the gardens previously. Key motives for visiting were identified as escape, stress reduction, enjoyment, safety, prior experience, and appreciation of natural beauty. Despite pandemic-induced visitation declines, the research reveals consistent high visitor satisfaction, with a significant majority expressing intentions to revisit and recommend the gardens. Regression analysis conclusively identifies facilities & infrastructure (as the predominant factor), the on-site environment (cleanliness, tranquillity, natural beauty), and visitors' escape/relaxation motive as key drivers of this satisfaction. However, critical deficiencies in essential services (e.g., restrooms, accessibility, first-aid, parking, and diversified marketing) were also identified, hindering optimal engagement. The study underscores the pivotal function of green spaces in public well-being during health crises and provides actionable recommendations for enhancing garden management and visitor experience. These insights are crucial for developing resilient urban planning and tourism strategies in response to future health crises.
The paper examines the financial market as a potential environment for the transformative power of the metaverse. The primary hypothesis is to construct future scenarios by investigating the metaverse's influence on the financial returns of companies involved in its development. The focus lies on analyzing the relationship between the Meta-verse Index (MVI) returns and the returns of metaverse-oriented companies, with the aim of predicting emerging socio-technological trends. A dataset comprising daily closing prices of the time series from 2019 to 2023 was collected, including the MVI and 47 metaverse-related assets classified into 13 business areas. A two-step methodological approach was adopted: 1) correlation network analysis and 2) graph embedding strategy performed on correlation networks. The results highlight that the current scenario, characterized by a strong connection between MVI, technologies, crypto currencies, and real estate, which defines the meta-economy and digital property, will play a pivotal role in the future. The forecasts emphasize the development of metaverse-native enterprises, the creation of new stock market indexes designed to assess metaverse performance, and the development of customized intellectual property for their business models.
In this study, we investigate novel identities involving the incomplete gamma function through the application of probabilistic techniques. By examining the distribution of order statistics derived from the gamma distribution, we establish integral identities that incorporate expressions of the incomplete gamma function. Incomplete Gamma function identities can be derived by integrating out order statistic densities. These findings provide deeper insight into the function's analytical structure and hold practical relevance. Notably, we leverage these results to construct bivariate gamma distributions, demonstrating their utility in statistical modeling.
In this article, a group sampling scheme for lot sentencing is developed under time-censoring when the lifetime of a product follows the odd-Perks-Lomax distribution. The test plans are constructed by limiting a linear combination of the producer and consumer risks from frequentist and Bayesian frameworks. Integer nonlinear programming is used to designate the optimal number of groups and acceptance limit. Several tables and figures are constructed to scrutinize the performance of the proposed testing strategies. The proposed optimal test plans outperform the traditional optimal two-point plan in terms of sample size. Furthermore, using prior information for defectives proportion increases the effectiveness of the proposed plans. A numerical example is provided to demonstrate the application of the introduced scheme.
Situations are often encountered, especially in the medical sciences, where observing each stage of an event is necessary and overlooking it might be risky for the well-being of an individual. Keeping the same very viewpoint, this article presents the analysis of real Modified Rankin score data with multiple responses from a Bayesian perspective using a polytomous logistic regression model. The study involves utilizing the Markov chain Monte Carlo technique for acquiring samples from the resulting posterior distribution. Finally, to check the scope of the model simplification, several covariates are tested against zero and then a comparison between the full model and the simplified model is proposed based on the deviance information criterion.
Missing data is a problem that often arises in a variety of real-world systems. The performance of classification algorithms operating on these systems would suffer as a result. Effective imputation approaches abound to tackle this issue in case of missing data with low dimensions. In addition, one of the most common methods for concurrently doing variable selection and coefficient estimation in high-dimensional data is the penalized regression technique. However, one of the most significant problems associated with high-dimensional data is that it often includes an enormous quantity of missing data, which means that conventional imputation methods may not adequately address it. This paper proposes the imputation of missing values with the adaptive elastic net as an extension of penalized techniques to enhance gene selection and impute missing values in high-dimensional data. The effectiveness of the proposed method is evaluated by applying it to high-dimensional datasets that are taken from real-world situations with varying numbers of features, sample sizes, and percentages of missing datasets. A comparison is made between the proposed approach and various imputation-penalized methods that are currently in use for high-dimensional data. The findings of the comparison experiments reveal that the proposed technique is superior to its rivals since it achieves a better value for classification accuracy, sensitivity, and specificity than its competitors.
This study explores the integration of trigonometric functions into traditional statistical models, focusing on the development of the Weibull Sine Generalized (WSG-G) family of distributions. A special case was formulated name Weibull Sine generalized exponential (WSG-E) distribution. This new distribution extends the baseline exponential distribution, accommodating heavier tails and outliers, thereby effectively modeling positively skewed data. Key statistics such as mean, variance, skewness, and kurtosis indicate the distribution's capacity to handle clustered data. A simulation study demonstrates the performance of Maximum Likelihood Estimation (MLE), revealing convergence in the mean squared error and root mean squared error for the parameter alpha with increasing sample sizes, although convergence is less evident for other parameters. The WSG-E distribution's applicability is further illustrated through its fitting of medical datasets on bladder cancer remission times and growth hormone deficiency in children, both characterized by extreme values. Overall, the WSG-E distribution proves to be a robust model for skewed data, and future research could extend this framework to additional continuous distributions.
Frequent injuries pose a problem in professional soccer that is being tackled with preventive measures. Consequently, injury prediction and prevention are also increasingly addressed from a statistical perspective. In a pilot study, several machine learning algorithms and conventional statistical approaches have been compared regarding their potential to predict time-loss non-contact lower-body injuries in professional youth soccer players, using data from a prospective cohort study with 56 players of which 22 were injured. The covariates considered here include basic soccer-related as well as neuromuscular and biomechanical features derived from physical testing. Lasso regularized logistic regression, naive Bayes, linear discriminant analysis, k-nearest neighbors, classification trees, random forests, XGBoost, and support vector machines are considered for binary classification and prediction of an injury occurrence. The prediction results from a cross-validated procedure are compared regarding multiple quality measures. Post Lasso logistic regression with a reduced penalty gives the best results with an accuracy of 0.625, a predictive likelihood of 0.593, and a Brier score of 0.228. The respective sensitivity and specificity are 0.773 and 0.529, with an AUC of 0.672. Moreover, an XGBoost model slightly outperforms the Lasso model in terms of accuracy (0.661), while for the other performance measures it is dominated by the Lasso. In addition to the specific results on the available injury data set, the proposed comparison procedure of several models for binary prediction provides a generally applicable analysis guideline. This roadmap can also be applied in other contexts where similarly structured small but rich data sets are available.
Partial Least Squares Structural Equation Modeling (PLS-SEM) is a powerful statistical approach that has become a mainstream method in many application areas. It offers flexibility in handling formative and reflective measurement blocks, enabling researchers to model relationships among observed and latent variables. The crucial step in this approach is the PLS-SEM algorithm, which involves computing the scores of latent variables by alternating between inner and outer estimation. The aims of the present paper are twofold. The first contribution shows that the computations used in the outer estimation are inappropriate for reflective blocks. The second contribution involves introducing an alternative algorithm to overcome this drawback by using a new strategy based on considering the true structure of reflective blocks. Numerical studies and empirical simulations are provided to illustrate the advantages of the proposed algorithm compared to the classical one.
The determinants and correlates of income distribution have received significant attention in economics and public policy literature over recent decades. Income distribution, representing the share of income received by each quintile or decile of a population expressed as a vector of nonnegative proportions that sum to one, is inherently compositional data. However, most research has traditionally used aggregate inequality measures, such as the Gini coefficient, as the dependent variable when modeling relationships with economic indicators. Unlike a compositional data analysis (CoDA) approach, this reliance on aggregate measures limits insights into the tradeoffs among income classes as inequality determinants change. To date, only one study has applied a logratio-based model to analyze the determinants of income inequality in the U.S., leaving substantial gaps in understanding the broader implications of CoDA in income studies. To address this, our study proposes a Dirichlet regression model for country-level income distribution, integrating relevant economic and development indicators. This model aims to identify key determinants of income inequality and assess their specific impacts on income shares across different income groups. The performance of our proposed model is compared against a traditional Gini-based model, highlighting its potential for more nuanced and comprehensive insights into income distribution dynamics.
Studies have proven that the ridge estimator proves itself as a highly desirable shrinkage tool for addressing multicollinearity issues. A widely used model called negative binomial regression model (NBRM functions ciently when count data contains overdispersion properties. Maximum likelihood estimator (MLE) produces coefficients whose variance becomes affected negatively by multicollinearity issues. The proposed paper introduces generalized ridge estimator to resolve the shortcomings of ridge estimator. Various approaches to estimate the shrinkage matrix have been developed. Monte Carlo simulation findings demonstrate that the proposed estimation technique produces superior MSE results than traditional MLE estimates and ridge estimates regardless of the selected shrinkage matrix estimation methodology. The estimating methods used for shrinkage matrices result different levels of performance enhancement.
The probability distribution is of great significance in probability theory, which is inherent in virtually all the branches of science. It is said to be used selectively in actuarial science with reference to insurance and finance, medicine, agriculture, demography and econometrics. However, the main contribution of the current research work is to propose a new distribution called as neutrosophic exponentiated power Lomax distribution or briefly NEPL. Several other mathematical characteristics that describe life survival and the related characteristics, such as hazard rate and functions and moment-generating functions and other tests of mean, variance, and standard deviation, asymmetry and kurtosis, have been built and analyzed. Monte Carlo method has been applied also to assess the efficiency of NEPL distribution estimate. Therefore, the results of the simulation carried out for this study reveal that the process of estimating with reasonable degree of accuracy is feasible only when the size of the sample is comparatively large. The existence of the premature infant staying time data has been utilised to illustrate the specific manner in which the elaborated NEPL distribution has been suggested for being applied. Based on the discussions of the previous sections, it can be deduced that the NEPL distribution is also general in terms of its applications because it can deal with all forms of data that is, it does not distinguish between certainty, probabilities of uncertainties, ambiguties or imprecisions.
In industrial contexts where managing customer service costs is critical, accurately predicting and analyzing these costs presents a significant challenge, particularly when dealing with zero-inflated count data. This study proposes a Two-Stage Machine Learning approach that extends the traditional hurdle model, offering enhanced flexibility and adaptability to complex data structures without compromising interpretability. Through a real-world case study in the cleaning service sector focused on one-time service purchases, the proposed method identifies key cost drivers and provides actionable insights into customer behavior. This research advances the field by presenting a highly effective method for analyzing zero-inflated data, outperforming popular models based on Poisson distribution. Simultaneously, it addresses practical business needs by supporting data-driven strategies to optimize operational resources and manage customer costs more effectively.
In this paper, we introduce a new two-parameter extension of the Teisisier distribution using the Topp-Leone distribution as a generator, namely ToppLeone Teissier distribution. The new model exhibits increasing, decreasing and bathtub shaped hazard rate functions. Several properties of the model are derived utilizing the Lambert W, the generalized integro-exponential and the incomplete generalized integro-exponential functions. Maximum likelihood and Bayesian procedures are used to estimate the model parameters. Lindley's approximation under squared error loss function is utilized for Bayesian computations. Moreover, a simulation study is carried out to analyze the performance of these estimators on the basis of mean squared error. The applicability of the proposed model is evaluated using two real data sets. Also, we highlight the neutrosophic approach on Topp-Leone Teissier distribution as a pathway to address issues related to indeterminate, vague, or uncertain data set. This enhancement integrates neutrosophic logic into the model parameters, providing a robust framework for addressing the inherent challenges of ambiguous data and thereby broadening the model's applicability to real-world scenarios involving incomplete or imprecise information.
In scenario analysis, collinearity is a big issue in analyzing such relationship as between the response variable and several explanatory variables. As for these difficulties, the linear regression model, often traditionally, offers a range of shrinkage estimators. One such estimator is the ridge estimator. Thus, in order to fit count data with over-dispersion, for the bell regression model, this paper presents an improvement of the new Ridge-type estimator. Judging from the Monte Carlo simulation and the application of the Bell regression model, it was noted that the proposed estimate yields on average a smaller mean squared error than the other candidate estimators.