
One of the most popular methods for assessing the distribution tail thickness parameter (tail index) of a two-parameter PD is the Hill estimator, derived from order statistics. For reliability evaluations and engineering design, an accurate assessment of the distribution’s tails is crucial. Since the original Hill estimator is not shift- and scale-invariant, it must be modified in practice to include at least shift and scale parameters. Note that in a real-world engineering application, distribution tails—such as the Weibull type—are thin rather than thick (heavy). Fortunately, Weibull-type distributions can be estimated using a modified Hill-type estimator. The original two-parameter Weibull-type distribution suffers from the drawback that it resembles the two-parameter PD; without incorporating a shift model parameter, it lacks a shift parameter and is not shift-invariant. By adding shift parameters, the basic two-parameter Weibull distribution is transformed into a three-parameter form. Even so, fitting distributions with only three-parameters does not provide sufficient flexibility for most engineering and design applications; in other words, they often do not fit the underlying raw data. In many facets of statistical inference for rare events, selecting the threshold for the tail distribution is crucial. A novel method for selecting tail distribution thresholds is presented in this study. This study introduces a tailor-made four-parameter Weibull distribution with added q -parameter that bridges it with PD and Extreme Value Theory (EVT) power parameters and thus enables Inverse Hill Statistic (IHS) extension from thick to thin tail distributions. We present a semi-analytical solution that improves accuracy of the design values and failure probability P_F estimates, when the underlying data sample is of limited size. Offshore wind speed dynamics from industrial applications in renewable energy were selected for comparison and validation of the proposed model parameter estimators. Accurate modeling of the ambient wind speed distribution is critical for estimating in situ wind energy potential.
The pivotal role of private corporate investment in facilitating India’s economic expansion is widely recognized. Nevertheless, the recent slowdown in corporate investment has posed a significant challenge to India’s sustained growth. This study investigates the firm-level determinants of private corporate investment in India using an unbalanced panel of manufacturing firms. It examines how sales growth, bank borrowing, borrowing costs, and internal cash flows influence investment behaviour across heterogeneous firms, including listed and unlisted companies of varying sizes. The analysis combines fixed-effects panel regressions, controlling for unobserved firm heterogeneity, with a panel vector autoregression (VAR) framework to capture dynamic interrelationships among key variables. The findings highlight the critical role of indicators like sales growth, bank borrowing, and internal cash flows in shaping private corporate investment, with implications for policy aimed at promoting investment-led economic growth. Moreover, machine learning techniques are employed to capture non-linearities and complex firm-level interactions, providing further evidence that financing costs and credit conditions are critical determinants of firms’ investment decisions.
The discrete point objects of a Dataset are transformed to continuous functions by the proposed method of Reverse Discretization. Each point is defined by its coordinates xr and yr and it carries the information mr (mass, price, velocity etc.). The information is treated as a field extending on the x-axis from -∞ to +∞. The field is expressed by the proposed function D_r(x)=m_rλ _r/π [1+λ _r^2(x-x_r)^2] which is continuous and symmetric about xr. It is easy to prove that -∞ ∞ ∫ D_r(x)dx=m_r , which is a fundamental property of the introduced field. It is also apparent that its intensity Dr(x) is controlled by the parameter λr. A large value of this parameter makes the intensity of the field highly concentrated around xr and quickly diminishing away from xr, tending asymptotically to zero. Thus, a Dataset of discrete data is transformed to a continuous function. If the Dataset has n data points of interest located within a space from X1 to X2 then outside this space, for a large λr, the intensity Dr(x) is negligeable and F12 =[ ∑ _r=1^nX_1X_2∫ D_r(x)dx] ≈ ∑ _r=1^n-∞ ∞ ∫ D_r(x)dx . The transformation of a Dataset of point-like discrete objects, described above, yields a continuous differentiable function that lends itself to the operations of Calculus and can have many applications. One application is the proposed New Regression Analysis. In the classical method, the Regression line is defined by n nodes with arbitrary x-coordinates, which remain constant, and the y-coordinates that are variable. Thus, the least-square-error (SSE) is a function of y, and it is optimized by differentiation with respect to y only. In the new Regression Analysis by Reverse Discretization, the nodes of the regression line have an additional degree of freedom and F12 is a function of x and y. With one more degree of freedom the optimization yields a better regression line. The obtained regression line does not simply indicate the trend of the Dataset but also provides an interpretation of the data by revealing the underlying generating law. In two examples given in this article the data points are distributed randomly around the segments of a line, which is the generating law. Only the xr- and yr- coordinates and the intensity mr of the information carried by each point were given in the analysis. The regression line obtained by this method revealed, with acceptable accuracy, the line that generated the data points. Thus, it provides an insight on the properties, and the trend of the data. The interpretive, predictive and forecasting potential of this method can find applications in several disciplines like Engineering, Statistics, Artificial Intelligence, Medicine, Economics, Marketing and Polling.
This paper introduces the smooth transition duration model, designed to model the dependence of duration on explanatory variables, allowing the duration time to vary with smooth transitions over different regimes. The proposed model is a generalisation of parametric survival regression models, and it makes it possible to detect nonlinear behaviour when the response of interest is the duration time until some event occurs. A Lagrange multiplier (LM) test is derived together with the maximum likelihood estimators of the smooth transition duration model. The practical use of the introduced model is exemplified by assessing the time between abnormal price increases in the electricity spot prices in Queensland, Australia. A deregulation process might have led to a change in the behaviour of the market participants, and the smooth transition duration model is used to detect and examine such possible transitions. The results show clear support for a gradual change in the appearance of abnormal price increases.
The traditional ranked set sampling (RSS) scheme can be viewed as an alternative to the simple random sampling (SRS) scheme in many lifetime scenarios. Over the years, numerous RSS-type methods have been proposed in the literature such as the ones with unequal sample sizes. This paper employs four sampling methods: the traditional RSS plan, two RSS-type plans with unequal sample sizes, and the SRS method to estimate the parameters of the exponentiated Shanker (E-Sh) distribution. The E-Sh distribution is particularly valuable for modeling lifetime phenomena due to its increasing or bathtub-shaped hazard rate function. Both classical and Bayesian frameworks are employed to derive point and interval estimates for the parameters. Since the Bayesian estimates seem to lack closed-form expressions, the Metropolis-Hastings within Gibbs algorithm is utilized to approximate these estimates. A simulation study is conducted to assess the performance of the different sampling strategies. The findings reveal that the traditional RSS method generally performs better than the others in classical estimation and Bayesian estimation under approximate non-informative priors. However, the RSS-type methods with unequal sample sizes prove competitive under Bayesian estimation with informative priors. A real data application involving diameter at breast height (DBH) data is also presented to illustrate the practical utility of the methods. The paper concludes with some final remarks.
This study models the impact of factors influencing the level of learning in semester system of education in Pakistan using the logistic regression model. The model is evaluated in classical and Bayesian frameworks. Maximum likelihood estimates of the model parameters are evaluated along with their standard errors and the significance of the estimate are reported therein. To keep the results using non-informative priors, we have used uniform prior to let the data speak for themselves. Marginal posterior distributions of the model parameters are also evaluated and discussed. The odd ratios of the response variable are computed to assess the impact of unit change in a particular explanatory variable or covariate on the learning probability of the students while keeping other covariates fixed. The goodness of the model is examined. It is observed that level of learning increases by an increase in satisfaction, marks in previous exams, class level, type of institution, regularity in taking classes, study hours per day and use of technology. Whereas demographic background, marital status, disease, residence mode, father’s qualification, family problems, and parents’ involvement in student’s study casts an adverse effect on the learning level. The model is proved to give good fit to the data.
The case presented refers to the Aspect entangled spin in photon pairs experiment. In the paper it is demonstrated that the conclusion of the experiment is based on a statistical flaw.
Symbolic data analysis can provide statistical inferences for macroscale data while preserving as much information as possible from microscale data. In this study, we focus on the symbolic interval-valued regression model. The microdata are reorganized into intervals by using the largest and smallest order statistics. Afterward, we develop innovative symbolic interval-valued regression models to construct the relationships between two or more intervals. Owing to the properties of order statistics, we maintain the natural order in which a higher value of the dependent variable is larger than its lower value. First, we develop a simple linear symbolic interval-valued regression model and derive the corresponding maximum likelihood estimators (MLEs). In addition, we describe the Fisher information matrix of the MLEs and show that they demonstrate asymptotic normality. Next, we extend the aforementioned model to a multiple linear symbolic interval-valued regression model, and the corresponding MLEs are again derived. Monte Carlo simulations and real data analysis confirm the validity of the proposed method.
Non-mydriatic retinal images, akin to neuroimaging, often suffer from artifacts and blurring that significantly inhibit the accurate diagnosis and early detection of neurodegenerative diseases. Current image enhancement methods, such as Robust Principal Component Analysis (RPCA), are commonly employed for early disease detection and medical diagnosis. However, these methods typically assume uniform singular value weights with existing nuclear norms, an assumption that may not hold because of the inherent variations in noisy image data. In addition, RPCA approaches primarily focus on global enhancement, failing to capture the fine details critical for precise analysis of retinal images. This limitation underscores the need for advanced and more adaptive method that can address these challenges and improve the accuracy of diagnostic tools in the biosciences, particularly for the early detection of diseases. To address these drawbacks, we propose a novel method that combines RPCA, the truncated weighted nuclear norm (TWNN), and the equalization of adaptive histograms (AHE). Unlike existing methods, the proposed approach enhances degraded retinal images by further applying Gaussian filtering to correct the edges and AHE used to detail the retinal images, preserve key structures, and improve contrast. The method is formulated as an optimization technique, incorporating Histogram Oriented Gradients (HOG) features, with the enhanced images used for diabetes prediction via machine learning, and the parameters are updated iteratively using ADMM. We conducted ablation studies to select the best model that predicts diabetes from the machine learning algorithms, including Support Vector Machine, Logistic Regression, and K-Nearest Neighbors, on retinal fundus images across all severity levels using the proposed method. Among these, SVM demonstrated the highest performance in predicting diabetes, from which we considered it for the classification purpose in this study. .One of the interesting aspect of this study is that, unlike the existing methods, we use the enhanced images produced by our approach for diabetes classification.The results of the study show that the proposed method substantially enhanced the quality of the images and improved prediction accuracy based on the public databases.
This paper develops an extended spherical fuzzy multi-criteria group decision-making (MCGDM) framework that integrates hybrid weighting with the total area method based on orthogonal vectors (TAOV). The main motivation is to address key limitations in existing MCGDM approaches by simultaneously accounting for decision-maker heterogeneity and criteria correlation, aspects that are typically overlooked in prior studies. To this end, the proposed model combines subjective and objective weights through a game-theoretic formulation, enabling a balanced representation of expert judgment and data-driven information. Furthermore, orthogonality among criteria is ensured via principal component analysis (PCA), thereby mitigating redundancy and enhancing interpretability within a unified mathematical structure. By integrating spherical fuzzy modeling with hybrid weighting and orthogonalization, the framework improves both the accuracy and robustness of group decision processes. The effectiveness and robustness of the proposed framework are illustrated through a real-world case study on capital-raising strategies, demonstrating its advantages over conventional spherical fuzzy MCGDM models.
This paper presents a comprehensive analysis of the bifurcation behavior of a two-dimensional predator-prey model described by a system of nonlinear difference equations. The study investigates the behavior of the system under different parameter regimes using both analytical and numerical techniques. Analytically, we employ a perturbation method to expand the system’s equations in a power series about the bifurcation point. This allows us to identify the types and scenarios of bifurcations that occur in the system, including period doubling, Neimark-Sacker, and strong resonance bifurcations. Numerical continuation methods implemented in MatcontM are used to validate the analytical results and gain further insights into the bifurcation structures.
Abstract Artificial intelligence is rapidly transforming financial decision-making across banking, lending, insurance, auditing, fraud detection, and customer-facing financial services. At the same time, its growing use has intensified ethical concerns related to fairness, accountability, transparency, privacy, trust, and human oversight. Although research on artificial intelligence (AI) in finance has expanded considerably, the literature remains fragmented across disciplinary and application-specific streams, limiting a consolidated understanding of its intellectual foundations and thematic development. This study provides a bibliometric and thematic review of research at the intersection of artificial intelligence, finance, and ethics. Drawing on Scopus-indexed journal articles and a PRISMA-guided screening process, a final sample of 338 articles published between 2000 and 2025 was analyzed using the bibliometrix package in R. The findings show that the field is young but rapidly expanding, particularly after 2019, with strong momentum in recent years. Intellectual structure analysis identifies foundational contributions centered on algorithmic fairness, accountability, explainability, and governance, while historiographic patterns reveal major developmental pathways in credit scoring, financial services, accounting and auditing, and generative AI. Conceptual and thematic analyses further show that the literature is organized into six interconnected clusters covering AI ethics and governance, algorithmic fairness, explainable AI, fraud detection, trustworthy AI, and human-in-the-loop financial decision-making. The study contributes a structured map of this emerging field and shows that AI in finance is increasingly understood not merely as a technical innovation, but as a socio-technical governance challenge requiring responsible design, institutional accountability, and sustained human oversight.
In this study, a new three-parameter Lomax distribution is proposed using the beta transformation technique. The new distribution is named Beta Transformed Lomax distribution. Some mathematical properties of the proposed distribution were derived, including moments, the moment-generating function, the mean residual life, quantile function, reliability function, failure rate, and order statistics. The model parameters were estimated using the maximum likelihood approach. The performance of the estimation method was evaluated using a simulation study. Two medical data sets were analyzed using a new distribution and the modeling performance was compared with other generalizations of Lomax distributions, and the results show that the new model performs better than others. Also, the Metropolis-Hastings approach within Bayesian estimation was used to obtain approximate Bayes estimates, and convergence diagnostics based on Markov Chain Monte Carlo techniques were conducted.
Likelihood-based inference for exponential-tail lifetime models can be unstable in finite samples, particularly under right-tail contamination. We develop a robust score-based inference procedure for the two-parameter Gompertz distribution by constructing estimating equations in which influence is bounded after standardisation by the local curvature of the likelihood, implemented via OPG-whitened score contributions and Huber weighting. Uncertainty is quantified using a Godambe (sandwich) covariance estimator, enabling Wald-type inference. Simulation results under clean sampling and controlled right-tail contamination show improved stability and more reliable coverage than maximum likelihood, with a moderate efficiency loss under correct specification.
Abstract Predicting unwanted pregnancies accurately may reduce abortion and population growth in a country. Almost half of India’s 48.1 million pregnancies were unintended (Lancet Global Health, 2018). This paper uses National Family Health Survey (NFHS-4, 2015-16) data to explore the prediction of unwanted pregnancy using various statistical and machine learning models. It uses under-sampling to improve the predictive power of the models, given the imbalanced distribution of unwanted pregnancy rates in the data. We have proposed a weighted variable to eliminate the effect of under sampling. Among the machine learning approaches tested, Random Forest stands out as the strongest for predicting unwanted pregnancies, delivering 80.35% accuracy and an AUC score of 0.86 on the test data. The model outputs suggest that unmet need for contraception, knowledge of contraception methods, age at first birth, wealth status, total children ever born, knowledge of ovulation cycle, place of residence, woman’s education, woman’s age, and marital duration are the strongest predictors to estimate the unwanted pregnancy. The trained model has also been validated and performed well on recent NFHS-5 data (2019-21). The model can be deployed by policymakers in the sub-regions of India in the future to predict the prevalence of unwanted pregnancy and run a customized campaign to reduce the events of unintended pregnancy through various medical and social interventions.
The idea commonly accepted is that probability distributions play a major role in reliability modeling, whereas classical models are often not flexible enough to accommodate the complex behavior of life data. This paper presents a more flexible modeling approach to reliability analysis through the introduction of the Inverse Gompertz Gumbel (IGoGum) distribution and Non-Homogeneous Poisson Process (IGoGum NHPP). The IGoGum-NHPP model, which is derived from the reciprocal transformation of the Gompertz-Gumbel distribution, accommodates skewness in positive and negative directions and allows for increasing and decreasing forms of hazard. The essential statistical properties of the model, including the survival and hazard functions, moments and order statistics, and the entropy of Bayesian risk, will be formulated along with parameter estimation through maximum likelihood, Bayesian, and hybrid Bayesian Neural Networks being optimized through Firefly Algorithms (BNN–FFA). Monte Carlo simulation and other model selection criteria (AIC, BIC, CAIC, HQIC, KS, MSE, RMSE) demonstrate that the proposed estimators perform very efficiently and to great effect. Applications to bladder-cancer remission time’s data and a data set regarding failures of diesel-engine turbochargers revealed that the IGoGum -NHPP model outperforms the others in terms of the fit and predictive capacity (invert Gompertz -Frankfurt -Vredenburg, Gompertz -Burr XII, and Gompertz -Lomax). Thus, the IGoGum family becomes a multifunctional and efficient tool to characterize lifetime modeling, reliability evaluation, and risk identification of complex systems.
This paper introduces a new and flexible family of distributions, namely the Gamma Topp-Leone-Heavy-Tailed-G distribution. Several statistical properties of the proposed family are derived, including the hazard rate function, quantile function, moments, moment generating function, Rényi entropy, distribution of order statistics and stochastic orderings. Parameter estimation is performed using various estimation techniques, including Maximum Likelihood (ML), Cramér-von Mises (CVM), Weighted Least Squares (WLS), and Least Squares (LS). In addition, the consistency and efficiency of the estimators are evaluated through a Monte Carlo simulation study. To demonstrate the flexibility and practical applicability of the proposed family, three real-world data sets from different application domains are analyzed.
Abstract Domestic violence is a deep-rooted societal issue impacting all aspects of women’s lives. The WHO reported that nearly 30% of women globally experience some form of violence. Early identification through a predictive model can lead to timely interventions. With this focus, the study aims to build a predictive machine-learning model to assess the incidence of domestic violence in India. Utilizing nationally representative data from the latest National Family Health Survey (NFHS-5), the study aims to identify and understand the patterns and key determinants of domestic violence in India. The statistical methodology involves exploratory data analysis, feature selection, model development and evaluation, and identifying feature importance. Seven different supervised classification machine learning algorithms are used, namely Logistic Regression, Support Vector Machine, Artificial Neural Networks, Random Forest, XG-Boost, Naive Bayes, and K-Nearest Neighbor. Results exhibit that the Logistic regression model is a more effective predictive model among all utilized machine learning models for assessing the prevalence of domestic violence in India. This finding contradicts the misconception that advanced machine learning algorithms consistently outperform the logistic regression model. Predictors like the partner’s control over the woman, alcohol consumption by the partner, the woman’s characteristics like their age, education, family history of violence, number of family members living together, and some of the socio-demographic predictors like region, caste, and wealth index are identified as major contributors to the incidence of domestic violence. We expect that the insights from the perception of machine learning will provide more impactful and technology-driven information to enhance interventions and policy initiatives aimed at eliminating domestic violence against women.
This study introduces a factorial extension of the Analysis of Means with Covariate (ANOMC), building on the classical ANOM framework by incorporating auxiliary covariate information. The proposed factorial ANOMC approach is designed for experiments involving two fixed factors and a continuous covariate, where traditional ANOM or factorial ANOM may lose efficiency. Six variants of the factorial ANOMC test are developed using regression and ratio-type estimators, and their performance is evaluated against the standard factorial ANOM test, which does not utilise covariate information. A comprehensive Monte Carlo simulation study is conducted to assess these tests under diverse conditions, including normal and non-normal error distributions, varying correlation structures, different sample sizes, multiple levels of each factor, and both homogeneous and heterogeneous variances. Performance is examined through empirical Type I error rates and statistical power. The findings show that while factorial ANOM maintains stable Type I error rates under ideal settings, several factorial ANOMC variants (i.e., ANOMC-MR1_F, ANOMC-MR2_F, ANOMC-MR4_F and ANOMC-Reg_F ) achieve improved detection capability, especially when regression estimators are used. Some ratio-based versions also perform well under specific correlation and distribution structures. Power increases noticeably for factorial ANOMC when sample sizes or factor-level combinations grow. Overall, the factorial ANOMC framework provides a more adaptable and informative alternative for multifactor experiments involving covariates.
A Block-Basu bivariate exponential (BBBE) distribution is one of the most popular and widely used absolutely continuous bivariate distributions. Later, Kundu and Gupta (Stat Method 7:464–477, 2010) obtained the Block-Basu bivariate Weibull (BBBW) distribution. Extensive work has been done on the BBBW model over the past several decades. Interestingly, it is observed that the BBBW model can be extended to the modified Weibull model. We call this new model as the Block-Basu bivariate modified Weibull (BBBMW) distribution. We consider the properties of the BBBMW distribution and provide the associated copula function. The BBBMW model has five unknown parameters and the maximum likelihood estimators (MLEs) cannot be obtained in closed form. To compute the MLEs directly, one needs to solve a five-dimensional optimization problem. We propose to use the EM algorithm for computing the MLEs of the unknown parameters. The proposed EM algorithm can be carried out by solving a two-dimensional optimization problem at each EM step. An extensive simulation is carried out, which demonstrates that the proposed EM algorithm performs quite well. A real data set is analyzed for illustrative purposes.