
In this study, we have derived the expression for Maximum Likelihood (ML) estimates and likelihood ratio test (LRT) for two Kumaraswamy populations under Joint Ranked Set Sampling (JRSS), Joint Modified Minimum Ranked Set Sampling (JMnRSS), Joint Modified Maximum Ranked Set Sampling (JMxRSS), and Joint Simple Random Sampling (JSRS). The performance of the ML estimates is evaluated using Root Mean Square Error (RMSE) and the bias criterion. LRT is conducted to test the equality of the first shape parameters of two Kumaraswamy populations when the other two shape parameters are known. The power of the LRT is determined to compare the test performance under the mentioned sampling schemes. This simulations for ML and LRT results are obtained using Monte Carlo simulations in R Studio. We found that all joint ranked sampling schemes demonstrate superior performance compared to JSRS in both ML estimation and LRT. However, the JRSS exhibited the best results in LRT when the other two shape parameters are known, while the JMnRSS excelled in ML estimation. Additionally, we provide an illustration using real-life lung cancer data to highlight these findings.
Traditionally, frailty models are built with the assumption that frailty influences the baseline hazard function in a multiplicative manner, known as a multiplicative frailty model. An alternate model available in the literature is the additive frailty model, in which the frailty variable is linearly related to the baseline hazard function. Although both models are popular, there is a need for a model that incorporates both multiplicative and additive models, especially in epidemiological data where neither the multiplicative nor the additive model adequately describes the data. This paper aims to address this gap by developing a family of models that include both multiplicative and additive frailty models under the inverse Gaussian frailty distribution. The paper also discusses the inference procedure for estimating model parameters using the MCMC method and applies the proposed model to real-life datasets.
T20 cricket is an exciting format characterized by explosive hitting and strategic play, engaging fans with each boundary and a wicket. A comprehensive dataset of T20 matches is analyzed to understand the factors affecting batsmen's performance in this highly dynamic format. Survival analysis approach is used to study the performance of the batsmen, measured in terms of 'number of runs' taken as the 'innings survival time'. In this context, dismissal of batsmen is taken as the 'event'. The dismissal may be due to getting Bowled, being Caught, LBW, Run out, Stumped or Hit wicket. These different forms of dismissal can be taken as 'competing risks' and this study specifically focuses on identifying factors associated with specific dismissals. In this process, Cumulative Incidence Function (CIF), Cause-Specific Hazard function (CSH), and Fine and Gray's Sub distribution Hazard function (SDH) were used. The results from this analysis offer insights into the game dynamics and aids in player's performance evaluation and strategic decision-making, such as, team composition, batting order and making the choice of 'batting first' or 'chasing'. Data for representative batsmen from the top ICC ranking of T20 game with specific inclusion and analysis was carried out using R Programming Language (R 4.5.0) with suitable packages.
This study proposes a novel statistical distribution called the power Gemeay distribution by modifying the Gemeay distribution. Various statistical properties like hazard rate function and its graphics, moments and related measures, order statistics, and incomplete moments are derived. Many estimation methods like the maximum likelihood method, Anderson-Darling estimation, right-tail Anderson-Darling estimation, and more are used to estimate the parameters. A simulation study is carried out to evaluate the efficiency of the estimation techniques. Based on the analysis of two real data sets, the novel distribution is more suitable than other existing models.
This essay explores the life and contributions of Andrei Nikolaevich Kolmogorov(1903-1987), one of the twentieth century's most influential mathematicians. Beginningwith the Borel-Kolmogorov paradox, we examine how Kolmogorov transformed probabilitytheory from a collection of informal methods into a rigorous mathematical framework. Wetrace his remarkable journey through the tumultuous Soviet era, his historic visit to theIndian Statistical Institute, Kolkata, and his profound contributions spanning probabilitytheory, turbulence, complexity theory, topology, and mathematical education. Kolmogorov'sideas continue to shape modern science, from stochastic modeling and statistical inference toturbulence theory and algorithmic information theory. This essay explores both the math-ematical contributions that reshaped probability theory and the historical context in whichthose ideas emerged
Count data appears in diverse field of study with a variety of patterns in certainfrequencies such as excessive occurrence of zeros and ones compared to other possible values.Here a new zero-one-inflated Poisson-Garima distribution and its applications is studied.After introducing the modified Poisson-Garima distribution, its various statistical propertiesand estimation of parameters are discussed. The estimation methods are then illustratedwith simulated samples and the proposed model was applied to two real data sets. It isdemonstrated that the new model outperformed its competitors like zero-one-inflated Poissondistribution and zero-one-inflated Poisson-Lindley distribution
Accurate detection of wheat spikes and reliable yield prediction are critical for optimizing crop production and resource management. This study presents an integrated framework for spike detection and yield estimation using pseudo-RGB images derived from hyperspectral data. A YOLOv8 model was trained on 1,050 images, achieving high precision, recall, and mean average precision values. The bounding boxes and masks generated byYOLOv8 were used to quantify spike count and spike area, while six vegetation indices were extracted from hyperspectral images acquired at the booting stage. Three multiple linear regression models were developed for yield prediction: one based on spike features, another on vegetation indices, and a third combining both. The combined model achieved the highest accuracy, with a five-fold cross-validation R2of 0.902 +/- 0.007, RMSE of 1.739 +/- 0.133 g,and MAE of 1.289 +/- 0.066 g. Compared with previous approaches, the proposed framework demonstrated improved performance, highlighting the value of integrating spike morphology and spectral data for yield prediction. Overall, the study shows that hyperspectral imaging can simultaneously provide morphological and physiological traits, reducing reliance on high-resolution RGB data in wheat phenotyping
This study introduces Conditional Dynamic Failure Extropy (CDFEX), a novel mea-sure for quantifying uncertainty in bivariate systems.CDFEX captures interdependencebetween components via the joint distribution.We develop nonparametric estimators forCDFE(X) and establish their asymptotic properties under mild conditions. Through simula-tions and real-world datasets, we demonstrate the robustness of the proposed estimators. Ourresults show thatCDFE(X) outperforms both univariate Dynamic Failure Extropy (DFEX)and Conditional Dynamic Cumulative Past Entropy, highlighting its potential to enhancereliability analysis in complex systems
A novel over dispersed count distribution is obtained by convolving independently distributed Poisson and transmuted geometric random variables. This distribution extends the Poisson-geometric, geometric, and transmuted geometric distributions. Essential statis-tical properties are analyzed. Maximum likelihood estimators of the unknown constants are derived using numerical optimization techniques and the EM algorithm. Extensive simulation studies evaluate the estimators performance under different conditions. Additionally, a flexible regression model built on the proposed distribution is formulated. Real-world applications in modeling over dispersed count data, both with and without covariates, demonstrate the model's relevance. The proposed model demonstrates desirable statistical properties and surpasses its closest competitors in empirical applications
A lot of ROC models have been derived in the area of classification to classify the subjects using many distributional assumptions based upon the nature of the data. Though, there are many models in the literature, still such models are of need current situation to address the different needs of the nature of data "dine to The Skewness or hon normal behavior of the data in reality. Therefore, this paper addresses such an ROC model which consists of the X Lindley (XL) distribution and this distribution has more flexibility than other one-parameter distributions. An attempt has been made to develop an ROC model in classification when the healthy population is at the higher side than the abnormal/diseased population. Further, simulations as well as real data set have been used to find out area under the curve. The simulations are done with various parameter values of the distribution, which will represent the better, moderate and worst case of classification in ROC analysis. The X Lindley distribution can be used quite effectively in analyzing the data in classification and which is easy to make computations even for a non statistician. Further, the properties of the XL ROC curve are verified mathematically and the likelihood ratio test is also proposed and the relation is established with the slope of the ROC curve.
Feature selection is a critical pre-processing step in machine learning. For supervisedproblems, class labels are used to identify important features. However, labeling or annotat-ing the data is labor-intensive and hence costly. Consequently, there is an abundance of un-labeled data and limited labeled data. Therefore, semi-supervised learning is very pertinentin this case. The problem of feature selection is equally relevant to semi-supervised learn-ing. In this research work, a novel semi-supervised method of feature selection,RepeatedSampledSemi-supervisedFeatureSelection(RSSFS), is proposed with a time complexity ofO(nd), which is significantly better than other existing algorithms. This follows an unbiasedand neutral strategy in producing a final feature set that has less redundancy and frequentlyoccurring feature components, thereby inducing better stability in the feature set. In thefirst step, a decision tree classifier is used for labeling the unlabeled portion of the data andaugmenting the training set. Repeated sampling is done from the pseudo-labeled portion, togenerate multiple augmented training sets. The topkfeatures are selected based on mutualinformation from each augmented training set. While choosing the features, it is ensuredthat the features are not redundant by applying a correlation coefficient based threshold.A voting-based approach is used to combine these multiple features into the final featureset. The proposed algorithm is compared with a) a benchmark, using the full feature set,and b) top k % supervised feature selection on the labeled portion. Comparing these threemethods across 18 datasets, it was found that RSSFS outperformed the supervised methodbased on F1 scores by 2.79% and the benchmark by 0.36%. Thus, the proposed algorithmwould prove impactful in applications where there is a plethora of unlabeled data comparedto labeled data
R-optimality has been proposed in the literature as an alternative to the widely usedD-optimality criterion, particularly when the goal is to construct rectangular confidenceregions. In this study, R-optimal designs are investigated for the special cubic model inmixture experiments. The optimality of the proposed designs is verified using the equivalencetheorem.
In this paper, reliability of a Phased Mission System (PMS) under internal and ex-ternal degradation has been analysed. The degradation path of a component in a phaseis taken as a linear combination of the internal degradation and a proportion of commonexternal degradation that influences each component and modelled using Wiener process.The PMS is modelled using cumulative exposure model with components' dependency andthat of phases modelled using different copulas: Gumbel-Hougaard, Clayton, and Frank.A numerical illustration is presented using an aircraft flight PMS with comparative studyamongst different copulas carried out. Sensitivity analyses are also undertaken to deter-mine which parameters are sensitive to small deviations from their respective true value sothat extra care may be taken by an engineer while setting the true value of each of suchparameters
In this article, we examine the probability of disaster in a stress-strength model where item strength follows the power function distribution and stress follows the Nakagami distribution. We compare simple random sampling (SRS) and ranked set sampling (RSS) approaches to assess their efficiency and accuracy in this context. Our study contributes to stress-strength modeling literature by introducing a novel application of ranked set sam-pling with Nakagami and Power function distributions. We also performed cost function optimization in this context.
Gene-gene and gene-environment interactions are widely believed to play significantroles in explaining the variability of complex traits. While substantial research exists inthis area, a comprehensive statistical framework that addresses multiple sources of uncer-tainty simultaneously remains lacking. In this article, we synthesize and propose extensionof a novel class of Bayesian nonparametric approaches that account for interactions amonggenes, loci, and environmental factors while accommodating uncertainty about populationsubstructure. Our contribution is threefold: (1) We provide a unified exposition of hier-archical Bayesian models driven by Dirichlet processes for genetic interactions, clarifyingtheir conceptual advantages over traditional regression approaches; (2) We shed light onnew computational strategies that combine transformation-based MCMC with parallel pro-cessing for scalable inference; and (3) We present enhanced hypothesis testing procedures foridentifying disease-predisposing loci. Through applications to myocardial infarction data, wedemonstrate how these methods offer biological insights not readily obtainable from standardapproaches. Our synthesis highlights the advantages of Bayesian nonparametric thinking ingenetic epidemiology while providing practical guidance for implementation. Our parallelC code for implementing our key Bayesian nonparametric model based on hierarchies ofDirichlet processes, is available athttps://github.com/Sourabh-Bhattacharya/HDP-REALDATA-CODE/blob/main/HDP_REALDATA.zip.
The North Atlantic Oscillation (NAO) index, a measure of sea-level atmospheric pressure variability, holds significant influence over weather patterns in North America and Northern Europe. A negative (positive) NAO value signifies increased cold air outbreaks and storm occurrences (reduced occurrences) in these regions. NAO, a product of multiple climate factors, demonstrates intricate dynamics with sea surface temperature (SST) and sea ice extent (SIE). In this study, we adopt a data-driven approach to explore the complex interplay between NAO, SST, and SIE, revealing a critical instability rooted in positive feedback loops among these climate variables. Our statistical machine learning methodology examines the impacts of melting Arctic SIE and rising SST on NAO, thereby understanding the weather patterns across the North Atlantic region. The skewness analysis yields a negative skewness in NAO across various time intervals -- daily, weekly, and monthly. This skewness, coupled with NAO's mean zero stationary nature, accentuates system instability. To capture these dynamics, we formulate a Bayesian Granger-causal dynamic linear model, which effectively updates the predictor-dependent variable relationship over time. The findings underscore an impending critical instability, indicative of more frequent occurrences of intensely cold climates in eastern North America and northern Europe, theory signifies a notable climate shift. By delving into the intricate feedback mechanisms of NAO, SST, and SIE, our study enhances our comprehension of climate variability, fostering a more informed perspective on the imminent climate changes that lie ahead.
In this article, we proposed a new estimator, termed the Modified Logistic Two-Parameter Estimator (MLTPE), and enhanced it by modifying its coefficients, yielding three variants: Modified Logistic Two-Parameter Estimator1 (MLTPE1), Modified Logistic Two-Parameter Estimator2 (MLTPE2), and Modified Logistic Two-Parameter Estimator3 (MLTPE3). These estimators are designed for logistic regression models in the presence of multicollinearity. Theoretically, we demonstrated the superiority of the MLTPE over existing estimators, including the Maximum Likelihood Estimator (MLE), Modified Almost Unbiased Ridge Logistic Estimator (MAURLE), and Logistic Two-Parameter Estimator (LTPE), in terms of mean square error (MSE). The superiority of the estimators is examined using a simulation study and a real-world example. In the simulation study, we varied the degree of correlation and sample size. The findings revealed that the efficacy of the estimators is significantly influenced by these factors. Furthermore, we evaluated the prediction performance of these estimators using balanced accuracy. The results suggested that the new estimators, MLTPE1, MLTPE2, and MLTPE3, outperformed the others slightly in terms of balanced accuracy, with MLTPE2 exhibiting superior performance regarding both scalar mean square error (SMSE) and balanced accuracy. Finally, we validated the simulation study using the myopia dataset, which produced satisfactory results.
The increasing complexity in the dimensionality of data throws new challenges in modelling. Matrix variate data with underlying subpopulations need more sophisticated models to make meaningful statistical inferences and clustering results. This work introduces the finite mixtures of matrix variate log-normal distributions to model the right skewed multi-modal matrix variate data, and its application to model-based clustering is discussed. In addition, an extended K-means algorithm is developed as an alternative to the ordinary K-means approach for clustering matrix variate data. Furthermore, it can also be a useful initialization approach for matrix variate finite-mixture models, which received much attention lately for modelling this kind of data. Using the suggested initialization approach, the Expectation-Maximization algorithm is employed to estimate the parameters. The ability of the proposed methodology is illustrated through simulations and real data studies.
Bayesian estimation is carried out for the scale parameter of Gumbel distribution under various loss functions. A new class of loss functions is introduced and an extensive Monte Carlo simulation study is carried out for the comparison of estimators under loss functions. Two different prior distributions for scale parameter are used. The posterior distributions have no closed form hence Lindley approximation as well as Markov Chain Monte Carlo methods are used for deriving estimates. The study is illustrated through a real life data.