Existence of a well defined and perfect sampling frame is the fundamental requirement of any sample survey. But there is enough evidence to support the fact that a perfect sampling frame that captures all the individual units of the population is rarely available, especially for a dynamic population where a constant movement of the units of the population is observed. In such cases, the sample collected can not be considered as a good representative of the population and the problem of incomplete frame arises. The results so drawn can immensely change the survey results and influence the legitimacy of the research. This study deals with the incomplete frame problem using ratio method of estimation, where the information on the auxiliary or ancillary variable is collected in the first phase and a second phase sample is then drawn to obtain estimates of the population mean of the characteristic under study and its mean square error up to first order approximation. Two different estimators, viz., combined ratio estimator and separate ratio estimators have been used and their efficiencies are compared. Further, the results are illustrated numerically with the help of Monte-Carlo simulations.
Survey sampling heavily relies on the availability of a robust sampling frame, which is often difficult to obtain, leading to undercoverage bias and misleading results. In the present era, an incomplete sampling frame poses a significant challenge, affecting survey results and findings. Literature offers widely used solutions to this pervasive issue that include multiple framework approach and post-survey adjustments. The predecessor-successor (PS) method proposed by Hansen et al. (1963) provides a way to uniquely identify the units not listed in the sampling frame by linking them to the existing units. Despite its potential to mitigate the bias arising from incomplete sampling frames, the P-S method was not explored extensively until it was later formulated mathematically by Singh (1983), Singh (1989) and others. Building on this legacy and addressing the limitations of simple random sampling in diverse populations, this paper introduces an estimator for the population mean within a stratified sampling framework. The properties of the proposed estimator are discussed and the findings are further supported by an empirical study as well as a simulation study. Journal of Statistical Research 2026, Vol. 60, No. 1, pp. 79-90.
This article deals with the statistical inference of the Modified-Weibull (M-W) distribution under left-censored data. The M-W model is an extension of the Weibull distribution, and it is a newly developed lifetime model. It is suitable for bath-shaped and increasing hazards. The necessary parameter estimation is discussed here for drawing inferences. The maximum likelihood estimation method is applied for the estimation of unknown parameters in the presence of left censoring. A simulation study is performed to monitor the accuracy of estimates in small samples. It is observed that proposed estimates have greater accuracy for small samples as well. Application of illustrative examples is given to study the estimation procedure under left censoring.
In this article, a new Modi-Exponential distribution is considered under Type-I and Type-II censoring. Some statistical and reliability characteristics of the distribution are specified. Statistical inferences of the Modi-Exponential parameters based on Type-I and Type-II censored data are discussed. The maximum likelihood estimation method is used for estimation of unknown parameters along with reliability characteristics. Asymptotic confidence intervals are also obtained for interval estimation. Monte-Carlo Simulation study is performed to study the performances of estimators. Further, two real data sets are provided to illustrate the results.. KEYWORDS :Censoring, Maximum likelihood estimation, Confidence bounds, Simulation.
Sampling frames play an important role in survey research as they lay the foundation on the basis of whichrepresentative sample drawn from a target population. However, in practice, sampling frames are often incomplete, meaningthat they do not fully capture the complete population of interest. In such cases, as a consequence of the incompleteness insampling frame, the sample drawn does not provide a good representation of the population and hence, the results infer tovague and misleading conclusions. The present study is an attempt to somehow uplift the efficacy of the estimator ofpopulation mean by taking into consideration, the additional information gathered from the units currently excluded in theexisting sampling frame along with mean square error of the proposed estimator. Optimum sample sizes are also obtainedusing suitable cost function under Neyman scheme. To compare the proposed estimator with the traditional ones, a simulationstudy has been done
Cigarette smoking is a preventable epidemic that is a leading cause of death. It increases the risk of coronary heart disease, stroke, lung cancer, chronic obstructive lung diseases etc., multifold. Smoking tobacco is not only injurious to oneself but also to those who are exposed second hand. Smoking induces endothelial dysfunction via inflammatory cytokines that can be quantified precisely. Cytokines can be leveraged as powerful predictive biomarkers for identifying risk of potential diseases. Current advances in biomarker research are providing substantive evidence of the roles of cytokines in disease. This is driving precision-based diagnosis and translational therapeutic interventions. Innovative machine algorithms (ML) are pioneering transformative changes in the field of medical research. This research implements the Neural Networks (NN) algorithm to classify smokers versus non-smokers using 63 cytokines as predictor features. In addition to the fact that NN is a generative algorithm, which makes it a very powerful tool to achieve the objective of this differentiation, techniques like cross validation and hyperparameter tuning improve the efficacy of the algorithm. The study identified the 10 most impactful predictor features that contributed to the classification and then used these to characterize smokers versus non-smokers. Primarily, the study constructed and investigated two classifiers, of which the first implemented NN using the entire set of 63 cytokines and the second using 10 most informative cytokines. The performance of the first classifier, implemented using 63 cytokines, evaluated by area under receiver operating characteristic (AUROC), was extremely good with an AUROC score of. 949 and 95% Confidence Interval (CI) (.923,.974). The second classifier that used the 10 most impactful cytokines with regard to the classification, demonstrated an exemplary performance, with an AUROC score of. 995 and a 95% CI (.991,1). The 10 most impactful cytokines from the aspect of smoker versus non-smoker differentiation, listed in order of importance, include: I-TAC, IL-22, IL-2R, IL-3, HGF, IL-18, G-CSF-CSF-3, MIF, SDF -1alpha, MMP-1. To gain a deeper understanding of the effect of smoking on cytokine levels, a 2-sample independent t test was performed, ascertaining the statistical significance of the 63 cytokine levels in smokers versus non-smokers. Machine Learning using biomarkers such as cytokines will enhance the ability to predict the advent of a disease and its outcome, and lead to novel treatment strategies.
Smoking is a major cause of premature and preventable death. Tobacco exposure has a detrimental effect on many organs and contributes to multiple diseases including chronic obstructive pulmonary disease (COPD), cardiovascular disease, cancer, and diabetes. Cytokines are inflammatory biomarkers that are mechanistically associated with smoking. Machine Learning algorithms allow for the quantitative assessment of the contributions of individual cytokines to tobacco-related diseases. The mapping of cytokines to disease can facilitate and direct treatment modalities. By the application of k Nearest Neighbor (k-NN) and Random Forest machine learning algorithms on 63 plasma cytokines we have demonstrated the classification of smoking. To ensure optimal results, performance improvement techniques such as k-fold cross validation and hyper parameter tuning are employed. Separability efficiency achieved by the models is evaluated using the Area Under the Receiver Operating Characteristic (AUROC) metric. The most significant cytokines that enabled the classification are identified and presented. The statistically significant difference for AUROC score of k-NN and Random Forest has been ascertained using the 2-sample independent t test. A reasonably good classification performance was achieved by k-NN algorithm with an AUROC metric of. 87, and a 95% CI of (.823,.917). Random forest exceeded k-NN algorithm's performance, with a perfect AUROC score of 1 and a 95% CI of (1,1). From among the ten most prominent cytokines that contributed to the classification, the ones common to both algorithms are: LIF, IL22, G-CSF/CSF-3, TRAIL. AUROC scores for k-NN and Random Forest are significantly different (p-value = 5.105e-16). The discovery and transference of biomarkers such as cytokines from the platform of molecular investigation to clinical practice, can facilitate precision medicine-based therapeutic interventions.
This paper presents a study on a new family of distributions using the Weibull distribution and termed as Modi-Weibull distribution. This Modi-Weibull distribution is based on four parameters. To understand the behaviour of the distribution, some statistical characteristics have been derived, such as shapes of density and distribution function, hazard function, survival function, median, moments, order statistics etc. These parameters are estimated using classical maximum likelihood estimation method. Asymptotic confidence intervals for parameters of Modi-Weibull distribution are also obtained. A simulation study is carried out to investigate the bias, MSE of proposed maximum likelihood estimators along with coverage probability and average width of confidence intervals of parameters. Two applications to real data sets are discussed to illustrate the fitting of the proposed distribution and compared with some well-known distributions.
Coronary artery disease (CAD) is a leading cause of mortality in the world. It is important to be able to proactively assess the risk of the disease, using novel biomarkers like cytokines that are indicators of inflammation in addition to traditional predictors of risk. Atherosclerosis, the primary cause of CAD, is an inflammatory disease involving cytokines. Identifying which cytokines are specifically altered can advance diagnosis and personalized treatment. Emerging research demonstrates that cytokines are transported on high density lipoproteins (HDL). Therefore, it is important to explore the roles of HDL-associated cytokines in vascular inflammation. Machine Learning (ML) algorithms are enhancing pioneering research from the standpoint of precision medicine. This technology can materially enable the translation of scientific research to clinical practice. In this study we implemented logistic regression and the derived regularized techniques using age and multidimensional cytokine biomarkers with the objective of identification of individuals “At Risk” for CAD. These techniques were further empowered by k-fold cross validation and hyper parameter tuning. Of the numerous algorithms investigated, the three most prominent ones, assessed based on area under receiver operating characteristic (AUROC) score are as follows: logistic regression, least absolute shrinkage, and selection operator (LASSO) regression with feature selection and ridge regression with feature selection. Logistic regression demonstrated an AUROC score of. 85 with a 95% Confidence Interval CI (.804,. 897), LASSO regression achieved a better AUROC score of. 875 with a 95% CI (.832,. 917) and finally ridge regression with feature selection exhibited the highest AUROC score of. 878 with a 95% CI (.837,. 92). The 2-sample independent t test proved that the three techniques were statistically significantly different from each other. With regard to the best classification demonstrated by ridge regression with feature selection, the most prominent biomarkers identified for the best classification achieved by ridge regression by feature selection, in the order of importance are as follows: Age, IL-7, RANTES, IFN-gamma, IL-3, GM-CSF, IL-15, IP-10, GCSF, IL-12. The identification and quantification of cytokines transported by HDL provide novel mechanistic insights that can inform the assessment of risk and therapeutic intervention in CAD.
In this paper, the authors have dealt with incomplete sampling frames and tried to use a weighted estimator to get refined and improved estimates of the population characteristics and their standard errors for simple random sampling cases (with and without replacement schemes). Further, considering the cost restrictions for surveys, the optimum values of the sample sizes have been obtained using a suitable cost function that has not been taken up in the past for such problems. Finally, a Monte Carlo simulation technique is applied to numerically illustrate the efficacy and applicability of the technique in real-life situations.
Presently, the role of cytokines in severe illness like COPD, cancer, cardiac disease associated with smoking is being explored to enable preemptive diagnosis and delivery of treatment interventions. We are investigating the connection between the elevation of inflammatory plasma cytokine in smokers versus nonsmokers. Disease indicator cytokines can be used to monitor the progression of disease which can help in the crucial task of prognosis and definitive diagnosis.Powerful and versatile Machine Learning algorithms can be leveraged to extract insights that cannot be obtained manually. We have applied Support Vector Machine (SVM) on 65 plasma cytokines and other traditional biomarkers to differentiate smokers and nonsmokers. To optimize the classification separability, we have used the following techniques: Principal component analysis (PCA), 10-fold cross validation and variable importance. The primary metric of evaluation is Area Under Receiver Operating Curve (AUROC), though we have additionally recorded and compared prediction accuracy across classifiers.The results are very promising. The AUROC classification accuracy achieved by SVM using the selected predictor feature variables is 89.2% with a 95%CI (85.4%,93.1%). The most prominent cytokines, contributing to the classification, in the order of importance are: I-TAC, Age, TG, G-CSF-CSF-3, MDCCCL22, Eotaxin-3, LIF, IL-2, Eotaxin-2, MIP-3alpha. The AUROC classification accuracy improved to 93% with a 95% CI (90.1%,99.5%) upon choosing the five most prominent cytokines.The versatile prowess of Machine Learning algorithms such as Support Vector Machine can translate pioneering molecular discoveries into actionable insights that can be applied in the field of translational and precision medicine to save life.
This paper suggests weighted ratio estimator of population mean based on the sample from incomplete frames. The bias and mean square error (MSE) of the proposed estimator are obtained up to the First Order of Approximation It is found more efficient than estimator of population mean under incomplete frame given by Agarwal and Gupta (2008). It has been illustrated with the help of hypothetical data as no data for incomplete frame is available in the literature.
Smoking is a major cause of cardiac and pulmonary disease, cancer, and other inflammation related diseases. Smoking impairs lipid and lipoprotein metabolism. The observed modification and reduction in levels of HDL in smokers has adverse effects on atheroprotective properties. It has been hypothesized that HDL transports inflammatory cytokines which accelerate tobacco-related diseases. To investigate the role of HDL in the transport of inflammatory cytokines and their detrimental effects on the immune response, it is paramount to compare cytokine levels in HDL for Smoker versus Nonsmoker groups. We isolated HDL from plasma using selected affinity immunosorption of apolipoprotein A-I-bearing lipoproteins, followed by quantitative ELISA of cytokines. We implemented a powerful stacked ensemble Machine Learning algorithm, namely Super Learner (SL) with base-learners: Decision Tree classifier, AdaBoost classifier, Bagging classifier, Extra Tree classifier, Logistic Regression and Random Forest classifier and meta learner: Logistic Regression. Prediction Accuracy metric was used to ascertain the separability efficacy of Smoker versus Nonsmoker based on cytokine levels. Super Learner composed of a Logistic Regression meta learner, achieved a 100% prediction accuracy, outperforming all the base learners. Machine learning-enabled Precision Medicine allows the investigation of the role of novel biomarkers such as HDL-transported cytokines which have a potential to generate valuable molecular insights. The discovery that cytokines are transported by HDL presents a new dimension in understanding inflammatory disorders and the potential for therapeutic intervention. The outstanding classification and prediction performance of Ensemble learning can be leveraged to revolutionize the biomarker discoveries, enabling insight that can lead to novel treatment modalities.
The paper suggests the improvement in estimator (Y) of population mean given by Agarwal & Gupta (2008) in case of incomplete sampling frame. Authors have used the auxiliary information (X) in terms of linear regression estimator. Bias and mean square error are obtained. It is shown theoretically and numerically that the proposed estimator is more efficient than the above mentioned estimator.
Background As per the 2017 WHO fact sheet, Coronary Artery Disease (CAD) is the primary cause of death in the world, and accounts for 31% of total fatalities. The unprecedented 17.6 million deaths caused by CAD in 2016 underscores the urgent need to facilitate proactive and accelerated pre-emptive diagnosis. The innovative and emerging Machine Learning (ML) techniques can be leveraged to facilitate early detection of CAD which is a crucial factor in saving lives. The standard techniques like angiography, that provide reliable evidence are invasive and typically expensive and risky. In contrast, ML model generated diagnosis is non-invasive, fast, accurate and affordable. Therefore, ML algorithms can be used as a supplement or precursor to the conventional methods. This research demonstrates the implementation and comparative analysis of K Nearest Neighbor (k-NN) and Random Forest ML algorithms to achieve a targeted “At Risk” CAD classification using an emerging set of 35 cytokine biomarkers that are strongly indicative predictive variables that can be potential targets for therapy. To ensure better generalizability, mechanisms such as data balancing, repeated k-fold cross validation for hyperparameter tuning, were integrated within the models. To determine the separability efficacy of “At Risk” CAD versus Control achieved by the models, Area under Receiver Operating Characteristic (AUROC) metric is used which discriminates the classes by exhibiting tradeoff between the false positive and true positive rates. Results A total of 2 classifiers were developed, both built using 35 cytokine predictive features. The best AUROC score of .99 with a 95% Confidence Interval (CI) (.982,.999) was achieved by the Random Forest classifier using 35 cytokine biomarkers. The second-best AUROC score of .954 with a 95% Confidence Interval (.929,.979) was achieved by the k-NN model using 35 cytokines. A p -value of less than 7.481e-10 obtained by an independent t-test validated that Random Forest classifier was significantly better than the k-NN classifier with regards to the AUROC score. Presently, as large-scale efforts are gaining momentum to enable early, fast, reliable, affordable, and accessible detection of individuals at risk for CAD, the application of powerful ML algorithms can be leveraged as a supplement to conventional methods such as angiography. Early detection can be further improved by incorporating 65 novel and sensitive cytokine biomarkers. Investigation of the emerging role of cytokines in CAD can materially enhance the detection of risk and the discovery of mechanisms of disease that can lead to new therapeutic modalities.
COVID-19 is now becoming a global issue and declared as pandemic by world health organization. This virus spread out from China to entire world. This paper performed a descriptive analysis of COVID-19 in special reference of India. Different attributes such as age, gender, travel history, communication type and current status are analyzed. Till now it can be stated from the study that age is not a significant factor that affect a person to be captured by this disease, also age attribute is normally distributed in current dataset. A significant relationship is found between gender (male and female) and Transmission Type (imported from other country or communicated from local) of the patients.
In this paper, a weighted-product estimator of population mean is proposed, when the sampling frame is incomplete. Bias, mean square error up to the first approximation is obtained and its efficiency is also discussed.
Kasturba Gandhi Balika Vidyalayas (KGBVs) for adolescent girls from backward society is one of the initiatives of Government of India to enhance educational status and quality of life among adolescent girls of underprivileged communities. In these institutes free education, food, clothes, books and free residential facilities are provided to adolescent girls. The present study was an attempt to assess the nutritional status of adolescent girls residing in KGBVs of two districts of Rajasthan i.e. Jaipur and Tonk. Nutritional anthropometry and dietary survey (food inventory method and 24 hour dietary recall) were carried out on 457 girls. The respondents were in the age group of 9-18 years belonging to Schedule Caste (28.2 per cent), Schedule Tribe (31.7 percent), Other Backward Caste (36.76 percent) and General (3.2 percent) category. Energy deficiency was found among the 18.35 percent girls while stunting (height for age less than - 2SD of WHO Z score) among 15.25 percent girls. Energy deficiency as well as stunting was observed higher among girls studying in class 6th as compared to class 8th. The food provided at KGBVs was basically cereal based. All food groups except cereals were available in inadequate quantities. Availability as well as intake of energy was adequate while that of micronutrients was below Recommend Dietary Allowances. A huge gap was observed between availability and intake of nutrients as well as food groups indicating possibility of pilferage of food stuffs.
Reliability modeling methods used to model combined Software/Hardware systems for the purposes of reliability estimation and allocation needs to accurately assess the interdependence between individual software elements and the hardware and on the platforms on which these software elements execute. The inclusion of hardware reliability into software reliability models increases the complexity of the models. This article provides modeling methods for inclusion of hardware into software reliability analysis.
In this paper an expert system using fuzzy logic has been proposed for evaluating performance of students of a regular degree course. This system explains the identification of three most important parameters attendance, internal assessment and external assessment for input of the result of a student and fuzzification of the parameters by fuzzy inference rules and then finding the final result as output and defuzzifying the output.