This study investigates the optimization of population mean estimation in stratified random sampling under fixed budget constraints. Although stratified sampling improves precision by dividing the population into homogeneous strata, data collection costs often vary across strata, making optimal resource allocation essential. We develop a cost-efficient allocation strategy that minimizes the variance of the estimated population mean under a linear budget constraint. The methodology incorporates auxiliary information and cost variations across strata to determine optimal sample sizes. Both real data and simulated data are used to evaluate the performance of the proposed allocation strategy. Analytical results, supported by simulation experiments, demonstrate that optimal allocation substantially improves estimation efficiency when compared with traditional proportional and equal allocation methods. Findings from the real dataset further confirm the practical applicability of the method in actual survey settings. Overall, the study provides valuable guidance for researchers and survey practitioners seeking to balance statistical accuracy with financial limitations in diverse real-world applications.
Lumpy Skin Disease (LSD), caused by the Lumpy Skin Disease Virus (LSDV), which is a part of the family of Poxviridae, subfamily Chordopoxviridae, and genus Capripoxvirus, is a rapidly spreading viral infection impacting mainly cattle globally, causing huge economic loss. This study explores the effectiveness of time series model Autoregressive Integrated Moving Average (ARIMA), and machine Learning models Decision Tree (DT), Random Forest (RF), and Support Vector Regression (SVR) in predicting LSDV infection rates across the 10 high-risk countries by LSDV, using data obtained from the Food and Agriculture Organization (FAO) EMPRES database, covering reported cases from 2006 to 2024. By using the evaluation metric Mean Squared Error (MSE), we identified the best-fitted model out of four distinct models for each country. Results revealed that the ARIMA model achieved the lowest MSE among all models for Thailand (MSE: 349.57), Bulgaria (4183.09), and North Macedonia (395.02). DT achieved the lowest MSE for Serbia (1.00) and Israel (32.27), while RF scored the lowest MSEs for Turkey (2014.25) and Greece (37.01). SVR achieved the lowest MSE scores in Albania (2.54), the Russian Federation (5.47), and Malaysia (818.50), indicating superior predictive performance in terms of error minimization. These results demonstrate that machine learning models, particularly SVR and DT, consistently achieve lower prediction errors than traditional ARIMA in several countries, highlighting their effectiveness in capturing complex, nonlinear outbreak patterns and their potential for improving disease surveillance and policy decision-making.
This paper addresses the problem of estimating the population mean of the study variable using auxiliary information on an auxiliary variable in stratified random sampling. A class of transformed ratio-product type estimators is introduced, with expressions for the biases and mean squared errors (MSEs) of the proposed class of estimators derived up to the first order of approximation. Some estimators are particular members of this introduced class, and many new estimators can be generated from the suggested family of estimators. Additionally, a linear cost function is introduced, and the MSEs and optimum values for the entire family of estimators are determined. The proposed family of estimators is theoretically compared with competing estimators, and conditions under which the suggested estimators are more efficient than others are also established. An empirical study is conducted to support the suggested family of estimators. The numerical illustration demonstrated the substantial practical utility of the theoretical findings in real-world applications. The estimator with the lowest MSE is recommended for various statistical applications and for analyzing large datasets in interdisciplinary research involving statistics and data science.
Dengue remains a significant public health challenge in India, motivating the use of statistical and machine-learning models to explore and compare forecasting approaches that can inform surveillance planning and epidemiological understanding. This study presents an exploratory, state-wise comparative assessment of classical Statistical, time series and machine-learning (ML) models applied to annual dengue incidence and deaths data in India from 2017 to 2024. We employed Naïve Method, Simple Exponential Smoothing (SES), ARIMA, Linear Regression (LR) and Support Vector Regression (SVR) based on temporal data. Model performance was evaluated using Root Mean Squared Error (RMSE), and the best-performing model for each state was selected based on the lowest RMSE value. Uttar Pradesh, Karnataka, Punjab, and Maharashtra are projected to report higher dengue incidences and deaths in 2025. Kerala consistently shows the highest CFR among all Indian States, with an average of 0.551
In this paper, an optimization framework for estimating the population mean in stratified random sampling is proposed by integrating an improved ratio estimator with a nonlinear integer programming (NLIP) and nonlinear travel cost function. Unlike classical approaches that assume linear cost function, the proposed framework explicitly incorporates nonlinear travel cost function into the sample allocation process, enabling the determination of an optimal allocation under a fixed cost. The sample sizes are determined by using two methods. The first method uses proportional allocation to determine sample sizes from each stratum, and the second uses a nonlinear integer programming to get optimum allocation from each stratum subject to a fixed cost constraint. The mean squared error (MSE) and percent relative efficiency (PRE) of the proposed estimator are derived and examined by comparing its performance with several well-known existing estimators. The practical performance of the proposed framework is evaluated by using both real and simulated datasets. The real dataset is based on the 2011 Census of India, taking district population as the primary (study) variable and the number of households as the secondary (auxiliary) variable, where all districts are divided into six homogeneous regions- North, East, Northeast, Central, West, and South. The simulated dataset are generated by using Python to replicate similar stratified structures. The empirical findings reveal that the proposed estimator consistently achieves a lower MSE and higher PRE in comparison to other existing estimators, indicating improved statistical efficiency while accounting for fixed survey costs. The study provides a unified optimization based approach for designing cost-effective and statistically efficient large-scale socio-economic and demographic surveys through optimal allocation under nonlinear cost constraints.
This paper examines improved estimation of shipment delivery times across four major regions of India—North, West, South, and East—using stratified sampling techniques to enhance estimation precision. Motivated by recent developments on optimal estimation under cost constraints in stratified sampling (Yadav et al. in Ann Data Sci 12:517–538, 2025), we propose a novel population mean estimator that incorporates shipment volume as an auxiliary variable alongside delivery time as the study variable. By effectively exploiting the correlation between shipment volume and delivery time, the proposed estimator achieves reduced estimation error and improved efficiency. The sampling properties of the estimator, including bias and mean squared error (MSE), are analytically derived to assess its theoretical performance. Comparative analysis with existing estimators demonstrates the superior efficiency of the proposed method in estimating expected delivery times. Extensive simulation experiments are conducted to evaluate robustness under heterogeneous regional delivery patterns. The results confirm consistent gains over traditional estimators, highlighting the practical relevance of the proposed approach for logistics optimization and informed supply chain decision-making.
This study presents an efficient method for estimating the population mean of a study variable using a Searls exponential ratio estimator. The approach is based on simple random sampling and leverages known population parameters of an auxiliary variable for enhanced estimation of population mean. Using large sample approximations up to the first order, the expressions for the bias and mean squared error of the proposed estimator are derived. To evaluate its performance, the estimator is compared conceptually and mathematically with several existing estimators. Theoretical results are further validated through both simulated and real-world data sets. Findings from these evaluations demonstrate that the proposed estimator consistently outperforms many competing estimators, particularly at higher population mean values. Therefore, it holds promise for improving population mean estimation across a range of practical applications.
This study introduces a new neutrosophic-type estimator for estimating the population mean in stratified sampling, motivated by the growing need for reliable statistical tools that can handle uncertainty, imprecision, and incomplete information. By combining auxiliary variables with neutrosophic theory, the proposed estimator is able to capture indeterminacy more effectively and produce meaningful interval-based estimates. We derive its bias and mean squared error (MSE) and compare its performance with several existing estimators. The results show that the new estimator consistently achieves lower MSE and much higher percentage relative efficiency (PRE), demonstrating clear improvements in accuracy. Numerical examples further confirm its advantages. The method is especially useful for data that naturally involve uncertainty or interval measurements, such as environmental indicators, financial data, sensor readings, and medical diagnostics. Overall, the study provides a more robust and efficient estimation approach for modern datasets where uncertainty cannot be ignored.
In survey sampling, post-stratification is widely used to improve estimation accuracy when stratification information becomes available only after data collection. This study proposes a new estimator for the population mean in post-stratified sampling that incorporates auxiliary information to enhance estimation efficiency. Building on ratio- and regression-type estimators, the proposed method combines auxiliary variables with post-stratum weights to reduce bias and Mean Squared Error (MSE). Theoretical properties of the estimator are derived using a first-order approximation. Its performance is evaluated through numerical illustrations based on two real populations and one simulated population, and comparisons are made with existing estimators. Under the considered settings, the proposed estimator exhibits lower MSE and higher Percent Relative Efficiency (PRE) than competing methods, indicating improved precision when reliable auxiliary information is available. The results suggest that the estimator is a useful alternative for population mean estimation in post-stratified sampling, with scope for further investigation under different sampling conditions and designs.
This article introduces an innovative adjusted estimator of ratio-type of the explanatory variable’s population mean, utilizing both traditional and unconventional auxiliary parameters. Frequently, data analysis encounters the challenge of outliers. To address this issue, a resilient measure of the auxiliary variable, unaffected by outliers is employed. We derive the bias and mean squared error (MSE) of the suggested estimator till the first-order approximation. An optimal value for the defining scalar is determined, yielding the minimum MSE for this optimal constant. The suggested estimator is scrutinized against the rival estimators. The efficiency conditions of the recommended estimator are obtained over the rival estimators. These efficiency conditions are substantiated through examination with a real dataset as well as simulated datasets, demonstrating an enhancement compared to alternative estimators.
This study proposes a new multivariate ratio-type estimator for improving the estimation of population means under stratified sampling, with particular application to child mortality and parental education data. The proposed estimator incorporates two study variables the number of children ever born and the number of deaths among children under five years of age and two auxiliary variables representing the educational levels of mothers and fathers. By effectively exploiting the relationships among the study and auxiliary variables, the proposed estimator achieves greater estimation precision. Its performance is evaluated theoretically using the Mean Squared Error (MSE) and Percentage Relative Efficiency (PRE) and is compared with several existing estimators. The theoretical results are validated through both an empirical study based on a real dataset and a simulation study using a synthetic non-normal stratified population. The findings demonstrate that the proposed estimator consistently yields lower MSE and higher PRE than the competing estimators, indicating its superior efficiency and robustness. The proposed methodology has practical applications in demographic and public health surveys, particularly in studies related to fertility, child mortality, and parental education. Overall, this work contributes to the advancement of multivariate estimation techniques in stratified sampling by providing a more efficient and reliable estimator for population mean estimation.
In many production and agricultural surveys, obtaining accurate auxiliary information at the initial stage of sampling is often expensive or infeasible. Stratified double sampling offers a practical and cost-effective solution by selecting a large first-phase sample to collect inexpensive auxiliary data, followed by a smaller second-phase sample to observe the main study variable. In this study, a new estimator and a family of improved estimators for the population mean are proposed under a stratified double sampling framework. The proposed class includes modified ratio, product, regression, and exponential-type estimators that efficiently utilize first-phase auxiliary information, stratum-specific characteristics, and optimally determined weights. Analytical expressions for bias and mean squared error (MSE) are derived up to the first-order approximation, and the optimal constants are obtained by minimizing the MSE. Theoretical and empirical results demonstrate that the proposed estimators achieve lower MSE and higher percentage relative efficiency (PRE) compared to several existing estimators. A practical application is illustrated in agricultural yield estimation, where satellite-based vegetation indices are used as auxiliary information in the first phase, while crop-cutting experiments are conducted on a smaller subsample. This approach substantially reduces survey cost while improving the precision and timeliness of production estimates.
In many practical survey situations, researchers need to estimate several population characteristics simultaneously while operating under limited resources. The motivation of this study arises from the need to obtain efficient estimates for multiple variables in stratified sampling while accounting for differences in measurement costs. In practice, the cost of measuring different variables may vary considerably, but many existing allocation methods do not adequately incorporate these variations, which may result in inefficient sampling designs. To address this issue, this study proposes a new multivariate stratified random sampling technique based on a family of estimators for compromise allocation. The proposed approach aims to improve the accuracy and efficiency of population mean estimation while controlling the overall measurement cost. The problem is formulated as an integer nonlinear multivariate stratified sampling model within a multi-objective mathematical programming framework using the proposed cost functions. The developed procedure is based on integer programming, and the resulting coefficients of variation are compared with those obtained from some existing compromise allocation methods. Numerical illustrations demonstrate that the proposed allocation technique provides improved efficiency. The proposed method is particularly useful in practical survey designs where budget, time, and effort must be carefully balanced with the requirement of achieving reliable statistical estimates.
This paper presents a new Searls type family of ratio estimators that use various auxiliary measures to estimate the population mean of a study variable under a simple random sampling without replacement framework. The study examines specific cases that include auxiliary data such as the median, quartile deviation, and coefficient of variation, maintains a first-order approximation for the introduced estimators' bias and Mean Square Error (MSE), and provides theoretical conditions for comparing their efficiency with existing estimators. Numerical analysis demonstrates that the proposed family of estimators are more effective than other ratio-type estimators, making them appropriate for use in a range of real-world scenarios including Agriculture, Biological Science, Commerce, Defence, Economics and Engineering, Forestry, Mathematical Sciences etc.
An estimator of the general class ratio, exponential, product type, was suggested to estimate the population mean (Y ) utilizing two auxiliary variables, given that the parameters of the population of the auxiliary variables under consideration have been identified. The bias and mean squared error (MSE) are derived up to the first order of approximation. The suggested estimator is then compared with the competing estimators of Y . One real data set and a simulation study are done to compare the efficiencies of various estimators. The study concludes that the proposed estimator outperforms the traditional mean estimator, standard ratio estimator, and many estimators suggested from time to time by various authors.
To improve the transformed ratio type estimators, this study uses new population parameters that are derived from extra information using a randomized response technique (RRT). Additionally, we suggest a modified family of powerful estimators for estimating the population mean of the sensitive variable in the presence of auxiliary data that are not sensitive. The bias and mean squared error (MSE), which are the primary statistical characteristics of the proposed estimator, have been determined up to the first order of approximation. We conduct theoretical comparisons among the contending estimators. Theoretical claims are supported by empirical evidence obtained from actual datasets. The suggested and competing estimators are further compared by analyzing their performances on a simulated data set. For a wide range of sensitive research applications, it is advisable to choose an estimator that possesses desirable sample properties and a minimized mean squared error (MSE).
The present article deals with a generalized class of estimators for estimating the population mean in sample surveys, employing various combinations of auxiliary variables and considering some values of characterizing constant alpha ranging from -1 to +1. The proposed estimator may be consider as an efficient extension to the work of Singh and Shukla (Metron, 45(1-2): 273-283, 1987), Bahl and Tuteja (Journal of information and optimization sciences, 12(1), 159-164, 1991) and Kadilar (Journal of Modern Applied Statistical Methods: Vol. 15 : Iss. 2 , Article 15, 2016). The sampling properties of the suggested estimators have been derived up to the first degree of large sample approximations. The suggested estimators are shown to have smaller mean squared errors than the existing exponential estimators considered in this paper. The percent relative efficiencies with respect to the usual mean estimator are calculated. An improvement has been shown over the existing exponential estimators through theoretical conditions as well as by a numerical and simulation study based on COVID-19 death in India.
Estimating the population mean using auxiliary information has been extensively explored within the classical framework. While point estimators offer simplicity, they provide only a single value without reflecting the uncertainty or variability inherent in real-world data. This shortcoming is particularly critical in high-precision applications and decision-making contexts. Moreover, classical estimators are often vulnerable to the influence of outliers, which can distort results and introduce significant bias. To address these limitations, this study introduces a novel estimator grounded in neutrosophic theory, which is specifically designed to handle imprecise, vague, and incomplete information-conditions frequently encountered in practical sampling scenarios. The proposed neutrosophic estimator incorporates auxiliary information expressed in neutrosophic terms, offering a more flexible and resilient estimation strategy. We derive expressions for the bias and mean square error of the proposed estimator under a first-order approximation. A comprehensive theoretical analysis, supported by simulation studies, confirms that the neutrosophic estimator consistently outperforms classical estimators in terms of accuracy and robustness, especially in uncertain environments. These findings underscore the potential of neutrosophic approaches as powerful alternatives to conventional estimation techniques in modern statistical inference.
In this article, we propose a new family of estimators for estimating the unknown population median of a study variable by utilizing auxiliary information under simple random sampling. The choice of the median, as opposed to the mean, is particularly advantageous in the presence of outliers or skewed distributions, where the mean may be unduly influenced. We derive the expressions for the bias and mean square error (MSE) of the proposed class of estimators up to the first order of approximation. Furthermore, we examine several notable subclasses within the proposed family and calculate their respective MSEs. To assess the efficiency and robustness of the proposed estimators, an empirical study is conducted using real-world data and benchmarked against existing estimators from the literature. The results of this empirical analysis demonstrate that the proposed estimators achieve lower MSEs, underscoring their practical relevance and effectiveness in survey sampling applications.
The Quest for the improved estimators of population mean is a continuous process to estimate it more efficiently and more closely to the true population mean of the study variable. There are various estimators of population mean in the literature, which estimate it efficiently, but still there is space for more efficient estimators. For improved estimation of population mean of primary variable, we suggest six different searls type ratio estimators utilizing the known auxiliary parameters. We derive the expressions for the biases and the mean squared errors of the introduced estimators for an approximation of degree one. The optimal values of the Searls characterizing constants, which minimize the mean squared errors of the introduced estimators, are obtained. The least value of the mean squared error of the proposed estimators for these optimal values of the characterizing scalars is also acquired. The efficiency of the proposed estimators is compared theoretically with the competing estimators of population mean through their mean squared errors. We derive the efficiency conditions of the proposed estimator for which it is more efficient than the estimators of population mean in competition. These efficiency conditions of the proposed estimators are verified through empirical data and the improvement over the competing estimators is calculated based on the minimum values of the mean squared errors. It is evident from the numerical study that the proposed estimator has the least mean squared error and the highest percentage relative efficiency in comparison to competing estimators. Thus, the proposed estimators are recommended for applications in different areas of applications.