
The assumption of normality is fundamental to many statistical methods, and violation of this assumption may result in biased interpretations and inferences. This paper evaluates the performance of seven formal normality tests for time series data including Shapiro-Wilk (SW), Kolmogorov-Smirnov (KS), Jarque-Bera (JB), Lilliefors (LF), Cramer-von Mises (CvM), Vasicek-Song (VS) and Anderson-Darling (AD) in various settings. A simulation study with symmetric and asymmetric distributions was carried out to find the detection rates, considering the most important time series characteristics, including trend, seasonality, and presence of outliers. The tests were evaluated based on their ability to correctly classify normal and non-normal data, and their performance was summarized using accuracy metrics. The results indicate that Cramer-von Mises, Lilliefors, Anderson-Darling and Vasicek-Song tests provide superior performance as compared to the other tests in various situations. Shapiro-Wilk test showed good performance for data with trend and seasonality and Kolmogorov-Smirnov test was most effective in detecting non-normality. Anderson-Darling, Lilliefors and Cramer-von Mises test demonstrated good performance with the presence of outliers. These findings suggest that one should select normality tests based on the specifics of the data. We suggest supplementing formal tests with graphical procedures to have a more detailed and reliable test of distributional assumptions.
The COVID-19 pandemic has presented unprecedented challenges to global healthcare systems, underscoring the critical need for accurate prediction of infection cases to facilitate effective resource allocation and decision-making. This study evaluates the performance of two widely used time-series forecasting models, ARIMA and LSTM in predicting COVID-19 infection trends. Using a dataset of daily infection cases spanning January 2020 to June 2020, both models were trained and evaluated. The results demonstrate that the LSTM model achieves superior performance compared to the ARIMA model, as evidenced by lower Mean Absolute Error (MAE) and Mean Squared Error (MSE). The LSTM models ability to capture complex patterns and non-linear relationships in the data contributes significantly to its enhanced predictive accuracy. These findings highlight the potential of LSTM models to deliver more reliable forecasts of COVID-19 infection cases, providing healthcare authorities with valuable insights to inform strategic planning and preparedness for future outbreaks.
Road accidents and subsequent injuries and deaths are at alarming stage in India and the same is true for almost all part of the world with some variations in the cases. Continuous expansions of road networks, disproportionate increase in urbanization, enormous state of motorization and many other micro factors. Rising accidental deaths result into large number of losses of life especially between the age group of 15–50, imposes a great concern to all stake holders (from policy makers to common people). Road accidents are multi-causal events mainly categorized by human error, road resource optimization, effective policy formulation. As such we are motivated to set our objectives to compute the hotspot and coldspots to develop deep understanding of related characteristics. The second objective is to support the policy makers in optimal resource utilization needed to combat and evolve an effective strategy ranging from mass awareness campaign to technological innovations and developing the infrastructure to meet pre-and-post accident related challenges. To achieve the computational efficacy and drawing in-depth inferences we have implemented the software's namely MS Excel, saTScan and Python. The finding of the current work highlights southern states and northern states have disproportionality higher rate. We found cities like Coimbatore, Aurangabad, Chandigarh, Bangalore, Delhi, Chennai and Hyderabad as the urban hotspots based on computed values of the statistical likelihood ratio and risk ratio. We found them as the most likelihood clusters (primary hotspots). We found time series model model Holt-Winter method is the best model when compared to other exponential model.
Estimating parameters accurately in the presence of uncertain and imprecise data is a key challenge in statistical analysis, particularly for complex models involving two populations. Fuzzy data provides a structured way to handle such uncertainties by effectively representing real-world ambiguity. While extensive research has been conducted on parameter estimation for single-population models using fuzzy data, extending these methods to dual populations remains a difficult task. This study addresses the issue by developing estimation techniques for two Weibull distributions that share a common scale parameter σ but have different shape parameters k 1 and k 2 , under fuzzy data conditions. We apply the expectation-maximization (EM) algorithm for Maximum Likelihood estimation and utilize a Bayesian approach with TK approximation for parameter estimation. To further refine Bayesian estimates, Gibbs sampling is employed to derive posterior distributions. Through Monte Carlo simulations on both simulated and real-world datasets, we evaluate the accuracy and robustness of our estimators, demonstrating their effectiveness in handling imprecise data. Additionally, asymptotic and HPD confidence intervals are also obtained. This research highlights the importance of reliable statistical methods for dual-population Weibull models, contributing to improved analytical precision across various domains.
Crude oil is the primary fuel and its price has a direct impact on oil exploration, exploitation, and other activities, as well as on the environment and on our economy. Hence, it is among the world's most abundant resources today. Crude oil is essential to the functioning of every modern economy. Considering crude oil's high volume of trading, speculators, analysts, and economists have vested interest in correctly projecting the commodity's future spot price. However, predicting such an apparently uncertain, economic environment is one of the primary challenges of econometric models. Oil price forecasts based on fundamental, technical, and time series analysis have been met with mixed success. This highlights the requirement for more refined methods of predicting future crude oil prices. This study uses a neural network to objectively foretell the price of crude oil. There are thirteen predictors and one dependent variable in this study. In this study Neural Networks (NN) and time series techniques are used to forecast time series data and out of these NN have been found to be the most proficient method. The split in the data is 70-30. 30% of the data is being used to verify the accuracy of the network's predictions. Speculating on the future cost of crude oil, requires the employment of feed forward and back propagation algorithms. Latest neural networks techniques are quite predictive as time series models are beaten by even the simplest of neural networks. The results of the investigation showed that the back propagation algorithm is superior in predicting the cost of crude oil. Hence, ANN can be used by financiers, and forecasters.
The estimation of finite population characteristics, particularly the mean, is critical in survey sampling, as it has a direct impact on the validity of sample data results. As the number and complexity of available data increases, so does the demand for more robust and precise estimators. Traditional estimators frequently fail to take full advantage of auxiliary information, which can be critical for enhancing estimating accuracy. This paper presents a novel log-ratio estimator that incorporates auxiliary information in a non-linear manner using logarithmic transformations of the study and auxiliary variables. The log-ratio estimator's performance is assessed against 9 different classical and modern estimators using two essential metrics: mean squared error (MSE) and percentage relative efficiency. Our empirical analysis includes a variety of applications, such as estimating the area under wheat based on cultivated area, estimating peppermint oil production based on field area, and three real-world datasets: breast cancer deaths, cancer deaths by gender (male and female), and brain tumor survival rates. To supplement these applications, a simulation study with a population size of 100,000 and a sample size of 1,000 was run up to 100,000 times to assess the estimators’ resilience and stability. The empirical and simulation results consistently reveal that the log-ratio estimator produces lower MSE and higher PRE values across all datasets, exhibiting greater accuracy and efficiency over traditional estimators. These findings suggest that the log-ratio estimator gives more trustworthy and efficient estimates, particularly when dealing with complicated data. As such, this estimator is a promising tool for future survey sampling, helping to enhance estimating methodologies capable of dealing with the challenges provided by modern, large-scale datasets.
This study addresses the challenge of estimating parameters for two logistic populations that share a common scale parameter but have different location parameters in the presence of fuzzy data. To handle these complexities, both Maximum Likelihood Estimation (MLE) and Bayesian methods are employed. Asymptotic confidence intervals are constructed using ML estimates. For Bayesian estimation, a conjugate prior is utilized, and Bayes estimators are approximated using Lindley’s method due to the lack of closed-form solutions. Furthermore, Approximate Bayesian Computation (ABC) and Markov Chain Monte Carlo (MCMC) techniques, including Hamiltonian Monte Carlo (HMC) and the Metropolis–Hastings (MH) algorithm, are utilized to sample from the posterior distributions and construct Highest Posterior Density (HPD) intervals. A detailed comparative analysis of MLE, Lindley’s approximation, ABC, HMC, and MH is conducted to assess their performance. The effectiveness of the proposed methodology is demonstrated using a real-world dataset under fuzzy conditions.
Generalized partially linear models (GPLMs) provide a versatile regression framework that blends parametric and nonparametric components, allowing flexible modeling of complex data structures. In binary response settings, particularly within logistic frameworks, verifying the independence between covariates and error terms is essential for ensuring model adequacy and validity. This paper develops a nonparametric diagnostic based on the Bergsma–Dassios measure of association, τ * , to assess the independence between the regressors ( X , W ) and the random error component in logistic GPLMs. Unlike traditional correlation measures, τ * captures broad classes of dependencies, including nonlinear and nonmonotonic associations, thus offering a powerful and robust diagnostic tool. Both complete data and missing-response scenarios are considered, where responses are missing completely at random (MCAR) or missing at random (MAR). Consistent and asymptotically efficient estimators for the parametric vector β and the nonparametric function m ( W ) are constructed under these settings. Theoretical properties of the proposed τ * -based test are established, including its asymptotic distribution and power against local alternatives. Simulation studies and real-data analyses further confirm the practical effectiveness and robustness of the proposed method, demonstrating its utility in semiparametric logistic regression with incomplete or potentially misspecified data.
Floods remain one of the most devastating natural disasters, causing significant human, economic, and ecological losses worldwide. Accurate prediction of flood peaks is therefore essential for effective water resource management, urban planning, and disaster mitigation. However, modeling flood peak data is challenging because such observations are nonnegative, highly skewed, and often exhibit heavy-tailed characteristics. Traditional symmetric probability models, such as the normal distribution, fail to capture this behavior. In contrast, the skewed chi-square distribution, particularly its scaled and noncentral variants, offers a mathematically flexible and practically interpretable framework for modeling positively skewed hydrological extremes. This paper develops a statistical formulation for flood peak prediction based on skewed chi-square modeling, including parameter estimation, return level analysis, and diagnostic validation. To complement simulation-based evaluation, an empirical study using annual flood-peak records from a Midwestern U.S. catchment (1950–2020) was conducted with data obtained from USGS and NOAA archives. The real data analysis confirms that the scaled noncentral chi-square distribution accurately captures the strong right skewness and heavy tails observed in hydrological extremes, outperforming traditional gamma and lognormal models. The proposed approach aims to bridge the gap between theoretical distributional modeling and real-world flood risk assessment.
Student evaluations of teaching (SETs) are a common measure of teaching quality in higher education, yet valid inference from such data remains challenging. Hierarchical dependencies among observations, variation in response styles, and measurement error in predictors complicate interpretation, while traditional analyses based on mean ratings often fail to distinguish teaching-related effects from correlated, non-teaching influences. This study introduces a Bayesian mixed-effects location–scale probit model as a model-assisted inferential framework for ordinal SET data. The model jointly estimates effects on the latent mean evaluation (location) and on latent response variability (scale), incorporating hierarchical random effects and correcting for measurement error in multi-item predictors. The framework is illustrated using four semesters of SET data from more than 5,000 students in psychology, pedagogy, and teacher education programs at a German university. The application is intended as a methodological demonstration rather than a population-level generalization. Within this context, didactic quality and lecturer likeability emerged as the strongest predictors of overall evaluations. An interaction effect suggested that comprehensibility receives higher ratings when lecturers are viewed as likeable. Beyond mean effects, substantial heterogeneity was observed at the student, lecturer, and course levels, along with systematic differences in rating precision linked to individual response tendencies. These findings highlight that robust inference from SET data requires models capturing both location and scale heterogeneity. More broadly, the proposed Bayesian approach demonstrates how location-scale modeling can improve the analysis of ordinal data with hierarchical dependencies and measurement error.
This study assesses the effectiveness of transmission models, including the Markov SIR, Gillespie Algorithm, and Reed-Frost Model, in simulating disease trends. Each model was estimated using daily infection and recovery counts and applied to project the progression of active cases. The Markov SIR model demonstrated a strong fit for COVID-19 transmission, particularly in modeling ongoing outbreaks. These findings emphasize its potential for forecasting the spread of infectious diseases, highlighting its adaptability for future outbreak monitoring and response efforts.
Non-destructive image-based classification of fruit varieties is essential for optimizing cost, nutritional benefits and supply chain management in food system. In this study we propose an interpretable machine learning framework for the multi-class classification of banana varieties using handcrafted image features extracted from colour, texture and geometric properties. A total of 195 banana images from five varieties were collected and expanded the dataset to 669 samples through rotational augmentation to enhance model generalization. Five supervised learning models, viz., k-Nearest Neighbours, Naïve Bayes, Random Forest, Decision Tree, and Support Vector Machine, were implemented using standard evaluation metrics. Among all models, Random Forest demonstrated the highest performance with an accuracy of 98.5 per cent and an MCC of 0.9814. Statistical validation was performed using one-way ANOVA and post-hoc Tukey HSD tests, which confirmed significant difference between model performances. Additionally, SHAP analysis provided insights into feature importance and model decision processes. The findings suggest ensemble learning models, especially Random Forest offer a compelling combination of accuracy and interpretability for agricultural classification tasks. The proposed approach enables applications in automated fruit sorting and mobile based advisory platform in smart agriculture.
Long Short-Term Memory (LSTM) neural network models have become the cornerstone for sequential data modeling in numerous applications, ranging from natural language processing to time series forecasting. Despite their success, the problem of model selection, including hyperparameter tuning, architecture specification, and regularization choice remains largely heuristic and computationally expensive. In this paper, we propose a unified statistical framework for systematic model selection in LSTM networks. Our framework extends classical model selection ideas, such as information criteria and shrinkage estimation, to sequential neural networks. We define penalized likelihoods adapted to temporal structures, propose a generalized threshold approach for hidden state dynamics, and provide efficient estimation strategies using variational Bayes and approximate marginal likelihood methods. Several biomedical data centric examples demonstrate the flexibility and improved performance of the proposed framework.
In the biology field of botany, leaf shape recognition is an important task. One way of characterising the leaf shape is through the centroid contour distances (CCD). Each CCD path might have different resolution, so normalisation is done by associating each contour to a circular density. Densities are rotated by subtracting the mean or mode preferred direction. Distance measures between densities are used to produce a hierarchical clustering method to cluster the leaves. We illustrate our approach with a motivating small dataset as well as a larger dataset.
This paper introduces a new calibration estimation technique for the estimation of the population distribution function (DF). We have introduced a new class of calibrated estimators using the non-linear constraints of an auxiliary variable in a simple random sampling design. Their performances have been assessed based on some real and artificially generated data sets under numerical and simulation studies. It has been found that the proposed estimators attain lower absolute relative bias (ARB), lower mean squared error (MSE) and higher percentage relative efficiency (PRE) against the usual unbiased, ratio, product, regression and GREG estimators. The results highlight the effectiveness of the proposed estimators, which may further encourage survey practitioners in their real-life applications.
Unsupervised learning is a major class of machine learning techniques where response information is missing or unavailable. Among these techniques, clustering plays a central role by grouping objects based on a chosen similarity measure. K-Means is one of the most established and widely used clustering methods, known for its simplicity and computational efficiency. For continuous data, K-Means performs well when the number of clusters ( K ) is known and correctly specified. However, it faces convergence and overfitting challenges when K is unknown. These issues stem from K-Means’ objective function, which monotonically decreases as the number of clusters increases—leading to a tendency to overfit. In this article, we propose an augmented K-Means algorithm that introduces a penalized version of the standard K-Means objective, designed to guard against overfitting and promote model parsimony when K is unknown. We establish key optimality properties of both the traditional K-Means loss function and the proposed penalty term. Extensive simulation studies on benchmark datasets demonstrate the improved performance of the proposed method, including accurate identification of the true number of clusters. Extensive simulation studies on benchmark datasets demonstrate the improved performance of the proposed method, including accurate identification of the true number of clusters. Additionally, we apply our approach to the clustering of globular galaxy datasets—an example of truly large-scale (“Big”) data—to further illustrate its effectiveness.
Count data sets in many real-world scenarios exhibit zero-inflation, characterised by an excessive zero counts. Zero-inflated negative binomial (ZINB) and hurdle negative binomial (HNB) models are commonly applied to handle this issue, particularly in the presence of overdispersion. The hurdle model distinguishes between zero and positive counts, applying a truncated count model to the non-zero values while the ZINB model posits two separate processes: one generating only zeros and the other producing counts that follow a negative binomial distribution. This study focuses on obtaining estimators of the parameters of both ZINB and HNB models using method of moments, method of moments with proportion estimator and maximum likelihood method. The estimators are compared with respect to relative bias and mean squared error (MSE). Further, approximate simultaneous T 2 confidence intervals for the parameters of these models are constructed by obtaining variance-covariance matrix using inverse of Fisher information matrix, inverse of Fisher information matrix based on profile likelihood obtained by eliminating the nuisance parameter π (inflation parameter) and using parametric bootstrap method. Monte-Carlo simulations are conducted to evaluate the performance of these methods, in terms of coverage probabilities and the length of the confidence intervals. The study analyses number of cases registered for violation of Essential Commodities Act by applying ZINB model.
Semiparametric regression models provide a powerful framework that combines the parametric and nonparametric paradigms, particularly effective for analyzing complex data structures. In practical scenarios, missing data is a pervasive issue that complicates statistical inference. This paper addresses semiparametric estimation when the response variable is subject to missingness under the Missing at Random (MAR) mechanism. We develop a kernel-based estimation strategy for the nonparametric component and employ partial regression methods—specifically, an adaptation of Robinson’s approach—to estimate the parametric part. The estimation procedure incorporates inverse probability weighting and nonparametric imputation to account for missing responses. Theoretical properties such as asymptotic bias, consistency, and variance are derived. The methodology is validated through two real-data analyses using the Abalone and Airfoil Self-Noise datasets, where missingness is artificially induced, demonstrating the effectiveness of the proposed strategy in preserving estimation accuracy. Our results underline the robustness and flexibility of semiparametric models in the presence of incomplete data.
The anatomical distribution of ischemic stroke in patients with atrial fibrillation (AF) is highly heterogeneous, yet its biological determinants remain poorly defined. This study aimed to identify clinical and molecular predictors of stroke topography and to develop explainable machine learning models to forecast lesion location. We retrospectively analyzed 500 patients with AF-related ischemic stroke, collecting clinical data and blood-based biomarkers including C-reactive protein (CRP), interleukin-6 (IL-6), B-type natriuretic peptide (BNP), D-dimer, and fibrinogen. Patients were stratified by stroke territory (anterior vs. posterior circulation) and hemispheric side (left vs. right). Two XGBoost models were built to predict each spatial pattern. Model performance was assessed using ROC analysis, and SHapley Additive exPlanations (SHAP) were applied for feature interpretation. Posterior circulation infarcts were associated with higher CRP and IL-6 levels (p < 0.0001), while BNP was elevated in left-sided strokes, and IL-6 in right-sided strokes. The circulation model achieved an AUC of 0.71, and the hemispheric model an AUC of 0.76. SHAP highlighted CRP and IL-6 as top predictors for posterior infarcts, and BNP for left-sided infarcts. These findings suggest that inflammatory and cardiac biomarkers carry spatially specific predictive value in AF-related stroke and support the use of interpretable models for personalized risk stratification.