
Chronic kidney disease (CKD) is a progressive disorder characterized by gradual loss of renal function over time. Accurate prediction of CKD in its early stages can help in timely diagnosis and clinical intervention. This study aims to develop a predictive model using machine learning (ML) techniques to identify CKD risk based on patient data. The dataset was sourced from Kaggle’s open Chronic Kidney Disease dataset repository. This study interprets the predictive capability of the selected model and highlights the importance of critical clinical features such as packed cell volume, hemoglobin, and serum creatinine in determining CKD risk. Received: November 25, 2025Revised: March 10, 2026Accepted: March 18, 2026
Wilson’s disease is a rare autosomal recessive disorder caused by impaired copper metabolism, resulting in toxic accumulation in vital organs. Owing to its progressive nature, early detection is essential to prevent severe clinical complications. This study proposes a combined machine learning framework for the classification of Wilson’s disease using clinical data. Five machine learning models were developed and evaluated, with model performance assessed primarily using accuracy metrics. Feature importance analysis was conducted to identify clinically relevant biomarkers that contribute to model predictions. Based on the most influential biomarkers, survival analysis was applied to examine disease progression patterns over time. The proposed approach supports data-driven clinical decision-making and highlights the potential of advanced analytical methods in improving early detection and long-term management of rare metabolic disorders.
This study presents a parametric survival analysis of data from 802 breast cancer patients, utilizing five-parametric distributions: exponential, Weibull, gamma, lognormal, and log-logistic. Model comparison was conducted using the Akaike Information Criterion (AIC) and Bayesian Information Criterion (BIC). The exponential distribution exhibited the weakest performance, as reflected by the highest AIC (1676.822) and BIC (1681.51), due to its assumption of a constant hazard rate that did not match the observed survival patterns. The Weibull model showed significant improvement (AIC = 1541.825, BIC = 1551.19), with a shape parameter of 2.13 (95% CI: 1.93-2.36) indicating an increasing hazard over time. Further enhancements were observed for the gamma (AIC = 1531.162, BIC = 1540.536), lognormal (AIC = 1530.82, BIC = 1540.194), and log-logistic (AIC = 1529.912, BIC = 1539.286) distributions, with the log-logistic model providing the best overall fit. Among the five-parametric models, the log-logistic distribution was selected as the best-fitting model, slightly outperforming the lognormal and gamma models.
Diarrhea disease remains a significant public health concern among children under 5 years of age, necessitating reliable forecasting tools to support timely intervention and healthcare planning. This study applies a combination of classical time-series models and supervised machine learning techniques to model and forecast monthly diarrhea cases using data obtained from childhood diseases records in Rivers State Hospital Management Board, in collaboration with agencies responsible for childhood disease surveillance. The dataset spans January 2015 to December 2024 and represents population-wide hospital-reported cases, collected under a rigorous sampling framework to ensure data credibility. All analyses were conducted using Python. Exploratory time-series analysis revealed persistent temporal dependence and seasonal fluctuations. Forecasting models implemented include FB Prophet with linear and logistic trend specifications [11, 12], ARIMA(1, 0, 1) [25], SARIMA(1, 0, 1) (1, 0, 1) [3, 26], Random Forest [13], Artificial Neural Network (ANN) [14], and a surrogate Long Short-Term Memory (LSTM) model [26]. The Prophet linear model assumes constant growth with additive monthly and yearly seasonality, while the logistic variant incorporates saturation effects through an externally defined carrying capacity. ARIMA and SARIMA models captured short-term persistence and seasonal shocks, whereas machine learning models were configured to learn nonlinear and temporal patterns in the data. Model performance was evaluated using Root Mean Square Error (RMSE) and Mean Absolute Percentage Error (MAPE). Results indicate that the ARIMA model achieved the best overall performance (RMSE = 124.17, MAPE = 21.49%), closely followed by the ANN (RMSE = 124.29, MAPE = 24.55%). The SARIMA model showed slightly inferior accuracy, suggesting limited seasonal dominance in the series. Among machine learning models, ANN outperformed Random Forest and LSTM, while the Prophet logistic model performed better than its linear counterpart but lagged behind ARIMA and ANN. The findings demonstrate that diarrhea incidence in Rivers State exhibits both short-term persistence and moderate nonlinear behavior. Linear time-series models and neural networks provide superior predictive accuracy, underscoring their suitability for disease surveillance and public health decision-making. Integrating statistical and machine learning approaches offers a robust framework for forecasting childhood diarrhea and optimizing healthcare resource allocation.
This study evaluates the effectiveness of the Yèrèlon IV strategy, a community-led intervention aimed at reducing HIV transmission among female sex workers (FSWs) in Bobo-Dioulasso, Burkina Faso. A prospective cohort of 305 female sex workers (FSWs) (258 HIV-positive and 47 HIV-negative at baseline) was followed over a two-year period. To assess longitudinal changes in HIV-related outcomes and to identify distinct behavioral and clinical trajectories, we applied a latent class linear mixed model (LCLMM). This approach allowed for the classification of participants into unobserved subgroups based on shared patterns of risk and response to intervention over time. Model selection was guided by Bayesian information criterion (BIC) and entropy values. Two classes were clearly identified: a high-risk group (10.2%) and a low-risk group (89.8%). The high-risk class was characterized by inconsistent condom use, higher psychoactive substance use, and less stable occupational status. Among HIV-negative FSWs, no seroconversions were recorded throughout the study period, indicating effective primary prevention. Among those living with HIV, a decrease in viral load was observed, reflecting improved treatment adherence and reduced potential for transmission. These findings highlight the capacity of the Yèrèlon IV strategy to both prevent new infections and reduce community-level infectiousness. The use of biostatistical modeling provided nuanced insights into heterogeneous risk profiles and intervention effects, supporting the relevance of differentiated, context-specific approaches to HIV prevention in key populations.
Within the medical community, multiple sclerosis (MS) is one of the uncommon diseases that is still untreatable and incurable. The objective of the study is to determine the impact of age and gender on several diagnostic features of MS patients. The diagnostic features include Varicella, initial symptom, mono- or poly-symptomatic, oligoclonal bands, visual evoked potential, and other MRI abnormalities. These features were statistically analyzed. Varicella shows a strong correlation with age (r = 0.89369) and a weak correlation with gender (r = 0.1341). According to the study, symptoms differ throughout age groups. Males typically experience motor symptoms, whereas females often experience sensory and motor symptoms. The poly-symptomatic record is more common in female patients than in male patients. A very modest connection was established between the maximum oligoclonal band and gender in the age group of 31-40 years. The VEP test indicates that the majority of positive cases are found in males compared to females. The study is unique because of its quantitative analysis and compilation of studies on the connections between gender, age, and various MS clinical characteristics. Additionally, the study aims to mathematically identify responsible variables that have not been fully analyzed before.
This paper thoroughly validates the proposed models designed to interpret anomalous drug diffusion in pharmacological compositions identified by delayed release concentrations. The paper targets the growing demand for realistic predictive models in drug delivery devices that display non-Fickian diffusion patterns in pharmaceutical research. The proposed models estimate key parameters influencing drug diffusion, such as transfer, elimination, and metabolism rate for several anomalous drugs, giving a good fit with the experimental data from the literature. Graphs have been plotted for the time-concentration profile of considered drugs with suited models. Further, statistical analysis is done, indicating the robustness of the nonlinear models.
Neuroinflammation critically shapes stroke outcomes, yet therapies often fail due to the biphasic nature of cytokines. Tumor Necrosis Factor-alpha ( ), for instance, drives acute injury (0-24h) but later supports repair. To address this paradox, we developed an Ordinary Differential Equation (ODE)-based model capturing interactions among , Interleukin- ( ), Interleukin-6 (IL-6), Interleukin-10 (IL-10), and microglial states ( pro-inflammatory, anti-inflammatory). The model simulates immune dynamics and therapeutic modulation over time. Simulations identified two optimal intervention windows: early suppression of (0-24h) reduced -driven damage, while delayed IL-10 enhancement ( ) promoted -mediated repair. The biphasic role of IL-6 - pro-inflammatory acutely, reparative later - was reproduced and validated against cytokine kinetics from tMCAo mouse models A combined regimen of inhibition in the acute phase and twofold IL-10 enhancement in the subacute phase reduced infarct volume by in silico. Although assuming uniform cytokine distribution, the model offers a framework for phased, precision-guided immunomodulation. These findings highlight the critical importance of timing in stroke immunotherapy and demonstrate the value of computational modeling for optimizing cytokine-based interventions.
Breast cancer is among the most common cancers in women worldwide, and outcomes improve with early detection. As machine learning enters routine care, data driven diagnostic systems may support earlier risk estimation. We present a compact pipeline that uses Principal Component Analysis for dimensionality reduction and Borderline-SMOTE for imbalance correction, followed by classification with Light Gradient Boosting Machine. Using the standardized Wisconsin Breast Cancer Diagnostic dataset, we retain 20 features to capture key variance while limiting redundancy and noise. Borderline-SMOTE is applied within each training fold to refine class boundaries. Performance is evaluated with stratified 10-fold cross validation and compared with seven alternatives: XGBoost, Support Vector Machines, Random Forests, Logistic Regression, Gaussian Naive Bayes, k Nearest Neighbor, and a Multilayer Perceptron. With 20 components, the proposed model attains accuracy 0.993, precision 1, recall 0.986, F10.993, and AUC 1.000 for distinguishing benign from malignant cases, outperforming baselines. These findings suggest that coupling dimensionality reduction, boundary focused resampling, and gradient boosted trees can enhance diagnostic performance and may inform clinical decision support.
Safe drinking water access remains challenging in sub-Saharan Africa, where groundwater is widely used without treatment. In M'Pody (Anyama municipality, C & ocirc;te d'Ivoire), this is amplified by climatic variability and inadequate sanitation, increasing microbiological health risks. This study characterizes contamination in well water and quantifies the influence of climatic factors (rainfall, temperature), sanitation proximity, and inter-indicator relationships on contamination probability and intensity. A mixed-effects hurdle modeling framework was applied to four indicators: Escherichia coli (E. coli), Enterococcus faecalis (E. faecalis), total coliforms (CT) and thermotolerant coliforms (CTH). Due to right-skewed concentrations and frequent zeros, a two-part model combined logistic regression for detection and generalized linear mixed models (GLMM) for conditional intensity on the log10(x + 1) scale. Rainfall increased detection probability, especially for E. faecalis, and intensity for E. coli, E. faecalis, and CTH, whereas higher temperatures were linked to lower concentrations. Proximity to septic pits had limited effects, except for increased E. coli within 15m. Predictive modeling showed CTH as a strong proxy for E. coli: CTH intensity predicts both presence and concentration, with a steep nonlinear risk gradient across its interquartile range. Compared with exploratory multivariate methods, the GLMM-hurdle framework provides robust, interpretable quantification of contamination drivers and indicator interdependencies. Well water in M'Pody fails to meet World Health Organization standards, highlighting the need for reinforced microbiological monitoring and targeted interventions during the rainy season.
This study used time series, differential equations, and probability theory as the foundation to analyze models of the spread of infectious and chronic diseases. The research focuses on the Susceptible-Infected-Recovered (SIR) model with the integration of time-varying parameters. Logistic regression and Bayesian models were also incorporated to assess the impact of genetic mutations on disease susceptibility. Additionally, time series models such as ARIMA were employed to study the progression of diseases like COVID-19, diabetes, and cardiovascular conditions. The aim of the research is to develop a comprehensive framework for understanding disease dynamics and supporting effective strategies for prevention, detection, and treatment, while enhancing public health initiatives and improving clinical decision-making.
Under-five mortality remains a crucial indicator of child health and overall development. This study utilizes the Box-Jenkins ARIMA model to examine historical patterns and predict future trends in under-five mortality. Temporal patterns were examined to predict future mortality figures using historical data, which was systematically tested for stationarity, autocorrelation, and model suitability. Among the various models assessed, (ARIMA 0, 1, 3) emerged as the best fit, determined by the lowest Akaike Information Criterion (AIC) and Bayesian Information Criterion values, as well as successful diagnostic checks. This model effectively captured the declining trend in under-five mortality while accommodating short-term fluctuations. Forecasts for the period from 2024 to 2033 project a gradual increase in under-five deaths. The 95% confidence intervals are getting wider, which means that long-term predictions are becoming less certain. While the decline in mortality has slowed, the forecast suggests a likely stabilization rather than a resurgence. These findings highlight the effectiveness of ARIMA modeling for monitoring mortality trends and emphasize the necessity for ongoing, targeted public health interventions to accelerate the reduction of preventable child deaths and achieve Sustainable Development Goal 3.2 by 2030.
The regulation of the human immune system by HIV infection is maintained by the CD4+ T cells and their decrease, severely impair the immune function of the body. Since statistical models capture biological variations, they can evaluate the decrease patterns of CD4+ T cells and enable predictive assessment for depletion of the same in HIV infected individuals. To understand the clinical evaluation, it is essential to understand the statistical behaviour of these immune cells and quantification of their uncertainty. The study proposes a predictive probability model for CD4+ T cell depletion in individuals with HIV infection that is formulated by the inverted Rayleigh distribution, and predictive probability distributions are derived under three different priors allowing a probabilistic forecast of CD4+ T cell depletion. The root mean square error was computed between the predictive probability values and the probability values based on maximum likelihood estimation to assess the closeness of predictive performance without parameters.
The classical Youden Index has been a widely used performance measure in statistics, assuming determinate observations in the dataset. However, when confronted with data containing uncertain, indeterminate, or unsure observations, its adequacy diminishes. To address this limitation, we propose the neutrosophic Youden index (NJ) as a generalization of the classical approach to accommodate neutrosophic or fuzzy observations. Inthis paper, we introduce two novel indices, namely, the neutrosophic Youden index and the neutrosophic weighted Youden index, specifically developed to handle fuzzy environments. Through comprehensive evaluation, we demonstrate the applicability and effectiveness of these indices in diagnostic testing. We present a practical application of the neutrosophic Youden index by comparing two diagnostic tests for cancer identification. Inaddition, we summarize the neutrosophic Youden index and confidence intervals for both tests and hypothesis testing, providing valuable insights for comparing two independent diagnostic tests. Furthermore, to showcase the potential of the neutrosophic weighted Youden index (NJw) in diagnostic test development, we employ it in a scenario involving two diagnostic tests for identifying pheochromocytoma. Alongside this, we present confidence intervals and hypothesis testing for these tests, showcasing the utility of the weighted approach under neutrosophy. In conclusion, our proposed NJ and NJw present suitable and robust options for comparative diagnostic analysis in medical science, biostatistics, and epidemiology. Particularly, the weighted neutrosophic Youden index emerges as a valuable tool when the consideration of sensitivity and specificity holds paramount importance in diagnostic test evaluation.
A stroke is a complex neurological condition marked by the sudden disruption of brain function due to vascular events such as hemorrhage or infarction. Rather than being a single disease entity, stroke emerges from a convergence of multiple risk factors and underlying health conditions. Recent estimates from the World Stroke Organization indicate that approximately 12.2 million individuals suffer a stroke each year, with around 6.5 million deaths attributed to it. Ischemic strokes account for more than 62% of these cases globally. Alarmingly, nearly 89% of stroke-related mortality and long-term disability occur in low- and middle-income countries, where preventive care and early intervention often remain limited. Several clinical indicators-including tobacco use, elevated glucose levels, cardiovascular complications, hypertension, and advanced age-have been consistently associated with poorer survival outcomes among stroke patients. Lifestyle modifications play a vital role in both recovery and prevention. This study explores the likelihood of stroke occurrence by analyzing patient data, identifying major risk factors, and examining the strength of association between these variables and stroke incidence.
Background: Malnutrition among children persists as a significant public health concern in India, contributing substantially to under-five mortality. Despite ongoing national and global efforts, undernutrition and micronutrient deficiencies persist. Maternal health, education, and socioeconomic status are key determinants of child nutrition. Understanding the spatial distribution and underlying determinants of malnutrition is critical for designing region-specific interventions. This study examines the determinants of under-five malnutrition and their spatial variations across Indian districts, employing data derived from the National Family Health Survey (NFHS-5). Methods: The analysis utilized data from the National Family Health Survey (NFHS-5), aggregated for 707 districts across India. Child nutritional status was evaluated based on WHO-recommended anthropometric indicators, namely stunting, underweight, and wasting. Spatial autocorrelation was assessed using Global Moran's I, and spatial dependence, along with key predictors, was examined through the Spatial Lag Model (SLM) and Spatial Error Model (SEM). Results: The analysis revealed that, on average, one-third of children were stunted (33.5%), nearly 30% were underweight (29.5%), and approximately one-fifth were wasted (18.6%). Significant positive spatial autocorrelation was observed for all indicators, with Moran's I values of 0.52, 0.65, and 0.44 (p-value < 0.001), confirming spatial clustering of malnutrition. Spatial regression analyses identified maternal undernutrition (BMI < 18.5 kg/m(2)) as the highest significant positive determinant (p-value < 0.001) across all indicators. Early maternal marriage (<18 years) and maternal anaemia were also positively associated, while maternal literacy and institutional delivery were protective. Underweight showed the highest spatial dependence (rho = 0.50; R-2 = 0.70). Conclusion: Under-five malnutrition in India exhibits strong spatial clustering, indicating that neighbouring districts share similar nutritional outcomes. Maternal undernutrition, early marriage, and socioeconomic deprivation are major drivers of child malnutrition. Targeted, region-specific strategies that focus on maternal nutrition, early marriage, and promoting maternal education and nutritional support are crucial for reducing regional disparities and achieving India's national nutrition goals.
Neuroinflammation critically shapes stroke outcomes, yet therapies often fail due to the biphasic nature of cytokines. Tumor Necrosis Factor-alpha (TNF-alpha), for instance, drives acute injury (0-24h) but later supports repair. To address this paradox, we developed an identified two optimal intervention windows: early suppression of reproduced and validated against cytokine kinetics from tMCAo volume by 32% in silico. Although assuming uniform cytokine distribution, the model offers a framework for phased, precisionguided immunomodulation. These findings highlight the critical interventions.
The existing t-test under classical statistics is used when all observations in the data are determinate. Incase of uncertain and imprecise observations, the existing t-test cannot be applied. In this paper, a t-test for two populations when variances are unknown and unequal in practice will be introduced when the data has imprecise observations. The test statistic will be introduced for imprecise observations. The testing procedure of the proposed t-test will be discussed considering the degree of uncertainty. The applications of the proposed t-test will be discussed using pet dogs fed data. From the data analysis, the proposed t-test was found to be more informative than the existing t-test under classical statistics.
This article analyzes malaria cases in River Nile State (Sudan) during the period 2018-2022. It concludes that Malaria prevalence rates in River Nile State (Sudan) were very similar across the study years (2018-2022), reaching 17%, 16%, 15%, 14%, and 13%, respectively. This is very high compared to the global malaria prevalence rate (3.1%), but close to the African average (16.4%). There is a significant difference in malaria cases across the years. Due to the increase in the number of cases in recent years, there is no significant difference in cases between months during each study year, meaning that cases are similar throughout the year. However, cases increase in winter and decrease in summer, cases increased for all months in 2022 compared to 2018, especially in April, except for June and July, where they decreased by 3% and 8%. Overall infections increased by 27.6% from 2018 to 2022 and the predictive values showed an increase in the number of infections over the next five years. The article recommends that there should be an increase in mosquito control measures, especially during the winter months. Mosquito nets be provided, especially to those living in rural areas. They must be educated about water conservation. Necessary treatment be provided to reduce the number of deaths and study the reasons that led to the sharp rise in April 2022.
This study aims to establish the fundamental concepts of multiple linear regression models involving qualitative predictor variables and to validate their associations using a Multilayer Feed Forward (MLFF) neural network. A simple guide is introduced to separate qualitative variables based on the number of classes, ensuring they meet the assumptions of multiple linear regression. The approach provides a basic and practical template for integrating qualitative predictors into applied linear models. Validation is carried out using the sum of square error and relative error obtained from the MLFF neural network. The low error values produced by the MLFF model highlight the effectiveness and superiority of the proposed methodology.