Childhood anemia affects around 40 We used DHS data from 16 countries across Africa, Asia, Latin America, the Caucasus, and the Middle East (n=68,856). We compared Logistic Regression, XGBoost, LightGBM, and TabPFN v2.6. Performance was assessed using AUC-ROC, Brier score, and ECE. Generalization was evaluated using leave-one-country-out (LOCO), reverse-LOCO, and few-shot settings. Subgroup analyses included sex, age, residence, maternal education, and wealth. Feature importance was estimated using SHAP. TabPFN outperformed classical models in low-data regimes (<200 samples), showing higher discrimination and better calibration. Across countries, it achieved the lowest Brier score (0.042) and ECE (0.203). Under full-data settings, AUC-ROC ranged from 0.59-0.76 with small between-model differences (≤ 0.05). LOCO performance was stable (0.58-0.69), driven by country context. Reverse-LOCO showed asymmetric transferability. Subgroup performance was consistent with no systematic demographic bias. SHAP identified child age, altitude, and height-for-age z-score as dominant predictors, followed by wealth and maternal education. Performance in childhood anemia prediction is driven more by population variation than model choice. TabPFN provides advantages in low-resource settings through improved discrimination and calibration, highlighting foundation models as promising tools for data-scarce global health prediction.
This study introduces a novel statistical model called the modified Fréchet-exponentiated exponential (MFrEE) distribution. The existing exponentiated exponential (EE) distribution, while useful for lifetime and reliability data, has limited flexibility in capturing diverse hazard shapes and may not adequately model extreme events or tail behavior. To address these limitations, the MFrEE distribution applies a modified Fréchet generator to the EE baseline, enhancing the model’s flexibility and robustness. Its survival and hazard functions, cumulative distribution function, and probability density function are derived, presented, and illustrated with plots for various parameter values. The study provides a comprehensive mathematical analysis of the distribution, deriving its moments, mean, variance, quantiles, and moment-generating function. Methodologically, the model is simulated using an accept–reject algorithm, and its parameters are estimated via maximum likelihood estimation (MLE). The performance of the estimators is assessed through Monte Carlo simulations using bias, mean squared error, and coverage probability (CP), with the CP results showing values close to the nominal 95% level across different parameter settings. Furthermore, the robustness and performance of the proposed method are evaluated using AIC, BIC, and AICc, demonstrating superior performance compared to baseline methods across three publicly available datasets. The study concludes by proposing this model as a significant contribution to probability theory and suggests two avenues for future research: applying the model to more real-world problems and using machine learning methods for parameter estimation to compare with the MLE approach used in this study.
IntroductionThe phrase “children ever born” refers to the total number of children a woman has during her lifetime, which is considered one of the three primary factors influencing a country's population size, composition, and structure. This study aimed to examine the spatial differences in the number of children ever born and related factors among women of reproductive age in rural Ethiopia. MethodsThis study utilized data from the 2019 Ethiopian Mini Demographic and Health Surveys, focusing on 5,934 rural women aged 15–49 years. Of the four count regression models considered, the zero-inflated Poisson regression model was identified as the most suitable for the data. Additionally, a spatial analysis was conducted to evaluate spatial dependencies across different zones in Ethiopia. ResultsIn Ethiopia, rural women typically have an average of 3.1 children throughout their lives. The distribution of the total number of children born was spatially clustered across different zones of Ethiopia (Moran's I=0.17). Notable hotspot areas were found in Shinile, Fik, Gode, Warder, Guji, Gurage, and West Harerge. Women who had their first child before turning 19 years old showed an IRR of 1.341 (IRR = 1.341), suggesting a higher likelihood of having more children than others. Conversely, women who practiced family planning (IRR = 0.961) compared to those who did not practice were less likely to have more children. DiscussionThe study, consistent with previous studies, shows that higher women’s education and wealth status, and use of family planning are associated with fewer children ever born, whereas experiences such as child mortality and early childbirth increase fertility, , highlighting the importance of improving reproductive health services, education, and socio-economic conditions to influence fertility patterns among rural Ethiopian women. ConclusionThe study, consistent with previous studies, shows that higher women’s education and improved wealth status, as well as the use of family planning, are associated with fewer children ever born, whereas experiences such as child mortality and early childbirth increase fertility, highlighting the importance of improving reproductive health services, education, and socio-economic conditions to influence fertility patterns among rural Ethiopian women.
Anemia remains a major public health concern in Ethiopia, particularly among women of reproductive age and children, with substantial variation across local administrative zones. Reliable small-area estimates of hemoglobin levels are essential for evidence-based health planning; however, direct survey estimates are often unreliable at disaggregated levels due to small sample sizes and high variability. This study applies a bivariate small area estimation (SAE) approach to improve the precision of hemoglobin level estimates for women and children by exploiting their inherent correlation. Data were obtained from the Ethiopian Demographic and Health Survey (EDHS) and auxiliary variables from the Population and Housing Census. The analysis employed the Bivariate Fay-Herriot (BFH) model to jointly model hemoglobin levels, allowing information sharing through both census covariates and the correlation between outcomes. Model performance was compared with the Univariate Fay-Herriot (UFH) model and traditional direct survey estimates. The results show that the BFH model provides more stable and precise estimates than both UFH and direct methods, demonstrating the benefits of borrowing strength across correlated health indicators. These findings highlight the value of bivariate modeling in enhancing the reliability of local-level health estimates and reducing uncertainty in survey-based measures. Accurate subnational hemoglobin estimates can support policymakers in designing targeted interventions for anemia reduction. The study encourages the use of multivariate SAE frameworks in future health research and national surveys to improve data-driven decision-making and monitoring of public health outcomes.
Gender-based violence can include sexual, physical, mental, and economic harm inflicted in public or in private. This violence also has a direct psychological effect, physical and financial consequences, and it has multiple underlying reasons, such as social, economic, cultural, political, and religious aspects. By applying multiple resampling techniques, this study aims to improve the precision and accuracy of supervised machine learning classifications of gender-based violence (GBV) using the SDHS dataset. The class imbalance between GBV-positive and GBV-negative instances makes it very challenging to produce reliable classification machine learning models. To address this issue, oversampling machine learning approaches, including synthetic minority over-sampling technique (SMOTE), adaptive synthetic (ADASYN), and random over-sampling (ROS), were employed to classify the GBV data in Somalia. The logistic regression (LR), decision tree (CART), random forest (RF), naïve Bayes (NB), k-nearest Neighbors (KNN), and support vector machine (SVM) methods were trained and evaluated. In addition, oversampling techniques were employed for improving the imbalanced datasets. Receiver operating characteristic curve (ROC) and the area under the curve (AUC) were used to assess each machine learning classifier and to compare performance on the original GBV dataset. Among the resampling techniques, SMOTE (RF = 0.992, CART = 0.969, and KNN = 0.957) outperformed ADASYN (RF = 0.912, CART = 0.910, and KNN = 0.876) and ROS (RF = 0.920, CART = 0.919, and KNN = 0.880) across almost all evaluation metrics. The classifiers that performed the best were random forest (RF) and classification and regression trees (CART), then k-nearest Neighbors. After resampling the imbalanced dataset, we may therefore conclude that the random forest (AUC = 0.972), CART (AUC = 0.969) and KNN (AUC = 0.957) machine learning classifiers are better at accurately classifying the k-nearest Neighbors dataset. In addition, compared to the other oversampling techniques, SMOTE was used to the machine learning classifiers to balance the imbalanced class distributions in favour of the minority class. In addition, SMOTE with the Mathews correlation coefficient (MCC) outperformed ADASYN and ROS resampling techniques. The MCC values for SMOTE reached their highest values (RF = 0.86, CART = 0.85, and KNN = 0.80), indicating strong overall predictive reliability of the machine learning models. Therefore, the findings of this analysis will assist government and non-government organizations in making policy decisions to GBV risks.
Survival analysis is a crucial statistical tool for evaluating the impact of risk factors on time-to-event outcomes, such as disease progression or death. In clinical monitoring of patients with chronic kidney disease (CKD), subjects are at risk for competing clinical outcomes, such as death or progression to end-stage renal disease, where one event may preclude or alter another from occurring. Standard survival analysis, which treats such events as independent censoring, yields biased estimates. A competing risks framework is therefore essential for valid inference. Time-varying covariates with measurement errors can introduce additional bias if unaccounted for. Furthermore, the demand for scalable and rapid estimation approaches in Bayesian inference for survival data with many features is critical in medical research. This paper proposes an Integrated Nested Laplace Approximation (INLA)-based Bayesian method to model competing risks in CKD while accounting for covariate measurement error. Bayesian cause-specific competing risks models that incorporate a mixed-effects covariate measurement error submodel were proposed. Various nonparametric and parametric baseline hazard distributions were evaluated. The proposed methods were illustrated using both a simulation study and real-world CKD data analysis. The simulation study revealed that the INLA approach provided nearly identical and accurate posterior parameter estimates while being more computationally efficient compared to a Markov-Chain Monte-Carlo (MCMC) approach. The simulation and application studies demonstrate that this study makes noteworthy contributions to the proper and efficient analysis of survival CKD data with competing risks and covariate measurement errors.
Hemoglobin, an iron-rich protein in red blood cells, enables oxygen delivery to all body tissues. The hemoglobin records in most countries are collected through surveys. Surveys often provide reliable estimates of target variables for the population at both national and regional administration levels. This study investigates the spatial distribution of hemoglobin levels among women and children across Ethiopian administrative zones. Because survey sample sizes are often insufficient to produce reliable direct estimates at the zonal level, small area estimation methods are used to improve the precision of estimates. Small area estimation provides reliable estimates at the local zones by combining survey, census, and spatial datasets. Therefore, we focus on geospatial small area estimation to provide reliable and precise hemoglobin estimates for women and children. Therefore, our prime objective is to provide reliable and precise estimates of hemoglobin levels for women and children using geospatial small area estimation by integrating the survey and census datasets. An area-level Fay-Herriot modeling framework was implemented to estimate hemoglobin levels across the studied zones. The results show that for both children and women, the Spatial Empirical Best Linear Unbiased Predictor (SEBLUP) consistently produces the lowest coefficients of variation (CVs), followed by the standard Empirical Best Linear Unbiased Predictor (EBLUP), while direct survey estimates display the highest variability. The comparison of CVs for hemoglobin estimates demonstrates that children’s mean CVs decrease from 13.64 (Direct survey estimates) to 11.29 (EBLUP) and then to 6.06 (SEBLUP). Similarly, women’s mean CVs decline from 13.48 (Direct survey estimates) to 10.71 (EBLUP) and then 6.37 (SEBLUP), confirming that SEBLUP provides the highest precision for both children and women, consistently producing the lowest CVs across all statistical measures. The results in the maps indicate the spatial disparities in hemoglobin levels for both women and children among the Ethiopian administrative zones. Policymakers will significantly benefit from these local area-level estimates, and researchers will gain methodological insights into the application of geospatial small-area estimation.
Objectives:Sub-Saharan Africa continues to experience the highest under-five mortality rates globally, contributing 29.7% of all under-five deaths despite a 60% global decline between 1990 and 2022. This study aims to analyze time to death among children under five in Somalia and identify the key factors influencing child survival. Study design:The data used in this study is a population-based cross-sectional survey using a multistage stratified cluster sampling design. Methods:In this study, 17,610 children under five from a Somalia 2020 demographic and health survey (DHS) were used. The accelerated failure time (AFT) model was used to analyze the time to death of under-five children. Survival time ratios (TR) and corresponding p-values were used to identify significant determinants of child survival. Results:Of the total 17,610 children, about 689 children (3.91%) experienced the event (death). Several AFT models were compared, and the Weibull AFT model was selected as the best fit. The results of the Weibull AFT model showed that significant factors that influence child survival include maternal age at the first birth, preceding birth interval, the number of children ever born, and regional disparities. Longer birth intervals (18-59 months) increased survival time for the children, while shorter or excessively long intervals reduced survival. Mothers aged 20-29 at first birth showed a 49.2% increase in survival time (TR = 1.492; p = 0.003), compared to the younger mother. The shape parameter (0.607) suggests a declining hazard rate over time. Conclusions:This study highlights critical maternal, familial, and regional factors that influence child survival in Somalia. Strengthening targeted interventions, particularly those promoting optimal birth spacing and supporting younger mothers, may substantially improve under-five children survival outcomes.
Lowering the rate of onset, progression, mortality, and morbidity remains a key goal in the management of Human Immunodeficiency Virus (HIV). Antiretroviral therapy (ART) is working effectively to enhance prognosis and quality of life among individuals living with HIV. Viral load suppression is a viral load of less than 1000 copies/ml in an individual receiving antiretroviral therapy (ART), indicating that ART is working effectively to stop the virus from multiplying. Despite significant progress in HIV treatment, achieving and sustaining viral load suppression remains an ongoing public health concern. Therefore, ART leads to a high chance of suppressing HIV viral load and undetectable levels of the virus to improve individual health and well-being from the disease. However, the prevalence of HIV remains high despite notable advances in increasing access to ART. The results of this study revealed that approximately 64.8
Breast cancer is the most frequently diagnosed cancer among women and persists as a societal problem worldwide. It remains a leading cause of cancer associated morbidity and mortality, specifically in low- and middle-income countries where access to timely diagnosis and treatment is often limited. This study aims to compare survival and classical machine learning models for predicting breast cancer survival in Ethiopia to identify approaches that balance predictive accuracy with interpretability. The study utilized retrospective data from 1164 women treated at Tikur Anbesa Specialized Hospital and Hiwot Fana Specialized University Hospital between 2019 and 2024. Methods like Kaplan-Meier estimation, Cox proportional hazards, random survival forests (RSF), DeepSurv, and classical machine learning (SVM, XGBoost, LGBM, and RF) classifiers were used with evaluation metrics such as AUC, C-index, and Integrated Brier Score (IBS). The Shapley additive explanation approach was used to ensure the interpretability of results from models such as RSF, DeepSurv, and random forests (RF). It allowed the identification of important predictors of breast cancer outcome by indicating consistent predictors across models. The findings demonstrated that random survival forest and random forest achieved the highest performance (C-index: 0.754; IBS: 0.091) and (0.729 ± 0.006), respectively, outperforming the other models under consideration. The Shapley Additive Explanations (SHAP) analysis for the RSF model showed that age, tumour size, metastasis, stage, comorbidities, and marital status as the most important predictors of breast cancer survival. Furthermore, the SHAP analysis for the RF model indicated that the higher age category (45 and above), metastasis status (M1), stage four, and larger tumour size contribute a strong influence on predictions. Among the machine learning models, the random forest algorithm effectively identifies the key predictors of breast cancer outcomes. For the survival analysis methods, the RSF offers robust capabilities for handling time-to-event data and censoring, making it well-suited for accurate survival prediction. By combining these approaches, we were able to gain clearer insights and better identify the key factors influencing breast cancer prognosis. This study highlights the value of data-driven methods in helping healthcare professionals identify high-risk patients with greater precision and take timely, informed actions to support their care.
BACKGROUND:Childhood immunization is an important part of public health efforts to reduce morbidity and mortality from diseases that can be prevented by vaccination. Sub-Saharan Africa (SSA) has the lowest childhood immunization coverage and the highest child mortality rate in the world. Therefore, this study aimed to assess community variation in childhood immunization and to identify determinant factors associated with childhood immunization using mixed effect count regression models. METHOD:This study used data from the 2012-2022 Demographic and Health Survey (DHS), which included 195,000 children between the ages of 12 and 23 months in 33 SSA countries. A various mixed-effects count regression model were employed to identify the variables associated with the prevalence of childhood vaccination. RESULT:In SSA, the mean average childhood immunization was 5.47 (95% CI = 5.46, 5.48), with an 8.38 variance. The mixed effect zero-inflated Poisson model fit the data the best, with the lowest DIC, AIC, and BIC values. The results of the model showed that working mothers, mothers with secondary or higher education (IRR = 1.119; 95% CI: 1.114, 1.125), rich wealth status (IRR = 1.077; 95% CI: 1.073, 1.082), having a health card (IRR = 1.265; 95% CI: 1.258, 1.271), being exposed to the media (IRR = 1.082; 95%CI: 1.077, 1.086), institutional delivery (IRR = 1.168; 95% CI: 1.163, 1.174), receiving eight or more ANC visits (IRR = 1.301; 95% CI: 1.288, 1.314), receiving vitamin A (IRR = 1.350; 95% CI: 1.344, 1.355), using a contraceptive (IRR = 1.187; 95% CI: 1.182, 1.192) and receiving Postnatal Care (PNC) (IRR = 1.075; 95% CI: 1.071, 1.080), were associated with a higher incidence of childhood immunization. While children in rural areas (IRR = 0.972; 95%CI: 0.968, 0.976) had a lower prevalence of childhood vaccinations than children in urban areas. CONCLUSION:SSA had a low coverage of childhood vaccination with significant disparities among countries. Therefore, it is essential to prioritize public health initiatives that target low-income households, rural mothers, uneducated parents, and those who have not utilized maternal health care services for the sake of increasing the coverage of childhood immunizations that in turn enhances the health of children. Furthermore, it is imperative to create policies and initiatives that tackle regional and national variations in childhood immunization rates and to actively work toward their implementation.
Inflation is a critical global issue and also a significant challenge in Ethiopia. Despite its profound impact on the economy, research on inflation volatility in Ethiopia remains limited and insufficient. This paper aims to address these gaps by employing BEKK (Baba, Engle, Kraft, and Kroner) and DCC (Dynamic Conditional Correlation) - GARCH (Generalized Autoregressive Conditional Heteroscedasticity) models and analyze the characteristics of inflation trends, which supports informed economic decision making. We focus on four key inflation indicators: the Consumer Price Index (CPI), the Non-Food Price Index (NFPI), the Food Price Index (FPI), and the Exchange Rate (ER), which were compiled from the National Bank of Ethiopia (NBE) from January 2010 to December 2020. The study confirms inflation volatility, supported by the ARCH effect and Ljung-Box Q(m) statistics, along with conditional heteroscedasticity tests. This study demonstrates that, unlike previous approaches that neglected dynamic correlations in inflation volatility, the DCC-GARCH model decisively outperforms the BEKK-GARCH model in both parameter estimation and forecasting accuracy, as evidenced by significantly better Akaike Information Criterion (AIC), Schwarz Bayesian Information Criterion (SBIC), and Hannan-Quinn Information Criterion (HQIC) metrics. Our findings revealed that the DCC (1,1) model effectively captured volatility clustering without being persistent or explosive, as the sum of coefficients (θ = 0.1794, β = 0.7023) is less than 1, confirming mean reversion. In contrast to previous studies, our approach provided a more robust understanding of inflation dynamics, identifying CPI and FPI as the most volatile indicators. The study reveals significant correlations among inflation indicators-CPI, FPI, NFPI, and ER indicating a cohesive inflationary pattern. The coefficients show that past volatility and shocks persistently influence current volatility, underscoring their interdependence. The forecast from the best model reveals substantial instability is observed in CPI and FPI returns. It suggests a sharp increase in FPI and a rise in ER. The better method captured inflation volatility more effectively than other competent models. The DCC-GARCH model offered deeper insights into volatility dynamics, revealing the shortcomings of earlier time series models in addressing inflation volatility.
Despite major policy reforms and improvements in healthcare coverage, maternal mortality remains a critical public health burden in Ethiopia. While progress has been made since the Millennium Development Goals era, the maternal mortality ratio (MMR) still exceeds national and global targets. This systematic review synthesizes evidence from the past decade (2015–2025) to describe the magnitude, determinants, and regional disparities of maternal mortality in Ethiopia, highlighting persistent challenges and future priorities. Following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) 2020 guidelines (registered in International Prospective Register of Systematic Reviews (PROSPERO)), we systematically searched PubMed, Scopus, Web of Science, Embase, Cochrane, and African Journal Online(AJOL), supplemented with grey literature from World Health Organization (WHO), United Nations International Children's Emergency Fund (UNICEF), and the Ethiopian Ministry of Health. Studies published in English between January 2015 and September 2025 was included. Data extraction followed standardized templates, and study quality was appraised using Joanna Briggs Institute (JBI) and Newcastle–Ottawa Scale (NOS) tools. Given methodological heterogeneity, a narrative synthesis approach was applied. A total of 61 studies met inclusion criteria, encompassing all Ethiopian regions. The pooled MMR was estimated at 366.6 maternal deaths per 100,000 live births, showing only modest progress from previous decades. The leading causes of maternal death were obstetric hemorrhage (29.6
BACKGROUND:Antenatal care (ANC) contacts, along with enhanced health facilities for delivery, are essential components of maternal and child healthcare, as these significantly contribute to both mothers and their newborn child's health. Antennal care contacts primarily help women maintain normal pregnancies by detecting pre-existing conditions and preventing complications that may arise during childbirth. This study intended to determine possible factors that affect both ANC contact and place of delivery among women in Ethiopia. METHODS:The 2019 Ethiopian Mini Demographic and Health Survey data were used for this study. A total weighted sample of 3,926 women nested within 68 zones was used. The bivariate multilevel logistic regression model was utilized to assess the association between antenatal care contact and place of delivery and determinant factors among reproductive-aged women in Ethiopia. RESULTS:In this study, 57% and 47.5% of women had no ANC contacts and home delivery respectively. Similarly, about 36.73% of women delivered at home and didn't utilize the recommended ANC contacts. Only 32.99% of women have both health facility delivery and at least four ANC contacts during their pregnancy. Women who reside in rural areas were 0.612 and 0.352 times less likely to have ANC and health facility delivery compared to women who reside in urban areas. Whereas, the estimated odds of women with higher education levels were 3.803 and 8.406 times the estimated odds of women with no education. CONCLUSION:A high proportion of women are still delivering their new child at home and still don't have at least four ANC contacts during their pregnancy. Women's age, women education level, marital status, wealth status, sex of household head, place of residence, and region were significant predictors of antenatal care visits and place of delivery simultaneously in Ethiopia. Although the country tried to maximize these services, it still requires expansion of health facilities media campaigns, and women's literacy to reduce maternal and newborn child mortality in Ethiopia.
Abortion is one of the leading causes of maternal death in developing countries, particularly in sub-Saharan Africa (sSA,). In this region, abortion is responsible for 38,000 maternal deaths, making the area with the highest rate of abortion-related mortality in the world. This study aimed to examine the prevalence and associated factors of induced abortion in 33 countries in the region. We used data from the most current Demographic and Health Surveys (DHS) conducted in 33 sSA countries between 2012 and 2022. A total 367,881 of women were included in the analysis. The Bayesian multilevel logistic regression model was used to determine the factors associated to induced abortion because of the hierarchical nature of the DHS data. The overall prevalence of induced abortion was 16.50
The child mortality rate is a leading factor in the well-being and development of a nation. It measures the quality of life for a given population. This study aimed to determine the effects of under-five child mortality in Ethiopia. The authors used a cross-sectional study design via the 2019 Ethiopian Demographic and Health Survey. For our study, we used 3837 births recorded by mothers in seven regions of Ethiopia. In this study, the author employed the Bayesian and classical logistic regression models. The study found that the household size, number of under-five children, Sex of child, twin, births in the last five years, and breastfeeding status are significant predictors of child mortality in Ethiopia. Consequently, governmental, non-governmental, and other concerned bodies should focus on targeted healthcare interventions for mothers and children by updating their health intervention policies. In addition, improved health services are needed for better health care for children and mothers. Education should be given to mothers during pregnancy and after birth. This helps improve health for mothers and children, along with addressing other risk factors.
Children worldwide can live lives free from various illnesses and disabilities due to vaccination. For instance, vaccination has eliminated smallpox, a deformative and frequently fatal illness that claimed an estimated 300 million lives in the twentieth century. However, due to a lack of access to immunization and other health services, 14.3 million infants in 2022 still did not receive their first dose of the Diphtheria-Tetanus-Pertussis (DTP) vaccine, and an additional 6.2 million received only a portion of the scheduled dose. This study aimed to assess prevalence and determinant factors of immunization among under-five children in Somalia using Somalia Health and Demographic Survey (SHDS) Data. The study design was cross-sectional, utilizing the SHDS 2020 data. A total of 3916 under-five children who fulfilled the inclusion criteria were included in this study. Count regression models were employed to explore factors associated with the number of vaccinations received per child. In this study, 9.14
BackgroundCommunity-based health insurance (CBHI) is a vital tool for achieving universal health coverage (UHC), a key global health priority outlined in the sustainable development goals (SDGs). Sub-Saharan Africa continues to face challenges in achieving UHC and protecting individuals from the financial burden of disease. As a result, CBHI has become popular in low- and middle-income countries, including Ethiopia. Therefore, this study aimed to identify the ML algorithm with the best predictive accuracy for CBHI enrollment and to determine the most influential predictors among the dataset.MethodsThe 2019 Ethiopian Mini Demographic and Health Survey (EMDHS) data were used. The CBHI were predicted using seven machine learning models: linear discriminant analysis (LDA), support vector machine with radial basis function (SVM), k-nearest neighbors (KNN), classification and regression tree (CART), and random forest (RF). Receiver operating characteristic curves and other metrics were used to evaluate each model’s accuracy.ResultsThe RF algorithm was determined to be the best machine learning model based on different performance assessments. The result indicates that age, wealth index, household members, and land usage all significantly affect CBHI in Ethiopia.ConclusionThis study found that RF machine learning models could improve the ability to classify CBHI in Ethiopia with high accuracy. Age, wealth index, household members, and land utilization are some of the most significant variables associated with CBHI that were determined by feature importance. The results of the study can help health professionals and policymakers create focused strategies to improve CBHI enrollment in Ethiopia.
Time-to-event data analysis without a well-defined time origin commonly occurs in observational studies that retrospectively collect survival endpoints. For instance, after enrolling participants who have or have not received a specific treatment, an event status can be observed for all participants; however, the start date of treatment is only observable for the treatment group. The corresponding time origin does not exist for the control group, resulting in missing survival time data. Complete-case analysis is often considered the standard approach, but it disregards information from all participants in the control group and does not allow us to compare their survival distributions. To address this challenge, we propose a novel semiparametric proportional hazards model by regarding these missing time origins as nuisance parameters. We approximate the risk sets as cumulative normal distributions to deal with these nuisance parameters and develop estimation and inference procedures for our proposed estimator. We study the asymptotic properties of this model and conduct the simulation studies to validate its finite sample property. Analysis of data from a recent SARS-CoV-2 seroprevaluence study illustrates the applicability of our methods. The proposed methods are implemented in the R package coxphm.
Antenatal care (ANC) utilization offers a wide range of interventions, such as education, counseling, screening, treatment, monitoring, and supporting the health of pregnant women, making it a significant opportunity for expectant mothers. This study aims to investigate the time to the first ANC contact among pregnant women and to identify associated factors by employing the Accelerated Failure Time (AFT) model using different frailty distributions. This study used Somalia's Health and demographic survey data. A sample of 3138 women of reproductive age (15-49 years) were included in the study and accelerated failure time (AFT) models with different frailty distributions were compared using information criteria to select the best model. Among the women included in this study, only 33.1% of them received their first ANC contacts within the recommended time during their pregnancy. A gamma frailty model with log-logistic as base-line distribution was found to be the best model for the time-to-first ANC utilization for our data. The final model, based on the log-logistic gamma frailty, identified marital status, mother's occupation, wanted pregnancy, region, parity, wealth index, education level of mother, persons deciding on mother health care, and media exposure are significant (p-value <0.05) predictors of time to the first ANC contact in Somalia. The final model evidenced a high degree of heterogeneity at an individual level regarding the time to the first ANC utilization in Somalia. The median time for the first ANC contact among pregnant women was 6.2 months. To ensure accurate analysis and better policy recommendation, different candidate models were compared, and the univariate gamma frailty model with a log-logistic baseline was found to be the most appropriate approach for analyzing time to the first ANC contact among pregnant women. Maternal and child health policies and initiatives must better focus on women's development and implement interventions aimed at increasing the timely initiation of prenatal care services. More specific policy measures, such as targeted educational campaigns, improved pregnancy services, and efforts to minimize regional disparities, should be prioritized as urgent intervention mechanisms.