To address the challenges associated with measuring and classifying household consumption (poverty) in developing countries, such as cost, time gaps, and inaccurate socio-economic data, this study suggests leveraging machine learning (ML) algorithms. We assessed the performance of various ML algorithms using data from 14,580 sample households from the Integrated Household Living Condition Survey (EICV5), considering 87 features. Among the 12 classifiers evaluated, multiple kernel support vector machines, eXtreme gradient boosting, and multinomial logit demonstrated the highest predictive accuracy, ranging between 86.6% and 88.5%. Notably, household food expenditure, the total number of children (<14 years) in the household, and household own food expenditures emerged as the most predictive features for consumption classification. Interestingly, including shock-coping strategies did not significantly improve prediction accuracy. The multiple kernel support vector machine consistently outperformed eXtreme gradient boosting and multinomial logit. These findings suggest that survey questions used to assess poverty in Rwanda could be streamlined, prioritizing important features, particularly those related to household food characteristics. This approach has the potential to address challenges associated with measuring and classifying household consumption in developing countries more effectively.
Background. The HIV epidemic varies significantly across different groups in the SSA region, and complicates designing effective general interventions. We aim at uncovering temporal trends in HIV prevalence disaggregated by age, sex, and country.Method. We determined HIV prevalence trends among males and females aged 15-49 years for surveys conducted from the years 2003-2007 (period 1) and 2013-2018 (period 2) in SSA. Countries were divided into three clusters based on their socio-behavioural characteristics, and age was categorized into ranges of 15 to 24 years, 25 to 34 years, and 35 to 49 years. A log-binomial regression model was employed in testing for a discrepancy in prevalence between these groups.Results. Swaziland had the highest increase in HIV prevalence between the two periods among females, with a 3.18% rise, followed by Lesotho, Ethiopia and Tanzania females with 2.96%, 2.18% and 0.05%, respectively. Men of ages 15 and 24 experienced an increase from 0.29 (95% CI 0.1-0.84%) to 0.32 (95% CI 0.15-0.68%), 0.36 (95% CI 0.18-0.73%) to 0.47 (95% CI 0.32-0.7%) between the two periods in Cote d'Ivoire and Rwanda, respectively. Females of age 25 to 34 in Ethiopia had an increase in the prevalence of (3.26 (95% CI 2.6-4.08%) to 4.25 (95% CI 3.47-5.21%)). Conclusions. There is a significant difference in HIV prevalence in the general population between the sexes and age categories. In general, females outperformed their male counterparts in every category.
AbstractArtificial intelligence (AI) is enabling organizations to address a range of real-world challenges in areas as diverse as global health, education and poverty alleviation.
In this research we use data from a number of different sources of satellite imagery. Below we describe and visualize various metrics of the datasets being considered. Satellite imagery is retrieved from Google earth which is supported by Data SIO (Scripps Institution of Oceanography), NOAA (National Oceanic and Atmospheric Administration), US. Navy (United States Navy), NGA (National Geospatial-Intelligence Agency), GEBCO (General Bathymetric Chart of the Oceans), Image Landsat, and Image IBCAO (International Bathymetric Chart of the Arctic Ocean). Using random sampling of spatial area in Kigali per target area, 342,843 thousands images were retrieved under the five categories: residential high income (78941), residential low income(162501), residential middle income(101401), commercial building, (67400) and industrial zone,(24400). For the industrial zone, we also included some images from Nairobi, Kenya industrial spatial area. The average number of samples for a category is 86929. The size of the sample per category is proportional to the size of the spatial target area considered per category. Kigali is located at latitude:-1.985070 and longitude:-1.985070, coordinates. Nairobi is located at latitude:-1.286389 and longitude:36.817223, coordinates.
Modelling, simulation, and forecasting offer a means of facilitating better planning and decision-making. These quantitative approaches can add value beyond traditional methods that do not rely on data and are particularly relevant for public transportation. Lagos is experiencing rapid urbanization and currently has a population of just under 15 million. Both long waiting times and uncertain travel times has driven many people to acquire their own vehicle or use alternative modes of transport. This has significantly increased the number of vehicles on the roads leading to even more traffic and greater traffic congestion. This paper investigates urban travel demand in Lagos and explores passenger dynamics in time and space. Using individual commuter trip data from tickets purchased from the Lagos State Bus Rapid Transit (BRT), the demand patterns through the hours of the day, days of the week and bus stations are analysed. This study aims to quantify demand from actual passenger trips and estimate the impact that dynamic scheduling could have on passenger waiting times. Station segmentation is provided to cluster stations by their demand characteristics in order to tailor specific bus schedules. Intra-day public transportation demand in Lagos BRT is analysed and predictions are compared. Simulations using fixed and dynamic bus scheduling demonstrate that the average waiting time could be reduced by as much as 80%. The load curves, insights and the approach developed will be useful for informing policymaking in Lagos and similar African cities facing the challenges of rapid urbanization.
Big data offers the potential to calculate timely estimates of the socioeconomic development of a region. Mobile telephone activity provides an enormous wealth of information that can be utilized alongside household surveys. Estimates of poverty and wealth rely on the calculation of features from call detail records (CDRs), however, mobile network operators are reluctant to provide access to CDRs due to commercial and privacy concerns. As a compromise, this study shows that a sparse CDR dataset combined with other publicly available datasets based on satellite imagery can yield competitive results. In particular, a model is built using two CDR-based features, mobile ownership per capita and call volume per phone, combined with normalized satellite nightlight data and population density, to estimate the multi-dimensional poverty index (MPI) at the sector level in Rwanda. This model accurately estimates the MPI for sectors in Rwanda that contain mobile phone cell towers (cross-validated correlation of 0.88).
Hundreds of organizations and analysts use energy projections, such as those contained in the US Energy Information Administration (EIA)'s Annual Energy Outlook (AEO), for investment and policy decisions. Retrospective analyses of past AEO projections have shown that observed values can differ from the projection by several hundred percent, and thus a thorough treatment of uncertainty is essential. We evaluate the out-of-sample forecasting performance of several empirical density forecasting methods, using the continuous ranked probability score (CRPS). The analysis confirms that a Gaussian density, estimated on past forecasting errors, gives comparatively accurate uncertainty estimates over a variety of energy quantities in the AEO, in particular outperforming scenario projections provided in the AEO. We report probabilistic uncertainties for 18 core quantities of the AEO 2016 projections. Our work frames how to produce, evaluate, and rank probabilistic forecasts in this setting. We propose a log transformation of forecast errors for price projections and a modified nonparametric empirical density forecasting method. Our findings give guidance on how to evaluate and communicate uncertainty in future energy outlooks.
This study evaluated methods for automated classification of rain events into groups of "high" and "low" spatial and temporal variability in offline and online situations. The applied classification techniques are fast and based on rainfall data only, and can thus be applied by, e.g., water system operators to change modes of control of their facilities. A k-means clustering technique was applied to group events retrospectively and was able to distinguish events with clearly different temporal and spatial correlation properties. For online applications, techniques based on k-means clustering and quadratic discriminant analysis both provided a fast and reliable identification of rain events of "high" variability, while the k-means provided the smallest number of rain events falsely identified as being of "high" variability (false hits). A simple classification method based on a threshold for the observed rainfall intensity yielded a large number of false hits and was thus outperformed by the other two methods.
Motivation: Clustering techniques are routinely applied to identify patterns of co-expression in gene expression data. Co-regulation, and involvement of genes in similar cellular function, is subsequently inferred from the clusters which are obtained. Increasingly sophisticated algorithms have been applied to microarray data, however, less attention has been given to the statistical significance of the results of clustering studies. We present a technique for the analysis of commonly used hierarchical linkage-based clustering called Significance Analysis of Linkage Trees (SALT). Results: The statistical significance of pairwise similarity levels between gene expression profiles, a measure of co-expression, is established using a surrogate data analysis method. We find that a modified version of the standard linkage technique, complete-linkage, must be used to generate hierarchical linkage trees with the appropriate properties. The approach is illustrated using synthetic data generated from a novel model of gene expression profiles and is then applied to previously analysed microarray data on the transcriptional response of human fibroblasts to serum stimulation.
Wind power forecasting techniques have received substantial attention recently due to the increasing penetration of wind energy in national power systems. While the initial focus has been on point forecasts, the need to quantify forecast uncertainty and communicate the risk of extreme ramp events has led to an interest in producing probabilistic forecasts. Using four years of wind power data from three wind farms in Denmark, we develop quantile regression models to generate short-term probabilistic forecasts from 15 min up to six hours ahead. More specifically, we investigate the potential of using various variability indices as explanatory variables in order to include the influence of changing weather regimes. These indices are extracted from the same wind power series and optimized specifically for each quantile. The forecasting performance of this approach is compared with that of appropriate benchmark models. Our results demonstrate that variability indices can increase the overall skill of the forecasts and that the level of improvement depends on the specific quantile.
There has been considerable recent research into the connection between Parkinson's disease (PD) and speech impairment. Recently, a wide range of speech signal processing algorithms (dysphonia measures) aiming to predict PD symptom severity using speech signals have been introduced. In this paper, we test how accurately these novel algorithms can be used to discriminate PD subjects from healthy controls. In total, we compute 132 dysphonia measures from sustained vowels. Then, we select four parsimonious subsets of these dysphonia measures using four feature selection algorithms, and map these feature subsets to a binary classification response using two statistical classifiers: random forests and support vector machines. We use an existing database consisting of 263 samples from 43 subjects, and demonstrate that these new dysphonia measures can outperform state-of-the-art results, reaching almost 99% overall classification accuracy using only ten dysphonia features. We find that some of the recently proposed dysphonia measures complement existing algorithms in maximizing the ability of the classifiers to discriminate healthy controls from PD subjects. We see these results as an important step toward noninvasive diagnostic decision support in PD.
In equation ( 18), the first line at the top of page 1328 inoccrectly reads) should be g(e t-1 ).
Tracking Parkinson's disease (PD) symptom progression often uses the unified Parkinson's disease rating scale (UPDRS) that requires the patient's presence in clinic, and time-consuming physical examinations by trained medical staff. Thus, symptom monitoring is costly and logistically inconvenient for patient and clinical staff alike, also hindering recruitment for future large-scale clinical trials. Here, for the first time, we demonstrate rapid, remote replication of UPDRS assessment with clinically useful accuracy (about 7.5 UPDRS points difference from the clinicians' estimates), using only simple, self-administered, and noninvasive speech tests. We characterize speech with signal processing algorithms, extracting clinically useful features of average PD progression. Subsequently, we select the most parsimonious model with a robust feature selection algorithm, and statistically map the selected subset of features to UPDRS using linear and nonlinear regression techniques that include classical least squares and nonparametric classification and regression trees. We verify our findings on the largest database of PD speech in existence (~6000 recordings from 42 PD patients, recruited to a six-month, multicenter trial). These findings support the feasibility of frequent, remote, and accurate UPDRS tracking. This technology could play a key part in telemonitoring frameworks that enable large-scale clinical trials into novel PD treatments.
BACKGROUND:Developing methods for understanding the connectivity of signalling pathways is a major challenge in biological research. For this purpose, mathematical models are routinely developed based on experimental observations, which also allow the prediction of the system behaviour under different experimental conditions. Often, however, the same experimental data can be represented by several competing network models.RESULTS:In this paper, we developed a novel mathematical model/experiment design cycle to help determine the probable network connectivity by iteratively invalidating models corresponding to competing signalling pathways. To do this, we systematically design experiments in silico that discriminate best between models of the competing signalling pathways. The method determines the inputs and parameter perturbations that will differentiate best between model outputs, corresponding to what can be measured/observed experimentally. We applied our method to the unknown connectivities in the chemotaxis pathway of the bacterium Rhodobacter sphaeroides. We first developed several models of R. sphaeroides chemotaxis corresponding to different signalling networks, all of which are biologically plausible. Parameters in these models were fitted so that they all represented wild type data equally well. The models were then compared to current mutant data and some were invalidated. To discriminate between the remaining models we used ideas from control systems theory to determine efficiently in silico an input profile that would result in the biggest difference in model outputs. However, when we applied this input to the models, we found it to be insufficient for discrimination in silico. Thus, to achieve better discrimination, we determined the best change in initial conditions (total protein concentrations) as well as the best change in the input profile. The designed experiments were then performed on live cells and the resulting data used to invalidate all but one of the remaining candidate models.CONCLUSION:We successfully applied our method to chemotaxis in R. sphaeroides and the results from the experiments designed using this methodology allowed us to invalidate all but one of the proposed network models. The methodology we present is general and can be applied to a range of other biological networks.
Wind power is an increasingly used form of renewable energy. The uncertainty in wind generation is very large due to the inherent variability in wind speed, and this needs to be understood by operators of power systems and wind farms. To assist with the management of this risk, this paper investigates methods for predicting the probability density function of generated wind power from one to ten days ahead at five U.K. wind farm locations. These density forecasts provide a description of the expected future value and the associated uncertainty. We construct density forecasts from weather ensemble predictions, which are a relatively new type of weather forecast generated from atmospheric models. We also consider density forecasting from statistical time series models. The best results for wind power density prediction and point forecasting were produced by an approach that involves calibration and smoothing of the ensemble-based wind power density.