Plant-based product replacements are gaining popularity. However, the long-term health implications remain poorly understood, and available methods, though accurate, are expensive and burdensome, impeding the study of sufficiently large cohorts. To identify dietary transitions over time, we examine anonymised loyalty-card shopping records from Co-op Food, UK. We focus on 10,626 frequent customers who directly replaced milk with alternative milk. We then use product nutritional information to estimate weekly nutrient intake before and after the transition. 83% who converted to alternative milk saw a fall in iodine (44%), calcium (30%) and vitamin B12 (39%) consumption, with 57% reducing iodine purchase by more than 50%. The decline is even higher for those switching dairy and meat products. Our findings suggest that dietary transitions - such as replacing milk with alternative milk - could lead to nutritional deficiencies, notably iodine, which, if not addressed, may represent a significant public health concern, particularly in countries which do not mandate salt iodisation.
Introduction & Background The shift towards plant-based diets remains on the rise. Several observational studies have suggested that adopting these diets can result in some fundamental nutrient deficiencies, such as iodine deficiency. This can be especially harmful to a developing fetus, leading to growth impairment and, in extreme cases, cretinism. Nonetheless, understanding of long-term health consequences of this shift remains a challenge, particularly regarding nutritional impact at broader population scales. Objectives & Approach Our study focuses on the effects of transitioning to plant-based diets on the purchasing and assumed intake of essential nutrients like iodine, calcium, and vitamin B12. We analysed anonymized shopping records of 10,626 loyal customers who switched from regular milk to alternative milk products. By matching the transaction data with nutritional information, we estimated the weekly nutrient purchases before and after the transition. Our data was collected from a national food retailer across the UK. Relevance to Digital Footprints Loyalty-card transactional logs held by retailers reflect a valuable lens into nutritional intake data. This data can provide insight into the potential impact of purchasing behaviours, such as the potential health effects of dietary changes at scale. Our approach leverages AI modelling accompanied by rigorous variable importance methods to uncover potentially hidden insights on the impact of nutritional shifts to plant-based goods. Results Results indicate that 83% of individuals deemed regular customers, who switched to plant-based milk, experienced a decrease in their purchases of iodine (44%), calcium (30%), and vitamin B12 (39%) from their normal purchase patterns at the retailer. Additionally, 57% of these individuals decreased their iodine purchases by more than 50%. The reduction in these nutrients is even more significant for those who switch to plant-based dairy and meat products. Conclusions & Implications Our research indicates that dietary changes, such as switching from purchasing regular milk to alternative milks, may lead to insufficient intake of essential dietary nutrients such as iodine. This represents a significant potential health concern for the public if not remediated, especially in countries that do not require salt to be fortified with iodine.
Introduction & Background In England, The Indices of Deprivation (IoD) are a widely used and referenced measure to assess local levels of deprivation across a range of domains, including health and disability. However, due to their complex nature and the number of inputs required to generate these measures, they are only updated infrequently. Typically every 4-5 years, with the most recent versions released in 2019 and 2015. This study expands on previous research looking at the feasibility of using digital footprint data, in the form of retail loyalty card transactions, to predict local deprivation. This work focuses specifically on the health and disability subdomain of IoD. Our hypothesis is that retail behaviour relating to food purchases and their associated nutritional content, can be used to predict health deprivation. Objectives & Approach The work utilises loyalty card data from a large UK grocery retailer. Anonymised geo-location data for loyalty card members was used to assign retail grocery transactions to individual Lower Layer Super Output Areas (LSOAs) for each of the ten quarters in the study period (July 2019 - December 2021). A nutritional lookup was developed to enable the nutritional content of food transactions to be assigned to each LSOA. A number of metrics based on categories of food sold and their nutritional content were developed and used in a Machine Learning model, based on a Random Forest classifier, to predict areas with high levels of health related deprivation. Relevance to Digital Footprints This study uses data derived from digital footprint data of grocery transactions. It demonstrates the potential for utilising digital footprint data as a proxy for traditional demographic data without the need for expensive, both in terms of time and cost, surveying to be performed. Results The random forest classifier was able to predict neighbourhoods (at the LSOA level) with the top 20% of health related deprivation. A high level of predictive power was identified (Overall accuracy 80%). SHAP (SHapley Additive exPlanations) and Model Class Reliance (MCR) were used to determine the importance of the input features. Areas with higher proportional spending on cigarettes and soft drinks and lower spending on fish, wine and fruit and vegtables were found to be associated with extreme levels of health deprivation. In terms of nutrition, two derived metrics, calories per pound spend and the obesogenicity of food purchased, were found to be important predictors of health deprivation. Conclusions & Implications Digital footprint data on grocery purchases have been shown to be highly effective at predicting areas of extreme health related deprivation at the LSOA level. Features related to proportional spend on food categories and proportions of nutrients associated with these purchases were identified as optimal for predicting health related deprivation. The number of calories per pound spent and, to a lesser extent, the proportion spent on cigarettes, in an LSOA was found to be the most important predictor of high levels of health related deprivation. The high level of predictive accuracy obtained offers the potential for using digital footprint data as a proxy for traditional deprivation measures. This could enable rapid and near real-time surveillance of areas with poor health outcomes compared to traditional approaches. This could allow early interventions to be put in place mitigating some of the negative impacts of health related deprivation.
Anthocyanins are a class of polyphenols that have received widespread recent attention due to their potential health benefits. However, estimating the dietary intake of anthocyanins at a population level is a challenging task, due to the difficulty of scaling dietary surveys. Further, there is limited evidence as to who regularly consumes anthocyanins, whether temporally, spatially, or culturally according to levels of socioeconomic deprivation. Leveraging a massive retail loyalty card dataset in the UK, we pair two years of real-world purchasing data for 619,524 regular shoppers and 207 million shopping baskets with anthocyanin estimates drawn from polyphenol databases. We subsequently analyse relative deprivation levels of the neighbourhoods in which shoppers reside, illustrating how anthocyanin intake varies according to affluence. Results indicate that deprivation is linked dramatically with both lower total intake of anthocyanins and lower breadth of dietary sources for them, potentially aggravating the incidence of diet-related diseases in the poorest sections of society.
Understanding and measuring the predictability of consumer purchasing (basket) behaviour is of significant value. While predictability measures such as entropy have been well studied and leveraged in other sectors, their development and application to very large multi-dimensional data sets present in the retailing sector are less common. While a small number of methods exist, we demonstrate they fail to accord with intuition, leading to the potential for misunderstandings between those who conduct the analysis and those who act on the insights. We delineate the requirements for such a measure in this domain to demonstrate these issues in context. A novel measure is then developed based on entropy to directly measure the predictability of basket composition. The measure is designated as bundle entropy (zero denotes a bundle’s total predictability, one the total unpredictability). We empirically compare the proposed bundle entropy against existing measures using two large-scale real-world transactional data sets, each including more than 2,000 households (frequent shoppers) over two years. First, we demonstrate how the proposed measure is the only measure that behaves according to the desired properties. Second, we show empirically that bundle entropy differs noticeably from the other measures. Finally, we consider some use case analyses and discuss the utility of the proposed measure in practice.
Variable Importance (VI) has traditionally been cast as the process of estimating each variable's contribution to a predictive model's overall performance. Analysis of a single model instance, however, guarantees no insight into a variables relevance to underlying generative processes. Recent research has sought to address this concern via analysis of Rashomon sets - sets of alternative model instances that exhibit equivalent predictive performance to some reference model, but which take different functional forms. Measures such as Model Class Reliance (MCR) have been proposed, that are computed against Rashomon sets, in order to ascertain how much a variable must be relied on to make robust predictions, or whether alternatives exist. If MCR range is tight, we have no choice but to use a variable; if range is high then there exists competing, perhaps fairer models, that provide alternative explanations of the phenomena being examined. Applications are wide, from enabling construction of 'fairer' models in areas such as recidivism to health analytics and ethical marketing. Tractable estimation of MCR for non-linear models is currently restricted to Kernel Regression under squared loss [7]. In this paper we introduce a new technique that extends computation of Model Class Reliance (MCR) to Random Forest classifiers and regressors. The proposed approach addresses a number of open research questions, and in contrast to prior Kernel SVM MCR estimation, runs in linearithmic rather than polynomial time. Taking a fundamentally different approach to previous work, we provide a solution for this important model class, identifying situations where irrelevant covariates do not improve predictions.
Continuous monitoring of ventilatory parameters such as tidal volume (TV) and minute ventilation (MV) has shown to be effective in the prevention of respiratory compromise events in hospitalized patients. However, the non-invasive estimation of respiratory volume in non-intubated patients remains an outstanding challenge. In this work, we present a novel approach to respiratory volume monitoring (RVM) that continuously predicts TV and MV in normal subjects. Respiratory flow in 19 volunteers under spontaneous breathing was recorded using respiratory inductance plethysmography and a temperature-based wearable sensor. Temperature signals were processed to identify features such as temperature amplitude and mean value, among others. The feature datasets were then used to train and validate three machine-learning (ML) algorithms for the prediction of respiratory volume based on temperature-related features. A model based on Random-Forest regression resulted in the lowest root mean-square error and was subsequently chosen to predict ventilatory parameters on subject test data not used in the construction of the model. Our predictions achieve a bias (mean error) in TV and MV of 16.04 mL and 0.19 L/min, respectively, which compare well with performance metrics reported in commercially-available RVM systems based on electrical impedance. Our results show that the combination of novel respiratory temperature sensors and machine-learning algorithms can deliver accurate and continuous estimates of TV and MV in healthy subjects.