Basal Cell Carcinoma (BCC) is the most common type of skin cancer, accounting for nearly 80
The BETTER4U project (Preventing Obesity through Biologically and bEhaviorally Tailored inTERventions for You) is a Horizon Europe initiative (GAP 101080117) dedicated to advancing our understanding of the multifaceted etiology of obesity. The project aims to move beyond current knowledge of obesity-related determinants by examining the system-level interactions between biological (including genetics), lifestyle behaviors (physical activity, nutrition, sedentary behaviors), and contextual factors, including social, economic, psychological, and environmental factors. BETTER4U advances from traditional approaches by exploring determinants embedded within interconnected systems and biological frameworks. It seeks to identify and integrate polygenic risks, omics-based markers, lifestyle and contextual factors using data from biobanks and previous landmark studies. This comprehensive approach enables the refinement of Artificial Intelligence (AI) algorithms, enhancing individualization in the design of tailored interventions. The project employs real-time behavioral monitoring tools and remote technologies to track individual behaviors and metabolic responses, allowing for the development of personalized obesity prevention strategies. By doing so, BETTER4U aims to bridge the gap between research and practical application, enabling the design of interventions that are biologically and behaviorally informed. The anticipated outcomes of BETTER4U include advanced AI models that support the development of sophisticated, targeted interventions for individuals at risk of overweight or obesity. These models will provide actionable insights into how personalized lifestyle adjustments in lifestyle behaviors, such as diet and physical activity, can effectively mitigate weight gain trajectories across the lifespan. Ultimately, the project seeks to empower individuals and support precision medicine approaches in obesity prevention and treatment through active participant engagement.
Advances in IoT technologies combined with new algorithms have enabled the collection and processing of high-rate multi-source data streams that quantify human behavior in a fine-grained level and can lead to deeper insights on individual behaviors as well as on the interplay between behaviors and the environment. In this paper, we present an integrated system that collects and extracts multiple behavioral and environmental indicators, aiming at improving public health policies for tackling obesity. Data collection takes place using passive methods based on smartphone and smartwatch applications that require minimal interaction with the user. Our goal is to present a detailed account of the design principles, the implementation processes, and the evaluation of integrated algorithms, especially given the challenges we faced, in particular (a) integrating multiple technologies, algorithms, and components under a single, unified system, and (b) large scale (big data) requirements. We also present evaluation results of the algorithms on datasets (public for most cases) such as an absolute error of 8-9 steps when counting steps, 0.86 F1-score for detecting visited locations, and an error of less than 12 mins for gross sleep time. Finally, we also briefly present studies that have been materialized using our system, thus demonstrating its potential value to public authorities and individual researchers.
Parkinson’s disease (PD) is a long-term neurode-generative disorder characterized by motor and non-motor symptoms, with tremor being a key indicator. Early detection of PD is vital for effective symptom management, yet current methods often struggle in real-world scenarios. This study addresses Parkinsonian tremor detection using smartphone-captured accelerometer data collected in-the-wild. We propose a novel approach that integrates contrastive pretraining within a pre-existing Multiple-Instance Learning (MIL) framework to leverage large-scale unlabeled data and enhance representation learning. Additionally, we extend the method to a federated learning framework for scalable and privacy-preserving deployment. Results confirm that the pretrained MIL model surpasses the baseline, with improved performance as the unlabeled dataset size increases. Furthermore, federated pretraining achieves comparable results to centralized training while maintaining privacy and addressing non-Independent and Identically Distributed (non-IID) data challenges.
Accurate monitoring of eating behavior is crucial for managing obesity and eating disorders such as bulimia nervosa. At the same time, existing methods rely on multiple and/or specialized sensors, greatly harming adherence and ultimately, the quality and continuity of data. This paper introduces a novel approach for estimating the weight of a bite, from a commercial smartwatch. Our publicly-available dataset contains smartwatch inertial data from ten participants, with manually annotated start and end times of each bite along with their corresponding weights from a smart scale, under semi-controlled conditions. The proposed method combines extracted behavioral features such as the time required to load the utensil with food, with statistical features of inertial signals, that serve as input to a Support Vector Regression model to estimate bite weights. Under a leave-one-subject-out cross-validation scheme, our approach achieves a mean absolute error (MAE) of 3.99 grams per bite. To contextualize this performance, we introduce the improvement metric, that measures the relative MAE difference compared to a baseline model. Our method demonstrates a 17.41% improvement, while the adapted state-of-the art method shows a -28.89% performance against that same baseline. The results presented in this work establish the feasibility of extracting meaningful bite weight estimates from commercial smartwatch inertial sensors alone, laying the groundwork for future accessible, non-invasive dietary monitoring systems.
In this paper, we propose a weakly supervised semantic segmentation approach for food images which takes advantage of the zero-shot capabilities and promptability of the Segment Anything Model (SAM) along with the attention mechanisms of Vision Transformers (ViTs). Specifically, we use class activation maps (CAMs) from ViTs to generate prompts for SAM, resulting in masks suitable for food image segmentation. The ViT model, a Swin Transformer, is trained exclusively using image-level annotations, eliminating the need for pixel-level annotations during training. Additionally, to enhance the quality of the SAM-generated masks, we examine the use of image preprocessing techniques in combination with single-mask and multi-mask SAM generation strategies. The methodology is evaluated on the FoodSeg103 dataset, generating an average of 2.4 masks per image (excluding background), and achieving an mIoU of 0.54 for the multi-mask scenario. We envision the proposed approach as a tool to accelerate food image annotation tasks or as an integrated component in food and nutrition tracking applications.
Prediabetes is a common health condition that often goes undetected until it progresses to type 2 diabetes. Early identification of prediabetes is essential for timely intervention and prevention of complications. This research explores the feasibility of using wearable continuous glucose monitoring along with smartwatches with embedded inertial sensors to collect glucose measurements and acceleration signals respectively, for the early detection of prediabetes. We propose a methodology based on signal processing and machine learning techniques. Two feature sets are extracted from the collected signals, based both on a dynamic modeling of the human glucose-homeostasis system and on the Glucose curve, inspired by three major glucose related blood tests. Features are aggregated per individual using bootstrap. Support Vector Machines are used to classify normoglycemic vs. prediabetic individuals. We collected data from 22 participants for evaluation. The results are highly encouraging, demonstrating high sensitivity and precision. This work is a proof of concept, highlighting the potential of wearable devices in prediabetes assessment. Future directions involve expanding the study to a larger, more diverse population and exploring the integration of CGM and smartwatch functionalities into a unified device. Automated eating detecting algorithms can also be used.
Transportation mode recognition (TMR) is a critical component of human activity recognition (HAR) that focuses on understanding and identifying how people move within transportation systems. It is commonly based on leveraging inertial, location, or both types of signals, captured by modern smartphone devices. Each type has benefits (such as increased effectiveness) and drawbacks (such as increased battery consumption) depending on the transportation mode (TM). Combining the two types is challenging as they exhibit significant differences such as very different sampling rates. This paper focuses on the TMR task and proposes an approach for combining the two types of signals in an effective and robust classifier. Our network includes two sub-networks for processing acceleration and location signals separately, using different window sizes for each signal. The two sub-networks are designed to also embed the two types of signals into the same space so that we can then apply an attention-based multiple-instance learning classifier to recognize TM. We use very low sampling rates for both signal types to reduce battery consumption. We evaluate the proposed methodology on a publicly available dataset and compare against other well known algorithms.
Monitoring aquatic vegetation, including both floating and emergent types, plays a crucial role in understanding the dynamics of freshwater ecosystems. Our research focused on the Lower Dniester Basin in Southern Ukraine, covering approximately 1800 square kilometers of steppe plains and wetlands. We applied traditional machine learning algorithms, specifically random forest and boosting trees, to analyze Sentinel-2 satellite imagery for segmenting aquatic vegetation into emergent and floating types. Our methodology was validated against detailed in-situ field measurements collected annually over a 5-year study period. The machine learning classifiers achieved an F1-score of 0.88 ± 0.03 in classifying floating vegetation, outperforming our previously suggested histogram-based thresholding methodology for the same task. While emergent vegetation and open water were easily identifiable from satellite imagery, the robustness and temporal transferability of our methodology included accurately delineating floating vegetation as well. Additionally, we explored the significance of various features through the Minimum Redundancy - Maximum Relevance algorithm. This study highlights advancements in aquatic vegetation mapping and demonstrates a valuable tool for ecological monitoring and future research endeavors.
Data-driven approaches for remote detection of Parkinson's Disease and its motor symptoms have proliferated in recent years, owing to the potential clinical benefits of early diagnosis. The holy grail of such approaches is the free-living scenario, in which data are collected continuously and unobtrusively during every day life. However, obtaining fine-grained ground-truth and remaining unobtrusive is a contradiction and therefore, the problem is usually addressed via multiple-instance learning. Yet for large scale studies, obtaining even the necessary coarse ground-truth is not trivial, as a complete neurological evaluation is required. In contrast, large scale collection of data without any ground-truth is much easier. Nevertheless, utilizing unlabelled data in a multiple-instance setting is not straightforward, as the topic has received very little research attention. Here we try to fill this gap by introducing a new method for combining semi-supervised with multiple-instance learning. Our approach builds on the Virtual Adversarial Training principle, a state-of-the-art approach for regular semi-supervised learning, which we adapt and modify appropriately for the multiple-instance setting. We first establish the validity of the proposed approach through proof-of-concept experiments on synthetic problems generated from two well-known benchmark datasets. We then move on to the actual task of detecting PD tremor from hand acceleration signals collected in-thewild, but in the presence of additional completely unlabelled data. We show that by leveraging the unlabelled data of 454 subjects we can achieve large performance gains (up to 9% increase in F1-score) in per-subject tremor detection for a cohort of 45 subjects with known tremor ground-truth. In doing so, we confirm the validity of our approach on a real-world problem where the need for semi-supervised and multiple-instance learning arises naturally.
Automatic dietary monitoring has progressed significantly during the last years, offering a variety of solutions, both in terms of sensors and algorithms as well as in terms of what aspect or parameters of eating behavior are measured and monitored. Automatic detection of eating based on chewing sounds has been studied extensively, however, it requires a microphone to be mounted on the subject's head for capturing the relevant sounds. In this work, we evaluate the feasibility of using an off-the-shelf commercial device, the Razer Anzu smart-glasses, for automatic chewing detection. The smart-glasses are equipped with stereo speakers and microphones that communicate with smart-phones via Bluetooth. The microphone placement is not optimal for capturing chewing sounds, however, we find that it does not significantly affect the detection effectiveness. We apply an algorithm from literature with some adjustments on a challenging dataset that we have collected in house. Leave-one-subject-out experiments yield promising results, with an F1-score of $0.96$ for the best case of duration-based evaluation of eating time.
The relation among the various causal factors of obesity is not well understood, and there remains a lack of viable data to advance integrated, systems models of its etiology. The collection of big data has begun to allow the exploration of causal associations between behavior, built environment, and obesity-relevant health outcomes. Here, the traditional epidemiologic and emerging big data approaches used in obesity research are compared, describing the research questions, needs, and outcomes of 3 broad research domains: eating behavior, social food environments, and the built environment. Taking tangible steps at the intersection of these domains, the recent European Union project "BigO: Big data against childhood obesity" used a mobile health tool to link objective measurements of health, physical activity, and the built environment. BigO provided learning on the limitations of big data, such as privacy concerns, study sampling, and the balancing of epidemiologic domain expertise with the required technical expertise. Adopting big data approaches will facilitate the exploitation of data concerning obesity-relevant behaviors of a greater variety, which are also processed at speed, facilitated by mobile-based data collection and monitoring systems, citizen science, and artificial intelligence. These approaches will allow the field to expand from causal inference to more complex, systems-level predictive models, stimulating ambitious and effective policy interventions.
1. Kyritsis K, Diou C, Delopoulos A. A Data Driven End-to-End Approach for In-the-Wild Monitoring of Eating Behavior Using Smartwatches. IEEE J Biomed Health Inform. 2021;25(1):22-34. doi:10.1109/JBHI.2020.2984907 CrossRef Google Scholar
Heart murmurs are abnormal sounds present in heartbeats, caused by turbulent blood flow through the heart. The PhysioNet 2022 challenge targets automatic detection of murmur from audio recordings of the heart and automatic detection of normal vs. abnormal clinical outcome. The recordings are captured from multiple locations around the heart. Our participation investigates the effectiveness of self-supervised learning for murmur detection. We train the layers of a backbone CNN in a self-supervised way with data from both this year's and the 2016 challenge. We use two different augmentations on each training sample, and normalized temperature-scaled cross-entropy loss. We experiment with different augmentations to learn effective phonocardiogram representations. To build the final detectors we train two classification heads, one for each challenge task. We present evaluation results for all combinations of the available augmentations, and for our multiple-augmentation approach. Our team's, Listen2YourHeart, SSL murmur detection classifier received a weighted accuracy score of 0.737 (ranked 13th out of 40 teams) and an outcome identification challenge cost score of 11946 (ranked 7th out of 39 teams) on the hidden test set.
Global-scale canopy height mapping is an important tool for ecosystem monitoring and sustainable forest management. Various studies have demonstrated the ability to estimate canopy height from a single spaceborne multispectral image using end-to-end learning techniques. In addition to texture information of a single-shot image, our study exploits multitemporal information of image sequences to improve estimation accuracy. We adopt a convolutional variant of a long short-term memory (LSTM) model for canopy height estimation from multitemporal instances of Sentinel-2 products. Furthermore, we utilize the deep ensembles technique for meaningful uncertainty estimation on the predictions and postprocessing isotonic regression model for calibrating them. Our lightweight model ( ${\sim }320{\mathrm {k}}$ trainable parameters) achieves the mean absolute error (MAE) of $1.29 ~{\mathrm {m}}$ in a European test area of $79 ~{\mathrm {km}}^{2}$ . It outperforms the state-of-the-art methods based on single-shot spaceborne images as well as costly airborne images while providing additional confidence maps that are shown to be well calibrated. Moreover, the trained model is shown to be transferable in a different country of Europe using a fine-tuning area of as low as ${\sim }2 ~{\mathrm {km}}^{2}$ with ${\mathrm {MAE}}=1.94 {\mathrm {m}}$ .
While automatic tracking and measuring of our physical activity is a well established domain, not only in research but also in commercial products and every-day lifestyle, automatic measurement of eating behavior is significantly more limited. Despite the abundance of methods and algorithms that are available in bibliography, commercial solutions are mostly limited to digital logging applications for smart-phones. One factor that limits the adoption of such solutions is that they usually require specialized hardware or sensors. Based on this, we evaluate the potential for estimating the weight of consumed food (per bite) based only on the audio signal that is captured by commercial ear buds (Samsung Galaxy Buds). Specifically, we examine a combination of features (both audio and non-audio features) and trainable estimators (linear regression, support vector regression, and neural-network based estimators) and evaluate on an in-house dataset of 8 participants and 4 food types. Results indicate good potential for this approach: our best results yield mean absolute error of less than 1 g for 3 out of 4 food types when training food-specific models, and 2.1 g when training on all food types together, both of which improve over an existing literature approach.
Parkinson’s disease (PD) is a neurodegenerative disorder with both motor and non-motor symptoms. Despite the progressive nature of PD, early diagnosis, tracking the disease’s natural history and measuring the drug response are factors that play a major role in determining the quality of life of the affected individual. Apart from the common motor symptoms, i.e., tremor at rest, rigidity and bradykinesia, studies suggest that PD is associated with disturbances in eating behavior and energy intake. Specifically, PD is associated with drug-induced impulsive eating disorders such as binge eating, appetite-related non-motor issues such as weight loss and/or gain as well as dysphagia—factors that correlate with difficulties in completing day-to-day eating-related tasks. In this work we introduce Plate-to-Mouth (PtM), an indicator that relates with the time spent for the hand operating the utensil to transfer a quantity of food from the plate into the mouth during the course of a meal. We propose a two-step approach towards the objective calculation of PtM. Initially, we use the 3D acceleration and orientation velocity signals from an off-the-shelf smartwatch to detect the bite moments and upwards wrist micromovements that occur during a meal session. Afterwards, we process the upwards hand micromovements that appear prior to every detected bite during the meal in order to estimate the bite’s PtM duration. Finally, we use a density-based scheme to estimate the PtM durations distribution and form the in-meal eating behavior profile of the subject. In the results section, we provide validation for every step of the process independently, as well as showcase our findings using a total of three datasets, one collected in a controlled clinical setting using standardized meals (with a total of 28 meal sessions from 7 Healthy Controls (HC) and 21 PD patients) and two collected in-the-wild under free living conditions (37 meals from 4 HC/10 PD patients and 629 meals from 3 HC/3 PD patients, respectively). Experimental results reveal an Area Under the Curve (AUC) of 0.748 for the clinical dataset and 0.775/1.000 for the in-the-wild datasets towards the classification of in-meal eating behavior profiles to the PD or HC group. This is the first work that attempts to use wearable Inertial Measurement Unit (IMU) sensor data, collected both in clinical and in-the-wild settings, towards the extraction of an objective eating behavior indicator for PD.