
Many intelligent systems are optimized for the average user. With personal cognitive systems, we envision AI systems that are highly personalized and that continuously adapt to a specific individual, reflecting an individual’s abilities, preferences, and emotions. In the context of the ERC project, "AI-Twins of Human Experience," we are investigating how to move beyond generic, universal AI tools to personalized digital cognitive systems that continuously learn from users’ experiences, interactions, and reactions. Our goal is to lay the foundation for digital systems that facilitate human-AI partnerships, amplifying cognition, creativity, and well-being. Physiology-driven learning and adaptation are key to realizing this vision.
Accurate prediction of future trajectories in unstructured environments is essential for safe autonomous vehicle operation, especially when pedestrians, cyclists, and motor vehicles interact in complex and unpredictable ways. Many existing prediction models predict based on a single class or rely on multiple separate graphs for multi-class prediction. This work introduces a lightweight trajectory prediction framework that represents all agents within a single Comprehensive Heterogeneous Graph, enriched with a directional interaction mask that highlights behaviourally important neighbours based on relative distance, velocity, and motion direction. A spatial-temporal design combining a Spatial-Temporal Graph Convolutional Network with a micro Temporal Convolutional Network allows efficient modeling of motion dynamics, while a Gaussian mixture model decoder enables probabilistic training and deterministic maximum likelihood prediction for simulation. Experiments on the Stanford Drone Dataset show improved accuracy over recent multi-class baselines, and qualitative visualizations confirm stable, socially coherent, and scene-consistent predictions suitable for real-time autonomous systems.
Digital Twins (DTs) have emerged as a critical technology for enhancing the management and optimization of buildings, enabling real-time monitoring, predictive maintenance, and energy efficiency through the integration of Internet of Things (IoT) sensors and advanced data analytics. In the context of smart buildings, in particular, DTs are essential in bringing the best architectural practices for real-time processing of sensor data, building information models, and supporting efficient operation, energy management, and occupant comfort. This paper critically examines recent advancements in applications of sensor-driven DTs for smart buildings, and highlights the main aspects of their technical solutions, including types of sensors, use-case problems, visualization approaches, as well as continuously present challenges in these applications.
Fully decentralised AI challenges the long-standing assumption that intelligence must reside in a centralised infrastructure. By allowing data to remain at its source and enabling models to evolve through peer-to-peer interactions, decentralised learning opens the door to AI systems that are privacy-preserving, sovereign, and deeply embedded in their operational environments. However, removing the center also reshapes the problem space. How does collective intelligence arise from the structure of the underlying communication graph? What guarantees of convergence and robustness can we provide in the presence of unreliable links, heterogeneous data, or even adversarial behavior? And how can meaningful knowledge emerge when local signals are noisy or imperfect? This keynote examines decentralisation not merely as a technical design choice, but as a conceptual shift. By embracing network structure, resilience, and collaboration as first-class elements, we can move toward AI systems that behave less like monolithic engines and more like adaptive, distributed ecosystems.
Digital twins are a growing field of research and are increasingly implemented as a virtual copy of physical systems. This paper presents an Urban Mobility Digital Twin, developed within the European AVANT project (IPCEI-CIS), to support citizens and city administrators in smart mobility planning. The platform integrates traffic forecasting, accident risk estimation, and interactive simulation through a map-based interface. The main system capabilities are made possible by two machine learning models and a simulator engine, that have been built or adapted for this use case and for which performance metrics are provided. This paper presents the results achieved with the proposed models and aims to serve as a reference for future implementations of urban mobility digital twins.
Deploying human pose estimation (HPE) pipelines in privacy-aware settings requires processing sensitive data close to the point of acquisition, as set out in regulations such as the AI Act. Edge devices can meet this requirement, but their limited resources modify the impact of privacy mechanisms, hardware capacity, and scene complexity on latency. This paper presents a framework that jointly analyzes latency and privacy in HPE pipelines executed on edge architectures. We consider two representative scenarios: i) all the pipeline stages on the acquisition device; ii) acquisition and obfuscation on the camera, with inference on a separate edge node. We present the framework by evaluating each scenario with deployments based on three state-of-the-art HPE models (MoveNet, YOLO-Pose, and OpenPose), three representative hardware configurations, and three levels of scene complexity based on the number of subjects and visual entropy. For each deployment, the framework estimates stage-level and end-to-end latency distributions, obtains the probability of completing the pipeline within task-specific deadlines, and the attainable privacy level. Our framework also generates reliability-privacy-cost maps, a decision support tool that relates improvements in timing and privacy to the cost of implementing or upgrading a deployment. A surveillance case study illustrates how the framework helps design or adapt privacy-preserving edge infrastructures.
This work explores whether the public log layer of an EVM compatible ledger can act as a minimal and globally reachable transport for encrypted communication. We present Verbeth, a proof of concept protocol that uses ordinary wallet keys and per message forward secrecy to send ciphertext directly through transaction logs. The design requires no auxiliary servers and offers strong censorship resistance and guaranteed final delivery, while revealing the economic and cryptographic limits of on chain messaging. We also compare it with established messaging protocols to clarify where its guarantees align and where they diverge. Looking ahead, the protocol can absorb hybrid or post quantum designs that remove harvest now decrypt later concerns and place its security profile on par with other messengers.
Non-invasive and widely accessible methods for Diabetes Mellitus (DM) screening and monitoring are crucial to address the limitations of conventional blood-based assays. Exhaled breath (EB) analysis provides a contactless, real-time alternative by detecting disease-associated changes in Volatile Organic Compounds (VOCs). This PhD research work introduces a compact Non-dispersive Infrared (NDIR)-based optical Electronic Nose (E-nose) for rapid VOCs detection. Preliminary laboratory tests using a general-purpose IR sensor and a MEMS IR emitter showed feasibility for detecting low VOCs concentrations. Expanding the system with nonspecific near-/mid-IR, MOS, and ambient sensors could further enhance performance, enabling AI-driven multi-modal strategies for DM-oriented breath analysis.
Post-stroke rehabilitation is a critical phase for stroke survivors constituting majority of recovery period. Gamification is a common approach taken towards rehabilitation enhancing patient engagement. This involves using customized and sophisticated sensing setups. In lower-middle income countries (LMICs), acquiring such setups is a challenge, exacerbated by financial constraints post acute-care period and a very low number of therapy centers within practical reach of patients. We are building an accessible & affordable rehabilitation suite targeting such population groups, with an integrated gaming suite as one of the offerings. In this demonstration, we present a 3D game aiming motor function of fingers, which is a critical upper body function towards an independent living. With Google’s Mediapipe running on a smartphone, hand landmarks are detected in real-time and transmitted to the device running the game, which can be a laptop/PC or another smartphone. This makes our solution accessible and can be used by the patients at home in a ubiquitous way without any special setup. Clinically relevant indexes are calculated and shown on screen in an intuitive manner providing sufficient performance insights to both patient and the therapist. The gaming suite is poised for clinical trials along with other assessment modules of our rehabilitation suite, which includes lower-body assessments as well.
This paper presents an Artificial Intelligence (AI)-powered Wi-Fi channel state information human sensing system for presence detection via breathing pattern recognition through a three-layer hierarchical architecture. The primary goal is to detect human presence near the device, using breathing detection as a discriminative feature rather than estimating breathing rate values. The system combines a two-stage pipeline (activity detection pre-filtering followed by breathing detection) with temporal user state tracking and semantic classification. Based on three features (Breathing-to-Noise Ratio, Subcarrier Agreement, and Nonlinearity), our activity detection model achieves 88.5% F1-score on test data, and the breathing detection model 96.47% F1-score. Operating at 20 Hz sampling rate with 8 sec processing windows, the system enables intelligent presence detection while maintaining cross-dataset robustness. Extensive validation across independent datasets demonstrates consistent high accuracy, making this approach suitable for realistic deployment in human-computer interaction and wellness monitoring applications.
This artifact releases the annotated dataset used in our study of object-level privacy risks in home interior images. It contains 279 images with manually highlighted objects, including participant-derived sensitivity labels, object categories, and privacy information-type classes. The dataset supports privacy inference analysis via Vision-Language Models (VLMs) by enabling comparative and extended studies of sensitive visual cues in domestic environments.
In this paper, we propose DSSP, an SLO-aware, multi-Domain-based Sensing-as-a-Service (SaS) Platform. DSSP enables a marketplace in which institutions/organizations who own and operate Independent Administration Domains (IADs) for sensing sell the sensory data collected from their respective IADs to Cloud Service Providers (CSPs), who in turn, make use of the sensory data to enable query tail-latency Service-Level-Objective (SLO) guaranteed Sensing-as-a-Service (SaS) at scale, while preserving data privacy and autonomy of control for individual IADs. At the core of DSSP is the design of a budget decomposition technique that translates: (a) a query tail-latency SLO into exact task response time budgets for sensing tasks of the query dispatched to individual IADs; and (b) the task budget for a task arrived at an IAD into exact subtask queuing deadlines for subtasks of the task dispatched to individual edge nodes in each IAD. This enables IADs to allocate their internal resources independently and accurately to meet the task budgets and hence, query tail-latency SLO. The performance and viability of DSSP is evaluated and verified by experiments in an on-campus testbed.
Reinforcement Learning (RL) is a method for learning policies to execute sequential tasks, and it is gaining popularity in several domains, amongst which are pervasive systems and robotics. There, learning optimal control policies goes along with another desiderata: robustness to contextual uncertainties in environment sensing and actuation. Causal reasoning can help by leveraging cause-and-effect relationships between sensors’ readings and actuators’ actions to rule out spurious correlations due to noise, measurement errors, and the like. In this paper, we propose a specific instantiation of causal RL to explicitly learn and model the causal value function of actions’ in pervasive and robotic tasks. We test our proposed framework across the robotic tasks provided by the MuJoCo suite, ranging from object manipulation to locomotion, in improving, specifically, value function estimation. We open-source our code at: https://github.com/Giovannibriglia/benchmarking_causal_rl.
Wi-Fi-based human activity recognition is promising but often limited by costly retraining and poor scalability. This work introduces a consensus-based framework for distributed CSI sensing, enabling robust and scalable activity recognition with minimal training and communication overhead. Transmitter–receiver pairs are ranked based on their short training momentum and allocated by a central coordinator (e.g., a router) to monitor specific locations. Experiments across three locations and twelve participants show that a location-aware consensus approach matches optimally placed solutions (F1 =0.98) while improving robustness and temporal stability during dynamic activity flows.
Maintaining low Age of Information (AoI) is essential for connected and autonomous vehicles, yet the Semi-Persistent Scheduling (SPS) procedure standardized for 5G NR-V2X sidelink employs a fixed sensing window that cannot adapt to changes in traffic load or vehicle density. This limitation increases resource-selection conflicts and degrades the timeliness required for cooperative perception and cooperative driving. This paper introduces a lightweight, fully distributed congestion-control mechanism that incorporates real-time congestion information into the SPS sensing process. By deriving statistics from neighboring transmission periodicity, the method dynamically adjusts the sensing-window span to reduce collisions and prevent excessive channel occupation. The approach is evaluated in a system-level framework combining microscopic mobility with a detailed wireless channel and SPS access model. Across diverse densities, mobility conditions, and communication ranges, the adaptive sensing-window mechanism consistently improves both mean and tail AoI and enhances the stability of the Resource Reservation Interval. It achieves substantial reductions in information staleness, including significant gains for short-range, safety-critical exchanges. Because it solely refines the sensing-window computation without modifying SPS signaling, the method offers a practical and readily deployable enhancement to NR-V2X Mode 2 congestion control.
This study uses depth-sensor technology to characterize visitors’ spatial and social behavior, including engagement in museum environments. Museums are complex spaces that serve as centers of learning and social interaction, attracting diverse populations. Traditional methods of studying visitor interaction with exhibits are often intrusive and costly, requiring specially trained personnel to analyze data. To address these challenges, we propose a system to automatically analyze skeleton data collected from an on-site deployed solution using a set of Kinect v2 sensors across exhibition spaces at the Musikinstrumenten-Museum in Berlin. The system captures detailed information on visitor-exhibition interaction through body joints as skeleton data, providing insights into how designated areas and object placement influence engagement. By using raw skeleton data to derive high-level features, such as dwell time, proximity, pointing gestures, and qualitative observations of group dynamics, we compute spatial and behavioral metrics of visitors while avoiding the storage of identifiable audio or video data and aiming to preserve visitor privacy. The depth camera system enables tracking and displaying visitor interactions, offering a non-intrusive, interactive way to measure attention and engagement. With this study, we demonstrate the feasibility of skeleton-based sensing for automated unobtrusive assessment of human behavior indoors.
Human Activity Recognition (HAR) underpins applications in healthcare, rehabilitation, fitness tracking, and smart environments, yet many existing approaches remain difficult to integrate into real applications due to their training requirements and computational overhead. This paper presents RAG-HAR, a training-free, retrieval-augmented HAR framework based on large language models (LLMs), implemented by the authors as a lightweight widget within an Android fitness application developed using the Flutter framework and Android Studio. The demonstration highlights the end-to-end integration of the RAG-HAR pipeline within a user-facing application, illustrating its operation under continuous sensor input and its suitability for practical health and fitness scenarios.
Atrial fibrillation (AF) is the most common cardiac arrhythmia and is associated with an increased risk of stroke and heart failure, making early detection critical for effective cardiac management. This paper presents a comparative and explainable AF detection framework. We use the state-of-the-art handcrafted features to train Kolmogorov– Arnold Networks (KAN), a deep neural network (DNN), and an AdaBoost-based shallow ensemble. The core contribution of this work is a multi-view explainability analysis combining internal KAN attribution, SHapley Additive exPlanations (SHAP) for neural models, and impurity-based feature importance from the shallow model. We propose a correlation-weighted edge score (CWES) refinement of the KAN attribution that improves agreement with a consensus ranking (CR). Five-fold cross validation of CWES and CR on 15 top ranked features, on PhysioNet 2017 single-lead ECG dataset, reveals 82.5% F1-score which is an improvement over basic KAN F1-score of 80.6%. Finally, we demonstrate and analyze the medical explainability of the top ranked features.
Mobility systems are evolving into large-scale cyberphysical infrastructures that continuously sense, learn from, and act upon human movement. From indoor positioning and activity recognition to ride sharing and shared urban services, these systems increasingly shape how people interact with buildings, campuses, and cities. As intelligence becomes pervasive, trust in sensing, learning, and data use emerges as the central challenge for sustainable deployment. This talk presents a vision for trustworthy cyber-physical mobility systems, arguing that trust must be designed holistically across the entire pipeline rather than addressed at isolated stages. We begin at the sensing layer, highlighting privacy conscious modalities such as LiDAR and wireless sensing, which enable rich spatial and behavioral understanding without directly capturing identifiable visual information. This reflects a broader shift toward responsible sensing as a foundational design principle. We then examine learning and collaboration in distributed mobility ecosystems, where federated learning and GenAI enable collective intelligence while keeping raw data at the source. While promising, such approaches are not inherently safe, as threats such as membership inference attacks expose subtle privacy leakages even in decentralized settings. These challenges reveal a fundamental utility–privacy trade-off that future mobility systems must explicitly manage. Finally, the talk highlights transparency and visualization as key enablers for making cyber-physical intelligence observable and interpretable, outlining a path toward mobility systems that are intelligent, resilient, and worthy of long-term trust.