
The cooperative tracking of Multi-Autonomous Underwater vehicles (AUV) has shown great potential in fields such as ocean environmental monitoring, marine resource exploration, and underwater security. However, limited underwater acoustic communication range, sparse deployment of sensor nodes, and environmental uncertainties often lead to incomplete target trajectory information and partially observable states, resulting in decreased tracking accuracy and unstable performance in complex scenarios. To solve the above problems, this paper proposes a Hierarchical Deep Reinforcement Learning framework, termed ASF-HDRL (Attention-guided high-level Switching with stage-conditioned Feature modulation), which decouples high- and low-level policies and incorporates a target estimation algorithm to enhance adaptability and robustness in complex tasks. Simulation results demonstrate that ASF-HDRL achieves superior cooperative tracking performance in challenging simulation environments, outperforming several mainstream baseline methods in terms of convergence speed and tracking accuracy.
Depression is a prevalent mental health disorder with significant consequences, making its early and accurate detection essential. Traditional diagnostic methods relying on self-report questionnaires are subjective, underscoring the need for objective, automated approaches. However, unimodal models often fail to capture the full complexity of depressive symptoms. To address this, we propose VTA-DepressNet, a multimodal deep learning architecture that integrates visual, audio, and textual features through an attention-driven fusion mechanism. The model utilizes a Conv-BiLSTM for visual features, a Conv-BiGRU for audio features, and a Transformer encoder with an attention mechanism for textual data. Experimental results on the DAIC-WOZ dataset with fivefold cross-validation demonstrate that VTA-DepressNet achieves an F1-score of 0.83, significantly outperforming unimodal baselines. To validate generalization, the model was further evaluated on the Extended DAIC (E-DAIC) dataset, achieving a competitive F1-score of 0.77. These results not only affirm the effectiveness of combining behavioral and linguistic cues but also highlight the potential of digital mental-health tools for early screening and intervention in both clinical and educational contexts. In practice, VTA-DepressNet could be integrated into telehealth platforms or school-based mental health monitoring systems to support timely and scalable depression assessment.
Emotion fluctuations during music listening are closely tied to a listener’s current emotional state, yet most recommendation systems rely on historical behavior data, which struggle with cold-start issues and real-time adaptability. In this work, we propose IMJP-Net, a framework that leverages smart glasses IMU signals as a privacy-preserving implicit feedback channel for emotion-aware music recommendation. To resolve the non-stationarity of head-worn dynamics, the system integrates an IMU Evolution Module (IEM) for temporal inertial encoding and an IMU-Music Mutual Attentive Alignment (IMMAA) mechanism for cross-modal feature fusion. Evaluated on 1164 valid trials ( N=30 ), IMJP-Net achieves 19.6 N=10 ) validates practical utility, with participants highlighting enhanced emotional resonance, as well as privacy and non-intrusiveness. These results demonstrate the potential of IMJP-Net for pervasive emotion-aware music recommendation.
The widespread adoption of mobile and wearable devices has made human activity recognition (HAR) a key component of pervasive computing, enabling applications in smart healthcare, ambient assisted living, and smart cities. However, centralized HAR models face privacy concerns and performance degradation due to data heterogeneity arising from varying user behaviors, sensor placements, device types, and environmental conditions. While federated learning (FL) offers a privacy-preserving alternative, it suffers from client drift under non-IID data distributions, limiting both personalization and generalization. In this work, we propose FedAli-PHAR, a novel personalized federated learning framework that extends the Alignment with Prototypes (ALP) layer using optimal transport (Sinkhorn-Knopp algorithm) to align embeddings with learnable local and global prototypes. The proposed approach mitigates client drift through distribution-level alignment, incorporates adaptive feature fusion via a Gated Linear Unit (GLU), and employs exponential moving average (EMA) updates for stable prototype evolution in time-series sensor data. During inference, local prototypes enable efficient on-device personalization without additional communication. Extensive experiments on heterogeneous HAR benchmarks (HHAR, RealWorld, and UCI-HAR under realistic non-IID partitioning) demonstrate that FedAli-PHAR consistently outperforms state-of-the-art personalized FL methods (FedAvg, FedProx, FedProto, MOON, and FedAli), achieving statistically significant improvements of 2
Traditional waste management systems often fails to ensure safety and may cause health hazards for the workers. Depending on manual sorting and fixed infrastructure, systems increasingly fall short of sustainability goals. Recently, deep learning (DL) has revolutionized trash management, from garbage categorization and sorting to intelligent bin monitoring, and environmental effect assessment. This survey comprehensively reviews state-of-the-art deep learning techniques applied across the waste management lifecycle. Existing surveys mostly analyze different sorts of waste classification and lacks many practical challenges of applying DL methods to waste management. To address this research gap, we adopt a DL perspective to analyze automated waste management and segregation works and present research challenges from the implementation perspectives. Internet of Things (IoT) plays a very important role in such works. The problems of existing benchmark datasets are also discussed. The survey reports experimental results on representative benchmark datasets that are publicly available to signify the type of experimentation that could be conducted. Recent advances and hence, possible future research directions in the context of DL-based automated waste segregation have also been articulated.
Mobile manipulation is a fundamental capability in embodied intelligence robotics. The growing demand for robust and generalizable manipulation in unstructured household environments has driven rapid progress in embodied intelligence platforms. However, achieving a seamless transfer across the real-to-sim-to-real cycle faces three key challenges, including costly high-fidelity simulation scenes reconstruction, the complexity of systematic strategy evaluation in simulation, and incompatible real-world deployments. To address these challenges, we develop BestMan, a scalable and seamless real-to-sim-to-real platform that bridges the gap between the simulation and the real world, enabling effective strategy development, integration, and deployment for household mobile manipulation. Specifically, we design a novel Automated Scene Generation (ASG) module to reconstruct realistic simulations from real observations. Then, we propose a simulation-guided task formalization and skill learning architecture that supports the flexible integration and large-scale evaluations of hybrid skill strategies in simulation. Finally, to enhance the real-world scalability, we develop a Hardware-agnostic and Unified Middleware (HUM) to ensure seamless and compatible sim-to-real transfer across heterogeneous mobile manipulators for real deployments. Experimental results demonstrate the superior performance of our proposed platform in establishing standardized benchmarks and facilitating promising research in the field of mobile manipulation.
With the popularization and development of intelligent driving, driving safety is becoming increasingly important. Driver monitoring can timely detect safety hazards and provide warnings, which is an important method to improve driving safety. Since the driver's gaze is a crucial aspect of driver monitoring, determining the driver’s gaze in 3D space has become very important. This paper proposes a 3D spatial eye tracking method based on a single camera with a single light source and applies it to a driver monitoring scenario to design and develop a prototype system based on eye tracking for driver monitoring. The system uses a SVM classifier to classify the driver’s intentions and gives the driver a voice warning based on the classification results to avoid possible hazards and improve driving safety. In the experiment, the proposed 3D spatial eye tracking method achieved horizontal and vertical error accuracy of 1.69° and 1.60° respectively. Moreover, the designed driver monitoring system achieved 76
The absence of haptic feedback in stylus-based interactions with large touchscreens often leads to a suboptimal user experience. To address this, we propose a novel fine-grained vibrotactile feedback system, leveraging a comprehensive hardware-software co-design. Our core contribution is FinViPen, a wireless stylus integrating a 6-DOF attitude sensor, thin-film pressure sensors, and voice coil actuators. FinViPen captures real-time writing parameters, including normal force and pressure vector, transmitting them via Bluetooth to a PC. The PC then extracts dynamic features such as velocity, 3-axis acceleration, and pressure vector from FinViPen’s positional and writing data. These features are fed into a modified DDSP algorithm to synthesize fine-grained audio waveforms. These waveforms are wirelessly transmitted via stereo FM modules back to FinViPen, which in turn drives its voice coil actuator to deliver precise vibrotactile feedback. Objective experiments demonstrate that the vibrotactile output of FinViPen, when simulating materials like whiteboard, wood, and paper, achieves high consistency (MSE < 0.08) with established haptic texture database (HaTT) waveforms and spectra. Furthermore, a user study with 11 participants across 6 whiteboard writing scenarios revealed that our proposed technology significantly enhances users’ positional awareness and immersion during writing, directly attributable to the fine-grained vibrotactile feedback.
Facial expression reconstruction technology offers considerable potential in areas like human-computer interaction, affective computing, and virtual reality. To tackle the privacy challenges and environmental constraints inherent in traditional camera-based systems, researchers have recently introduced ear-worn devices as a viable solution. Nevertheless, these methods still demand enhancements, particularly in aspects such as design appeal and energy efficiency, to achieve broader applicability and practical use. This paper introduces a system called IMUFace. It uses inertial measurement units (IMUs) embedded in wireless earphones to detect subtle ear movements caused by facial muscle activities, allowing for covert and low-power facial reconstruction. A user study involving 12 participants was conducted, and a deep learning model named IMUTwinTrans was proposed. The results show that IMUFace can accurately predict users’ facial landmarks with a precision of 2.21 mm, using only five minutes of training data. The predicted landmarks can be utilized to reconstruct a three-dimensional facial model. IMUFace operates at a sampling rate of 30 Hz with a relatively low power consumption of 58 mW. The findings validate the feasibility of IMUFace and highlight potential directions for further research towards its practical adoption in mobile environments.
Data quality and budget are two major concerns in large-scale urban Mobile Crowdsensing (MCS) technologies. Traditional MCS research primarily measures data quality based on sensing coverage, without fully considering the importance of the data for downstream tasks. As a result, in Sparse Mobile Crowdsensing (SMCS), existing studies often focus on selecting the subregions most valuable for data inference, overlooking the ultimate goal of data collection, which is to support subsequent decision-making. With the rise of Embodied AI, a core challenge is how to actively collect key information that is most helpful for decision-making under conditions of sparse data and high costs. For example, in urban air quality monitoring, insufficient sensor deployment may cause critical polluted regions to go unobserved, undermining public health decisions. Similarly, sparse traffic data can lead to errors in autonomous driving systems, such as flawed route planning or hazard detection failures. As a feasible data collection paradigm for Embodied AI, existing SMCS subregion selection methods often ignore the relevance of data to tasks, thereby affecting decision-making in Embodied AI. To address this, we design a Sparse Mobile Crowdsensing active perception framework that selects the most valuable subregions for decision-making, considering the context of downstream tasks. Furthermore, existing SMCS subregion selection methods usually focus only on selecting the optimal subregion, whereas Embodied AI requires selecting a set of subregions. Choosing only the optimal subregion may overlook information overlap between subregions, leading to wasted data collection costs and reducing the perception efficiency of Embodied AI. This paper proposes an active acquisition strategy that takes into account parameter uncertainty and probabilistic margins, enabling the selection of an optimal set of subregions. Since this is an NP-hard problem, we approximate the proposed strategy using a greedy algorithm and provide performance guarantees for this approximation through theoretical proofs. Finally, we conduct experiments on three real-world datasets to validate the effectiveness of our proposed method. Experimental results show that when the data missing rate exceeds 60
Gesture-based biometric systems are emerging as a promising approach for secure, natural, and intuitive human–computer interaction in pervasive environments. However, unimodal gesture-based systems are inherently limited by modality-specific vulnerabilities, such as sensitivity to noise, occlusions, and reduced robustness under variability in execution, which can compromise security and reliability. While surface electromyography (EMG) and 3D skeletal motion capture have individually been explored, their systematic multimodal fusion remains under-investigated, despite its potential to enhance robustness and biometric security. In this work, we present a novel multimodal dataset that synchronously records 8-channel EMG signals from forearm muscles together with 3D hand skeleton data from a Leap Motion Controller. The dataset comprises multiple participants performing three distinct gestures (wave, fist, thumbs-up) enabling systematic evaluation across authentication and recognition tasks. Experimental results demonstrate that for person authentication, unimodal classifiers based on EMG and skeleton data achieve accuracies of 95.6
Pervasive computing environments are designed to operate efficiently in diverse contexts, delivering personalized recommendations at any time, anywhere, and for any purpose. However, because they handle sensitive personal data, privacy becomes a critical concern. With personalized recommendation systems becoming integral to daily life, achieving a balance between robust privacy protection and system effectiveness presents substantial challenges. This survey offers a comprehensive analysis of privacy-preserving methodologies developed for pervasive recommender systems (PRS). We rigorously review the recent state of research, identifying key privacy challenges inherent in recommendation frameworks, particularly those arising from malicious attacks such as data breaches, unauthorized access, and adversarial manipulations. We examine existing techniques and solutions designed to mitigate these threats, such as encryption, anonymization, federated learning, blockchain, and differential privacy, evaluating their strengths and limitations in pervasive environments. Additionally, we discuss how improving context awareness through emerging technologies such as edge/fog computing, blockchain, and federated learning can provide promising pathways toward decentralized, distributed computing models. Finally, we outline future research directions that aim to develop robust and scalable solutions that protect user privacy without compromising the performance of recommendation algorithms.
Mobile healthcare systems increasingly require intelligent, AI-driven mechanisms to deliver personalized and reliable services within dynamic and heterogeneous environments. The current research paper in-troduces PRISM-HS, an Intelligent and Ontology-Driven AI Framework for Personalized Mobile Healthcare, designed for intelligent service com-position and real-time adaptation. PRISM-HS leverages an OntoUML-based conceptual model enhanced with the SNOMED CT medical ontology to formally represent users’ profiles, contextual parameters, and service semantics. Through integrating Artificial Intelligence techniques, including Natural Language Processing (NLP) and semantic reasoning, the framework transforms unstructured user inputs into formal representations and ensures that generated workflows are contextually relevant and clinically valid. For execution, PRISM-HS adopts Business Process Execution Language (BPEL), providing a standardized orchestration layer that converts semantically validated requests into deployable healthcare processes. Designed for pervasive and mobile computing environments, PRISM-HS dynamically adapts to patients’ status, device constraints, and environmental conditions. Evaluation across multiple realistic healthcare scenarios revealed that PRISM-HS outperforms state-of-the-art approaches in terms of precision, recall, execution time, and success rate, achieving a high personalization accuracy (91
Emotion recognition is a critical technology in the field of affective computing. Among various carriers of emotional information, electroencephalogram (EEG) signals are widely studied due to their ease of acquisition and resistance to deception. However, raw EEG signals suffer from low signal-to-noise ratio and high dimensionality, necessitating effective feature extraction strategies to fully utilize their rich emotional information. Additionally, the directional interactions among brain neurons are not adequately characterized by traditional graph neural network (GNNs), which rely on spectral graph convolution and are limited to undirected graph structures. To address these challenges, this paper proposes a Temporal-Frequency Fusion Multi-Scale Diffusion Convolution Network (TFF-MDCN). First, a temporal-frequency fusion module based on a gating mechanism is developed to dynamically adjust the weights of time and frequency domain features, enhancing EEG signal representation. Second, a graph diffusion convolution mechanism is introduced to overcome the limitations of spectral graph convolution in handling directed graphs, enabling better capture of directional interactions between electrodes. Finally, a self-attention-based multi-scale fusion module is incorporated to capture dependencies across electrodes at varying diffusion steps, facilitating enhanced generalization. Under subject-dependent protocol, TFF-MDCN achieved accuracies of 93.41
Variations in deployment parameters such as new users or sensor locations are a major concern in sensor-based human activity recognition (HAR), accounting for a major source of performance drop from in-lab experiments to in-the-wild deployment. While fine-tuning a pre-trained model using data acquired during deployment can reduce such performance loss, it often leads to catastrophic forgetting. The hardware constraints of edge devices also limit the options of continual learning techniques. To address these challenges, we introduce COOL, a continual online on-device learning method leveraging Kolmogorov-Arnold Networks (KANs). Our method exploits the inherent plasticity of KANs. COOL is evaluated through two HAR scenarios (utilizing bio-impedance and Inertial Measurement Unit (IMU) signals separately) to demonstrate its performance in addressing both the catastrophic forgetting issue and concept drift issue caused by new targets (users and sensor locations). A significant average overall performance improvement of around 6.82
The trend of technology becoming more widely accepted in higher education is leading to the need for smart systems that can track and improve the efficiency of reading in English for college students. Traditional assessment techniques are slow to be adapted to the changing needs of the student and the process is time-consuming and not very effective in giving personalized feedback, as they rely mostly on grading by an assessor and being performed manually. This study presents a deep learning–monitoring system that is based on Monarch Butterfly Optimized Intelligent Capsule Networks (MBO-Int-CapsNet) to deliver an accurate assessment of reading efficiency and bolster it, as well. The uniqueness of the suggested tactic is the integration of the Monarch Butterfly Optimization (MBO) with the Capsule Networks (CapsNet) where the hyperparameters are automatically set and the routing of the dynamically increased traffic is made more efficient. Contrary to the standard deep learning approaches, MBO-Int-CapsNet ensures that the spatial and hierarchical relationships among the features are maintained, hence allowing the reading-related attributes to be captured with higher precision. While MBO is improving global search, avoiding local optima and speeding up convergence, the CapsNet structure is providing robustness against variations in orientation, position and scale of the speech features. The collection of spoken English recordings from the college-level learners is the dataset, and each of them is assessed by the experts of the field who give standardized proficiency scores depending on the criteria of pronunciation accuracy, fluency, intonation, and overall reading comprehension. Data preprocessing highlights the use of singular value decomposition (SVD) based matrix completion that eliminates the sparsity completely, and the extraction of Mel-frequency cepstral coefficients (MFCCs) that serve phonetic and prosodic pattern representation. The proposed MBO-Int-CapsNet model has shown its exceptional performance by generating minimum MAE and RMSE, achieving 95
The widespread issue of counterfeit pharmaceutical drugs presents a significant global challenge, especially for regions that are resource-constrained and have weak regulations and fragmented verification systems. Conventional methods, such as bar codes, holograms, and centralized databases have inherent weaknesses with regard to limited traceability, susceptibility to tampering, and an absence of trust by end-users. This paper proposes a comprehensive blockchain-based drug verification framework that incorporates smart contracts, role-based access, and a hybrid on/off-chain architecture, all of which may be used to verify a drug’s chain of custody, in real-time and in a tampered-proof way. The framework consists of a permissioned blockchain for secure logging, smart contract validation logic to provide autonomous validation logic, and Internet of Things (IoT). In order to provide the theoretical foundation for the latency of verification, authenticity accuracy, and validation probability, a mathematical model has been proposed. The framework testbed has been implemented using Hyperledger Fabric in a simulated drug supply chain that includes manufacturers, distributors, and pharmacies. Validation of over 10,000 synthetic transactions has been achieved at the cryptographic level. The performance of the proposed model was evaluated in comparison to four, state-of-the-art blockchain verification models using four distinct metrics such as, latency, throughput, false detections, and computational complexity. Generally speaking, the results indicate that the proposed framework reached lower latency at nearly 45
Personality traits play a pivotal role in shaping human behavior and decision-making, and their accurate prediction has garnered significant research interest. Traditionally, personality prediction has relied on self-reported questionnaires; however, advancements in technology have enabled alternative, indirect methods. Users’ interactions within digital environments generate behavioral footprints that can be leveraged for applications in psychology and human-computer interaction. Smartphones, as ubiquitous tools, provide rich data that reveal behavioral patterns, some of which are predictive of personality traits. This paper utilized smartphone data, specifically Call Detail Records (CDR) and Mobile Internet Usage (MIU) logs, to predict Big Five (OCEAN) personality traits. Data collection was facilitated through an Android application installed by 67 voluntary participants, who also completed an online personality questionnaire. Over an average of 30 days, behavioral features were extracted and used in linear regression and Gaussian process models for prediction. To address limited sample size, data augmentation techniques were employed to generate synthetic data, allowing further evaluation using diverse machine learning methods. This study highlights several strengths, including the extraction of novel features, the combined use of MIU and CDR, the higher granularity of collected logs, and diverse predictive methods. The results indicate that all OCEAN traits were predicted with acceptable accuracy, with MIU features outperforming CDR. Gaussian process methods demonstrated superior performance compared to linear regression. On augmented dataset, machine learning methods achieved remarkable accuracy (RMSE 0.1). These findings validate the proposed framework, offering a robust approach for personality prediction and implications for interdisciplinary applications in psychology and telecommunication.
Federated learning (FL) has become a widely adopted paradigm for privacy-preserving model training. However, traditional FL protocols heavily rely on data transmission between clients and servers over the wide-area network (WAN), which is often constrained and unreliable, leading to high communication costs and slow convergence. To address these issues, we propose a LAN-aware FL (LanFL) protocol that efficiently leverages the local-area network (LAN) capacity. By enabling frequent model aggregation within the same LAN, LanFL significantly reduces the need for global aggregation over the WAN, thereby speeding up the training process. However, due to the unique challenges presented by LAN environments, effectively utilizing LAN resources while maintaining the original performance of FL is not straightforward. To overcome this, LanFL incorporates several key techniques: LAN-aware hierarchical aggregation, intra-LAN device topology construction, and inter-LAN heterogeneous bandwidth coordination. We also provide theoretical analysis to derive the convergence bound. Extensive real-world experiments are conducted and the experimental results show that LanFL can significantly accelerate FL training, save WAN traffic, and reduce monetary cost while preserving the model accuracy.
The impact of productive AI on higher education garners significant attention among Chinese scholars. However, existing studies often lack empirical evidence of its actual usage by domestic users. Through investigating the respective variables that explain Chinese college students’ intention and behavior to use generative AI, the current research contributes new knowledge regarding this phenomenon. The following findings are from our analysis of the data according to the Technology Acceptance Model (TAM) and the AISAS Consumer Behavior Analysis Model: (1) Subjective cognition and experience sharing positively shape college students’ perceptions of the ease of use of generative AI. Perceived ease of use fully mediates the influence of subjective cognition and experience sharing on perceived usefulness. (2) There is a very high intention among college students to use generative AI, but variations exist against various demographic characteristics. (3) Intentions affect the generative AI use in terms of the perceived utility and user-friendliness among college students. This study clarifies the inherent processes through which subjective perception and experience of domestic college students regarding generative AI influence their intention toward adoption.