
Federated learning in mobile edge systems faces a key challenge: how to train models efficiently while still allowing slow devices to participate. Many existing methods either drop slow devices to speed up training, which hurts fairness, or wait for all devices, which wastes the computing power and time of faster ones. CLAMP addressed this through dynamic depth selection, allowing each client to train an adjusted number of model layers based on their capabilities. However, CLAMP's minimum depth aggregation creates a bottleneck. The server only aggregates the layers trained by all clients, which means that updates from deeper layers trained by faster devices are ignored. This strategy leaves the model capacity unused and limits the convergence speed. We propose Mask-Aware CLAMP, which replaces minimum-depth aggregation with layer-wise mask-aware aggregation. MA-CLAMP aggregates each model layer independently using only the clients that trained it, subject to minimum participation thresholds that ensure stability. This approach preserves deeper layer updates from capable clients while maintaining CLAMP's adaptive depth selection and straggler handling. MA-CLAMP does not require changes to client-side training. Experiments across MNIST, Fashion-MNIST, and CIFAR-10 with 100 heterogeneous clients show that MA-CLAMP achieves faster convergence on moderately complex tasks, improved resource efficiency in computation and energy use, maintains or improves model accuracy across all datasets, and enhances slow-device inclusion rates. MA-CLAMP provides a more effective solution for production federated learning systems where model quality, resource efficiency, and fair client participation are all essential.
Since the invention of automobiles, engineers have spent endless efforts to enhance their safety, operability, and fuel efficiency. The rapid development of electric vehicles (EVs) has transformed them into movable computing hubs, equipped with ample sensing, computing capabilities, and battery power. However, the placement of all the vehicle's sensors within its structure poses limitations to its sensing capability, consequently restricting potential applications. Meanwhile, unmanned aerial vehicles (UAVs or drones) offer aerial-perspective sensing capability from flexible viewpoints, processing which in real time can complement cars' sensing and enable novel applications. In this paper, we introduce our vision of how the real-time collaboration of ground and aerial vehicles can enable innovative mobility solutions. We discuss the system architectures, potential usage scenarios, and enabling technologies for supporting drone-car collaboration. Furthermore, we examine the latest technical advancements in the field and shed light on the promising opportunities that lie ahead in the future. Overall, this article provides a comprehensive understanding of the potential and significance of real-time collaboration between drones and cars in revolutionizing future mobility, improving fuel efficiency, and enhancing driving safety.
To provide timely outdoor emergency rescue, it is necessary to assess the human injury severity (HIS) in real time and proactively send a distress signal when an injury is detected. In the past years, there have been many studies on HIS assessment. However, these studies commonly use bulky medical equipment for human physiological measurements, making them inappropriate to be applied outdoors. Given that Photoplethysmogram (PPG) sensors, which are small in size and low in power cost, can measure a series of human physiological indicators, many studies have embedded PPG sensors into wearable devices, e.g., smartwatches, to enable human health monitoring. Based on PPG technology, a series of human physiological indicators such as heart rate, respiratory rate, oxygen saturation and blood pressure can be measured over long periods of time. This article aims to investigate the way of using a wearable PPG device to assess HIS for outdoor activity scenarios. First, the feasibility of using PPG-derived data to assess HIS is discussed. Then, a database used for HIS assessment research is built. Next, the PPG-based human physiological features used for HIS assessment are extracted. Later, six HIS assessment models, including LR, LDA, Linear_SVM, RBF_SVM, RF and CatBoost, were investigated. Experimental results show that the CatBoost model demonstrates overall optimal performance for assessing mild injury cases, while the RBF_SVM model exhibits superior performance for assessing moderate and severe injury cases.
Data spaces provide governed ecosystems for sovereign data exchange among organizations. However, current implementations rely on static, accept-or-reject contract mechanisms that limit agreement flexibility and constrain ecosystem adoption through suboptimal or failed negotiations. This work presents an automated and bilateral negotiation framework for multi-issue contracts in data spaces, aligned with the Dataspace Protocol (DSP) and the Open Digital Rights Language (ODRL) standards. The framework comprises a negotiation engine that implements iterative, time-dependent concession strategies, utility-based evaluation, and opponent modeling, together with a translator module that provides bidirectional mapping between ODRL policy constructs and the negotiation logic. The framework design is grounded in a practitioner needs assessment that identifies current negotiation gaps, confirms demand for automated support, and informs the selection of negotiation issues and the human-in-the-loop approval mechanism. The framework is integrated with the Eclipse Dataspace Components (EDC) and validated in an urban data-sharing scenario. A systematic evaluation compares the framework against baseline strategies, analyzes the effect of concession pace on the efficiency-welfare tradeoff, and characterizes framework robustness across varying degrees of preference overlap between parties.
Deciding which colors to use for a grayscale film without any guidance can lead to a myriad of colorization outputs, with some more believable than others. Unlike cinema, the lighter burden of colorizing photos allows human text guidance to be used. This paper attempts to take advantage of recent advances in large language models and neuro-symbolic artificial intelligence (AI) to extend human-guided image colorization to the video domain. The instantiation of this methodology we call " RAGCol++ : R etreival A ugemented G eneration based automatic video Col orization using Semantic Similarity Search and Probabilistic Grounded Knowledge". This work improves upon the original RAGCol paper by expanding the size and quality of COL-KG, adding efficiencies to the automatic video colorizer and incorporating semantic similarity search as opposed to the original Cypher query-based search. Our system achieves an average improvement of 3% over the previous state-of-the-art L-CAD + BVD across the DAVIS and Videvo datasets when using PSNR, SSIM, FID, FVD, and CDC as the metrics. This result is also substantiated by our user study, where RAGCol++ was preferred 56% of the time.
Machine learning models risk reproducing social biases, so algorithmic decision-making must incorporate principles that prevent discrimination. Post-processing methods, such as threshold optimization, can support this goal, but striking an appropriate trade-off between predictive performance and group equity metrics remains challenging. The recent application of Differential Item Functioning (DIF) in model selection sheds light on its potential as a promising yet unexplored approach to threshold tuning in more unbiased machine learning systems. Building on this premise, this work introduces DIF-PP, a fairness-aware post-processing method that draws on concepts from Item Response Theory (IRT) and DIF for optimizing decision boundaries. DIF-PP represents these thresholds as test items, derives classification characteristic curves through IRT, and uses DIF to identify the most impartial cutoff point. Experimental results with 18 datasets show that DIF-PP consistently outperforms existing methods when we analyze group fairness metrics and predictive performance simultaneously. By combining IRT and DIF, our proposal effectively mitigates discriminatory effects in binary classification, marking a significant advance toward the development of responsible artificial intelligence solutions.
Motivated by providing a model that fits real-world scenarios, we develop a tractable and realistic analytical framework by carrying out an analysis on quantifying the performance of SWIPT-assisted underlaid D2D networks. The analysis is conducted under two key constraints-imperfect channel state information (CSI) and tiered power allocation for cellular users-which are of critical importance in real-world communications and are frequently understudied in the literature. Specifically, we first derive the expected received energy at a D2D receiver (DRU) from its associated transmitter (DTU), assuming only statistical CSI availability at the receiver. Our stochastic geometry-based analysis reveals the impact of CSI uncertainty and, to our knowledge, constitutes the first such treatment for SWIPT-enabled D2D energy harvesting. Then, we extend the analysis to a stratified power model, where cellular users transmit at tier-specific power levels to capture device heterogeneity. Closed-form expressions for ergodic harvested energy are derived, signifying the impact of the tiered structure under imperfect CSI. Finally, we validate the analytical results through Monte Carlo simulations and discuss their practical implications for system design, including tier planning and user association strategies. The joint consideration of imperfect CSI with stratified power allocation enhances the tractability, usability, and accuracy of the model with respect to real-world SWIPT-enabled D2D deployments.
Personalized news recommendation has become an essential technology for online news services. For effective personalized news recommendation, it is ideal to utilize various information, such as the title and body text. However, some news services do not retain rich information such as the body text, and only titles may be available. In this paper, we propose a news recommendation framework that utilizes data augmentation with ChatGPT. By inputting our prompt and news title into ChatGPT, we extend the information in news articles to supplement the news content feature. In particular, we focus on two directions of title extension: (1) user direction and (2) content direction. In the proposed framework, ChatGPT infers and outputs (1) the target audience of the news article in the user direction, and (2) the categories of the news article in the content direction. In addition, to further enrich the extended information, we introduce a title-similarity-based augmentation module. Evaluation experiments on a real-world news service dataset show that the proposed framework outperforms conventional methods by up to 1.65% in AUC (area under the ROC curve). These findings highlight the importance of extending the content features of news articles from their titles and utilizing them for recommendation through various prompting strategies. Furthermore, experiments across multiple datasets confirmed that the title-similarity-based augmentation module works well in some cases, but does not always select articles that contribute to recommendations when titles contain little information or are excessively short. We have provided our code and GPT-generated data to enable other researchers to reproduce our findings. (1)
Real-time tasks, especially control tasks, can often tolerate occasional missed deadlines due to robust algorithms. The weakly-hard model offers an approach for specifying the maximum number tolerable deadline misses m(i ) within a sequence of K(i )executions. Research has shown that utilizing the weakly-hard model can significantly reduce the over-provisioning typically required in real-time system design. This led to the development of various scheduling algorithms and schedulability analyses in recent years. However, existing state-of-the-art analyses have limitations: they do not scale with larger values of K-i and focus solely on job-kill as a system level action. We propose a new job-level fixed priority scheduling algorithm that overcomes these limitations. Our approach first considers the traditional job-kill method but then extends to a skip-next job strategy in case of deadline miss. The schedulability analysis our algorithm scales with K-i , reducing computational time by up to 100 times compared to existing approaches.
Accurate traffic forecasting is fundamental for intelligent transportation systems, directly influencing congestion management, safety, and emissions. A central challenge is the non-stationary nature of traffic flow, where relationships between sensors shift dynamically due to incidents, demand changes, and propagation waves. Many existing graph-based models rely on static graphs or computationally heavy dynamic mechanisms, limiting their suitability for real-time deployment. We propose ST-Hybrid, a spatio-temporal model that incorporates a lightweight state-conditioned dynamic graph learner. The module updates connectivity in real time using adaptive node embeddings and sparse top- k neighborhood selection, offering a practical balance between flexibility and efficiency. On PeMSD8, it obtains a mean absolute error (MAE) of 15.19 and root mean square error (RMSE) of 24.18, ranking second among recent dynamic-graph approaches. It maintains competitive accuracy on the larger PeMSD4 network (MAE 20.16, RMSE 32.17) and demonstrates robust performance on PeMSD3 (MAE 15.58, RMSE 25.93) and PeMSD7 (MAE 21.58, RMSE 34.36). Sensor-level analyses reveal that most residual error is concentrated in a small subset of volatile sensors, suggesting that targeted refinement may yield greater improvements than further global architectural complexity. Importantly, ST-Hybrid maintains low inference latency (approximately 15 ms), demonstrating its suitability for large-scale, real-time traffic forecasting applications.
The problem of decentralized multi-robot patrol has previously been approached primarily with hand-designed strategies for minimization of "idleness" over the vertices of a graph-structured environment. These approaches often involve heuristic utility functions and complex inter-agent coordination mechanisms, which may introduce unnecessary complication or performance degradation due to inefficient utility function formulations. To examine this, we present two lightweight learned neural network-based strategies and show that they outperform existing strategies in both idleness minimization and against an intelligent intruder model, as well as presenting an examination of robustness to communication failure. We also present a minimal regression of these strategies, and show that performance comparable to leading literature strategies is achievable with an extremely simple controller. Our results indicate important considerations for strategy design and analysis of patrol systems, which we discuss in depth.
Urban traffic networks evolve continuously as cities add and relocate sensors, yet most prediction pipelines must be retrained from scratch to exploit these new data sources. Such retraining is costly in computing, carbon, and engineering effort. We introduce DESPT—Distance-Enhanced Spatial Pre-Training—a sustainable framework that learns spatial embeddings once and reuses them whenever new sensors are added. DESPT couples a graph module whose adjacency weights decay with real driving distance and a time-series contrastive encoder that is pre-trained on historical streams, then frozen. At inference time, the frozen embeddings are concatenated with live readings, allowing any downstream forecaster to generalize immediately to unseen sensors with negligible extra energy. Extensive experiments on three benchmarks—METRLA, PeMS-BAY, and a two-year Hague network—show that DESPT cuts mean absolute error by 3.4%, 3.7%, and 9.2%, respectively, relative to strong graph baselines. A runtime analysis confirms that DESPT amortizes computation across deployments, delivering accurate, low-footprint forecasts that align with the sustainability goals of modern smart-city infrastructure.
Medical image processing (MIP) can make the medical diagnosis more accurate and efficient. However, the operating process of MIP involves many computationally-intensive tasks, which results in a long computational time for MIP. To optimize the processing time of MIP, many research works use the heterogeneous computing architecture (HCA) to accelerate MIP. These works are effective. However, most of them focus on single-accelerator HCA (SA-HCA), which commonly consists of one CPU plus one kind of accelerators, but pay less attention to the multi-accelerator HCA (MA-HCA) such as CPU/FPGA/NPU or CPU/GPU/FPGA/NPU. Considering different accelerators are suitable for executing different tasks, a MA-HCA can theoretically achieve better computing performance than a SA-HCA. Therefore, this paper aims to design a MA-HCA dedicated to MIP to enhance the acceleration performance of MIP. First, a series of representative MIP algorithms are selected, and the work of accelerating these algorithms by using Huawei's Ascend neural processing unit (NPU) is investigated. Then, the NPU acceleration mechanism is integrated with the FPGA and GPU acceleration strategies to build a MA-HCA dedicated to MIP, and a unified programming model for this MA-HCA is also designed. To evaluate the performance of the proposed MA-HCA, the execution time and energy cost of various MIP algorithms on different SA-HCAs and MA-HCA are measured, and the results show that the MA-HCA presented in this paper can significantly improve the acceleration performance while still maintaining low power consumption.
In recent years, TLC NAND flash memory has emerged as the dominant storage medium, primarily due to its high storage density, cost-effectiveness, and widespread adoption in various applications. However, despite these advantages, TLC NAND flash memory suffers from significant reliability challenges, especially when data is retained for an extended period. One of the most critical issues is retention errors, which tend to accumulate over time, eventually surpassing the error correction capabilities of ECC (Error Correction Code). This problem is particularly severe in the early stages of usage, when the memory has undergone fewer program/erase (P/E) cycles and is highly vulnerable to retention errors caused by lateral charge migration between adjacent cells. In this paper, we propose a reliability-based coding strategy to suppress the lateral charge migration and minimize retention errors during the early lifespan of TLC NAND flash memory. By leveraging an optimized coding approach, the proposed method effectively enhances data reliability without introducing significant computational overhead. Experimental results demonstrate that our strategy significantly outperforms previous methods, providing a substantial reduction in retention errors while maintaining efficient storage performance.
With the increasing integration of renewable energy sources, AC microgrids have become an essential medium for renewable energy utilization. However, the rapid expansion of microgrids necessitates efficient methods to evaluate their steady states. This paper proposes an efficient, non-iterative linear power flow analysis method for droop-controlled islanded AC microgrids. The frequency-dependent bus admittance matrix is approximated by linearization to indicate the power change under a frequency deviation, and then a closed-form formula is derived to ensure efficient and non-iterative power flow solving. Simulation results show that the proposed method can 1) calculate the steady state of the droop-controlled AC microgrids with an acceptable accuracy and 2) address the convergence issues of the existing methods and greatly enhance the computational efficiency, achieving a speedup of 12 times for a large stacked microgrid with 8,500 buses when compared with an exiting method.
Video anomaly detection is crucial for applications like surveillance and autonomous systems. Traditional methods often rely solely on visual cues, missing valuable contextual data. This paper presents Perceptual and Analytical Representation Learning (PEARL), a novel method that combines perceptual (raw sensory input) and analytical (higher-level context) modalities. Specifically, we integrate visual information with object tracking data, along with the tracking data-specialized normalization method, DOT-Norm, leveraging ID switching to capture high-level contexts of abnormal movements. We evaluate early- and late-fusion strategies to enhance anomaly detection, particularly for irregular movements marked by frequent track ID switches. Our approach, tested on the UCSD-Ped1 dataset, outperforms the state-of-the-art by improving precision (+0.082), recall (+0.104), F1 score (+0.149), and AUC (+0.053). These findings highlight the potential of integrating analytical tracking data with perceptual video frames in a multimodal learning approach for anomaly detection, paving the way for future applications and research where knowledge-driven analytical modalities are crucial.
Due to the vast amount of reviews available on the Web, in the past decades, a growing share of work has focused on sentiment analysis. Aspect-based sentiment classification is the subtask that seeks to detect the sentiment expressed by the content creators towards a defined target within a sentence. This paper introduces three novel unsupervised attentional neural network models for aspect-based sentiment classification, and tests them on English restaurant reviews. The first model employs an autoencoder-like structure to learn a sentiment embedding matrix where each row of the matrix represents the embedding for one sentiment. To improve the model, a target-based attention mechanism is included that de-emphasizes irrelevant words. Last, a redundancy and a seed regularization term constrain the sentiment embedding matrix. The second model extends the first by including a Bi-LSTM layer in the attention mechanism to exploit contextual information. The third model further adapts a Left-Center-Right separated neural network with Rotatory attention structure from the supervised realm to an unsupervised setting. Although all three models construct meaningful sentiment embeddings, experimental results indicate that the inclusion of the Bi-LSTM in the attention mechanism leads to a more precise attention mechanism and, thus, better predictions. The best model, i.e., the second, outperforms all investigated unsupervised and weakly supervised algorithms for aspect-based sentiment classification from the literature.
User Experience (UX) is an essential factor in today's business, and so large companies have adopted a process of continuous UX monitoring. In this process, design experts run user tests and analyze UX performance indicators, which are usually obtained from standardized questionnaires performed with users. Thus, this process is expensive, as it requires paying test participants, and allocating the time to manually analyze the collected data. Previously, we have defined interaction effort as a metric for the dynamic performance of web elements, based on measures of the user interaction, and we presented UX-Analyzer as a tool to visualize the interaction effort. UX-Analyzer allows UX experts or other team members to evaluate web pages automatically and in a transparent way. Visualizing the interaction effort may provide a good indication towards the level of UX of an online system. It may also be used in the context of an A/B testing approach that instead of revenue or conversions, it compares the UX of alternative versions of a website. In this paper we show how UX-Analyzer can perform in real settings. We present a case study with 152 participants that show its applicability, and we also add a qualitative study with design professionals to assess its adoption and usefulness. Additionally, we have extended the related work, and describe some new features that improve the tool's flexibility.
A federated network can be seen as a collection of network domains, typically under different authorities, who collaborate and share network resources to enable the provision of end-to-end multi-domain services with possibly performance guarantees. To provide these services, domains need to disclose a compact portrayal of their topology. The interconnection of all exposed abstracted topologies serves as input to compute the appropriate paths that are needed to support an end-to-end service. Domains basically resort to a topology abstraction that reduces to a mesh of abstract links that connect to their border nodes. Such an abstraction voluntarily limits the information disclosed to other domains, but, on the other hand, leads to inefficient resource usage. In order to enable effective collaborations between domains, in this paper, we propose to enrich the topology aggregation exposed by domains by including additional abstract topology constructs, e.g. abstract non-border nodes, domain-level slices, etc. We then revisit the virtual link embedding problem to include the proposed aggregation. Our evaluations on real and random topologies show significant gains in terms of admission ratios and resource usage.
We study how decentralized multi-agent systems adapt on local data in dynamic environments. In this extension of our earlier work, we build on our original setting of a driverless transportation system for a manufacturing scenario with Autonomous Guided Vehicles (AGVs) to include not only disturbances, but also perception and communication errors that hinder the information sharing of agents. We analyze how well the decentralized MAS can correct these errors, where corrections come from, how the errors impact the performance, and consider the communication efficiency among the MAS population.