Federated Learning (FL) enables collaborative model training across decentralized clients without sharing private data. However, FL suffers from biased global models due to non-IID and long-tail data distributions. We propose \textbf{FedSM}, a novel client-centric framework that mitigates this bias through semantics-guided feature mixup and lightweight classifier retraining. FedSM uses a pretrained image-text-aligned model to compute category-level semantic relevance, guiding the category selection of local features to mix-up with global prototypes to generate class-consistent pseudo-features. These features correct classifier bias, especially when data are heavily skewed. To address the concern of potential domain shift between the pretrained model and the data, we propose probabilistic category selection, enhancing feature diversity to effectively mitigate biases. All computations are performed locally, requiring minimal server overhead. Extensive experiments on long-tail datasets with various imbalanced levels demonstrate that FedSM consistently outperforms state-of-the-art methods in accuracy, with high robustness to domain shift and computational efficiency.
Since the release of GPT2-1.5B in 2019, the large language models (LLMs) have evolved from specialized deep models to versatile foundation models. While demonstrating remarkable zero-shot ability, the LLMs still require fine-tuning on local datasets and substantial memory for deployment over the network edges. Traditional first-order fine-tuning techniques require significant GPU memory that exceeds the capacity of mainstream hardware. Besides, the LLMs have been expanded beyond text generation to create images, audio, video, and multi-modal content, necessitating careful investigation of efficient deployment strategies for large-scale foundation models. In response to these challenges, model fine-tuning and model-compression techniques have been developed to support the sustainable growth of LLMs by reducing both operational and capital expenditures. In this work, we provide a comprehensive overview of prevalent memory-efficient fine-tuning methods for deployment at the network edge. We also review state-of-the-art literature on model compression, offering insights into the deployment of LLMs at network edges.
Large language models (LLMs) have demonstrated exceptional proficiency in language understanding. However, when LLMs align their outputs with deceptive and/or misleading prompts, the generated responses could deviate from the de facto information. Such observations are known as fawning hallucinations, where the model prioritizes alignment with the input's implied perspective over accuracy and truthfulness. In this work, we analyze fawning hallucinations in various natural language processing tasks and tailor the so-termed contrastive decoding method for fawning-hallucination mitigation. Specifically, we design two paradigms to generate corresponding deceptive and/or misleading inputs for the consistent fawning hallucinations induction. Then, we propose the collaborative contrastive decoding (CCD) to handle the fawning hallucinations across different tasks in LLMs. By contrasting the deviation in output distribution between induced and transformed neutral inputs, the proposed CCD can reduce reliance on deceptive and/or misleading information without requiring additional training. Extensive experiments demonstrate that the proposed CCD can effectively mitigate fawning hallucinations and improve the factuality of the generated responses over various tasks.
Facial expressions convey human emotions and can be categorized into macro-expressions (MaEs) and micro-expressions (MiEs) based on duration and intensity. While MaEs are voluntary and easily recognized, MiEs are involuntary, rapid, and can reveal concealed emotions. The integration of facial expression analysis with Internet-of-Thing (IoT) systems has significant potential across diverse scenarios. IoT-enhanced MaE analysis enables real-time monitoring of patient emotions, facilitating improved mental health care in smart healthcare. Similarly, IoT-based MiE detection enhances surveillance accuracy and threat detection in smart security. Our work aims to provide a comprehensive overview of research progress in facial expression analysis and explores its potential integration with IoT systems. We discuss the distinctions between our work and existing surveys, elaborate on advancements in MaE and MiE analysis techniques across various learning paradigms, and examine their potential applications in IoT. We highlight challenges and future directions for the convergence of facial expression-based technologies and IoT systems, aiming to foster innovation in this domain. By presenting recent developments and practical applications, our work offers a systematic understanding of the ways of facial expression analysis to enhance IoT systems in healthcare, security, and beyond.
We investigate robust federated learning under ad-versarial settings where a subset of participating devices may behave in a Byzantine manner, submitting arbitrarily corrupted updates to disrupt global model training. To jointly address the critical challenges of device bottlenecks, slow convergence, and Byzantine threats, we propose Byrd-Katyusha-a Byzantine-resilient, memory-efficient federated learning algorithm that integrates Katyusha's momentum-based acceleration, SVRG-style variance reduction, and robust aggregation. Byrd-Katyusha strategically distributes the computational burden between server and workers, achieving fast convergence and robustness while maintaining an 0 (1) memory footprint per device. The algorithm leverages corrected mini-batch gradients and Krum aggregation to mitigate the impact of adversarial updates, while avoiding the per-sample memory overhead and parallelism limitations of prior SAGA-based methods. Extensive experiments on benchmark datasets such as IJCNNI and CIFAR-IO demonstrate that Byrd-Katyusha consistently outperforms existing baselines in terms of convergence speed, accuracy, and robustness to adversarial behavior.
Scalable video coding (SVC) and power-division multiplexing (PDM) are crucial technologies for video streaming in Internet of Vehicles (IoV) networks due to their robustness over time-varying channels. However, the correlation of video tiers and channel fluctuations induces significant complexity in the joint design of SVC and PDM in IoV networks. To reduce such complexity, we propose a novel transmission strategy for SVC-based video streaming in PDM technology by facilitating cross-layer control across the application (APP), data link, and physical (PHY) layers. Specifically, we propose a concise framework where the modulation order, the power control, and the prioritization of video tiers are jointly controlled. Moreover, we leverage the network abstraction layer units (NALUs) within an arbitrary group-of-pictures (GoPs) to improve the overall video data rate. The formulated joint control of the modulation order, the transmission power, and video-tier prioritization is proven to be a quasi-concave problem. Consequently, our proposed cross-layer optimization-based SVC video transmission strategy could efficiently utilize the prioritized video tiers in the APP layer, power control in the data-link (DL) layer, and modulation orders in the PHY layer. Numerical results are used to demonstrate performance improvements over the benchmarks.
Recently, scalable video coding (SVC) has gained significant recognition in mobile video streaming because it can adapt bitstreams to time-varying transmission conditions. However, the coding performance of SVC, which is determined by its coding structure, has not been thoroughly studied. To address this issue, we propose analyzing the redundancy, reduction, distortion, and mutuality of video information within the video coding processes. This analysis facilitates the development of a novel information-theoretical framework for quantifying coding performance, which includes an information theory (IT)-based quantification method and a graphical representation system. The representation system accurately delineates the coding reference structure for encoding each video frame, while the proposed method utilizes mutual information to quantify the achievable coding performance of SVC under the delineated structure. To demonstrate the significance of our research, we apply the proposed framework to encode a basic coding unit, showcasing its effectiveness in improving SVC schemes. Consequently, our framework not only provides an efficient approach for quantifying the coding performance of SVC but also serves as an invaluable tool for optimizing SVC in various applications.
Personalized Federated Learning (PFL) targets client-specific models under heterogeneous and limited data. However, conventional methods often use heuristic or data-size-based averaging and overlook the true contributions of client updates. We propose a contribution-oriented PFL framework that quantifies client contributions via gradient alignment and prediction discrepancy for informed aggregation. We further develop a parameter-wise personalization mechanism for adaptive local updates and a mask-aware momentum optimizer for stable training. Preliminary results on CIFAR10 validate its effectiveness.
An aptamer sensor for the rapid and sensitive recognition of oxytetracycline (OTC) was developed by modifying a screen-printed carbon electrode with Cu-based metal organic framework (MOF), electrochemically reduced graphene oxide (ErGO), and gold nanoparticles (AuNPs). The Cu-MOF had a large surface area and accommodated AuNPs and ErGO. The electrodeposition of AuNPs and the immobilization of ErGO improved the conductivity to amplify the sensor response. The developed sensor had an extremely broad linear range (0.1-105 ng/ mL) and a very low limit of detection (0.03 ng/mL), and it was highly stable, specific, and reproducible. When used to detect OTC in milk and pork samples, the recovery ranged from 87.0 % to 110.2 %.
Communication remains a central bottleneck in large-scale distributed machine learning, and gradient sparsification has emerged as a promising strategy to alleviate this challenge. However, existing gradient compressors face notable limitations: Rand-K discards structural information and performs poorly in practice, while Top-K preserves informative entries but loses the contraction property and requires costly All-Gather operations. In this paper, we propose ARC-Top-K, an All-Reduce-Compatible Top-K compressor that aligns sparsity patterns across nodes using a lightweight sketch of the gradient, enabling index-free All-Reduce while preserving globally significant information. ARC-Top-K is provably contractive and, when combined with momentum error feedback (EF21M), achieves linear speedup and sharper convergence rates than the original EF21M under standard assumptions. Empirically, ARC-Top-K matches the accuracy of Top-K while reducing wall-clock training time by up to 60.7%, offering an efficient and scalable solution that combines the robustness of Rand-K with the strong performance of Top-K.
Federated Learning (FL) has gained significant attention for enabling privacy preservation and knowledge sharing by transmitting model parameters from clients to a central server. However, with increasing network scale and limited bandwidth, uploading complete model parameters has become increasingly impractical. To address this challenge, we leverage the high informativeness of prototypes—feature centroids representing samples of the same class—and propose Federated Prototype Momentum Contrastive Learning (FedPMC). At the communication level, FedPMC reduces communication overhead by using prototypes as carriers instead of full model parameters. At the local model update level, to mitigate overfitting, we construct an expanded batch sample space to incorporate richer visual information, design a supervised contrastive loss between global and real-time local prototypes, and adopt momentum contrast to gradually update the model. At the framework level, to fully exploit the sample's feature space, we employ three different pre-trained models for feature extraction and concatenate their outputs as input to the local model. FedPMC supports personalized local models and utilizes both global and local prototypes to address data heterogeneity among clients. We evaluate FedPMC alongside other state-of-the-art FL algorithms on the Digit-5 dataset within a unified lightweight framework to assess their comparative performance. The code is available at https://github.com/zhy665/fedPMC
With the extraordinary success of generative artificial intelligence, large pretrained models (LPMs) have been widely used to achieve human-level performance. Despite the one-shot capability, it is always preferred to fine-tune the LPMs for domain-specific downstream tasks. Therefore, the federated learning system is leveraged to fine-tune the large pretrained models enabling concurrrently use multiple distributed clients as well as their local datasets. While the first-order fine-tuning methods suffer from high computational and memory costs due to the backward propagation, we are motivated to propose a federated zeroth-order fine-tuning method with only forward propagation. Moreover, we also leverage differential privacy to further preserve the data privacy of local clients. Experimental results illustrate that our proposed federated zeroth-order method can reduce the memory and retain a similar testing accuracy over the state-of-the-art benchmarks.
We investigate robust federated learning, where a group of workers collaboratively train a shared model under the orchestration of a central server in the presence of Byzantine adversaries capable of arbitrary and potentially malicious behaviors. To simultaneously enhance communication efficiency and resilience against such adversaries, we propose a Byzantine-resilient Nesterov-accelerated federated learning (Byrd-NAFL) algorithm. Byrd-NAFL seamlessly integrates Nesterov's momentum into the federated learning process alongside Byzantine-resilient aggregation rules to achieve fast and safe convergence against gradient corruption. We establish a finite-time convergence guarantee for Byrd-NAFL under non-convex and smooth loss functions with relaxed assumptions on the aggregated gradients. Extensive numerical experiments validate the effectiveness of Byrd-NAFL and demonstrate the superiority over existing benchmarks in terms of convergence speed, accuracy, and resilience to diverse malicious attacks.
Personalized federated learning (PFL) addresses a critical challenge of collaboratively training customized models for clients with heterogeneous and scarce local data. Conventional federated learning, which relies on a single consensus model, proves inadequate under such data heterogeneity. Its standard aggregation method of weighting client updates heuristically or by data volume, operates under an equal-contribution assumption, failing to account for the actual utility and reliability of each client's update. This often results in suboptimal personalization and aggregation bias. To overcome these limitations, we introduce Contribution-Oriented PFL (CO-PFL), a novel algorithm that dynamically estimates each client's contribution for global aggregation. CO-PFL performs a joint assessment by analyzing both gradient direction discrepancies and prediction deviations, leveraging information from gradient and data subspaces. This dual-subspace analysis provides a principled and discriminative aggregation weight for each client, emphasizing high-quality updates. Furthermore, to bolster personalization adaptability and optimization stability, CO-PFL cohesively integrates a parameter-wise personalization mechanism with mask-aware momentum optimization. Our approach effectively mitigates aggregation bias, strengthens global coordination, and enhances local performance by facilitating the construction of tailored submodels with stable updates. Extensive experiments on four benchmark datasets (CIFAR10, CIFAR10C, CINIC10, and Mini-ImageNet) confirm that CO-PFL consistently surpasses state-of-the-art methods in in personalization accuracy, robustness, scalability and convergence stability.
Recent advancements have introduced federated machine learning-based channel state information (CSI) compression before the user equipments (UEs) upload the downlink CSI to the base transceiver station (BTS). However, most existing algorithms impose a high communication overhead due to frequent parameter exchanges between UEs and BTS. In this work, we propose a model splitting approach with a shared model at the BTS and multiple local models at the UEs to reduce communication overhead. Moreover, we implant a pipeline module at the BTS to reduce training time. By limiting exchanges of boundary parameters during forward and backward passes, our algorithm can significantly reduce the exchanged parameters over the benchmarks during federated CSI feedback training.
To address the growing challenge of energy efficiency in next-generation coordinated multipoint (CoMP) communication systems, this article develops a green CoMP optimization framework that integrates renewable energy harvesting, smart grid interactions, and real-time power control. We formulate a stochastic long-term weighted sum-rate maximization problem, incorporating transmit covariance variables and joint channel-aware precoding. To enable online implementation, we convert the time-averaged problem into an equivalent per-slot formulation and design an online dynamic beamforming and energy management (ODBEM) algorithm. The proposed ODBEM integrates three synergistic mechanisms: 1) dual-driven energy pricing; 2) Lyapunov drift-plus-penalty scheduling; and 3) momentum-based energy smoothing. We further conduct rigorous convexity and Karush-Kuhn-Tucker optimality analysis to ensure algorithmic correctness and convergence. Simulation results demonstrate that ODBEM outperforms baseline strategies in both throughput and energy cost, confirming its effectiveness for sustainable and adaptive CoMP transmission.
With the rapid development of autonomous driving and edge computing, vehicular edge computing (VEC) has become an emerging paradigm that allows vehicles with abundant computational resources to work as edge nodes. By introducing vehicles as infrastructures, VEC has the potential to improve users' quality of experience and decrease operator's deployment expenditure, especially for hot spots. In this article, a novel VEC-based resource reservation framework is designed to handle the time-varying computation requests. To articulate realistic scenarios, the online durations of provider vehicles (PVs) are assumed to be different. Besides, the PVs will not always be online to wait for the reservation assignment for the limited revenue, i.e., the PVs are dynamic and the computational resource reservation points (CRRPs) are static. In this way, dynamic matching is leveraged to model the interaction between the PVs and CRRPs. To prevent the CRRPs from manipulating their preferences for better partners, a strategy-proof and stable resource reservation algorithm is proposed to ensure all CRRPs are truthful during the resource reservation procedure. Finally, numerical simulation results are presented to validate the proofness, truthfulness, and performance of our proposed resource reservation algorithm.
Enhancing the sensitivity of targeted substance detection is crucial, yet prior research seldom incorporates more than two signal amplification methods. In this study, we introduced a novel triple signal amplification strategy for tetracycline detection using an aptasensor. This strategy integrates a graphene and multi-walled carbon nanotubes composite (GO-MWCNTs), Exonuclease I (Exo I), and a hybrid DNA-gold nanoparticle (AuNPs)horseradish peroxidase (HRP) system. The GO-MWCNTs serve as a conductive carrier, boosting electron transfer for initial signal amplification. Exo I, targeting single-stranded DNA, facilitates target recovery and secondary signal amplification. The gold nanoprobe, through specific base pairing, binds to the tetracycline aptamer's complementary chains on the electrode surface. Horseradish peroxidase's catalytic action then generates a robust electrochemical signal, culminating in three-stage signal amplification. This optimized approach achieved a low detection limit of 3.3 x 10-4 ng mL- 1, with a range from 1 x 10-3 ng mL- 1 to 1 x 103 ng mL- 1. Notably, the aptasensor demonstrated high selectivity, repeatability, stability, and reliability. These findings offer a promising reference for developing effective aptasensors in antibiotic detection.
With the rapid expansion of Internet of Things (IoT) applications, federated learning (FL) has emerged as a critical technology for managing distributed datasets. However, Byzantine faults pose significant security threats to the practical deployment of FL in IoT contexts. This research aims to evaluate the impact of these threats on FL implementations in IoT environments. Specifically, we conducted experiments within the Federated Stochastic Gradient Descent (FedSGD) framework using the CIFAR-10 dataset to evaluate the resilience of FL against various Byzantine attacks. These experiments involved testing seven types of Byzantine attacks and assessing the accuracy and effectiveness of nine different Byzantine defense mechanisms. We introduced a novel performance metric, $P_{BD1}$ , which enhanced our ability to comprehensively evaluate the effectiveness of these defenses. Our results indicate that while most defense mechanisms exhibit varying degrees of effectiveness, DnC (Divide-and-Conquer) and ClippedClustering emerged as the most promising techniques. Additionally, the new performance metric P B D 1 has proven to be a valuable tool for comprehensive evaluation.
The next generation of communication is envisioned to be intelligent communication, that can replace traditional symbolic communication, where highly condensed semantic information considering both source and channel will be extracted and transmitted with high efficiency. The recent popular large models such as GPT4 and the boosting learning techniques lay a solid foundation for the intelligent communication, and prompt the practical deployment of it in the near future. Given the characteristics of "training once and widely use" of those multimodal large language models, we argue that a pay-as-you-go service mode will be suitable in this context, referred to as Large Model as a Service (LMaaS). However, the trading and pricing problem is quite complex with heterogeneous and dynamic customer environments, making the pricing optimization problem challenging in seeking on-hand solutions. In this paper, we aim to fill this gap and formulate the LMaaS market trading as a Stackelberg game with two steps. In the first step, we optimize the seller's pricing decision and propose an Iterative Model Pricing (IMP) algorithm that optimizes the prices of large models iteratively by reasoning customers' future rental decisions, which is able to achieve a near-optimal pricing solution. In the second step, we optimize customers' selection decisions by designing a robust selecting and renting (RSR) algorithm, which is guaranteed to be optimal with rigorous theoretical proof. Extensive experiments confirm the effectiveness and robustness of our algorithms.