The widespread use of image acquisition technologies in broadcasting and multimedia services, together with advances in facial recognition, has raised serious privacy concerns. In live news or streaming interviews, audiences may legitimately view participants, yet automated systems can capture and recognize their identities without consent, creating significant legal and ethical risks. Face de-identification, which refers to the process of concealing or replacing personal identifiers, has therefore emerged as an effective means to protect the privacy of facial images. A significant number of methods for face de-identification have been proposed in recent years. In this survey, we provide a comprehensive review of state-of-the-art face de-identification methods, categorized into three levels: pixel-level, representation-level, and semantic-level techniques. We systematically evaluate these methods based on two key criteria, the effectiveness of privacy protection and preservation of image utility, highlighting their advantages and limitations. Our analysis includes qualitative and quantitative comparisons of the main algorithms, demonstrating that deep learning-based approaches, particularly those using Generative Adversarial Networks (GANs) and diffusion models, have achieved significant advancements in balancing privacy and utility. Experimental results reveal that while recent methods demonstrate strong privacy protection, trade-offs remain in visual fidelity and computational complexity. This survey not only summarizes the current landscape but also identifies key challenges and future research directions in face de-identification.
With the rapid development of artificial intelligence, large language models (LLMs) have made remarkable advancements in natural language processing. These models are trained on vast datasets to exhibit powerful language understanding and generation capabilities across various applications, including chatbots, and agents. However, LLMs have revealed a variety of privacy and security issues throughout their life cycle, drawing significant academic and industrial attention. Moreover, the risks faced by LLMs differ significantly from those encountered by traditional language models. Given that current surveys lack a clear taxonomy of unique threat models across diverse scenarios, we emphasize the unique privacy and security threats associated with four specific scenarios: pre-training, fine-tuning, deployment, and LLM-based agents. Addressing the characteristics of each risk, this survey outlines and analyzes potential countermeasures. Research on attack and defense situations can offer feasible research directions, enabling more areas to benefit from LLMs.
This paper presents a robust yet efficient Grey Wolf Optimizer-Support Vector Machine algorithm, termed K-GWO-SVM, for the analysis of ECG signals in smart healthcare systems, aiming to improve classification accuracy and computational efficiency. The proposed model introduces three main contributions: (1) the use of GWO to automatically search for the optimal hyperparameters of SVM tailored to each dataset, (2) a mini-batch strategy guided by K-means clustering to improve the efficiency and convergence of GWO by selecting representative subsets of data, and (3) an enhanced regulation function integrated into GWO that prevents premature convergence by improving the balance between exploration and exploitation. A convergence study is conducted to demonstrate the influence of mini-batch size on both classification accuracy and computational efficiency, showing that using mini-batches as small as 10% of the training data significantly improves computational efficiency without compromising classification accuracy. The K-GWO-SVM framework is evaluated on two benchmark datasets: WESAD for emotion recognition and MIT-BIH Arrhythmia for cardiac classification. The proposed model achieves 99.02% accuracy on WESAD with over a 90% reduction in computational time (10% mini-batch), and 100% accuracy on MIT-BIH with over a 50% reduction in computational time (50% mini-batch), validating its effectiveness, robustness, and suitability for deployment in resource-constrained smart healthcare environments.
Emerging immersive services, such as virtual reality (VR), are featured by multi-modal data streams (audio, video, and haptic). The fusion of distinct service characteristics introduces a significant challenge to wireless communications. This paper considers hybrid VR video and haptic services, where VR video segments and haptic data packets are transmitted in two different time scales. To optimize resource utilization, we employ dynamic time division duplexing (TDD) to address the asymmetry in uplink/downlink (UL/DL) traffic. We formulate a dual time-scale optimization problem that minimizes overall bandwidth while satisfying both the round-trip delay of VR service and the reliability and latency requirements of haptic service. To address this problem, we propose a hierarchical deep reinforcement learning (DRL) framework: 1) deep deterministic policy gradient (DDPG) optimizes total bandwidth at the beginning of each video segment’s transmission; 2) a diffusion-based actor-critic (Diffusion-AC) algorithm determines the UL time resource ratio at each time slot, which is significantly shorter than the transmission duration of a video segment. Simulation results show that the proposed algorithm can reduce the overall bandwidth usage by 20% compared with baseline methods. In addition, it provides greater adaptability across diverse scenarios and converges faster than the baseline methods.
Federated learning (FL) over wireless networks suffers from significant training latency and degraded convergence due to unreliable wireless transmission, especially under blocked propagation environments. Although reconfigurable intelligent surfaces (RISs) can improve communication reliability, existing wireless FL studies rarely characterize the trade-off between learning convergence and communication delay under modulation-dependent transmission errors. In this paper, we consider a wireless FL system operating under RIS-assisted blocked-link propagation scenarios, and focus on adaptive modulation and sub-channel allocation for convergence-latency aware communication design. By characterizing the effect of symbol errors on uploaded local gradients, we derive a convergence-related upper bound that reveals the impact of symbol error rate (SER) on FL loss decay. Based on this result, we formulate a joint convergence-latency optimization problem, which is cast as a mixed-integer nonlinear programming (MINLP) problem, and solve it using a low-complexity hybrid alternating optimization framework. Extensive experiments on MNIST, CIFAR-10, and Speech Commands show that the proposed scheme consistently achieves faster convergence and higher test accuracy than existing adaptive communication schemes, especially in complex tasks and challenging wireless scenarios.
The vision for 6G aims to enhance network capabilities, supporting an intelligent digital ecosystem where artificial intelligence (AI) is a key. However, the expansion of 6G raises critical security and privacy concerns due to the increased integration of IoT devices, edge computing, and AI. This survey provides a comprehensive overview of 6G protocols with a focus on security and privacy, identifying risks that have not been experienced in preceding 5G systems, and presenting mitigation strategies. While many vulnerabilities from earlier generations persist, the introduction of AI/ML introduces novel risks like model inversion and malicious manipulation of AI. Vulnerabilities in emerging personal IoT networks and autonomous vehicles are also underscored, where falsified command signaling or privacy leakage can pose safety and ethical concerns. The survey also discusses the transition toward lattice-based, post-quantum encryption standards, and identifies limitations in current security frameworks and calls for new, dynamic approaches tailored to 6G's complexity. Close collaboration among stakeholders, including governments, industry, and researchers, is indispensable to developing robust standards, secure architectures, and risk assessment frameworks that address AI, quantum threats, and privacy at scale.
Deep Learning (DL) powered by Deep Neural Networks (DNNs) has revolutionized various domains, yet understanding the intricacies of DNN decision-making and learning processes remains a significant challenge. Recent investigations have uncovered an interesting memorization phenomenon in which DNNs tend to memorize specific details from examples rather than learning general patterns, affecting model generalization, security, and privacy. This raises critical questions about the nature of generalization in DNNs and their susceptibility to security breaches. In this survey, we present a systematic framework to organize memorization definitions based on the generalization and security/privacy domains and summarize memorization evaluation methods at both the example and model levels. Through a comprehensive literature review, we explore DNN memorization behaviors and their impacts on security and privacy. We also introduce privacy vulnerabilities caused by memorization and the phenomenon of forgetting and explore its connection with memorization. Furthermore, we spotlight various applications leveraging memorization and forgetting mechanisms, including noisy label learning, privacy preservation, and model enhancement. This survey offers the first-in-kind understanding of memorization in DNNs, providing insights into its challenges and opportunities for enhancing AI development while addressing critical ethical concerns.
Reconfigurable intelligent surface (RIS) has emerged as a promising technology for future 6G wireless communications. However, the passive nature of RIS and the high-dimensional cascaded channels pose significant challenges for channel estimation (CE), particularly in practical scenarios where decomposition dictionaries cannot be predefined. This paper proposes a novel three-stage joint CE and positioning (JCEP) framework for RIS-assisted communication systems. It first performs the initial CE based on a predefined row dictionary that exploits the structural properties of cascaded channels, and then conducts positioning based on the initial CE results. Finally, it refines the CE results by incorporating the positioning output to construct customized column dictionaries. The framework employs a unitary approximate message passing sparse Bayesian learning (UAMP-SBL) based channel estimator that adapts to both initial and CE refinement stages. For positioning, we design a graph attention network (GAT) to achieve robust positioning performance in dynamic environments. Furthermore, in the CE refinement, we introduce a location-aware dictionary design that leverages position priors to reduce computational overhead. Additionally, we employ meta-learning to enable rapid adaptation to new environments. Extensive simulations show that our framework achieves superior performance in CE and positioning accuracy with low complexity.
In this article, we explore federated customization of large models and highlight the key challenges it poses within the federated learning framework. We review several popular large model customization techniques, including full fine-tuning, efficient fine-tuning, prompt engineering, prefix-tuning, knowledge distillation, and retrieval-augmented generation. Then, we discuss how these techniques can be implemented within the federated learning framework. Moreover, we conduct experiments on federated prefix-tuning, which, to the best of our knowledge, is the first trial to apply prefix-tuning in the federated learning setting. The conducted experiments validate its feasibility with performance close to centralized approaches. Further comparison with three other federated customization methods demonstrated its competitive performance, satisfactory efficiency, and consistent robustness.
In recent years, the widespread adoption of location based services (LBS) across diverse mobile applications has accentuated the urgent need to safeguard users' location privacy. To address this concern, the dummy location selection (DLS) al gorithm, grounded in the k-anonymity criterion, has been exten sively studied as a countermeasure against adversaries leveraging query probability information. However, its efficacy diminishes when confronted with the hybrid Retrospect attack introduced in this work. This novel attack method strategically exploits the spatial continuity between adjacent locations, augmented by temporal insights derived from continuous queries spanning an entire movement trajectory. To mitigate this threat, we propose an anti-propagating spatial attack algorithm for individual queries, followed by an anti-Retrospect algorithm designed to intelligently select plausible dummy locations that adhere to stringent privacy requirements. The effectiveness of both our proposed attack and defense mechanisms is rigorously validated through evaluations on two real-world datasets, benchmarked against state-of-the-art methodologies.
Collecting multidimensional user data is essential for personalized services, yet it poses significant privacy risks. While privacy regulations like the GDPR and CPRA advocate for data minimization, attribute correlations can inadvertently amplify unintentional information disclosure, leading to correlation-induced information leakage (CIL). Although data collectors often possess rich prior knowledge of these correlations, existing Local Differential Privacy (LDP) mechanisms are inadequate for effectively leveraging this information to reduce CIL. In this paper, we propose CoP, a coordinated perturbation mechanism designed to mitigate CIL in multidimensional data collection while preserving utility. Unlike traditional LDP approaches, CoP explicitly incorporates prior distribution knowledge to coordinate the perturbation process across attributes. By optimizing the perturbation strategy based on known correlations, CoP achieves a better privacy-utility trade-off. Extensive evaluations across both synthetic and real-world datasets demonstrate that CoP significantly outperforms state-of-the-art LDP mechanisms in reducing disclosure while preserving analytical accuracy.
Federated learning (FL) faces significant challenges from modality heterogeneity, which motivates multimodal federated learning (MFL) to leverage complementary modalities across decentralized clients for improved performance. However, modality imbalance introduces a new attack surface, making MFL more vulnerable to membership inference attacks (MIAs), an issue that remains largely unexplored. In this work, we present the first systematic study of MIAs against MFL and propose a modality-aware attack framework. We show that multimodal models are inherently more susceptible to MIAs due to heterogeneous modality contributions, and existing attacks are suboptimal as they treat multimodal parameters as a whole. By performing MIAs on individual modalities, we find that (i) attacking the dominant modality achieves comparable accuracy with lower overhead, and (ii) different modalities expose distinct membership patterns. To identify members with different patterns, we propose a modality-aware framework that exploits cross-modal performance gaps to adaptively select attack modalities and calibrate inference results. Experiments on three datasets show our approach outperforms baselines across multiple metrics.
With the growing need for collaborative machine learning across institutions holding sensitive data, ensuring data privacy without compromising model performance has become an important challenge. This work introduces secure federated learning algorithms that use encryption and masking techniques to protect the privacy of data during collaborative model training. Three federated learning algorithms were developed: one for vertical federated learning and two combining horizontal and vertical data partitioning. The proposed algorithms are designed such that participating clients communicate only with the server, even when data exchange between clients is required. This exchange occurs through the server with the help of encryption and masking. The performance of the algorithms, evaluated in terms of accuracy and loss, shows competitive results. The accuracy remains unchanged compared to the centralised scenario for the vertical federated learning algorithm and one of the combined federated learning algorithms, and it remains highly competitive with the other combined federated learning algorithm. The privacy analyses conducted as part of this work demonstrate no risk of data leakage ensuring that no party involved can infer sensitive information.
The integration of the Internet of Things (IoT) with Federated Learning (FL) offers a transformative approach to addressing the challenges of massive data processing and privacy preservation in distributed systems. As a decentralized machine learning paradigm, FL enables model training on distributed datasets while safeguarding data privacy, making it well-suited for IoT applications. However, the performance of wireless FL systems is often constrained by limited communication resources and the mobility of participating clients, which can disrupt efficient model training and convergence. In this paper, we propose a novel mobility-aware FL scheduling strategy that leverages interpretable machine learning to enhance resource allocation in wireless networks. A more effective and fair resource allocation strategy can be achieved by dynamically adjusting the weight of the model quality and the communication quality of the training participants. We evaluate the proposed strategy against traditional scheduling methods in both single and multi-base station scenarios. Simulation results reveal that our approach significantly enhances overall learning efficiency by prioritizing high-value local models. Furthermore, for mobile clients, we identify an optimal range of average speed and participant numbers that maximizes the performance of wireless FL systems, offering practical insights for real-world deployments.
Federated learning (FL) is a promising distributed machine learning framework for mobile networks, where an aggregation server produces a global model by aggregating the local models from clients over multiple rounds. Usually, selecting a subset of clients instead of all improves communication efficiency. Therefore, several heuristic strategies for model selection and aggregation have been proposed to improve the global model accuracy or speed up model convergence. However, the effectiveness of these strategies lacks theoretical backing. To investigate this, an empirical comparison study that systematically and quantitatively evaluates existing heuristic strategies is conducted in this paper. Specifically, a FL prototype including nine model selection and aggregation strategies is developed. Experiments with three levels of non-IID data settings on this prototype reveal a trade-off between convergence stability and global model accuracy. Notably, selecting local models with max parameter entropy achieves an excellent balance between model accuracy and convergence stability when handling non-IID data. These findings contribute to a better understanding of heuristic model selection and aggregation strategies, offering valuable guidance for future FL development.
Multiplex graphs represent diverse real-world interactions among entities, where multiple relationship types coexist within the same set of entities. These graphs introduce privacy risks, as data collectors can exploit cross-layer dependencies to infer hidden and sensitive connections. In this work, we propose a C2P-M framework that identifies and protects critical connections while preserving the structural information in multiplex graphs. Unlike conventional methods for single-layer graphs that perturb all edges uniformly, C2P-M selectively protects critical connections, maintaining the analytical usability of the graph. To achieve this, we introduce the multiplex $p$p-cohesion model, which incorporates new score functions that account for both intra-layer and inter-layer dependencies, enabling precise identification of critical connections for each vertex. For privacy protection, our method protects the identified critical connections, leveraging an adaptive Randomized Response (RR) mechanism to ensure $\varepsilon$epsilon-Local Differential Privacy (LDP). We formally prove that C2P-M satisfies $\varepsilon$epsilon-LDP. Extensive experiments on eight real-world multiplex graph datasets demonstrate that C2P-M significantly outperforms baseline privacy-preserving methods, achieving a better privacy-utility trade-off.
This paper investigates the finite-blocklength performance of reconfigurable holographic surfaces (RHS) for ultra-reliable low-latency communication (URLLC). A physics-consistent RHS model is established, and the information-theoretic dispersion is derived in closed form using the Mellin transform method, we derive a closed-form probability density function of the RHS-induced channel gain. Then, we apply second-order Taylor expansion and saddle point approximation to obtain analytical expressions for the mutual information and unconditional information variance, which explicitly capture the amplitude-induced anisotropic variance. Finally, leveraging the Berry-Esseen theorem, the closed-form achievability and converse bounds are established which quantify the impact of RHS design parameters, blocklength n , and average error probability & varepsilon; on the rate. Analytical and simulation results demonstrate that RHS reduces the variance of the effective channel power by 35-50% in required blocklength compared with reconfigurable intelligent surfaces (RIS) under identical resource budgets. The findings identify RHS as a physically scalable and mathematically tractable architecture for short-packet 6G communication.
The proliferation of fifth-generation (5G) communication technologies and the Internet of Things (IoT) has led to massive distributed datasets across numerous user devices. Traditional centralized machine learning methods suffer from large communication overhead and damage of data privacy. Federated Learning (FL) offers a decentralized approach, allowing collaborative local model training and global aggregation, thus preserving data privacy and reducing communication costs. However, implementing FL in wireless networks is complicated due to limited wireless resources, increasing model complexity, and user mobility. In this paper, we address the impact of user mobility in hierarchical federated learning (HFL) systems and propose a resource allocation and user scheduling strategy to minimize energy consumption while maintaining learning performance. We design a comprehensive model that considers user mobility, wireless communication, and computing resources. Using the Lyapunov optimization method, we transform the long-term optimization problem into manageable subproblems, enabling efficient resource allocation and user selection. Our proposed Low Cost Scheduling Algorithm (LCSA) achieves an O(1N) convergence rate, balancing local-edge divergence and improving overall convergence. Experimental results testify that our algorithm significantly reduces energy consumption while achieving high test accuracy compared to baseline methods, highlighting the positive effects of mobility on system performance.