Off-the-shelf RDBMS typically expose only the query execution plan ( QEP ) of an SQL query, without presenting information about representative alternative query plans ( AQP s) considered during plan selection in a user-friendly manner. Providing easy access to representative AQP s is valuable in database education, as it helps learners understand the plan choices made by a query optimizer, one of the several important components related to the topic of relational query processing. In this paper, we present a novel problem called informative plan selection problem ( tips ) which aims to discover a set of k informative AQP s from the underlying plan space so that the plan informativeness of the set is maximized. Specifically, we explore two variants of the problem, batch TIPS and incremental TIPS , to cater to diverse learners. Due to the computational hardness of the problem, we present an approximation algorithm to address it efficiently while providing theoretical guarantees for the results. An extensive experimental study, including feedback from real-world learners and a three-year in-class evaluation of academic outcomes, demonstrates the effectiveness of our solutions for database education.
Online social networks (OSN) such as Twitter, Facebook, etc. have an overall user base of more than 5 billion as of today. Traditional centralized OSN users’ data and content are stored in centralized servers, which has the risk of data leakage and privacy violation. Distributed OSN (DOSN) address the single-point-of-failure and user data privacy concerns faced by centralized OSN by enabling the operation of network infrastructures and services without centralized ownership or control. However, DOSN face privacy protection issues. The group key agreement (GKA) is an important method to construct secure channels to protect the secure communication of network group members. Asymmetric GKA methods allow external members to securely communicate with group members without having to join the group. However, it is worth noting that existing AGKA schemes rely on bilinear pairing, resulting in a high computational overhead. Meanwhile, considering scenarios where external users join, or group members leave, we design a dynamic asymmetric group key agreement (DAGKAwP) scheme based on Schnorr batch multi-signature that does not depend on bilinear pairing. During the key generation phase, the members generate self-authenticating public-private key pairs to resist malicious public key attacks. For group key agreement, new hash functions are embedded in Schnorr signatures to generate aggregated public keys, and this scheme supports external member addition and internal member exit. In group message encryption, sender anonymity and message non-repudiation are realized. The security comparison with related DAGKA schemes reveals that the DAGKAwP scheme offers more comprehensive security. Performance evaluations suggest that this scheme offers computational efficiency and lower communication overhead than related Dynamic AGKA schemes.
The space-air-ground integrated network has emerged as a critical enabler to achieve high-capacity 6 G communications. However, frequent handovers between satellites and gateways, along with unbalanced gateway traffic, significantly degrade the overall transmission capacity. To address these issues, this paper proposes a balanced satellite-ground scheduling architecture based on the analytic hierarchy process, called AHP-BSA. First, three spatiotemporal parameters (interconnection time $R(t)$, transmission capacity $C(t)$, and propagation delay $D_{p}$) are defined. The rationality and consistency of AHP-BSA is proved using the spatiotemporal parameters. In the parameter calculation phase of AHP-BSA, this paper further proves the relationship between $R(t)$ and outage probability, as well as between $D_{p}$ and transmission efficiency. Then, a parameter optimization algorithm is designed in AHP-BSA. These spatiotemporal parameters are normalized, weighted, and incorporated into the parameter optimization algorithm. Through stability-aware adjustment, it adjusts scheduling decisions in response to link dynamics and real-time fluctuations. Simulation results confirm that AHP-BSA outperforms existing methods in transmission capacity, time complexity, and long-term traffic balance.
With the rapid advancement and widespread application of the graph neural network (GNN), the collaborative graph learning, in which multiple parties collaboratively construct a GNN model using their respective graph data, has attracted increasing attention. However, this paradigm also raises significant privacy concerns, as both nodes and edges may contain sensitive personal information, while existing privacy-preserving schemes often come at the cost of degraded model performance or substantial system overhead. Therefore, this paper proposes an efficient and privacy-preserving collaborative learning framework on vertically partitioned graph data, dubbed Plog. Specifically, we first design a decomposition algorithm to split the sparse adjacency matrix into the summation of multiple independent permutations, which are lightweight, parallelizable, and well-suited for secure multi-party computation. Building on this, a weighted oblivious batch permutation protocol is carefully customized based on correlated randomness to securely and efficiently compute adjacency matrix multiplications, addressing the core efficiency bottleneck in GNN inference and training. The selective security of Plog is formally verified under the ideal-real paradigm. Extensive experimental results on three real-world datasets demonstrate that compared to the state-of-the-art scheme, Plog can reduce online communication rounds by 46% and achieve a 1.73 & times; speedup in the overall inference and training time.
Traditional Proof-of-Work (PoW) consensus secures blockchains by incurring significant energy and resource wastage on solving valueless puzzles. The Proof-of-Useful-Work (PoUW) paradigm emerges to repurpose this effort for tasks of societal or scientific value. However, existing PoUW approaches face a critical trade-off between the work's practical utility and its verifiability, as efficiently validating complex computations without full re execution remains a major open problem. This paper proposes Proof-of-Useful-Federated-Work (PoUFW), a novel consensus protocol that resolves this dilemma by recasting the “work” as Federated Learning (FL) model aggregation. Specifically, PoUFW mandates that miners not only perform the aggregation, but also generate a cryptographic proof of its correctness. To achieve this, we propose a novel proof mechanism that builds upon the Kate Zaverucha-Goldberg (KZG) polynomial commitment scheme to generate a concise, non-interactive proof. This mechanism enables validators to rapidly verify the aggregation's integrity without re-executing the computation or accessing the underlying model. Consequently, PoUFW transforms the computational overhead of consensus into valuable, verifiable machine learning outcomes, while simultaneously establishing a decentralized, incentivized framework for FL. Rigorous theoretical analysis substantiates the protocol's security and correctness. Extensive simulations further demonstrate that PoUFW achieves high verification efficiency alongside robust security, confirming its practical feasibility and scalability. This work provides a robust theoretical and empirical foundation for building blockchain systems with improved resource efficiency and societal value.
Vehicle trust management is closely related to the identity security of intra-domain vehicles and has garnered widespread attention. However, there is a significant conflict between vehicle trust and identity privacy, which has led to the emergence of novel trust link attack issues. To address the problem, this paper proposes a vehicle identity trust management algorithm based on differential privacy, called DITDP. In DITDP, we first design an identity trust evaluation model based on D-S evidence, which is used to quantify vehicle trust values, including the direct trust, the recommended trust, and the aggregated trust. Then, we design the dynamic trust evaluation and identity privacy protection modules in DITDP. The dynamic trust evaluation module corrects conflict between vehicle trust components and can dynamically update vehicle trust values in the time domain. The identity privacy protection module achieves indistinguishable trust while ensuring the availability of trust value by using Differential privacy. Finally, the simulations verify that the DITDP algorithm performs well in terms of identity protection ability, trust availability, and other aspects.
The massive growth of data has brought vigorous vitality to the Internet of Things (IoT). It has also brought new challenges, such as confidentiality privacy protection and redundant data transmission. Concerning this regard, data aggregation serves as an efficient technique to minimize the transmission frequency among massive objects in smart grid (SG). By aggregating a large amount of encrypted data from distributed IoT devices, the proposed framework enables efficient and privacy-preserving analytics at the edge. With the advent of the postquantum era, a good aggregation scheme must provide quantum resistance while ensuring the secure aggregation of ciphertext power data. However, the excessive overhead limits anti-quantum algorithms from being widely used in SG, where resource devices are limited. Therefore, it is an important part of the current private data security aggregation technology to find a low cost and lightweight inverse quantum algorithm to achieve user data security aggregation. In this article, we propose an improved NTRU-based cryptosystem with multidimensional coding, referred to as multidimensional coding NTRU (MC-NTRU), and use the lattice batch signature technique, which improves the efficiency of the scheme while satisfying the anti-quantum attack. Based on these, we design the multidimensional auditable lattice-based privacy-preserving data aggregation scheme (MA-PPDA) for privacy data on resource-limited IoT devices, such as SGs. In addition to this, the scheme achieves fault tolerance (FT) of the scheme by adding zeros and random numbers to the user data. The comparative study against the existing approaches demonstrates that the proposed scheme not only adheres to critical security aspects, including user privacy, data confidentiality, integrity, and authenticity, but also decreases both communication and computational burdens on the system. This makes our scheme particularly apt for IoT environments characterized by constrained device resources.
Federated learning enables decentralized clients to collaboratively train models without sharing local data. However, heterogeneous client distributions often induce client drift and hinder convergence. This paper proposes FedCDWA, a decoupled hierarchical federated distillation framework. FedCDWA decouples client-side personalized distillation from server-side mutual distillation to mitigate distillation-induced optimization conflicts. It further adopts Hierarchical Wasserstein Aggregation to aggregate prototypes without restrictive parametric assumptions while preserving intra-class structure and inter-class geometry. To achieve finer-grained feature alignment, Prototype–Variance Dual Alignment matches feature means and variances in the feature space. We prove convergence guarantees for FedCDWA. Experiments on three datasets demonstrate that FedCDWA consistently improves both global and personalized accuracy across heterogeneity levels, with smaller performance degradation under more severe heterogeneity.
Digital Twins (DT) are increasingly becoming a core infrastructure for large-scale Unmanned Aerial Vehicle (UAV) swarm orchestration in low-altitude intelligent networks. By constructing a physical-virtual closed loop, UAVs frequently need to send multicast command messages based on DTs to numerous UAVs via insecure wireless links. Signcryption ensures secure data exchange by simultaneously guaranteeing confidentiality and authentication, thus better matching UAV performance. However, most state-of-the-art signcryption schemes optimize the receiver through aggregation or batch verification, while the cost of sending the same command to multiple receivers still increases rapidly with swarm size, creating a bottleneck of receiver-friendly but sender-costly. To overcome this limitation, we propose a digital twin-enabled secure multicast signcryption protocol (DTSM), which enables multicast messages to be delivered in signcryption form, allowing only designated UAVs to decrypt, authenticate, and execute them. We first construct a three-layer digital twin architecture, which further illustrates the role of DTs in collaborative tasks within UAV swarms. Then, we design a group key distribution mechanism based on polynomial interpolation and a secure multicast signcryption protocol, embedding the multicast key into a Lagrange polynomial, allowing each designated UAV to recover the key locally. Formal proofs under the random oracle model and protocol formal verification using Scyther demonstrate that the proposed DTSM achieves confidentiality, unforgeability, authentication, and resistance to replay attacks, while informal analysis indicates that it also supports conditional anonymity and traceability. Evaluation results show that the proposed DTSM significantly reduces the communication and computational overhead at the sending side compared with related schemes, making it suitable for large-scale receiving scenarios.
Similarity-based vector search facilitates many important applications such as search and recommendation but is limited by the memory capacity and bandwidth of a single machine due to large datasets and intensive data read. In this paper, we present CoTra, a system that scales up vector search for distributed execution. We observe a tension between computation and communication efficiency, which is the main challenge for good scalability, i.e., handling the local vectors on each machine independently blows up computation as the pruning power of vector index is not fully utilized, while running a global index over all machines introduces rich data dependencies and thus extensive communication. To resolve such tension, we leverage the fact that vector search is approximate in nature and robust to asynchronous execution. In particular, we run collaborative vector search over the machines with algorithm-system co-designs including clustering-based data partitioning to reduce communication, asynchronous execution to avoid communication stall, and task push to reduce network traffic. To make collaborative search efficient, we introduce a suite of system optimizations including task scheduling, communication batching, and storage format. We evaluate CoTra on real datasets and compare with four baselines. The results show that when using 16 machines, the query throughput of CoTra scales to 9.8-13.4x over a single machine and is 2.12-3.58x of the best-performing baseline at 0.95 recall@10.
The application of vehicular ad hoc networks (VANETs) in intelligent transportation systems (ITS) enables the collection and dissemination of traffic event information, thereby helping to improve transportation efficiency and security. However, due to the open nature of VANETs, authenticated nodes are not completely trustworthy, posing challenges in determining the authenticity of events. Existing studies overlook the impact of node density on event validation performance, resulting in insufficient reliability in low-node-density scenarios and a lack of detailed quantitative assessments of attack resistance. To address this issue, this article proposes an attack-resistant and adaptive node density event validation scheme called AR-NDA. In AR-NDA, the roadside unit (RSU) can assess the trustworthiness of opinions on event authenticity from various nodes using the designed trustworthiness assessment model. Based on the assessment results, the RSU can reasonably leverage the trustworthiness of nodes’ opinions while considering other relevant factors to ultimately infer the authenticity of the event. The effectiveness of the AR-NDA scheme is validated through extensive experiments under varying node densities and proportions of malicious nodes. The scheme is capable of coping with attacks from malicious nodes, maintaining an event validation accuracy of over 0.8 even under the extreme conditions of low node density and high proportions of malicious nodes.
Shuffler-based differential privacy (shuffle-DP) is a privacy paradigm providing high utility by involving a shuffler to permute noisy report from users. Existing shuffle-DP protocols mainly focus on the design of shuffler-based categorical frequency oracle (SCFO) for frequency estimation on categorical data. However, numerical data is a more prevalent type and many real-world applications depend on the estimation of data distribution with ordinal nature. In this paper, we study the distribution estimation under pure shuffle model, which is a prevalent shuffle-DP framework without strong security assumptions. We initially attempt to transplant existing SCFOs and the naïve distribution recovery technique to this task, and demonstrate that these baseline protocols cannot simultaneously achieve outstanding performance in three metrics: 1) utility, 2) message complexity; and 3) robustness to data poisoning attacks. Therefore, we further propose a novel single-message adaptive shuffler-based piecewise (ASP) protocol with high utility and robustness. In ASP, we first develop a randomizer by parameter optimization using our proposed tighter bound of mutual information. We also design an Expectation Maximization with Adaptive Smoothing (EMAS) algorithm to accurately recover distribution with enhanced robustness. To quantify robustness, we propose a new evaluation framework to examine robustness under different attack targets, enabling us to comprehensively understand the protocol resilience under various adversarial scenarios. Extensive experiments demonstrate that ASP outperforms baseline protocols in all three metrics. Especially under small ε values, ASP achieves an order of magnitude improvement in utility with minimal message complexity, and exhibits over threefold robustness compared to baseline methods.
Shared-nothing geo-distributed SQL databases, such as CockroachDB, are increasingly vital for enterprise applications requiring data resilience and locality. However, we encountered significant performance degradation at the customer side, especially when their deployments span multiple data centers over a Wide Area Network (WAN). Our investigation identifies the bottleneck in the performance of the Distributed Hash Join (Dist-HJ) algorithm, which is contingent upon a crucial balance between communication overhead and computational load. This balance is severely disrupted when processing skewed data from real-world customer workloads, leading to the observed performance decline. To tackle this challenge, we introduce Bala-Join, an adaptive solution to balance the computation and network load in Dist-HJ execution. Our approach consists of the Balanced Partition and Partial Replication (BPPR) algorithm and a distributed online skewed join key detector. The former achieves balanced redistribution of skewed data through a multicast mechanism to improve computational performance and reduce network overhead. The latter provides real-time skewed join key information tailored to BPPR. Furthermore, an Active-Signaling and Asynchronous-Pulling (ASAP) mechanism is incorporated to enable efficient, real-time synchronization between the detector and the redistribution process with minimal overhead. Empirical study shows that Bala-Join outperforms the popular Dist-HJ solutions, increasing throughput by 25%-61%.
Trusted Execution Environments (TEEs) provide an efficient foundation for ciphertext computation through hardware-enforced isolated execution, enabling secure multi-party data processing. Integrating TEEs with cryptographic techniques allows sensitive data to be processed without exposing the plaintext, thereby supporting collaborative computation in untrusted environments. These lightweight and efficient security mechanisms have driven widespread interest and research into TEE-based ciphertext computation schemes. However, existing studies often overlook the impact of participating entity behaviors on overall security, lacking a systematic characterization of the complex threat landscape introduced by TEEs. To address this gap, we present a survey of TEE-assisted computation over ciphertext from this behavioral perspective, providing an in-depth analysis of the limitations of current trust models. To comprehensively characterize the complex threats introduced by TEEs, we propose a security framework driven by these entity behaviors. By defining the capabilities and goals of each entity, this framework systematizes the threat models of TEE-assisted ciphertext computation systems into three core trust assumptions: the honest, the semi-honest, and the malicious assumption. Based on this framework, we systematically classify TEE-assisted ciphertext computation protocols across typical scenarios. We evaluate their design features and security under different trust assumptions, and briefly discuss potential open challenges in this domain.
Federated learning (FL) enables the collaborative model training while preserving data privacy, making it particularly attractive for large-scale Internet of Things (IoT) systems. However, in practical deployments, data collected by distributed clients is often nonindependent and identically distributed (Non-IID), which amplifies the vulnerability of FL to poisoning attacks. Among them, label-flipping attacks (LFA) are especially stealthy, as they can induce targeted misclassification without noticeably affecting overall accuracy, posing serious risks to safety-critical IoT applications. In this article, we propose NeuDFL, a lightweight and robust defense framework against LFA under Non-IID settings. Unlike many existing defenses that primarily rely on the gradient analysis over the full-model update space or auxiliary clean datasets, NeuDFL exploits lightweight classwise parameter statistics extracted from the final fully connected layer (FCL). By leveraging the cumulative and task-aligned nature of model parameters, NeuDFL enables reliable identification of attacked classes and filters malicious clients via an adaptive statistical threshold, improving robustness to data heterogeneity while incurring low computational overhead. Extensive experiments on multiple datasets demonstrate that NeuDFL offers an effective and efficient defense against LFA, providing a robust solution for FL in complex real-world environments.
Modern Vision-Language Models (VLMs) pose significant individual-level privacy risks by linking fragmented multimodal data to identifiable individuals through hierarchical chain-of-thought reasoning. However, existing privacy benchmarks remain structurally insufficient for this threat, as they primarily evaluate privacy perception while failing to address the more critical risk of privacy reasoning: a VLM’s ability to infer and link distributed information to construct individual profiles. To address this gap, we propose MultiPriv, the first benchmark designed to systematically evaluate individual-level privacy reasoning in VLMs. We introduce the Privacy Perception and Reasoning (PPR) framework and construct a bilingual multimodal dataset with synthetic individual profiles, where identifiers (e.g., faces, names) are linked to sensitive attributes. This design enables nine challenging tasks spanning attribute detection, cross-image re-identification, and chained inference. We conduct a large-scale evaluation of over 50 open-source and commercial VLMs. Our analysis shows that 60\% of widely used VLMs can perform individual-level privacy reasoning with up to 80\% accuracy, posing a significant threat to personal privacy. MultiPriv provides a foundation for developing and assessing privacy-preserving VLMs.
Deploying outsourced graph neural network (GNN) inference services in the cloud is gaining widespread application across various fields, such as fraud detection and social network analysis. Cloud servers utilize outsourced model to analyze the graph data of data owners, enabling data owners to enjoy high-quality GNN inference services. However, this approach leads to privacy concerns regarding GNN models, graph data and inference results. To address the privacy issues, some privacy-preserving GNN inference schemes have been proposed. But the existing schemes are only applicable to graph convolutional network and not to graph attention network (GAT) with stronger expressive power. Therefore, we propose a secure GAT inference scheme (SecGAT) for outsourcing scenarios. First, we represent the Beaver triple-based multiplication process as a two-phase multiplication, which allows us to combine specific algorithms to optimize the communication overhead. Then, we design a graph data encryption method to protect the privacy of outsourced graph data. Finally, we propose a series of customized algorithms for secure GAT inference. Based on the proposed building blocks, we construct a complete GAT inference process. Rigorous security analysis and extensive evaluations demonstrate the effectiveness of our scheme. By comparing the core algorithms, our scheme can improve computational efficiency by more than 20% and reduce communication overhead by 20%-40% compared to existing schemes.
As the field of personal lending grows exponentially, the risk of credit default has become an issue of growing significance. Meanwhile, responsible and fair personal credit evaluation in safe environments is a key area of research in the field of financial technologies. Even the current strategies have a few problems in the real-world application: the imbalanced classes of credit data creates low detectability of delinquent cases, a centralized scoring solution does not provide credible verification methods, which makes them easily manipulated, and machine learning model decision-making is not typically transparent.In resolving these problems, this study introduces a verifiable and interpretable model of personal credit assessment that combines ensemble learning, blockchain-based data notarization, and zero-knowledge proofing. To begin with, a credit assessment model using soft-voting ensemble, with SMOTE oversampling is developed, to integrate the complementary advantages of random forest and gradient boosting classifier and thus would be useful in enhancing the detection of delinquent users. Second, a zero-knowledge proof system developed under the Groth16 protocol and integrated with blockchain notarization to make the verifiability and the immutability of the scoring outcomes accessible to protect data privacy. Third, the entire SHAP feature attribution vectors are hashed with the hash-digest algorithm SHA-256 and stored on-chain, providing end-to-end auditability over scoring results to the underlying rationale decision and raising decision transparency and regulatory acceptability.In the experiments conducted on UCI Germany credit dataset, the suggested method proves to be more efficient than the baseline models in major measurements including full term and Recall. The Recall was improved by approximately 30 percentage points and the AUC increased from 0.8938 to 0.9024. Both on-chain audit and zero-knowledge proof verification rates are 100%, and SHAP notarization completion is 100%, which show general benefits of the framework in predictive accuracy, privacy protection, model explainability, and result verifiability. The work presents a practicable technical route of the personal credit assessment frameworks within the financial regulation realms.
Local differential privacy (LDP) provides strict privacy guarantee in a distributed environment. Recent studies demonstrated that LDP protocols are vulnerable to data poisoning attacks where an attacker can manipulate the perturbed result on the local side and send bogus data to skew the final estimate on the server. Unfortunately, existing attack detections do not create an effective attack indicator and rely on particular characteristics of LDP protocols. As a result, they typically exhibit limited detection performance. In this paper, we use log-likelihood as the attack indicator and propose a chain-style detection to enhance the detection effectiveness, in which the attack impact could propagate along the chain and exhibit clear anomaly signal even under stealthy attack scenarios. The experimental results show that our detection consistently outperforms the existing methods. Using four datasets containing categorical and numerical data separately, our detection achieves an F1 score exceeding 96% in most cases. It even remains above 0.9 under stealthy attack settings, outperforming the state-of-the-art detection by up to 0.25.