
This paper proposes a UAV-assisted cooperative perception framework for vehicular networks, ensuring the timely transmission of image feature extraction results from vehicles to the RSU, thereby enabling cooperative perception such as HD map construction and object-level tracking in smart cities. We analyze the success probability that vehicles can complete image feature extraction and transmit the result before leaving the RSU’s communication range with the assistance of UAVs. The simulation results demonstrate that the proposed algorithm outperforms benchmark schemes in terms of success probability.
State-of-the-art deep learning models rely on large GPU clusters and various parallelism strategies, which in turn depend on collective communication (CC) operators to synchronize data. While vendor libraries (e.g., NCCL, RCCL) provide standard CC algorithms, they often suffer from bandwidth bottlenecks in imbalanced topologies. Recent synthesis-based methods improve performance but face three key limitations: poor scalability due to the combinatorial explosion of scheduling space, lack of support for multistage execution, and suboptimal communication throughput. We propose Canvas, a scalable and near-optimal CC scheduling framework that addresses these challenges. Canvas introduces: (1) Hierarchical synthesis to decompose the global scheduling problem into tractable subproblems for scalability. (2) Collective decomposition to enable structured, multi-stage algorithm generation. (3) Cross-micro-batch pipeline scheduling to parallelize communication across micro-batches and maximize link utilization. Evaluations show that Canvas achieves up to 1.98× bandwidth speedup over TACCL and 3.56× over TE-CCL, and synthesizes algorithms for 512-GPU topologies within 1.77 hours, whereas TACCL fails to produce results within 24 hours.
In recent years, significant progress has been made towards scalable network control-plane verification. Yet, operators are still hesitant to deploy such systems. We argue that this reluctance is in part due to a semantic gap between operators reasoning about routing states and verifiers exploring the space of environments. Indeed, operators express the specification in terms of behavior of routing states, while verifiers usually rely on solvers to find specific environments that violate the specification. This semantic gap prevents users from guiding these solvers to directly explore routing states that violate the specification, or to search for states that are most relevant or likely.In this paper, we present a new approach for flexible control-plane verification. Instead of relying on rigid off-the-shelf solvers, we design a novel backtracking algorithm to directly explore the space of routing states. This enables users to guide the exploration according to the specification and domain-specific knowledge from operators. This algorithm paves the way for novel use cases, ranging from finding relevant (e.g., likely) counterexamples to performing verification of probabilistic specifications.
Mobile-cloud collaborative Convolutional Neural Network (CNN) inference enables the efficient execution of CNN models by offloading partial inference workloads from mobile devices to the cloud. Although model partitioning for collaborative inference has been extensively studied, most existing approaches assume reliable mobile-cloud transmission, which often breaks down in real-world wireless environments with packet loss. In such scenarios, incomplete feature transmission can result in a significant drop in inference accuracy. In this paper, a joint scheduling approach is proposed to address this challenge, leveraging a search-based algorithm to determine both the model partition layer and redundancy level, to balance inference accuracy and latency under packet loss conditions. The proposed method is evaluated in a real-world mobile-cloud environment. Results show that it reduces the inference latency by up to 30.3% compared to the non-redundant baseline with the same accuracy threshold.
This paper offers a solution to generate concrete goals for automating the fulfillment of user intents in the computing continuum. The computing continuum guarantees a flexible infrastructure for services at the cost of more complex handling. Our proposed method innovates the state of the art, helping build performative automated strategies through the translation of service owner intents into concrete targets. We improve on existing intent-based systems by offering support for multi-domain infrastructures. Furthermore, we go beyond current computing continuum management solutions, offering full automation by generating concrete targets for the continuum of automated agents. We achieve that through a multi-agent system built on Large Language Models (LLMs) that translates high-level business intents into executable Reinforcement Learning (RL) environments. By leveraging infrastructure representations in the form of Knowledge Graphs, the framework identifies which system components require adaptation and estimates the target metric values needed to fulfill the intent. We evaluate it on a realistic use case with promising results. We can deploy a fully working RL agent to manage network and computing resources, achieving a success rate higher than 80% after preliminary training.
Time-sensitive networking (TSN) has been widely used in industrial automation and automotive applications by precisely opening and closing the gates of packet queues. However, both clock drift and network congestion will likely result in timing misalignment when the clock synchronization protocol gPTP (802.1 AS) is used. This timing misalignment may, in turn, cause failure in forwarding packets at scheduled times and hence unexpected delay jitters, or even miss application deadlines. To address this acute problem, we propose a novel TSN scheduling algorithm, called SDT-TSN (Synchronization-Deviation-Tolerant TSN), to ensure that packets can still arrive on time and be transmitted deterministically even in the presence of inexact time synchronization. First, we formalize the linear constraint model of flow scheduling to maximize the tolerance of inexact time synchronization. Then, we propose an optimal algorithm based on SMT (Satisfiability Modulo Theories) and a fast heuristic algorithm to solve the packet scheduling problem under inexact network synchronization. SDT-TSN is the first to derive the maximum tolerable time-synchronization deviation. Finally, we evaluate SDT-TSN, demonstrating its capability of eliminating packet-forwarding failures due to the commonly-used/assumed constant time-synchronization deviation and increasing the tolerable synchronization deviation from 140µs to 480µs.
With the rapid development of cloud-network convergence, alert data processing has become a critical task in intelligent cloud-network operations and maintenance. Enterprise threat detection systems serve as a vital defense against cyberattacks. However, due to well-known challenges such as massive unlabeled data, false alerts, and lack of interpretability, threat detection remains a formidable challenge. To address these issues, we propose the Enterprise Threat Detection Explainer (ETDE), an interpretable graph neural network (GNN)-based framework designed to automatically identify attackers from intrusion prevention system (IPS) alerts. ETDE extracts relevant subgraph structures and features from security attribute graphs to model attack patterns, while providing comprehensive explanations to help security teams rapidly pinpoint threats and accelerate incident response. Evaluations on real-world operational data demonstrate that ETDE significantly reduces the manual workload for security analysts, enhancing both efficiency and reliability in threat detection.
In-band network telemetry (INT) enables real-time network monitoring by embedding telemetry data into packets. The advent of programmable switches further enhances the flexibility of INT by enabling dynamic customization of telemetry collection at the hardware level. However, the practical application of INT is hampered by significant challenges arising from data noise (due to packet loss, delay, and measurement inaccuracies) and network dynamics (such as changing INT paths and feature requirements). These issues severely degrade the performance of machine learning models used for analyzing INT data, hindering the accurate capture of spatio-temporal network characteristics. This paper presents Planner, a novel generative graph learning framework designed to address these limitations. Planner enables the collection of network features at various levels of granularity on programmable switches. Crucially, it constructs dynamic graphs representing evolving INT paths and employs a hybrid Graph Neural Network (GNN) and Recurrent Neural Network (RNN) architecture to effectively learn spatial and temporal dependencies. Furthermore, Planner incorporates variational inference to generate robust latent representations, mitigating the detrimental effects of noise and instability in INT data. We have implemented a testbed prototype of Planner using Intel Tofino ASIC switches. Extensive experiments demonstrate the performance superiority and robustness of Planner over the baseline methods, achieving a 23.2% improvement in F1 score.
The vastness of the IPv6 address space has led to the common practice of allocating prefixes to end users rather than individual addresses. Users can assign these prefixes as a single subnet or divide them into multiple subnets for different purposes. Allocation strategies vary significantly in terms of prefix granularity, and identifying the actual granularity of subnet assignments is crucial for improving measurement efficiency, accuracy, and for better IPv6 network management. However, no existing method can discover IPv6 subnets at an Internet-Wide scale.To this end, we propose SubRecon, an Internet-Wide IPv6 subnet discovery system. SubRecon consists of two key phases: subnet delimitation and target expansion. In the subnet delimitation phase, we perform a systematic scan across the entire IPv6 address space without relying on any existing seed dataset. This phase adopts a top-down approach, probing prefixes from the shortest to the longest in a hierarchical manner. At each level, we recursively refine prefixes and discard sub-prefixes that do not meet the convergence condition. This pruning strategy eliminates redundant probes in unallocated regions, significantly reducing the search space and improving probing efficiency. To further improve coverage, the target expansion phase leverages the active address dataset as an auxiliary input. It identifies active addresses not covered by previously discovered subnets, expands them into new candidate target prefixes, and performs another round of subnet delimitation. This helps enhance the completeness and coverage of the final discovered subnet set. Experimental results show that SubRecon discovers 8,381,974 IPv6 subnets across 14,147 autonomous systems, and the resulting subnet list can serve as high-quality input for topology discovery. Additionally, during the subnet discovery process, SubRecon identifies a large number of last-hop router interfaces, discovering approximately 10 million more than the current state-of-the-art methods.
Power line communication (PLC) is emerging as the preferred connectivity solution to enlarge WiFi’s communication range for mobile clients. However, we have to consider the mobility in power-line assisted WiFi for a seamless mobile experience and wide connectivity. To this end, in this paper, we exploit the power line backbone to achieve seamless mobility and wide connectivity for mobile clients in a power line-assisted wireless transit network. To do so, we introduce TurboGrid that characterizes the power-line backbone and allows the network infrastructure to make performance-impacting decisions on roaming between PLC adapters. We first characterize the power line channel over time and space in the powerline infrastructure. Then, we show that the existing WiFi handover approaches cannot be directly applied when integrating power lines to assist WiFi communication. Therefore, we propose to integrate the link quality on the power line backbone with link quality over the air for seamless mobile client switching. The systematic evaluation and implementation of commodity WiFi AP and PLC adapters show the efficacy of our TurboGrid during the mobile client’s mobility.
With the increasing deployment of Low Earth Orbit (LEO) satellites (e.g., Starlink), rate adaptation has been widely applied in satellite communication systems to rapidly adapt changing satellite-ground channel. However, the state-of-the-art rate adaptation solutions still do not approach the Shannon capacity, since they always operates in discrete steps (e.g., switching between fixed modulation orders and coding rates), leading to abrupt performance deterioration when the channel fluctuates between two thresholds. We propose a new plug-and-play rate adaptation, called ERA-LEO integrated to satellite communication systems without any extra hardware modification. For adapting smoothly to changes in channels, we first design a probabilistic constellation shaping enabled system to provide a continuous and more granular adjustment of the constellation by tuning the probability distribution of symbols. Second, to avoid the frequent feedback, used to select a target probability distribution based on the channel condition, we establish a theoretical model to demonstrate that the SNR of the return link can reliably predict that of the forward link, and present a lightweight prediction mechanism based on neural network. We also implement ERA-LEO prototype using FPGA-based software defined radio platforms. The extensive evaluations demonstrate ERA-LEO achieves an average gain of 2.19 dB compared to state-of-the-art baselines and a 43.08% improvement in network throughput.
As industries increasingly embrace digital and intelligent transformation, enterprises face significant challenges in the containerization upgrades for their Artificial Intelligence(AI) applications and dynamic service migration and management in Cloud-Edge continuum(CEC). This paper presents CloudSkin, an innovative platform designed to realize streamlined, seamless and intelligent service migration in Cloud-Edge Continuum by integrating advanced containerization techniques and AI-driven orchestration capabilities. Our approach, leveraging intelligent algorithms for service migration, can seamlessly transit services between cloud and edge environments, ensuring optimised resource allocation and reducing service latency to assure quality of service(QoS). CloudSkin has been enabled in a Mobility Usecase, empowering Cellnex businesses undergoing digital transformation to achieve higher operational efficiency. The experimental results show that compared to traditional reactive service migration, using intelligent proactive service migration can provide better migration detection, improving F1-Score up to 23.5%, and reducing 28.9% the service running time where the service latency violates SLA.
Modern data-intensive applications require adaptive, goal-driven infrastructures, yet the growing complexity of the compute continuum makes it increasingly difficult for existing approaches to autonomously manage data placement in accordance with evolving user intents and real-time system states. In this paper, we propose a modular Intent-Based Agentic AI Framework that can accurately capture high-level user objectives, reason over distributed resource conditions, and automate data placement in a scalable and interpretable manner. Unlike static rule-based systems, our framework supports interactive user-agent interactions and adaptive forecasting to recommend and apply optimal storage and caching strategies across distributed cloud-edge environments. Our prototype implementation showed an accurate response to user inputs at 88%, and achieved objectives in around 30 seconds. This implementation and results demonstrate the feasibility and efficacy of the proposed solution.
Lightweight anonymity protocols provide a well-balanced anonymity and performance by encrypting and decrypting only packet headers under the active and local adversary threat model. Among them, PHI and dPHI are promising in universally providing relationship anonymity. However, when overlaid onto IP, they are susceptible to the router skipping attack, where honest routers are skipped by malicious routers. Although this attack poses a significant threat to anonymity, its prevention is challenging due to the lack of path integrity in these protocols. To address this limitation, this paper integrates path validation into dPHI. This integration is non-trivial, as anonymity and path validation are inherently contradictory requirements. This paper designs and implements pPHI, a novel protocol, and analyzes pPHI in terms of security and performance.
Multi-tenant GPU clusters are designed to concurrently run multiple distributed ML training workloads. However, frequent data transfers among GPUs via collective operations can slow down training, as collectives from different tenants compete for network bandwidth. Recent work (e.g., CASSINI) has considered collective coordination to prevent network contention, but primarily focused on static job-level optimizations at deployment time, oblivious to runtime network conditions and the specific traffic pattern of each workload. In this paper, we present Symphony, an application-layer solution that dynamically coordinates collective operations across tenants at runtime. Symphony integrates seamlessly with existing clusters with minimal modifications to the collective communication library and includes a lightweight online scheduling mechanism that requires no advance information about the workloads or their collectives. We evaluate Symphony using both a real GPU testbed implementation and trace-driven simulations. Specifically, using realistic ML workloads in our testbed, we observe improvements of up to 13.2% in average communication time and 9.6% in training time compared to state-of-the-art solutions.
Decentralized satellite federated learning (DSFL) enables LEO satellites to collaboratively train models without relying on terrestrial infrastructures. In DSFL, concurrent model updates for multiple services increase the amount of computation on the satellite, which can lead to longer total training time. This paper proposes a computation-efficient DSFL framework to construct the overlay network for services to minimize the total training time. Simulation results demonstrate that the proposed scheme outperforms the baseline algorithms in terms of total training time.
Placing workloads across the Edge and Cloud can be performed through containerization. However, conditions in Edge nodes are very different from Cloud, as environmental factors such as temperature, humidity, voltage provisioning or dust, heavily affect execution performance and device health. Reliability is an important factor when deciding to deploy such load onto Edge devices, indicating their capability to achieve a desired Quality of Service (QoS). For this, research on Edge computing must focus on how to monitor, estimate and manage devices, in a distributed, autonomous and reliable manner. Our current efforts on performance analysis for devices under "wild conditions" are moving towards integrating reliability into orchestrator systems. Here we present our vision and roadmap for expanding Cloud orchestration towards the Edge with technologies capable of providing knowledge about node environmental conditions and mitigate its impact. Through Out-Of-Band telemetry, we can retrieve temperature and power consumption variables from node components, indicating its health and estimating its reliability given external stress factors. In particular, using intent-based orchestration for containerized platforms, such factors can be used for enforcing reliability as a key-performance indicator. The current work in progress focuses on industrial and commercial scenarios, e.g., Edge computing for urban mobility, with road-side nodes performing AI-based Video-Analytics, exposed to uncontrolled weather conditions. The principal objective is to achieve an Edge network orchestration that takes into account node health in an automatic manner for reducing operational costs such as energy consumption, device repair and replacement, while maintaining QoS in the Edge.
Effective networking is crucial for the efficient and cost-effective training of artificial intelligence (AI) models. Remote Direct Memory Access (RDMA) has gained widespread adoption due to its low latency, high throughput, and minimal CPU overhead. However, current RDMA implementations remain constrained by underlying Equal-Cost Multi-Path (ECMP) forwarding mechanisms, hindering the full exploitation of parallel paths within distributed training data centers. Furthermore, the rapid evolution of network interface bandwidth necessitates the re-evaluation of existing multi-path approaches due to ineffective scaling to high-speed network environments.This study introduces LIBRA, a novel multi-path transmission method for RDMA within high-speed network environments, enabling efficient utilization of abundant network paths in modern data centers. LIBRA integrates segment routing with packet-level multi-path spraying. Deterministic path assignment eliminates the need for sender-side path detection, simplifying load-balancing design and enhancing overall system efficiency. Furthermore, novel algorithms are proposed to decouple multi-path load balancing from congestion control (CC). These algorithms identify network congestion positions and implement tailored traffic control strategies through fine-grained control mechanisms. The feasibility of the hardware implementation for LIBRA was experimentally validated using FPGA and Tofino-based prototypes. Large-scale simulations demonstrate that LIBRA effectively and robustly leverages the rich network multi-paths within data centers, yielding a 2~4× reduction in tail latency for large message transmission scenarios. Our simulation code will be released as open source at the time of publication.
With the explosive growth in the scale and complexity of large language models (LLMs), there is an urgent need to extend training and inference workloads from within a single data center to across multiple data centers. However, this also introduces new challenges for network transport protocols. To address these issues, we propose SRCC (Sub-RTT Congestion Control), a method designed for inter-datacenter networks. Specifically, SRCC introduces a flowset-based mechanism along with shared node tables, enabling Datacenter Interconnect (DCI) switches to be aware of the path status of each flow. By leveraging information shared among different flows, SRCC can accurately adjust the sending rate at a sub-RTT timescale, thereby significantly improving network performance. Building on this approach, we design detailed mechanisms to address the following challenges: (1) applying INT technology in wide-area networks; (2) acquiring INT information with low overhead; and (3) achieving precise congestion window adjustments under sub-RTT perception.We conducted large-scale simulations using NS3, and the experimental results show that our scheme reduces the average FCT slowdown by 44.17% and 53.86% compared to HPCC and DCTCP, respectively.