
Concurrent live virtual-machine (VM) migration is a network and service management problem: the controller must decide not only which VM should move, but also the path, reserved bandwidth rate, and migration mode used while other migrations consume the same hosts, links, and switch queues. This paper presents CAREM, Counterfactual Action Reduction and Evaluation for Migration. CAREM expands each pending VM-to-destination request into feasible path/rate/mode actions, predicts migration time, downtime, energy, and SLA risk with calibrated uncertainty, removes only same-request actions that are certifiably dominated, and passes the retained actions to a residual-state global scheduler. The analysis shows why request-level scoring collapses within-request control structure, when set-wise certified reduction is locally admissible under coverage and measurement-error assumptions, and why fiber-collapsed baselines can incur continuation regret when local choices affect later scheduling. In a 128-host fat-tree simulator, CAREM improves the scalar objective by 22% over the matched collapsed baseline, reducing makespan by 17%, downtime-tail risk by 24%, and energy by 6%. Under common-testbed comparisons with representative SDN, concurrency-aware, SLA-aware, policy-aware, hybrid-mode, and adaptive-placement baselines, CAREM achieves the lowest scalar objective while retaining compatibility with recent hybrid RL/ML mode-selection and placement layers. A 32-host KVM/Open vSwitch prototype confirms lower migration time and downtime with control-loop latency inside the planning interval.
Generative Artificial Intelligence (GAI) has been attracting a massive and rapidly expanding user base worldwide in recent years, resulting in enormous demand for inference requests that cannot be handled efficiently by centralized cloud-based architectures. Edge computing presents a promising approach to mitigate these challenges by leveraging the power of numerous edge devices to better provide GAI services to the users. In this work, we develop a novel two-stage approach to dynamically allocate GAI models and assign user requests to the best edge nodes. In the first stage, we model a joint optimization problem to minimize the expected processing time and decide the optimal model allocation based on predicted user demands. In the second stage, we develop an efficient online approach to assign requests to edge nodes when they arrive, as well as to reallocate GAI models when necessary. Moreover, to address the complexity of the optimization problems in this stage, we leverage Lyapunov optimization framework and Benders decomposition methods to efficiently solve the problems, thereby enabling the proposed approach to quickly adapt to the dynamics of the system. Extensive simulations are conducted to evaluate the performance of the proposed approach and investigate the impacts of important parameters. Simulation results show that the proposed approach can reduce the total processing time by up to 43% with very short running time.
The increasing demand for higher throughput and lower latency in modern applications has driven the evolution of IEEE 802.11 with the introduction ofWi-Fi 7. The new Extremely High Throughput (EHT) amendment enhances performance with Multi-Link Operation (MLO), which allows Wi-Fi stations equipped with multiple radios to transmit and receive concurrently over different frequency links, improving channel access opportunities, reducing latency, and enhancing overall throughput and reliability. However, in scenarios with irregular or high traffic loads across links, MLO can degrade performance due to contention imbalance and inefficient medium access coordination. In particular, TCP flows are highly sensitive to such conditions, as increased contention and collisions can trigger congestion control mechanisms that reduce sending rates and impair throughput. This work investigates the impact of MLO on TCP traffic and proposes a Reinforcement Learning (RL)-based contention control strategy to improve load distribution in congested multi-link environments. The RL-based method dynamically adapts the channel access aggressiveness of stations according to network conditions, mitigating contention imbalances across links and achieving TCP performance improvements while maintaining or, in some cases, improving UDP performance. Simulation results show that the proposed AI-driven contention control provides a flexible and effective solution for traffic balancing in multi-link Wi-Fi 7 networks.
Cloud platforms have become a core component of the Internet because most services and products rely on them to host their backends. Estimating and understanding the latency experienced when accessing those cloud platforms is a challenge of growing importance that has not been sufficiently studied. To address this relevant matter, we conducted a three-month measurement campaign, collecting traceroute data every 30 min across 256 source–destination probe pairs. Our specific goal is to analyze whether the current network is able to provide adequate performance for emerging applications and services. We use this dataset to evaluate the performance of forecasting algorithms when predicting cloud latency from both temporal and spatial perspectives, and we leverage post-hoc explainability methods to identify the drivers affecting latency. Several prior studies provide public cloud-latency datasets, but these datasets are generally analyzed in isolation. To close this gap and provide a cross-dataset comparative analysis of cloud-latency measurements, we analyzed the related publicly available datasets and applied a common forecasting and explainability workflow to compare their findings. Our analysis reveals that operators do not require complex methods to predict latency and that distance and a few other simple features are sufficient to achieve operationally accurate predictions. We find latency to be remarkably stable from the user’s perspective, both over the duration of the campaign and across hours of the day, which contrasts with previous findings, and we show that the specific path traversed has a significant impact on latency.
The rise of network function virtualization (NFV) technology has enabled virtual network functions (VNF) and service function chains (SFCs) to develop into standard paradigms for service delivery. The uncertain traffic brought about by edge computing has made it a key issue to figure out how to deploy VNFs for network load balancing. However, traditional methods are limited to SFC embedding solutions and resource management and pay less attention to traffic. To address the above issue, this paper takes into account the spatio-temporal characteristics of network traffic and the traffic scheduling after VNF deployment to solve the load balanced VNF deployment problem. We formalize the problem into an NP-hard nonlinear integer programming problem, which will be solved with the proposed algorithm adaptive VNF deployment and load balanced traffic scheduling (AdpVDLTS). AdpVDLTS divides network time into large and small time slots to operate on traffic and VNFs simultaneously, and achieves load balancing through traffic prediction and collaboration with VNF deployment. Compared with the excellent existing algorithms, AdpVDLTS can maintain more stable load balancing, higher throughput, lower deployment cost, and lower latency. In addition, the effectiveness of the traffic prediction algorithm is proved by ablation experiments.
In modern video streaming systems, clients typically make independent bitrate decisions without full knowledge of network conditions, often leading to inefficient resource usage and unstable quality. Network-assisted adaptive streaming is rapidly gaining importance in Network and Service Management (NSM), as operators strive to deliver consistently high Quality of Experience (QoE) under dynamic traffic conditions. This paper introduces FROG, a novel, real-time Software-Defined Networking (SDN)-assisted framework designed to coordinate multi-server and multipath HTTP Adaptive Streaming (HAS) using standardized metrics carriage: Common Media Client Data (CMCD) and Common Media Server Data (CMSD). FROG employs an innovative two-stage optimization workflow in which an initial Linear Programming (LP) model rapidly determines feasible bandwidth bounds, server selection, and path capacities, thereby transforming the remaining optimization into a sequential layer-selection process for coordinated quality selection and flow allocation. This decomposition enables sub-second optimization at the scale of thousands of users, while buffer-aware client feedback is integrated to proactively prevent stalls and maintain system stability. Experiments on an emulated SDN testbed demonstrate that FROG achieves QoE comparable to that of the optimal MILP solution on tractable instances while virtually eliminating video stalls, and significantly reduces quality oscillations, outperforming state-of-the-art network-assisted approaches by up to 2.32× under playback-driven evaluation scenarios. Scalability experiments with up to 4,000 clients further demonstrate sub-second optimization runtimes, confirming the practicality of FROG for large-scale deployments.
This paper presents a software-defined Time-Sensitive Networking (TSN) architecture that implements the IEEE 802.1Qcr Asynchronous Traffic Shaper (ATS) using Extended Berkeley Packet Filter (eBPF) technology within Linux-based TSN bridges. By moving traffic shaping logic to the kernel level, our solution eliminates the need for dedicated hardware and enables dynamic, programmable control of frame filtering, metering, and queuing. A Software-Defined Networking (SDN) controller complements the design, providing centralized orchestration of TSN behavior through standardized interfaces and a unified network view. We implement the ATS scheduling model to compute and enforce per-stream eligibility times, supporting time-aware scheduling of concurrent streams within the same priority class. This enables deterministic traffic delivery, which is critical for industrial automation and control. Our approach allows seamless integration into existing infrastructures and aligns with the flexibility objectives of Industry 5.0. Performance evaluations demonstrate accurate scheduling behavior under heterogeneous traffic conditions and quantify the delay introduced by a TSN bridge for multiple coexisting streams.
Multiple Radio Access Technology (multi-RAT) environments provide a promising foundation for service-aware communication in intelligent transportation systems (ITS) and smart cities. However, traffic steering (TS) in highly mobile and ultra-dense vehicular networks remains challenging due to dynamic network conditions, heterogeneous RAT capabilities, varying vehicle requirements, packet loss, latency, and frequent ping-pong RAT switching. In this context, we propose Holistic Intelligent Traffic Steering (HITS), a proactive bi-level TS management framework for multi-RAT vehicular networks. HITS integrates centralized network-wide guidance with local vehicleside decision-making. At the central level, a Graph Convolutional Network–Long Short-Term Memory (GCN–LSTM) model captures holistic spatio-temporal network dynamics and evaluates RAT optimality. At the local level, a State-Action-Reward- State-Action (SARSA) reinforcement learning agent performs adaptive, vehicle-specific RAT selection using local observations and central-level optimality guidance. Results show that HITS achieves up to 6.5% higher average throughput, reduces packet loss ratio by more than 30.2%, lowers latency by nearly 12.2%, and reduces the ping-pong RAT switching rate by over 24% compared with baseline and state-of-the-art (SoTA) TS approaches.
Distributed Denial of Service (DDoS) attacks threaten service continuity in next-generation networks, where autonomous mitigation must suppress attack traffic without unnecessarily degrading legitimate service. This paper presents Sentinel, a neuro-symbolic co-evolutionary framework for DDoS-aware network and service management. Sentinel combines a Proximal Policy Optimization (PPO) defender with a symbolic safety layer that repairs unsafe actions, enforces domain-specific constraints, and supports interpretable rule crystallization. Training further uses a Hall-of-Fame archive of historical attackers and benchmark checkpoint selection to mitigate late-stage co-evolutionary degradation. We evaluate Sentinel over 2,000 generations and 10 independent seeds across benign, mild-attack, strong-attack, chaos/flash-crowd, and held-out ICMP flood scenarios. Results show that Sentinel Pareto-dominates the ShieldOnly ablation on mild, strong, and chaos scenarios under a joint service-quality, leakage, and outage criterion, and Pareto-dominates unshielded PPO on zero-shot ICMP. In the ICMP setting, Sentinel eliminates severe outages in the evaluated scenario (0 vs. 145.7 for PPO) and reduces leakage by 8.6 percentage points through symbolic ICMP overrides. Benchmark checkpointing reduces strong-attack leakage by 8.1 percentage points and severe outages by 98.3 steps relative to the final checkpoint. Rule crystallization achieves 97.97% held-out accuracy using a depth-4 decision tree over four interpretable traffic features. Strict SLA compliance is not achieved under attack conditions, and all results are simulation-based within the modeled traffic and attack distributions.
Microservice applications commonly rely on service-mesh circuit breakers to maintain resilience under dynamic workloads. In platforms such as Istio, circuit breaking is enforced through connection and request-queue limits, making performance highly sensitive to parameter choices that static configurations cannot adapt across traffic phases. This paper proposes an SLO-aware reinforcement learning controller that adjusts Istio circuit-breaker parameters from runtime telemetry. The controller operates in the control plane through standard APIs and requires no changes to application code or data-plane components. A compact state representation and bounded incremental action space support stable, interpretable online adaptation. The controller is trained and evaluated on a real single-node Kubernetes cluster: training uses sustained heavy open-loop overload, while evaluation uses stochastic non-stationary moderate and heavy open-loop workloads. Results show that under heavy overload the controller substantially reduces tail latency relative to unprotected deployments, tuned static circuit breakers, and recent heuristic and adaptive baselines, while matching the strongest baseline on success rate. Under milder load it preserves high success within the latency SLO. These results show that reinforcement learning can provide effective adaptive traffic management for service-mesh environments.
Although 5G networks are more efficient in terms of power consumption to traffic ratio, efforts still need to be made to further increase energy efficiency not only for the radio part, but also with respect to the growing cloud component with edge servers. Consolidation of traffic workloads onto shared infrastructures is a key feature of cloud computing to reduce energy consumption, and logical functionality placement plays a key role in this regard. Here, in the cloud RAN context, we propose a unified and energy-aware logical placement of 5G E2E functionalities, i.e., distributed units (DUs), centralized units (CUs), and user plane functions (UPFs), together with traffic routing. The placement problem is formulated as a large-scale integer linear program and solved using a column generation-based decomposition technique, complemented by an efficient heuristic to ensure tractability and improved scalability. The model captures key network and cloud (compute) resources, jointly optimizing the placement of DU, CU, and UPF components, along with traffic routing, to minimize energy consumption while maintaining low latency and high Quality of Service (QoS). Numerical results, based on an open Montreal traffic dataset, demonstrate that the proposed column generation algorithm achieves near-optimal solutions, while the heuristic approach offers significantly better scalability with consistently strong performance. The proposed methods reduce energy consumption by up to 14% and maintain low-latency service delivery. Furthermore, the results highlight that static, peak-time-based placement strategies can lead to inefficiencies throughout the day, emphasizing the importance of accounting for broader temporal traffic patterns.
Modern decentralized applications frequently employ one-to-many messaging over dynamic recipient groups, a paradigm we term Selected Group Broadcasting (SGB). Though one-to-one traffic models (e.g., M/M/1 queue) are well studied, traffic characterization for one-to-many connectivity remains underdeveloped. In this paper, we investigate the delay distribution of one-to-many traffic over an SGB in a wide-area communication system. Depending on the underlying network input, the SGB problem naturally decomposes into two distinct paradigms: SGBs (for Sequential inputs) and SGBp (for Parallel inputs). Asymptotic analysis reveals that SGBs and SGBp are governed by fundamentally different stochastic dynamics, necessitating dedicated modeling approaches for each. This paper focuses exclusively on the SGBs paradigm, providing a rigorous analytical treatment and establishing an operatortheoretic performance framework. For first-order statistics of SGBs, we derive exact closed-form solutions under exponential inter-ACK interval (ACK Interval) inputs. By exploiting the rational Laplace-domain structure of the forward operator, we extend the model to precisely compute output distributions for any non-negative distribution input. Furthermore, we introduce an equivalent SGBs analysis scheme via the inverse operator, enabling the transformation of continuous commit-time distributions into physically interpretable ACK Interval profiles. For second-order statistics, we establish a formalized functional hypothesis under exponential ACK Intervals, validated via simulation. We verify the derived probability density functions with simulated delays; the results are confirmed via Kolmogorov-Smirnov tests.
Deep learning-based network intrusion detection systems (NIDSs) in the Industrial Internet of things (IIoT) are inevitably prone to producing misclassified samples. Correction models can identify and correct these samples to reduce the total error rate (TER) of NIDSs. Existing correction models fail to account for the uneven distribution of misclassified samples in the prediction intervals of NIDSs due to IIoT data imbalance, re-sulting in over-correction and increased TERs. Given this, we propose a novel correction model called Correction Forest for correcting the misclassified samples of NIDSs targeting imbalanced IIoT network traffic dataset. Correction Forest adopts a generation-correction strategy. The generation process divides the output values of NIDSs with imbalanced dataset into fine-grained bins, and then generates novel misclassification features for each bin using a balanced hybrid Random Forest. The correction process calculates feature importance score for misclassification features and corrects misclassification samples by a K-Nearest Neighbors (KNN) -based algorithm. Evaluated on 15 imbalanced IIoT datasets with varying malicious sample ratios, Correction Forest significantly outperforms 4 state-of-the-art models. Under the 1.25% setting of NSL-KDD, Correction Forest improves F1-Score from 0.3144 to 0.7345, an absolute gain of 0.4201. On the TON_IoT at 12.5%, it achieves a maximum R.TER of 0.3750 among the state-of-the-art models.
Intelligent network operations increasingly rely on structured anomaly knowledge to support anomaly analysis, alert correlation, and root-cause investigation. However, under emerging and few-shot anomaly scenarios, anomaly-related knowledge graphs are often incomplete, which limits their operational value. To address this issue, this paper studies the problem of few-shot anomaly knowledge completion and proposes a Neighbor-Enhanced Knowledge Graph Completion model (NEKC). NEKC employs a similarity-aware neighbor selection mechanism to retain semantically relevant neighbors while introducing diversity constraints to avoid representation bias caused by overly homogeneous neighborhoods. An attention mechanism is further used to dynamically weight neighbor entities, and a Transformer-based encoder is adopted to capture contextual dependencies for task-specific relation representation. To evaluate the proposed method, experiments are conducted on the generic few-shot knowledge graph completion benchmark NELL and on a constructed Network Anomaly Knowledge Graph (NAKG) derived from public anomaly-related knowledge sources. The results show that NEKC achieves moderate overall improvements on NELL, whereas its advantages are more evident on NAKG compared with representative baseline methods in few-shot link prediction and knowledge completion tasks. These results indicate that NEKC can effectively improve the completeness of anomaly knowledge graphs and provide semantic support for downstream anomaly analysis in intelligent network operations.
Paxos consensus protocol is widely used in distributed systems, yet their performance can degrade under heterogeneous workloads and deadline-constrained requests. Traditional priority-based Paxos extensions rely on static scheduling policies that are unable to adapt to dynamically changing urgency. This paper proposes a deadline-aware scheduling framework for Paxos based on the Shortest Remaining Processing Time (SRPT) discipline. The proposed approach dynamically prioritizes requests according to their expected completion behavior and their likelihood of meeting assigned deadlines, enabling preemptive scheduling decisions at the consensus leader. An analytical queueing model is developed to characterize the mean waiting time, mean residence time, and mean response time under SRPT scheduling. Using these results, a probabilistic measure of deadline satisfaction is derived and employed to guide scheduling decisions. Analytical and numerical evaluations demonstrate that the proposed framework significantly improves deadline satisfaction while preserving the delay-optimal properties of SRPT, making it well suited for latency-sensitive Paxos deployments.
Onion messages (OMs) in the Lightning Network (LN) enable private communication between nodes through onion routing. While they support important functionalities such as static invoices and asynchronous payments, they can also be exploited for spam. To counter this, the Basis of Lightning Technology Specifications (BOLTs) recommend implementing rate limiting on OM forwarding. Limiting the rate, however, opens the door for an adversary to degrade the OM service by flooding the network—anonymously, thanks to onion routing. In this paper, we analyze and quantify the impact of DoS attacks on the OM service. Following the approach suggested in the literature, we study a policy in which per-peer forwarding limits and forwarding-node selection probabilities are proportional to publicly observable channel-capacity weights. Our analysis shows that OMs are most resilient against DoS attacks when the honest nodes’ public-capacity weights are uniformly distributed. However, the distribution of these weights in LN is highly non-uniform, as demonstrated in prior empirical studies and confirmed by our analysis of three network snapshots spanning 2022–2026. To improve resilience under such skewed conditions, we propose restricting forwarding-node selection to a carefully selected subset of nodes. For the evaluated setting l = L = 3, at a common adversarial public-capacity weight of approximately 14.30 BTC—equivalent to about USD 1 million at the reference BTC/USD rate used in our evaluation—the top-target per-connection adversary matches the strong optimal aggregate-budget adversary in each snapshot. Their OM failure probabilities range from 40.9% to 56.7% when the full network is available for forwarding-node selection. Restricting selection to the optimized subsets reduces these probabilities to 3.6%–3.9%, with corresponding analytical upper bounds of 6.2%–6.9%.
In the era of 5G and beyond, the hard-isolated channels enabled by time slot cross-connects in metro transport network (MTN) effectively meet the demands of emerging network services for low latency, low jitter, flexible bandwidth, and secure isolation. However, the dynamic arrival and departure of tenant virtual network request (VNRs) lead to resource fragmentation within the MTN transport network, resulting in inefficient resource utilization. To mitigate network resource fragmentation, we formulate the MTN dynamic channel orchestration problem and propose a fragmentation-aware MTN dynamic channel orchestration method. This method comprises two key components: a greedy graph-reconstruction-based channel mapping algorithm and a fragmentation-aware channel reconfiguration algorithm. The former optimizes MTN channel resource allocation to achieve static channel orchestration, while the latter, leveraging a simulated annealing-based channel reconfiguration strategy, dynamically adjusts channel allocations based on the static orchestration results, thereby reducing fragmentation levels. Compared to existing approaches, under varying network load conditions, the proposed channel mapping algorithm reduces the consumption of network time slot resources by 16.3% -20.6%, while the channel reconfiguration algorithm significantly lowers the levels of network fragmentation by 48.7% -79.5%, and reduces the running time by 91.3%–94.2% compared with the baseline.
Machine learning (ML)–based intrusion detection systems (IDS) frequently degrade when deployed across heterogeneous networks due to domain shifts in traffic composition and monitoring configurations. Conventional domain adaptation (DA) methods mitigate this issue by aligning source and target distributions, but they often rely on retaining source-domain data at deployment—an impractical requirement that undermines operational scalability and reusability. To address this gap, we propose TRANSFA-IDS (Transformer Source-Free Adaptation for IDS), a lightweight source-free adaptation framework that recalibrates a source-trained IDS using only target traffic data. TRANSFA-IDS converts tabular flow records into structured RGB image embeddings and employs a compact Vision Transformer with a Deep Support Vector Data Description (Deep-SVDD) head to learn transferable normal representations. At deployment, adaptation is performed by fine-tuning only the last transformer block on a small target buffer, realigning target representations without retraining or access to source data. Experiments on cross-dataset transfer between CIC-IDS-2018 and UNSW-NB15 show that TRANSFA-IDS achieves AUROC of 0.9177 and 0.9071 in the two transfer directions, reduces target-domain benign false positives by over 60% relative to the same source-pretrained model deployed without source-free adaptation, and adapts substantially faster than supervised and unsupervised DA baselines while using at most 20% of the target-domain data. These results indicate that source-free adaptation can achieve both strong detection performance and a practical deployment-oriented design, with cross-benchmark evidence of scalable adaptation across heterogeneous network environments.
Low Earth Orbit (LEO) satellite constellations are emerging as an important platform for distributed dataflow execution in space-terrestrial integrated networks. Existing studies largely treat routing and processing separately, while next-generation LEO systems are expected to process and transform data in transit by leveraging on-board computing and software-defined infrastructures. However, jointly optimizing routing and in-network processing in dynamic LEO satellite networks remains challenging because of time-varying connectivity, limited on-board resources, and bandwidth constraints. In this paper, we formulate the Dynamic LEO In-network Processing Dataflow Optimization (DLIDO) problem, which aims to maximize the throughput of processed dataflows by jointly optimizing routing paths and processing-resource allocation over a dynamic flow network. We present an approximation algorithm with a proven (1−ϵ) approximation guarantee for 0 < ϵ ≤ 0.5, providing near-optimal throughput under dynamic processing and communication constraints. To further improve efficiency and practicality, we develop a 2-walk based iterative heuristic algorithm that substantially reduces runtime while maintaining strong empirical performance, and in some regimes provably optimal behavior. Extensive evaluations on realistic LEO network topologies show that both algorithms significantly outperform existing approaches in throughput and adaptability, highlighting a promising direction for dataflow-aware scheduling and optimization in dynamic satellite systems.