Cloud-hosted microservices enable scalable and flexible delivery of online services, but unfortunately, also introduce new attack surfaces. In particular, lateral movement attacks exploit interconnected microservice chains and compromise target service without defenders having time to discover and remediate the breach. In this paper, we explore ways to deploy honeypots as a deception-based defense against lateral movement attacks targeting a microservice oriented architecture. We consider a cost metric corresponding to the expected number of lateral movement steps to perform so that the attack succeed. Using a Markov model, we describe the performance of exploration strategies adopted by an attack to reach its target. The model describes two types of attackers, namely a naive attacker with no knowledge on the system, as well as an experienced attacker with partial information about the distance to travel to reach the target. We evaluate our framework and simulate an attack against a vulnerable microservices chain in a Kubernetes cluster. Evaluations confirm that the honeypot-based defense strategy significantly increases the average time involved to reach the target, providing a baseline for sizing honeypots deployment that effectively slows down attacks.
The Kelly or proportional allocation mechanism is a simple and efficient auction-based scheme that distributes an infinitely divisible resource proportionally to the agents bids. When agents are aware of the allocation rule, their interactions form a game extensively studied in the literature. This paper examines the less explored repeated Kelly game, focusing mainly on utilities that are logarithmic in the allocated resource fraction. We first derive this logarithmic form from fairness-throughput trade-offs in wireless network slicing, and then prove that the induced stage game admits a unique Nash equilibrium NE. For the repeated play, we prove convergence to this NE under three behavioral models: (i) all agents use Online Gradient Descent (OGD), (ii) all agents use Dual Averaging with a quadratic regularizer (DAQ) (a variant of the Follow-the-Regularized leader algorithm), and (iii) all agents play myopic best responses (BR). Our convergence results hold even when agents use personalized learning rates in OGD and DAQ (e.g., tuned to optimize individual regret bounds), and they extend to a broader class of utilities that meet a certain sufficient condition. Finally, we complement our theoretical results with extensive simulations of the repeated Kelly game under several behavioral models, comparing them in terms of convergence speed to the NE, and per-agent time-average utility. The results suggest that BR achieves the fastest convergence and the highest time-average utility, and that convergence to the stage-game NE may fail under heterogeneous update rules.
In the multi-resource Kelly mechanism, players obtain a share of each resource in proportion to the bids they place on it. They thus engage in a non-cooperative game where they distribute their budgets across multiple resources. In this paper, we study the repeated variant of this game under standard no-regret algorithms, namely Online Gradient Descent (OGD) and Dual Averaging (DA) algorithms. More specifically, we investigate an additive utility framework with heterogeneous valuations across resources, where each resource-specific utility can be either logarithmic or linear. In this setting, we prove uniqueness of the Nash equilibrium. Moreover, we prove convergence of OGD and DA in the repeated game to this unique Nash Equilibrium. Extensive numerical simulations validate the theoretical results and measure convergence speed across different settings.
Join-the-shortest queue (JSQ) and its variants have often been used in solving load balancing problems. The aim of such policies is to minimize the average system occupation, e.g., the customer's system time. In this paper, we extend the load balancing setting to include constraints that may be imposed, e.g., due to the communication network. First, we cast the problem in the framework of constrained MDPs: this permits us to address both action-dependent constraints, such as, e.g, bandwidth limitation, and state-dependent constraints, such as, e.g., minimum queue utilization. Hence, unlike the state-of-the-art approaches in load balancing, we derive new policies that satisfy the constraints while minimizing system occupancy. Extensive numerical simulations have evaluated their performance under various system settings.
Extensive-form games with perfect information admit at least one Nash equilibrium. The backward induction algorithm identifies in linear time a Nash equilibrium of the game, called subgame perfect. We introduce an extension of the backward induction algorithm which is the first to identify all the outcomes of pure Nash equilibria of the game in linear time with respect to the size of the game.
In edge computing systems, autonomous agents must make fast local decisions while competing for shared resources. Existing MARL methods often resume to centralized critics or frequent communication, which fail under limited observability and communication constraints. We propose a decentralized framework in which each agent solves a constrained Markov decision process (CMDP), coordinating implicitly through a shared constraint vector. For the specific case of offloading, e.g., constraints prevent overloading shared server resources. Coordination constraints are updated infrequently and act as a lightweight coordination mechanism. They enable agents to align with global resource usage objectives but require little direct communication. Using safe reinforcement learning, agents learn policies that meet both local and global goals. We establish theoretical guarantees under mild assumptions and validate our approach experimentally, showing improved performance over centralized and independent baselines, especially in large-scale settings.
The Kelly mechanism is a proportional allocation auction widely adopted in decentralized resource allocation systems to share an infinitely divisible resource among competing agents. We analyze the sequential game it induces when agents have α -fair utilities and behave strategically. Our main result proves that synchronous best-response updates drive bids to the unique Nash equilibrium at a linear rate for α∈{0,1,2} . Extensive simulations reveal that best-response dynamics reach equilibrium significantly faster than previously proposed no-regret learning algorithms.
5G and beyond networks with Massive Multiple Input Multiple Output (M-MIMO) systems benefit from beamforming technology allowing to efficiently allocate resources and manage interference at beam level resolution. This paper proposes a lightweight Inter-Cell Interference Coordination (ICIC) solution that is formulated as a dynamic optimization problem. First, an interference graph is constructed, with nodes representing cells, and the edges - the mutual interference produced by a group of beams serving interfered users per couple of adjacent cells. The ICIC is based on orthogonal resource allocation for a limited set of users associated to the edges. The decision to activate the ICIC per edge of the graph is taken by a low complexity Multi-Armed Bandit (MAB) algorithm to avoid possible performance degradation associated to the ICIC. System-level simulations demonstrate the performance gain brought about by the proposed Interference Management (IM) solution.
The Kelly or proportional allocation mechanism is a simple and efficient auction-based decentralized resource allocation scheme that distributes an infinitely divisible resource proportionally to the agents' bids. When agents are aware of the allocation mechanism, their interactions form a game. The properties of its Nash equilibria are well understood under the simplifying assumption of unbounded budgets. In this paper, we analyze the game in a more realistic budget-constrained setting, motivated by its optimality in terms of the liquid price of anarchy (LPoA). Specifically, we establish a sufficient condition for the uniqueness of the Nash equilibrium and design a distributed sequential learning procedure that provably converges to the equilibrium. In particular, our sufficient condition holds when the payoff functions of the agents are of the proportional fair type in the allocated fraction. Finally, extensive numerical experiments shed light on the interplay between the heterogeneity of the payoff functions and the agents' budgets.
The outcomes of extensive-form games are the realisation of an exponential number of distinct strategies, which may or may not be Nash equilibria. The aim of this work is to determine whether an outcome of an extensive-form game can be the realisation of a Nash equilibrium, without recurring to the cumbersome notion of normal-form strategy. We focus on the minimal example of pure Nash equilibria in two-player extensive-form games with perfect information. We introduce a new representation of an extensive-form game as a graph of its outcomes and we provide a new lightweight algorithm to enumerate the realisations of Nash equilibria. It is the first of its kind not to use normal-form brute force. The algorithm can be easily modified to provide intermediate results, such as lower and upper bounds to the value of the utility of Nash equilibria. We compare this modified algorithm to the only existing method providing an upper bound to the utility of any outcome of a Nash equilibrium. The experiments show that our algorithm is faster by some orders of magnitude. We finally test the method to enumerate the Nash equilibria on a new instances library, that we introduce as benchmark for representing all structures and properties of two-player extensive-form games.
With the advent of big data applications, coflow scheduling has become a cornerstone for the engineering of traffic in datacenters. Minimizing the average weighted Coflow Completion Times (CCT) is a crucial step to minimize the execution time of jobs running in distributed computing frameworks. In this paper, we present a new σ-order coflow scheduling solution, ONE-PARIS, an online semi-clairvoyant and semi-distributed implementation suitable to minimize the weighted CCT in production environments. We achieves this through ONE-PARIS scheduler for ordering coflows and a decentralized resource allocation mechanism, called Sync-Rate, enabling to respect the order of priority of coflows provided by ONE-PARIS and ensuring efficient synchronization between flows of the same coflow in order to free up bandwidth for low-priority flows. Extensive simulations on both synthetic and real traffics show that our proposed coflow scheduler outperforms other state-of-art schemes.
Datacenter networks commonly facilitate the transmission of data in distributed computing frameworks through coflows, which are collections of parallel flows associated with a common task. Most of the existing research has concentrated on scheduling coflows to minimize the time required for their completion, i.e., to optimize the average dispatch rate of coflows in the network fabric. Nevertheless, modern applications often produce coflows that are specifically intended for online services and mission-crucial computational tasks, necessitating adherence to specific deadlines for their completion. In this paper, we introduce $\mathtt {WDCoflow}$ , a new algorithm to maximize the weighted number of coflows that complete before their deadline. By combining a dynamic programming algorithm along with parallel inequalities, our heuristic solution performs at once coflow admission control and coflow prioritization, imposing a $\sigma$ -order on the set of coflows. With extensive simulation, we demonstrate the effectiveness of our algorithm in improving up to $3\times$ more coflows that meet their deadline in comparison the best SoA solution, namely $\mathtt {CS\rm{-}MHA}$ . Furthermore, when weights are used to differentiate coflow classes, $\mathtt {WDCoflow}$ is able to improve the admission per class up to $4\times$ , while increasing the average weighted coflow admission rate.
The deployment of large antenna arrays in Massive Multiple Input Multiple Output (M-MIMO) systems substantially improves mobile networks’ performance. Nevertheless, the increase in performance brought by additional antennas is often counterbalanced by an increase in Power Consumption (PC) caused by the use of supplemental hardware resources supporting additional Radio Frequency (RF) channels. In M-MIMO networks, switching on/off RF channels and muting the associated antennas according to load conditions is known to be an efficient Energy Saving (ES) mechanism. This paper introduces a new RF channels switch on/off solution to maximize the Energy Efficiency (EE) under Quality of Service (QoS) constraints. The proposed solution is based on a MAB algorithm, which appears simpler than state of the art approaches and can be easily implemented in real systems. The algorithm leverages the quasiconcave shape of the EE metric to sequentially select the optimal antenna array configuration - i.e., the number of RF channels in both azimuth and elevation - from a predefined set of configurations. Extensive system-level simulations demonstrate that the proposed algorithm achieves a significant EE gain over a baseline solution.
The average coflow completion time (CCT) is the standard performance metric in coflow scheduling. However, standard CCT minimization may introduce unfairness between the data transfer phase of different computing jobs. Thus, while progress guarantees have been introduced in the literature to mitigate this fairness issue, the trade-off between fairness and efficiency of data transfer is hard to control. This paper introduces a fairness framework for coflow scheduling based on the concept of slowdown, i.e., the performance loss of a coflow compared to isolation. By controlling the slowdown it is possible to enforce a target coflow progress while minimizing the average CCT. In the proposed framework, the minimum slowdown for a batch of coflows can be determined in polynomial time. By showing the equivalence with Gaussian elimination, slowdown constraints are introduced into primal-dual iterations of the CoFair algorithm. The algorithm extends the class of the sigma-order schedulers to solve the fair coflow scheduling problem in polynomial time. It provides a 4-approximation of the average CCT w.r.t. an optimal scheduler. Extensive numerical results demonstrate that this approach can trade off average CCT for slowdown more efficiently than existing state of the art schedulers.
Modern portable devices can execute increasingly sophisticated AI models on sensed data. The complexity of such processing tasks is data-dependent and has relevant energy cost. This work develops an Age of Information markovian model for a system where multiple battery-operated devices perform data processing and energy harvesting in parallel. Part of their computational burden is offloaded to an edge server which polls devices at given rate. The structural properties of an optimal policy for a single device-server system are derived. They permit to define a new model-free reinforcement learning method specialized for monotone policies, namely Ordered Q-Learning, providing a fast procedure to learn the optimal policy. The method is oblivious to the devices' battery capacities, the cost and the value of data batch processing and to the dynamics of the energy harvesting process. Finally, the polling strategy of the server is optimized by combining this policy improvement technique with stochastic approximation methods. Extensive numerical results provide insight into the system properties and demonstrate that the proposed learning algorithms outperform existing baselines.
The deployment of RISs in future 6G networks is expected to substantially improve mobile network coverage. This paper introduces a new cross-layer low-complexity scheme for the online optimization of RIS-assisted communication systems. It jointly combines BS and RIS configuration and fair UEs’ scheduling which is critical for high-performance deployments. A RIS beam synthesis method is especially proposed for RIS configuration. The proposed solution embeds two nested control loops: i) a fast control loop working at the OFDMA slot scale and consisting in a standard UEs proportional fair scheduler, and ii) a slow control loop operating at the OFDMA frame scale which adapts the RIS’ configuration to the UEs’ spatial distribution and maximizes the UEs’ aggregated performance. The slow control loop is based on an online stochastic approximation algorithm whose convergence to the optimal restpoint is proved. In a reference scenario, the proposed scheduler achieves a gain of 47% in mean spectral efficiency for NLOS UEs over a baseline scheme.
With the uptake of intelligent data-driven applications, edge computing infrastructures necessitate a new generation of admission control algorithms to maximize system performance under limited and highly heterogeneous resources. In this paper, we study how to optimally select information flows which belong to different classes and dispatch them to multiple edge servers where applications perform flow analytic tasks. The optimal policy is obtained via the theory of constrained Markov decision processes (CMDP) to take into account the demand of each edge application for specific classes of flows, the constraints on computing capacity of edge servers and the constraints on access network capacity. We develop DRCPO, a specialized primal-dual Safe Reinforcement Learning (SRL) method which solves the resulting optimal admission control problem by reward decomposition. DRCPO operates optimal decentralized control and mitigates effectively state-space explosion while preserving optimality. Compared to existing Deep Reinforcement Learning (DRL) solutions, extensive results show that it achieves 15% higher reward on a wide variety of environments, while requiring on average only 50% learning episodes to converge. Finally, we further improve the system performance by matching DRCPO with load-balancing in order to dispatch optimally information flows to the available edge servers.
In fog computing customers' microservices may demand access to connected objects, data sources and computing resources outside the domain of their fog provider In practice, the locality of connected objects renders mandatory a multi-domain approach in order to broaden the scope of resources available to a single-domain fog provider. We consider a scenario where assets from other domains can be leased across a federation of cloud–fog infrastructures. In this context, a fog provider aims to minimize the quantity of external resources to be rented to satisfy the applications' demands while meeting their requirements. We first introduce a general framework for the deployment of applications across multiple domains owned by multiple cloud–fog providers. Hence, the resource allocation problem is formulated in the form of an integer linear program. We provide a novel heuristic method that explores the resource assignment space in a breadth-first fashion to ensure that locality constraints are met. Extensive numerical results evaluate deployment costs and feasibility of the proposed solution demonstrating that it outperforms the standard approaches adopted in the literature.
Many cloud service providers (CSPs) offer an on-demand service with a small delay. Motivated by the reality of cloud ecosystems, we study non-interruptible services and consider a differentiated service model to complement the existing market by offering multiple service level agreements (SLAs) to satisfy users with different delay tolerance. The model itself is incentive compatible by construction. Two typical architectures are considered to fulfill SLAs: (i) non-preemptive priority queues and (ii) multiple independent groups of servers. We leverage queueing theory to establish guidelines for the resultant market: (a) Under the first architecture, the service model can only improve the revenue marginally over the pure on-demand service model and (b) under the second architecture, we give a closed-form expression of the revenue improvement when a CSP offers two SLAs and derive a condition under which the market is viable. Additionally, under the second architecture, we give an exhaustive search procedure to find the optimal SLA delays and prices when a CSP generally offers multiple SLAs. Numerical results show that the achieved revenue improvement can be significant even if two SLAs are offered. Our results can help CSPs design optimal delay-differentiated services and choose appropriate serving architectures.
David Starobinski合作论文数Laboratory of Networking and Information Systems4