An increasing number of cloud services are operated globally, where the service data are frequently replicated across geographically distributed datacenters to improve service quality and reliability. Such replication generates many one-to-many bulk data transfers over inter-datacenter networks from one datacenter to many receiver datacenters. To provide end-users with guaranteed services, these data transfers are usually required to be completed within designated deadlines. Despite the exponential growth in data demand, there has been little work on guaranteeing deadlines for one-to-many transfers, which is the subject of this paper. This paper proposes a centralized admission control coupled with a scheduling algorithm, named deAdline-Guaranteed transfEr (AGE), to guarantee the deadline of admitted data transfers and utilize the network capacity efficiently. The key idea is to flexibly select the source datacenter for receiver datacenters and allow the remaining receivers to obtain a replica from either the original source or the other receivers that have already received a copy. By jointly allocating the source for receivers and the bandwidth and routing paths for every data transfer, AGE maximizes the number of deadline-satisfied transfers. Our simulations show that compared to the state-of-the-art, AGE guarantees the deadline for up to 70 percent more transfers, achieves at least 2× higher network throughput, and reduces the completion time up to 80 percent.
As applications become more distributed to improve user experience and offer higher availability, businesses rely on geographically dispersed datacenters that host such applications more than ever. Dedicated inter-datacenter networks have been built that provide high visibility into the network status and flexible control over traffic forwarding to offer quality communication across the instances of applications hosted on many datacenters. These networks are relatively small, with tens to hundreds of nodes and are managed by the same organization that operates the datacenters which make centralized traffic engineering feasible. Using coordinated data transmission from the services and routing over the inter-datacenter network, one can optimize the network performance according to a variety of utility functions that take into account data transfer deadlines, network capacity consumption, and transfer completion times. In this dissertation, we study techniques and algorithms for fast and efficient data transfers across geographically dispersed datacenters over the inter-datacenter networks. We discuss different forms and properties of inter-datacenter transfers and present a generalized optimization framework to maximize an operator selected utility function. Next, in the several chapters that follow, we study, in detail, the problems of admission control for transfers with deadlines and inter-datacenter multicast transfers. For the admission control problem, our solutions offer significant speed up in the admission control process while offering almost identical performance in the total traffic admitted into the network. For the bulk multicasting problem, our techniques enable significant performance gain in receiver completion times with low computational complexity, which makes them highly applicable to inter-datacenter networks.
Bulk transfers from one to multiple datacenters can have many different completion time objectives ranging from quickly replicating some k copies to minimizing the time by which the last destination receives a full replica. We design an SDN-style wide-area traffic scheduler that optimizes different completion time objectives for various requests. The scheduler builds, for each bulk transfer, one or more multicast forwarding trees which preferentially use lightly loaded network links. Multiple multicast trees are used per bulk transfer to insulate destinations that have higher available bandwidth and can hence finish quickly from congested destinations. These decisions–how many trees to construct and which receivers to serve using a given tree–result from an optimization problem that minimizes a weighted sum of transfers’ completion time objectives and their bandwidth consumption. Results from simulations and emulations on Mininet show that our scheduler, Iris, can improve different completion time objectives by about 2.5 × .
Several organizations have built multiple datacenters connected via dedicated wide area networks over which large inter-datacenter transfers take place. Since many such transfers move the same data from one source to multiple destinations, using multicast forwarding trees can reduce bandwidth needs and improve completion times. However, using a single forwarding tree per transfer can lead to poor performance as the slowest receiver dictates the completion time for all receivers. Using multiple forwarding trees per transfer alleviates this concern-the average receiver could finish early; however, if done naively, bandwidth usage would also increase and it is apriori unclear how best to partition receivers, how to construct the multiple trees and how to determine the rate and schedule of flows on these trees. This paper presents QuickCast, a first solution to these problems. Using simulations on real-world network topologies, we see that QuickCast can speed up the average receiver's completion time by as much as $10\times$ while only using $1.04\times$ more bandwidth; further, the completion time for all receivers also improves by as much as $1.57 >$ faster at high loads. Thereby, while some implementation challenges remain, we advocate using a cohort of forwarding trees.
Datacenters provide cost-effective and flexible access to scalable compute and storage resources necessary for today's cloud computing needs. A typical datacenter is made up of thousands of servers connected with a large network and usually managed by one operator. To provide quality access to the variety of applications and services hosted on datacenters and maximize performance, it deems necessary to use datacenter networks effectively and efficiently. Datacenter traffic is often a mix of several classes with different priorities and requirements. This includes user-generated interactive traffic, traffic with deadlines, and long-running traffic. To this end, custom transport protocols and traffic management techniques have been developed to improve datacenter network performance. In this tutorial paper, we review the general architecture of datacenter networks, various topologies proposed for them, their traffic properties, general traffic control challenges in datacenters and general traffic control objectives. The purpose of this paper is to bring out the important characteristics of traffic control in datacenters and not to survey all existing solutions (as it is virtually impossible due to massive body of existing research). We hope to provide readers with a wide range of options and factors while considering a variety of traffic control mechanisms. We discuss various characteristics of datacenter traffic control, including management schemes, transmission control, traffic shaping, prioritization, load balancing, multipathing, and traffic scheduling. Next, we point to several open challenges as well as new and interesting networking paradigms. At the end of this paper, we briefly review inter-datacenter networks that connect geographically dispersed datacenters, which have been receiving increasing attention recently and pose interesting and novel research problems. To measure the performance of datacenter networks, different performance metrics have been used, such as flow completion times, deadline miss rate, throughput, and fairness. Depending on the application and user requirements, some metrics may need more attention. While investigating different traffic control techniques, we point out the tradeoffs involved in terms of costs, complexity, and performance. We find that a combination of different traffic control techniques may be necessary at particular entities and layers in the network to improve the variety of performance metrics. We also find that despite significant research efforts, there are still open problems that demand further attention from the research community.
Inter-datacenter networks connect dozens of geographically dispersed datacenters and carry traffic flows with highly variable sizes and different classes. Adaptive flow routing can improve efficiency and performance by assigning paths to new flows according to network status and flow properties. A popular approach widely used for traffic engineering is based on current bandwidth utilization of links. We propose an alternative that reduces bandwidth usage by up to at least 50% and flow completion times by up to at least 40% across various scheduling policies and flow size distributions.
Several organizations have built multiple datacenters connected via dedicated wide area networks over which large inter-datacenter transfers take place. This includes tremendous volumes of bulk multicast traffic generated as a result of data and content replication. Although one can perform these transfers using a single multicast forwarding tree, that can lead to poor performance as the slowest receiver on each tree dictates the completion time for all receivers. Using multiple trees per transfer each connected to a subset of receivers alleviates this concern. The choice of multicast trees also determines the total bandwidth usage. To further improve the performance, bandwidth over dedicated inter-datacenter networks can be carved for different multicast trees over specific time periods to avoid congestion and minimize the average receiver completion times. In this paper, we break this problem into the three sub-problems of partitioning, tree selection, and rate allocation. We present an algorithm called QuickCast which is computationally fast and allows us to significantly speed up multiple receivers per bulk multicast transfer with control over extra bandwidth consumption. We evaluate QuickCast against a variety of synthetic and real traffic patterns as well as real WAN topologies. Compared to performing bulk multicast transfers as separate unicast transfers, QuickCast achieves up to 3.64× reduction in mean completion times while at the same time using 0.71× the bandwidth. Also, QuickCast allows the top 50% of receivers to complete between 3× to 35× faster on average compared with when a single forwarding multicast tree is used for data delivery.
Long flows contribute huge volumes of traffic over inter-datacenter WAN. The Flow Completion Time (FCT) is a vital network performance metric that affects the running time of distributed applications and the users' quality of experience. Flow routing techniques based on propagation or queuing latency or instantaneous link utilization are insufficient for minimization of the long flows' FCT. We propose a routing approach that uses the remaining sizes and paths of all ongoing flows to minimize the worst-case completion time of incoming flows assuming no knowledge of future flow arrivals. Our approach can be formulated as an NP-Hard graph optimization problem. We propose BWRH, a heuristic to quickly generate an approximate solution. We evaluate BWRH against several real WAN topologies and two different traffic patterns. We see that BWRH provides solutions with an average optimality gap of less than 0.25%. Furthermore, we show that compared to other popular routing heuristics, BWRH reduces the mean and tail FCT by up to 1.46× and 1.53×, respectively.
Large cloud companies manage dozens of datacenters across the globe connected using dedicated inter-datacenter networks. An important application of these networks is data replication which is done for purposes such as increased resiliency via making backup copies, getting data closer to users for reduced delay and WAN bandwidth usage, and global load balancing. These replications usually lead to network transfers with deadlines that determine the time prior to which all datacenters should have a copy of the data. Inter-datacenter networks have limited capacity and need be utilized efficiently to maximize performance. In this report, we focus on applications that transfer multiple copies of objects from one datacenter to several datacenters given deadline constraints. Existing solutions are either deadline agnostic, or only consider point-to-point transfers. We propose DDCCast, a simple yet effective deadline aware point to multipoint technique based on DCCast and using ALAP traffic allocation. DDCCast performs careful admission control using temporal planning, uses rate-allocation and rate-limiting to avoid congestion and sends traffic over forwarding trees that are carefully selected to reduce bandwidth usage and maximize deadline meet rate. We perform experiments confirming DDCCast's potential to reduce total bandwidth usage by up to 45% while admitting up to 25% more traffic into the network compared to existing solutions that guarantee deadlines.
Datacenters are the main infrastructure on top of which cloud computing services are offered. Such infrastructure may be shared by a large number of tenants and applications generating a spectrum of datacenter traffic. Delay sensitive applications and applications with specific Service Level Agreements (SLAs), generate deadline constrained flows, while other applications initiate flows that are desired to be delivered as early as possible. As a result, datacenter traffic is a mix of two types of flows: deadline and regular. There are several scheduling policies for either traffic type with focus on minimizing completion times or deadline miss rate. In this report, we apply several scheduling policies to mix traffic scenario while varying the ratio of regular to deadline traffic. We consider FCFS (First Come First Serve), SRPT (Shortest Remaining Processing Time) and Fair Sharing as deadline agnostic approaches and a combination of Earliest Deadline First (EDF) with either FCFS or SRPT as deadline-aware schemes. In addition, for the latter, we consider both cases of prioritizing deadline traffic (Deadline First) and prioritizing regular traffic (Deadline Last). We study both light-tailed and heavy-tailed flow size distributions and measure mean, median and tail flow completion times (FCT) for regular flows along with Deadline Miss Rate (DMR) and average lateness for deadline flows. We also consider two operation regimes of lightly-loaded (low utilization) and heavily-loaded (high utilization). We find that performance of deadline-aware schemes is highly dependent on fraction of deadline traffic. With light-tailed flow sizes, we find that FCFS performs better in terms of tail times and average lateness while SRPT performs better in average times and deadline miss rate. For heavy-tailed flow sizes, except for tail times, SRPT performs better in all other metrics.
Using multiple datacenters allows for higher availability, load balancing and reduced latency to customers of cloud services. To distribute multiple copies of data, cloud providers depend on inter-datacenter WANs that ought to be used efficiently considering their limited capacity and the ever-increasing data demands. In this paper, we focus on applications that transfer objects from one datacenter to several datacenters over dedicated inter-datacenter networks. We present DCCast, a centralized Point to Multi-Point (P2MP) algorithm that uses forwarding trees to efficiently deliver an object from a source datacenter to required destination datacenters. With low computational overhead, DCCast selects forwarding trees that minimize bandwidth usage and balance load across all links. With simulation experiments on Google's GScale network, we show that DCCast can reduce total bandwidth usage and tail Transfer Completion Times (TCT) by up to 50% compared to delivering the same objects via independent point-to-point (P2P) transfers.
Datacenters provide the infrastructure for cloud computing services used by millions of users everyday. Many such services are distributed over multiple datacenters at geographically distant locations possibly in different continents. These datacenters are then connected through high speed WAN links over private or public networks. To perform data backups or data synchronization operations, many transfers take place over these networks that have to be completed before a deadline in order to provide necessary service guarantees to end users. Upon arrival of a transfer request, we would like the system to be able to decide whether such a request can be guaranteed successful delivery. If yes, it should provide us with transmission schedule in the shortest time possible. In addition, we would like to avoid packet reordering at the destination as it affects TCP performance. Previous work in this area either cannot guarantee that admitted transfers actually finish before the specified deadlines or use techniques that can result in packet reordering. In this paper, we propose DCRoute, a fast and efficient routing and traffic allocation technique that guarantees transfer completion before deadlines for admitted requests. It assigns each transfer a single path to avoid packet reordering. Through simulations, we show that DCRoute is at least 200 times faster than other traffic allocation techniques based on linear programming (LP) while admitting almost the same amount of traffic to the system.
Datacenter-based Cloud Computing services provide a flexible, scalable and yet economical infrastructure to host online services such as multimedia streaming, email and bulk storage. Many such services perform geo-replication to provide necessary quality of service and reliability to users resulting in frequent large inter- datacenter transfers. In order to meet tenant service level agreements (SLAs), these transfers have to be completed prior to a deadline. In addition, WAN resources are quite scarce and costly, meaning they should be fully utilized. Several recently proposed schemes, such as B4 [1], TEMPUS [2], and SWAN [3] have focused on improving the utilization of interdatacenter transfers through centralized scheduling, however, they fail to provide a mechanism to guarantee that admitted requests meet their deadlines. Also, in a recent study, authors propose Amoeba [4], a system that allows tenants to define deadlines and guarantees that the specified deadlines are met, however, to admit new traffic, the proposed system has to modify the allocation of already admitted transfers. In this paper, we propose Rapid Close to Deadline Scheduling (RCD), a close to deadline traffic allocation technique that is fast and efficient. Through simulations, we show that RCD is up to 15 times faster than Amoeba, provides high link utilization along with deadline guarantees, and is able to make quick decisions on whether a new request can be fully satisfied before its deadline.
In this paper, we study the problem of leader selection in the presence of selfish nodes in mobile ad-hoc networks (MANETs). In order to encourage selfish nodes to assume the leadership role, some form of incentive mechanism is required. We present an incentive-based leader selection mechanism in which incentives are incorporated in the form of credit transfer to the leader, which motivates nodes to compete with each other for assuming the leadership role. The competition among nodes is modeled as a series of one-on-one incomplete information, alternating offers bargaining games. Furthermore, we propose an efficient algorithm in order to reduce the communication overhead imposed on the network for selecting a new leader, in case the current leader is disconnected from the network due to reasons such as mobility and battery depletion. Simulation results show that the proposed mechanism not only increases the overall lifetime of a network, but also decreases the amount of imposed communication overhead on the network in comparison with the traditional leader selection algorithms.
The cooperative arrangement of Device-to-Device (D2D) communications-enabled devices for sharing distributed resources brings many new opportunities. In this paper, we use the term Mobile Cloud for representing a mobile peer-to-peer network where devices themselves are resource providers. One of the most important problems required to be addressed before real implementation of such mobile clouds is adopting a proper service discovery scheme. In this paper, we propose a task-based self-organizing model as our management framework for mobile clouds. In the proposed model, the service discovery task is handled by a node selected as the leader. In order to select the most qualified node as the leader, we propose a leader selection mechanism which is based on multi-player first-price sealed-bid auction. Incentives in the form of credit transfer from service discovery requesting nodes to the leader are considered to encourage nodes to assume the leadership role. Furthermore, we propose a secure transaction model between service discovery requesting nodes and the leader for each service discovery. The simulation results demonstrate that our proposed leader-based service discovery model balances the energy consumption of cloud participants, and increases the overall lifetime of mobile clouds.
We provide a novel solution for Resource Discovery (RD) in mobile device clouds consisting of selfish nodes. Mobile device clouds (MDCs) refer to cooperative arrangement of communication-capable devices formed with resource-sharing goal in mind. Our work is motivated by the observation that with ever-growing applications of MDCs, it is essential to quickly locate resources offered in such clouds, where the resources could be content, computing resources, or communication resources. The current approaches for RD can be categorized into two models: decentralized model, where RD is handled by each node individually; and centralized model, where RD is assisted by centralized entities like cellular network. However, we propose LORD, a Leader-based framewOrk for RD in MDCs which is not only self-organized and not prone to having a single point of failure like the centralized model, but also is able to balance the energy consumption among MDC participants better than the decentralized model. Moreover, we provide a credit-based incentive to motivate participation of selfish nodes in the leader selection process, and present the first energy-aware leader selection mechanism for credit-based models. The simulation results demonstrate that LORD balances energy consumption among nodes and prolongs overall network lifetime compared to decentralized model.
Task-based self-organizing algorithms have been proposed as a solution for the management of mobile ad-hoc networks (MANETs). Such algorithms assign tasks to nodes by sequentially selecting the best nodes as leaders. Since the correct functionality of such algorithms depends on the truthfulness of participant nodes, any misbehavior can disrupt the operation of the network due to absence of a centralized monitoring unit. In this paper, we consider the problem of malicious behavior in leader-based MANETs. Current solution analyzes the behavior of the leader by means of some checker nodes. This approach is vulnerable since a malicious checker can ruin the character of a normal behaving leader, and declare it as a malicious one. We propose an efficient self-organizing mechanism which can detect a malicious behaving leader, while protecting a normal behaving leader from being declared as a malicious node. In addition, our mechanism is designed with no constraints on leader selection algorithm, which makes it applicable to any kind of leader-based network. We also provide several simulation results, showing that our proposed mechanism is efficient and effective even for large MANETs.
Agent-Based Modeling and Simulation (ABMS) is a simple and yet powerful method for simulation of interactions among individual agents. Using ABMS, different phenomena can be modeled and simulated without spending additional time on unnecessary complexities. Although ABMS is well-matured in many different fields such as economic, social, and natural phenomena, it has not received much attention in the context of mobile ad-hoc networks (MANETs). In this paper, we present ABMQ, a powerful Agent-Based platform suitable for modeling and simulation of self-organization in wireless networks, and particularly MANETs. By utilizing the unique potentials of Qt Application Framework, ABMQ provides the ability to easily model and simulate self-organizing algorithms, and then reuse the codes and models developed during simulation process for building real third-party applications for several desktop and mobile platforms, which substantially decreases the development time and cost, and prevents probable bugs that can happen as a result of rewriting codes.
Cauligi S. Raghavendra合作论文数Electrical Engineering and Computer Science;Viterbi School of Engineering at the University of Southern California2