Network design is the process of dimensioning IP capacity over an optical network infrastructure to satisfy a given set of demands and reliability constraints. Specifically, we consider the problem of hose-based cross-layer network design, which seeks to find a minimum cost design that is able to route demand for all hose traffic matrices under all specified failure states. While most network design problems are solved as Mixed Integer Programs, a commercial solver can become intractable due to the scale of today's networks. We demonstrate how the classic Benders decomposition algorithm can be applied and improved for this problem and discuss practical implementation aspects. We showcase a horizontally scalable distributed framework to leverage the decomposable problem structure and solve millions of linear programs in a distributed manner, thereby making the network design problem tractable. In contrast to the conventional approach where failure states and traffic matrices are planned sequentially, the Benders algorithm finds global optimal designs across all traffic matrices and failure states. This leads to network designs with improved solution quality and reliability, with 20--30% less IP capacity and spectrum consumption, 50% less link augments and up to 20x faster runtime that enables design for hyper scale networks in a matter of hours.
This paper presents Meta's Production Wide Area Network (WAN) Entitlement solution used by thousands of Meta's services to share the network safely and efficiently. We first introduce the Network Entitlement problem, i.e., how to share WAN bandwidth across services with flexibility and SLO guarantees. We present a new abstraction entitlement contract, which is stable, simple, and operationally friendly. The contract defines services' network quota and is set up between the network team and services teams to govern their obligations. Our framework includes two key parts: (1) an entitlement granting system that establishes an agile contract while achieving network efficiency and meeting long-term SLO guarantees, and (2) a large-scale distributed run-time enforcement system that enforces the contract on the production traffic. We demonstrate its effectiveness through extensive simulations and real-world end-to-end tests. The system has been deployed and operated for over two years in production. We hope that our years of experience provide a new angle to viewing WAN network sharing in production and will inspire follow-up research.
Meta has a large scale backbone infrastructure supporting services with varying QoS requirements. As part of backbone network planning, a capacity plan that differentiates between different classes of services in terms of availability guarantees is generated and scheduled for deployment. Deployment progress is measured traditionally in terms of volumes of capacity deployed. Our work provides insights into the shortcomings of capacity volume driven deployments. We provide a methodology to rank the contribution of each entity pending deployment towards our network performance goals and use this metric to prioritize deployments helping higher classes of services meet their network guarantees earlier in the deployment schedule. By enabling QoS awareness in backbone deployments, we are able to demonstrate a 67% reduction of risk exposure period for high priority services.
This paper presents Meta's Production Wide Area Network (WAN) Entitlement solution used by thousands of Meta's services to share the network safely and efficiently. We first introduce the Network Entitlement problem, i.e., how to share WAN bandwidth across services with flexibility and SLO guarantees. We present a new abstraction entitlement contract, which is stable, simple, and operationally friendly. The contract defines services' network quota and is set up between the network team and services teams to govern their obligations. Our framework includes two key parts: (1) an entitlement granting system that establishes an agile contract while achieving network efficiency and meeting long-term SLO guarantees, and (2) a large-scale distributed run-time enforcement system that enforces the contract on the production traffic. We demonstrate its effectiveness through extensive simulations and real-world end-to-end tests. The system has been deployed and operated for over two years in production. We hope that our years of experience provide a new angle to viewing WAN network sharing in production and will inspire follow-up research.
Many modern Internet services and applications operate on top of geographically distributed datacenters, which replicate large volumes of business data and other content, to improve performance and reliability. This leads to high volumes of continuous network traffic between datacenters, constituting the bulk of traffic exchanged over the Wide-Area Networks (WANs) that connect datacenters. Operators heavily rely on the efficient and timely delivery of such traffic, as it is key for the performance of distributed datacenter applications, and bulk inter-datacenter traffic also plays an important role for network capacity planning. In this article, we discuss the unique and salient features of bulk across-datacenter transfers and extensively review recent technologies and existing research solutions proposed in the literature for optimization of such traffic. Moreover, we discuss several challenges and interesting directions that call for substantial future research efforts.
Flow routing over inter-datacenter networks is a well-known problem where the network assigns a path to a newly arriving flow potentially according to the network conditions and the properties of the new flow. An essential system-wide performance metric for a routing algorithm is the flow completion times, which affect the performance of applications running across multiple datacenters. Current static and dynamic routing approaches do not take advantage of flow size information in routing, which is practical in a controlled environment such as inter-datacenter networks that are managed by the datacenter operators. In this paper, we discuss Best Worst-case Routing (BWR), which aims at optimizing the tail completion times of long-running flows over inter-datacenter networks with non-uniform link capacities. Since finding the path with the best worst-case completion time for a new flow is NP-Hard, we investigate two heuristics, BWRH and BWRHF, which use two different upper bounds on the worst-case completion times for routing. We evaluate BWRH and BWRHF against several real WAN topologies and multiple traffic patterns. Although BWRH better models the BWR problem, BWRH and BWRHF show negligible difference across various system-wide performance metrics, while BWRHF being significantly faster. Furthermore, we show that compared to other popular routing heuristics, BWRHF can reduce the mean and tail flow completion times by over 1.5× and 2×, respectively.
We list some Software Defined Networking (SDN) products that support the Group Table ALL feature which can be used to forward incoming packets to multiple outgoing ports via packet replication.
Inter-datacenter networks connect dozens of geographically dispersed datacenters and carry traffic flows with highly variable sizes and different classes. Adaptive flow routing can improve efficiency and performance by assigning paths to new flows according to network status and flow properties. A popular approach widely used for traffic engineering is based on current bandwidth utilization of links. We propose an alternative that reduces bandwidth usage by up to at least 50% and flow completion times by up to at least 40% across various scheduling policies and flow size distributions.
Datacenters provide cost-effective and flexible access to scalable compute and storage resources necessary for today’s cloud computing needs. A typical datacenter is made up of thousands of servers connected with a large network and usually managed by one operator. To provide quality access to the variety of applications and services hosted on datacenters and maximize performance, it deems necessary to use datacenter networks effectively and efficiently. Datacenter traffic is often a mix of several classes with different priorities and requirements. This includes user-generated interactive traffic, traffic with deadlines, and long-running traffic. To this end, custom transport protocols and traffic management techniques have been developed to improve datacenter network performance. In this tutorial paper, we review the general architecture of datacenter networks, various topologies proposed for them, their traffic properties, general traffic control challenges in datacenters and general traffic control objectives. The purpose of this paper is to bring out the important characteristics of traffic control in datacenters and not to survey all existing solutions (as it is virtually impossible due to massive body of existing research). We hope to provide readers with a wide range of options and factors while considering a variety of traffic control mechanisms. We discuss various characteristics of datacenter traffic control, including management schemes, transmission control, traffic shaping, prioritization, load balancing, multipathing, and traffic scheduling. Next, we point to several open challenges as well as new and interesting networking paradigms. At the end of this paper, we briefly review inter-datacenter networks that connect geographically dispersed datacenters, which have been receiving increasing attention recently and pose interesting and novel research problems. To measure the performance of datacenter networks, different performance metrics have been used, such as flow completion times, deadline miss rate, throughput, and fairness. Depending on the application and user requirements, some metrics may need more attention. While investigating different traffic control techniques, we point out the tradeoffs involved in terms of costs, complexity, and performance. We find that a combination of different traffic control techniques may be necessary at particular entities and layers in the network to improve the variety of performance metrics. We also find that despite significant research efforts, there are still open problems that demand further attention from the research community.
Cauligi S. Raghavendra合作论文数Electrical Engineering and Computer Science;Viterbi School of Engineering at the University of Southern California2