TCP congestion control methods seriously and unnecessarily harm performance of network transmissions when used in dedicated clusters and grids. We present a simple method in which congestion control can be disabled under appropriate circumstances while still addressing fairness issues and avoiding congestion collapse. We discuss a Linux-based implementation of this “Rude TCP”1 and demonstrate the performance benefits this change provides over one and ten gigabit Ethernet links in Linux – showing a factor of 20 improvement in some cases.
Media companies (and other organizations with large amounts of digital content) require prompt broadcast of extremely large files from a single source to a collection of geographically dispersed destinations. Due to the high cost of terrestrial networks of sufficient bandwidth, satellite networks are commonly used for such transfers. However, current satellite transfers rely on expensive error correction via forward error correction and whole-file retransmission. This paper presents a new, hybrid solution combining the advantages of satellite and terrestrial networks to provide cost-effective reliable file transfer. Specifically, we propose a new peer-to-peer scheme exploiting fast terrestrial networks and multiple receivers to recover from high loss rates (5% or more) in near real-time (latency < 400ms). This solution is efficient, robust under variable packet loss and connectivity, user tunable, scales well, and doubles bandwidth compared to existing approaches. The system has been validated via extensive simulations using a terrestrial network based on the AT&T common backbone core network
Editors’ Message 1–2 Standards for Grid Computing: Global Grid Forum Charlie Catlett 3–7 Characterizing Grids: Attributes, Definitions, and Formalisms Zsolt Németh and Vaidy Sunderam 9–23 Mapping Abstract Complex Workflows onto Grid Environments Ewa Deelman, James Blythe, Yolanda Gil, Carl Kesselman, Gaurang Mehta, Karan Vahi, Kent Blackburn, Albert Lazzarini, Adam Arbree, Richard Cavanaugh and Scott Koranda 25–39 Ninf-G: A Reference Implementation of RPC-based Programming Middleware for Grid Computing Y. Tanaka, H. Nakada, S. Sekiguchi, T. Suzumura and S. Matsuoka 41–51 Simulation Studies of Computation and Data Scheduling Algorithms for Data Grids Kavitha Ranganathan and Ian Foster 53–62 Automatic Flow-Control Adaptation for Enhancing Network Performance in Computational Grids Wu-chun Feng, Mark K. Gardner, Michael E. Fisk and Eric H. Weigle 63–74 Design, Implementation, and Evaluation of the Remos Network Monitoring System Bruce Lowekamp, Nancy Miller, Roger Karrer, Thomas Gross and Peter Steenkiste 75–93 Instructions for Authors 95–98
Design, Implementation, and Evaluation of the Remos Network Monitoring System Bruce Lowekamp, Nancy Miller, Roger Karrer, Thomas Gross and Peter Steenkiste 75–93
With the advent of computational Grids, networking performance over the wide-area network (WAN) has become a critical component in the Grid infrastructure. Unfortunately, many high-performance Grid applications only use a small fraction of their available bandwidth because operating systems and their associated protocol stacks are still tuned for yesterday's WAN speeds. As a result, network gurus undertake the tedious process of manually tuning system buffers to allow TCP flow control to scale to today's WAN Grid environments. And although recent research has shown how to set the size of these system buffers automatically at connection set-up, the buffer sizes are only appropriate at the beginning of the connection's lifetime. To address these problems, we describe an automated and lightweight technique called dynamic right-sizing that can improve throughput by as much as an order of magnitude while still abiding by TCP semantics.
We present a new twist to the Beowulf cluster - the Bladed Beowulf. In contrast to traditional Beowulfs which typically use Intel or AMD processors, our Bladed Beowulf uses Trans-meta processors in order to keep thermal power dissipation low and reliability and density high while still achieving comparable performance to Intel- and AMD-based clusters. Given the ever increasing complexity of traditional supercomputers and Beowulf clusters; the issues of size, reliability power consumption, and ease of administration and use will be "the" issues of this decade for high-performance computing. Bigger and faster machines are simply not good enough anymore. To illustrate, we present the results of performance benchmarks on our Bladed Beowulf and introduce two performance metrics that contribute to the total cost of ownership (TCO) of a computing system - performance/power and performance/space.
In this paper, we present a novel twist on the Beowulf cluster - the Bladed Beowulf. Designed by RLX Technologies and integrated and configured at Los Alamos National Laboratory, our Bladed Beowulf consists of compute nodes made from commodity off-the-shelf parts mounted on motherboard blades measuring 14.7" /spl times/ 4.7" /spl times/ 0.58". Each motherboard blade (node) contains a 633 MHz Trans-meta TM5600/spl trade/ CPU, 256 MB memory, 10 GB hard disk, and three 100-Mb/s Fast Ethernet network interfaces. Using a chassis provided by RLX, twenty-four such nodes mount side-by-side in a vertical orientation to fit in a rack-mountable 3U space, i.e., 19" in width and 5.25" in height. A Bladed Beowulf can reduce the total cost of ownership (TCO) of a traditional Beowulf by a factor of three while providing Beowulf-like performance. Accordingly, rather than use the traditional definition of price-performance ratio where price is the cost of acquisition, we introduce a new metric called ToPPeR: total price-performance ratio, where total price encompasses TCO. We also propose two related (but more concrete) metrics: performance-space ratio and performance-power ratio.
With the advent of computational grids, networking performance over the wide-area network (WAN) has become a critical component in the grid infrastructure. Unfortunately, many high-performance grid applications only use a small fraction of their available bandwidth because operating systems and their associated protocol stacks are still tuned for yesterday's WAN speeds. As a result, network gurus undertake the tedious process of manually tuning system buffers to allow TCP flow control to scale to today's WAN grid environments. And although recent research has shown how to set the size of these system buffers automatically at connection set-up, the buffer sizes are only appropriate at the beginning of the connection's lifetime. To address these problems, we describe an automated and lightweight technique called dynamic rightsizing that can improve throughput by as much as an order of magnitude while still abiding by TCP semantics.
While tcpdump is an invaluable monitoring tool that has held up remarkably well for over a decade, it is showing its age. Network speeds have recently outstripped the ability of ‘stock’ tcpdump running on commodity hardware to keep up with the network, rendering it incapable of monitoring traffic at gigabit-per-second (Gbps) speeds. Tests over Gigabit Ethernet showed that tcpdump could monitor and record traffic at speeds no greater than 250 Mbps with O(ms) time granularity. To achieve monitoring at Gbps speeds and O(ns) time granularity with commodity parts, we present TICKET – the Traffic Information-Collecting Kernel with Exact Timing. TICKET combines efficient commodity-based hardware and software in an architecture that hides disk latency and bandwidth.
Rather than painful, manual, static, per-connection optimization of TCP buffer sizes simply to achieve acceptable performance for distributed applications, many researchers have proposed techniques to perform this tuning automatically. This paper first discusses the relative merits of the various approaches in theory, and then provides substantial experimental data concerning two competing implementations-the buffer autotuning already present in Linux 2.4.x and "dynamic right-sizing." The paper reveals heretofore unknown aspects of the problem and current solutions, provides insight into the proper approach for different circumstances, and points toward ways to further improve performance.
We present results from computations on Green Destiny, a 240-processor Beowulf cluster which is contained entirely within a single 19-inch wide 42U rack. The cluster consists of 240 Transmeta TM5600 667-MHz CPUs mounted on RLX Technologies motherboard blades. The blades are mounted side-by-side in an RLX 3U rack-mount chassis, which holds 24 blades. The overall cluster contains 10 chassis and associated Fast and Gigabit Ethernet switches. The system has a footprint of 0.5 meter2 (6 square feet), a volume of 0.85 meter3 (30 cubic feet) and a measured power dissipation under load of 5200 watts (including network switches). We have measured the performance of the cluster using a gravitational treecode N-body simulation of galaxy formation using 200 million particles, which sustained an average of 38.9 Gflops on 212 nodes of the system. We also present results from a three-dimensional hydrodynamic simulation of a core-collapse supernova.
Virtually all network applications requiring reliable end-to-end communication depend on TCP. Unfortunately, the performance of any stock TCP is abysmal over wide-area networks (WANs) and even over local-area networks (LANs) with very high-bandwidth links. Currently, network researchers manually optimize TCP buffer sizes to achieve acceptable performance over a given connection. Unfortunately, this manual optimization requires changes to the kernel on both end hosts involved in the network connection (changes that are only effective for connections between these two hosts). Furthermore, because two administrative domains must be coordinated to perform this optimization, this process can be tedious and time consuming. To address these problems, this paper illustrates the benefits of a new technique called dynamic right-sizing. This technique dynamically and automatically determines the best buffer size, and hence flow-control window size in TCP. Our simulation study shows that dynamic right-sizing can improve the performance of flows by two orders of magnitude over stock TCP implementations that have static flow-control windows.
Until recently, the average desktop computer has been powerful enough to saturate any available network technology; but this situation is rapidly changing with the advent of Gigabit Ethernet (GigE) and similar technologies. While CPU speeds have improved by over 50% per year since the mid-1980s (or roughly doubling every 1.6 years) [3], network speeds have improved by nearly 100% per year from 10-Mb/s Ethernet in 1988 to 6.4 Gb/s HiPPI-6400/GSN [10] in 1998! Network speeds have finally surpassed the ability of the computer to fi ll the network pipe, and this situation will get dramatically worse due to the above trends as well as slowly increasing I/O bus speeds. While the average I/O bandwidth of a PC is expected to increase from 1.056 Gb/s (32-bit, 33-MHz PCI bus) to 4.224 Gb/s (64-bit, 66-MHz PCI bus) over the next 12-18 months, the widespread availability of HiPPI-6400/GSN (6.4 Gb/s) [10] this year and 10GigE (10 Gb/s) [6] next year will far outstrip the ability of a computer to fill the network. Our experiments already demonstrate that a PC can no longer fully utilize available network bandwidth. With the default Red Hat Linux 6.2 OS running on dual processor 400-MHz PCs with Alteon AceNIC GigE cards on a 32-bit, 33-MHz PCI bus, the peak bandwidth achieved by TCP is only 335 Mb/s. With an 83% increase in CPU speed to 733 MHz, the peak bandwidth only increases by 25% to 420 Mb/s. These bandwidth numbers can be improved by about 5% by increasing the default send/receive buffer sizes from 64 KB to 512 KB, by another 10% by using interrupt coalescing, and by slightly more with further enhancements to the system set-up as described below. Unfortunately, TCP still utilizes only half the available bandwidth in the best case between two machines. This implies that file or web servers wishing to fully utilize their available gigabit-per-second bandwidth will have trouble doing so. Furthermore, remote visualization projects of large data sets will be bound by TCP/IP stack performance, as described below. To create an environment that delivers maximum bandwidth, we removed all processes from the Red Hat Linux 6.2 OS except for the kernel, init, and a shell, turned off virtual memory and interrupt coalescing, and set the benchmarking software as a real-time process. We then configured a single machine to run the tests over its loopback interface (thus not touching the PCI bus or network at all). In this best case, we found that the machine spends a majority of time in the protocol stack rather than transmission into the network and results in only 485 Mb/s on the 400-MHz machine and 590 Mb/s on the 733-MHz machine. Despite this optimal (albeit unrealistic) environment, these numbers demonstrate that it is impossible for TCP running with Linux on this hardware to achieve gigabit speeds between machines. Thus, even if the hardware-based bottleneck at the host interface is alleviated, e. g., the future InfiniBand [1], the software-based bottleneck at the ho st interface, e. g., TCP/IP, will persist. Additional problems that work to sabotage achieving high bandwidth include TCP's flow-and congestion-control mecha-nisms and the use of 1500-byte maximum transmission units (MTUs) in Ethernet. For example, TCP Reno generally viewed as the ubiquitous TCP has been shown to induce bursty and chaotic behavior in network traffic [5, 7, 9, 2], and its congestio n-control mechanism never uses more than 75% of the available link bandwidth on average [11], e. g., under 750Mbps on a gigabit Ethernet link. Other versions of TCP begin to address this problem, but it is still an open issue. In addition, with TCP's default flow-control window, connections with large bandwi dth-delay products cannot keep the network pipe full and suffer significant performance degradation when compared to the pe rformance on LANs. Finally, because of Ethernet's 1500-byte MTU (note: a few non-standard implementations support "jum bo packets" at 9000 bytes), a fully utilized 10-Gb/s Etherne t line would produce a minimum of 830,000 packets per second. If the network interface card (NIC) does not provide interrupt coalescing, each packet would interrupt the host and cause a jump to the interrupt handler. As 2000 cycles per interrupt is the minimum for Linux 2.2.17 to switch from a user process to the interrupt handler and back, a 1.6-GHz CPU would be required merely to handle the interrupt without doing any other processing such as receiving a packet! While jumbo packets and interrupt coalescing may improve delivered bandwidth, there are problems with both. Jumbo packets can only effectively be used in a local-area network; wide-area networks will likely fragment them to 1500 bytes or drop them entirely if we set the "don't fragment" bit, whic h we would want to do to avoid the cost of fragmentation and reassembly in high-performance systems. Furthermore, in the worst case, a packet may cross an old link that only guarantees the Internet standard 576-byte transmission unit [8]. Jumbo packets also induce blocking on networks that do not allow out-of-band data (such as Ethernet), where a small high priority packet from one connection may be enqueued in a switch behind several large jumbo packets from another connection and have to wait. Switched networks that either have a smaller MTU (e. g. ATM at 53 bytes) or support out-of-band data (e. g. Myrinet) do not suffer from this problem. Interrupt coalescing forces us to make a tradeoff between bandwidth and latency. That is, by amortizing the cost of an interrupt over several packets, we can increase delivered bandwidth but the latency increases for each individual packet. Other related approaches which offload work to the NIC are not scalable. For example, some programmers have tried to
Abstract: Computational grids such as the Information Power Grid [1], Particle Physics Data Grid [2], and Earth System Grid [3] depend on TCP to provide reliable communication between nodes across a wide-area network (WAN). Of the available TCP implementations, TCP Reno and its variants are the most widely deployed; however, Reno's performance in computational grids is mediocre at best. Due to conflicting results in the evaluation of TCP implementations [4,5,6,7,8,9,10,11,12,13], we present a detailed simulation study that unifies the conflicting results and demonstrates the limitations of earlier work. We focus on the two most debated versions of TCP - Reno and Vegas. Using real traffic distributions, we show that Vegas performs well over modern high-performance links and better than Reno with the proper selection of the Vegas parameters alpha and beta. Our results exhibit ways to significantly enhance the performance of distributed computational grids that rely on TCP.
Null ciphers are some of the oldest cited examples of modern steganography, and are some of the few steganographic algorithms that use either synthetic or immutable carriers. In contrast, the vast majority of today's steganographic algorithms use mutable carriers where the embedding process requires modifying the carrier in some way. The main deficiency with mutating the carrier during the embedding process is that the algorithms will leave some sort of signature. In this paper we explore the idea of embedding data without changing the carrier by mapping the message onto the carrier instead of making modifications to the carrier. We explore algorithms that use Variable Interval Symbol Aggregation (VISA) for both text and binary data. We study these variable interval algorithms in terms of several quantitative measures and show that these algorithms, often cited as classic examples of steganography, share many characteristics with encryption algorithms.
Michael S Warren合作论文数Los Alamos National Laboratory3
Adam Arbree合作论文数Cornell University2
R. Schlichting合作论文数Software Systems Research Department at AT&T Labs-Research1
Andrew A. Chien合作论文数 Department of Computer Science, University of Illinois at Urbana-Champaign;Department of Computer Science, The University of Chicago;Department of Computer Science and Engineering, University of California, San Diego1
Matti Hiltunen合作论文数AT&T Labs Research — Leading Invention, Driving Innovation1