As Deep Learning (DL) models grow larger and more complex, training jobs are increasingly distributed across multiple Computing Units (CU) such as GPUs and TPUs. Each CU processes a sub-part of the model and synchronizes results with others. Communication among these CUs has emerged as a key bottleneck in the training process. In this work, we present SiPAC, a Silicon Photonic Accelerated Compute cluster. SiPAC accelerates distributed DL training by means of two co-designed components: a photonic physical layer and a novel collective algorithm. The physical layer exploits embedded photonics to bring peta-scale I/O directly to the CUs of a DL optimized cluster and uses resonator-based optical wavelength selectivity to realize hardware multi-casting. The collective algorithm builds on the hardware multi-casting primitive. This combination expedites a variety of collective communications commonly employed in DL training and has the potential to drastically ease the communication bottlenecks. We demonstrate the feasibility of realizing the SiPAC architecture through 1) an optical testbed experiment where an array of comb laser wavelengths are shuffled by a cascaded ring switch, with each ring selecting and forwarding multiple wavelengths to increase the effective communication bandwidth and hence demonstrating the hardware multicasting primitive, and 2) a four-GPU testbed running a realistic DL workload that achieves 22% system-level performance improvement relative to a similarly-sized leaf-spine topology. Large scale simulations show that SiPAC achieves a 1.4× to 5.9× communication time reduction compared to state-of-the-art compute clusters for representative collective communications.
The diversity of workload requirements and increasing hardware heterogeneity in emerging high performance computing (HPC) systems motivate resource disaggregation. Resource disaggregation allows compute and memory resources to be allocated individually as required to each workload. However, it is unclear how to efficiently realize this capability and cost-effectively meet the stringent bandwidth and latency requirements of HPC applications. To that end, we describe how modern photonics can be co-designed with modern HPC racks to implement flexible intra-rack resource disaggregation and fully meet the bit error rate (BER) and high escape bandwidth of all chip types in modern HPC racks. Our photonic-based disaggregated rack provides an average application speedup of 11% (46% maximum) for 25 CPU and 61% for 24 GPU benchmarks compared to a similar system that instead uses modern electronic switches for disaggregation. Using observed resource usage from a production system, we estimate that an iso-performance intra-rack disaggregated HPC system using photonics would require 4× fewer memory modules and 2× fewer NICs than a non-disaggregated baseline.
We present a silicon photonic architecture for accelerating collective communications in distributed deep learning. We demonstrate a 22% job completion time improvement in a small-scale testbed and 1.4 to 5.9× improvement in large-scale simulations.
Broadband access services in the era of 5G and beyond are driving the evolution of optical access networks from a fiber-to-the-home/ fiber-to-the-building infrastructure to a more common access platform, compatible with both wired and wireless broadband services. Accordingly, a flexible optical access network plays a pivotal role in catering to various service requirements. This work proposes a reconfigurable converged optical fronthaul network for wireless and wired services using a 4◊4 microring resonator (MRR) based silicon photonic (SiP) switch fabric. Reconfigurable service selection is demonstrated for wavelength selective, multicast and space switching scenarios. The reconfiguration of digital data and digital radio-over-fiber (DRoF) services for cloud radio access networks (C-RAN) is demonstrated.
Designing efficient interconnects to support high-bandwidth and low-latency communication is critical toward realizing high performance computing (HPC) and data center (DC) systems in the exascale era. At extreme computing scales, providing the requisite bandwidth through overprovisioning becomes impractical. These challenges have motivated studies exploring reconfigurable network architectures that can adapt to traffic patterns at runtime using optical circuit switching. Despite the plethora of proposed architectures, surprisingly little is known about the relative performances and trade-offs among different reconfigurable network designs. We aim to bridge this gap by tackling two key issues in reconfigurable network design. First, we study how cost, power consumption, network performance, and scalability vary based on optical circuit switch (OCS) placement in the physical topology. Specifically, we consider two classes of reconfigurable architectures: one that places OCSs between top-of-rack (ToR) switches—ToR-reconfigurable networks (TRNs)—and one that places OCSs between pods of racks—pod-reconfigurable networks (PRNs). Second, we tackle the effects of reconfiguration frequency on network performance. Our results, based on network simulations driven by real HPC and DC workloads, show that while TRNs are optimized for low fan-out communication patterns, they are less suited for carrying high fan-out workloads. PRNs exhibit better overall trade-off, capable of performing comparably to a fully non-blocking fat tree for low fan-out workloads, and significantly outperform TRNs for high fan-out communication patterns.
The expected halt of traditional technology scaling is motivating increased heterogeneity in high-performance computing (HPC) systems with the emergence of numerous specialized accelerators. As heterogeneity increases, so does the risk of underutilizing expensive hardware resources if we preserve today’s rigid node configuration and reservation strategies. This has sparked interest in resource disaggregation to enable finer-grain allocation of hardware resources to applications. However, there is currently no data-driven study of what range of disaggregation is appropriate in HPC. To that end, we perform a detailed analysis of key metrics sampled in NERSC’s Cori, a production HPC system that executes a diverse open-science HPC workload. In addition, we profile a variety of deep-learning applications to represent an emerging workload. We show that for a rack (cabinet) configuration and applications similar to Cori, a central processing unit with intra-rack disaggregation has a 99.5% probability to find all resources it requires inside its rack. In addition, ideal intra-rack resource disaggregation in Cori could reduce memory and NIC resources by 5.36% to 69.01% and still satisfy the worst-case average rack utilization.
This paper for the first time reports an O-band micro-ring resonator-based switch-and-select silicon photonic switch fabric with bent couplers. An average crosstalk ratio of below -40dB is achieved with a 3dB passband of 43.6GHz.
We introduce an optically interconnected disaggregated architecture for GPU resources and demonstrate a 3 x increase in GPU utilization and up to 73.2% acceleration of application runtime for distributed machine learning workloads. © 2022 The Author(s)
We propose a reconfigurable optical switching architecture for shared-memory CPU-GPU systems. System-level simulations show improved execution time and energy efficiency up to 34% and 25% respectively compared to a static point-to-point architecture for specific application sets. © 2021 The Author(s)
Future HPC and datacenter systems are expected to contain an increasingly heterogeneous set of compute and memory resources as a strategy to preserve performance scaling in the long term. This heterogeneity, combined with the low average resource usage of today's systems, motivate resource disaggregation to allow applications to pool and compose no more than the fine-grain resources they require. In this paper, we start by motivating resource disaggregation by observing average utilization and rate of change of memory bandwidth and latency in NERSC's Cori. We then perform an analytical analysis that quantifies if today's photonic links and switches would meet key metrics to minimize overhead, and what that overhead will be. Finally, we perform preliminary experiments to demonstrate that even this minimal overhead penalizes application performance in some cases. Our study motivates future work on more aggressive photonic links and switches in order to make resource disaggregation more attractive in future systems.
The scaling trends of deep learning models and distributed training workloads are challenging network capacities in today’s datacenters and high-performance computing (HPC) systems. We propose a system architecture that leverages silicon photonic (SiP) switch-enabled server regrouping using bandwidth steering to tackle the challenges and accelerate distributed deep learning training. In addition, our proposed system architecture utilizes a highly integrated operating system-based SiP switch control scheme to reduce implementation complexity. To demonstrate the feasibility of our proposal, we built an experimental testbed with a SiP switch-enabled reconfigurable fat tree topology and evaluated the network performance of distributed ring all-reduce and parameter server workloads. The experimental results show up to 3.6× improvements over the static non-reconfigurable fat tree. Our large-scale simulation results show that server regrouping can deliver up to 2.3× flow throughput improvement for a 2× tapered fat tree and a further 11% improvement when higher-layer bandwidth steering is employed. The collective results show the potential of integrating SiP switches into datacenters and HPC systems to accelerate distributed deep learning training.
The explosive growth in data analytics is driving an intensely growing need for compute performance. We will review our ARPA-E PINE ENLITENED experimental testbed and simulation results to motivate performance advantages of disaggregation through the use of flexible photonic interconnect networks for distributed deep learning applications.
This paper proposes optical network interconnects as a key enabler for building high-bandwidth ML training clusters with strong scaling properties. Our design, called SiP-ML, accelerates the training time of popular DNN models using silicon photonics links capable of providing multiple terabits-per-second of bandwidth per GPU. SiP-ML partitions the training job across GPUs with hybrid data and model parallelism while ensuring the communication pattern can be supported efficiently on the network interconnect. We develop task partitioning and device placement methods that take the degree and reconfiguration latency of optical interconnects into account. Simulations using real DNN models show that, compared to the state-of-the-art electrical networks, our approach improves training time by 1.3-9.1x.
Free-space optical communication is a line-of-sight wireless communication scheme, which is preferred for its number of prime advantages over radio frequency wireless communication, such as no spectrum licensing, large bandwidth, inherent security, electromagnetic compatibility/electromagnetic interference immunity etc. Moreover, free-space optical communication also benefits from low-cost installation and maintenance. It has been studied for the next generation access networks, inter-building connections, ground-to-unmanned aerial vehicle links, underwater communication applications, inter-satellite links, deep space links etc. Among various detection approaches utilized in free-space optical communication, coherent detection can achieve the best sensitivity in a bandwidth-limited condition, effectively demodulate optical multilevel coded signals to attain high spectral efficiency, offer excellent background noise rejection. However, such an attractive free-space optical communication suffer from waveform distortion, scintillation, phase fluctuations etc. after transmission in atmospheric channels. Its link losses are almost dependent on atmospheric effects and climatic conditions. In this article, we present an up-to-date survey on coherent free-space optical communication, the atmospheric turbulent effects especially the impacts of turbulence in free-space optical links, and countermeasures against such impairments.