As Deep Learning (DL) models grow larger and more complex, training jobs are increasingly distributed across multiple Computing Units (CU) such as GPUs and TPUs. Each CU processes a sub-part of the model and synchronizes results with others. Communication among these CUs has emerged as a key bottleneck in the training process. In this work, we present SiPAC, a Silicon Photonic Accelerated Compute cluster. SiPAC accelerates distributed DL training by means of two co-designed components: a photonic physical layer and a novel collective algorithm. The physical layer exploits embedded photonics to bring peta-scale I/O directly to the CUs of a DL optimized cluster and uses resonator-based optical wavelength selectivity to realize hardware multi-casting. The collective algorithm builds on the hardware multi-casting primitive. This combination expedites a variety of collective communications commonly employed in DL training and has the potential to drastically ease the communication bottlenecks. We demonstrate the feasibility of realizing the SiPAC architecture through 1) an optical testbed experiment where an array of comb laser wavelengths are shuffled by a cascaded ring switch, with each ring selecting and forwarding multiple wavelengths to increase the effective communication bandwidth and hence demonstrating the hardware multicasting primitive, and 2) a four-GPU testbed running a realistic DL workload that achieves 22% system-level performance improvement relative to a similarly-sized leaf-spine topology. Large scale simulations show that SiPAC achieves a 1.4× to 5.9× communication time reduction compared to state-of-the-art compute clusters for representative collective communications.
We present a silicon photonic architecture for accelerating collective communications in distributed deep learning. We demonstrate a 22% job completion time improvement in a small-scale testbed and 1.4 to 5.9× improvement in large-scale simulations.
Adeep neural network based equalizer is proposed to mitigate the intersymbol interference observed in next generation high speed passive optical network (PON) links. The DNN based equalizer is shown to outperform the best known conventional equalizer, the maximum likelihood sequence estimator (MLSE) both in back-back and through fiber experiments. To reduce the hardware complexity of DNN based equalizer for PON systems, we investigate the use of embedded parallelization within a DNN structure having multiple symbol outputs from one DNN. We further investigate using a classification output stage with cross entropy cost to perform joint decision on multiple symbol outputs and demonstrated that the sensitivity gain of such scheme over regression output. To understand the complexity of hardware implementation, the fixed-point DNN based equalizers are developed and implemented in FPGA. The impact of fixed-point resolution on the receiver sensitivity and hardware resource utilization in FPGA implementation is analyzed and reported in detail. We show that a reduction of over 40% in LUTs (look up table) utilization is possible by reducing the DNN's weight resolution from 8-bit to 4-bit while incurring a small penalty in receiver sensitivity.
We introduce an optically interconnected disaggregated architecture for GPU resources and demonstrate a 3 x increase in GPU utilization and up to 73.2% acceleration of application runtime for distributed machine learning workloads. © 2022 The Author(s)
We demonstrate the thermal control of cascaded micro-ring DWDM filters using a single photodiode. The streamlined implementation maintains stable operation of the 8-ring bus with less than 0.1dB BER power penalty on an 8x10Gb/s link. (C) 2022 The Author(s)
The scaling trends of deep learning models and distributed training workloads are challenging network capacities in today’s datacenters and high-performance computing (HPC) systems. We propose a system architecture that leverages silicon photonic (SiP) switch-enabled server regrouping using bandwidth steering to tackle the challenges and accelerate distributed deep learning training. In addition, our proposed system architecture utilizes a highly integrated operating system-based SiP switch control scheme to reduce implementation complexity. To demonstrate the feasibility of our proposal, we built an experimental testbed with a SiP switch-enabled reconfigurable fat tree topology and evaluated the network performance of distributed ring all-reduce and parameter server workloads. The experimental results show up to 3.6× improvements over the static non-reconfigurable fat tree. Our large-scale simulation results show that server regrouping can deliver up to 2.3× flow throughput improvement for a 2× tapered fat tree and a further 11% improvement when higher-layer bandwidth steering is employed. The collective results show the potential of integrating SiP switches into datacenters and HPC systems to accelerate distributed deep learning training.
The explosive growth in data analytics is driving an intensely growing need for compute performance. We will review our ARPA-E PINE ENLITENED experimental testbed and simulation results to motivate performance advantages of disaggregation through the use of flexible photonic interconnect networks for distributed deep learning applications.
This paper proposes optical network interconnects as a key enabler for building high-bandwidth ML training clusters with strong scaling properties. Our design, called SiP-ML, accelerates the training time of popular DNN models using silicon photonics links capable of providing multiple terabits-per-second of bandwidth per GPU. SiP-ML partitions the training job across GPUs with hybrid data and model parallelism while ensuring the communication pattern can be supported efficiently on the network interconnect. We develop task partitioning and device placement methods that take the degree and reconfiguration latency of optical interconnects into account. Simulations using real DNN models show that, compared to the state-of-the-art electrical networks, our approach improves training time by 1.3-9.1x.
We demonstrate SiP switch-enabled server regrouping using bandwidth steering for performance improvement in distributed deep learning training in a Fat-tree testbed. Our proposed SiP switch control scheme enables scaling to large-scale datacenter and HPC systems.
We propose and experimentally demonstrate bandwidth steering using silicon photonic switches within a HyperX topology-based high-performance computing environment. Results running deep learning applications on a physical testbed show 20% improvement in execution time.
As bandwidth requirements and integration of photonic components in computing systems increase, the optical micro-ring resonator are becoming an important building block for dense, high-bandwidth interconnects. Ring resonators are small in size and can operate at data rates up to 60Gb/s NRZ,(1) making them well-suited for integrating many rings operating at different wavelengths into a single device. These devices are a promising solution for complex interconnected systems, such as chip-to-chip interconnects.(2) However, using these rings can be challenging as they are sensitive to fabrication and temperature variations and need constant tuning to lock them to their assigned optical wavelengths.(3) This tuning is commonly done by inserting embedded heaters in or above the ring. Existing techniques that tune the rings by the optical power coupled into the rings require extensive characterization and spectrum analysis, or out-of-band signalling, to account for rings drifting across optical channels and fabrication variations. In this work, we use four cascaded micro-rings implemented in a silicon photonic device operating in the C-band. Each ring taps a single wavelength from a bus waveguide. The remaining light in the bus waveguide is then fed into a photodiode, which is used to monitor the tuning of all rings. We show an algorithm that, based only on information about the design of the chip, tunes the rings to the exact desired optical channel. With this algorithm the use of more complex techniques to directly measure a ring's coupled wavelength can be avoided, reducing system complexity.
Deep learning has been revolutionizing many aspects of our society, powering various fields including computer vision, natural language processing, and activity recognition. However, the scaling trends for both datasets and model size are constraining system performance. Variability of memory requirements can lead to poor resource utilization. Reconfigurable photonic interconnects provide scalable solutions and enable efficient use of disaggregated memory resources. We propose a photonic switched optically connected memory system architecture that tackles the memory challenges while showing the functionality of optical switching for deep learning models. Our proposed system architecture utilizes a “lite” (de)serialization scheme for memory transfers via optical links to avoid network overheads and supports the dynamic allocation of remote memories to local processing systems. In order to test the feasibility of our proposal, we built an experimental testbed with a processing system and two remote memory nodes using silicon photonic switch fabrics and evaluated the system performance. The optical switching time is measured to be 119 μs and an overall 2.78 ms latency is achieved for the end-to-end reconfiguration. The collective results and existing high-bandwidth optical I/Os show the potential of integrating the photonic switched optically connected memory to state-of-the-art processing systems.
We explore a novel, silicon photonics-based approach to build a high bandwidth rack designated for machine learning training. Our goal is to scale state-of-the-art ML training platforms, such as NVIDIA’s DGX and Intel’s Gaudi, from a handful of GPUs in one platform to 256 GPUs in a rack while maintaining Tbps communication bandwidth. Our design, called TeraRack, leverages the emergence of silicon photonics technology to achieve Tbps bandwidth in/out of the GPU chip. TeraRack enables accelerating the training time of popular ML models using (i) a scheduling algorithm that finds the best wavelength allocation to maximize the throughput between communicating nodes; and (ii) a device placement algorithm that partitions ML models across nodes to ensure a sparse and local communication pattern that can be supported efficiently on the interconnect. We build a small prototype with FPGA boards and a 10 mm × 10 mm silicon photonics chip. Simulation results show that TeraRack’s performance on realistic ML training workloads is equivalent to a full-bisection 256×1.2Tbps electrical fabric at 6× lower cost, enabling faster model/data parallel training.
We present an optically connected disaggregated system architecture that supports dynamic resource allocation, leveraging the optical spatial switching capability of silicon photonic microring-based interconnects and demonstrate compute nodes on-request access to remote memory and PCIe-based resources.
We develop an automated wavelength locking of microring resonators for routing optical signals in unicast and multicast modes. The locking algorithm utilizes the photo-conductive effect of the integrated microheaters for tapless monitoring of the optical power coupled to the microring.
We develop a robust and scalable solution for control and feedback of silicon photonic circuits used for optical unicast and multicast. A single-wire DAC and ADC feedback architecture is evaluated with 20Gb/s PAM-4 data streams.
We present a scalable Software-Defined-Networking (SDN) control-plane to integrate Silicon Photonics (SiP) with conventional Ethernet/InfiniBand networks and simultaneously perform packet and circuit switching. Experimental evaluations demonstrate this unique solution with 224 microseconds control plane latency for data-center and high-performance-computing platforms.
We present an autonomous SDN network architecture that leverages the spatial and wavelength switching capabilities of silicon photonics microring-based circuits for self-adaptive bandwidth steering. These functionalities are seamlessly integrated and demonstrated in a datacom testbed.
We demonstrate significant performance improvements for data centers and high performance computing systems through dynamic bandwidth steering enabled by silicon photonic switches on a physical testbed. Experimental running the GTC benchmark on the testbed showed >30% reduction in total execution time.
Luca Carloni合作论文数Department of Computer Science, The Fu Foundation School of Engineering and Applied Science, Columbia University1
Giuseppe Di Guglielmo合作论文数VLSI Design and Education Center, The University of Tokyo1