BCube, one of the representative server-centric data center networks, is of high practical value in high-performance computing clusters and large model distributed training. During the long-term operation of data centers, component failures have become commonplace, which easily lead to network partitions and service interruptions. Therefore, fault tolerance serves as the core foundation for guaranteeing service continuity and stable operation of data centers. Cycle structures are a critical topological foundation for network fault tolerance, as they provide redundant communication paths, enable deadlock-free parallel communication, and facilitate network load balancing. The capability of constructing valid cycles in faulty networks directly determines the resource utilization efficiency and service availability of data centers. In this work, we first develop a theoretical framework for cycle fault tolerance analysis in BCube and derive the exact value of its cyclic connectivity to evaluate the network’s capability to preserve cycles under node failures. On this basis, we propose a two-stage fault-tolerant cycle construction algorithm adapted to the structural characteristics of BCube. The first stage identifies the largest fault-free connected component containing cycles to define the effective region for cycle construction. In the second stage, multiple strategies are integrated to efficiently generate valid cycles within the component, while maximizing the coverage of available fault-free nodes. Experiments conducted on BCube networks of different scales under diverse fault scenarios verify that our algorithm achieves superior performance in terms of running efficiency and node coverage, as well as excellent scalability and fault robustness.
更多
查看译文
关键词
Fault tolerance,BCube data center network,cyclic connectivity,cycle construction