This tutorial paper presents a data center exchange (Data Center Xchange, DCX) architecture for all-photonics networks-as-a-service in distributed data center infrastructures, enabling the creation of a virtual large-scale data center by directly interconnecting geographically distributed data centers in metropolitan areas. In contrast to existing vendor-driven optical networking approaches, the proposed architecture adopts an operator-driven and open digital twin paradigm, leveraging cloud-native transponder architectures and open tools/interfaces such as GNPy and CMIS/TAI, and a user–carrier collaborative control framework. In particular, the cloud-native architecture enables operators to flexibly develop, deploy, and manage their own control and automation functions across transponders and controllers using container-based software components. Key requirements for such an architecture in the era of AI are identified: support for low-latency operations, scalability, reliability, and flexibility within a single network architecture; the ability to add new operator-driven automation functionalities based on an open networking approach; and the ability to control and manage remotely deployed transponders connected via access links with unknown physical parameters. We propose a set of technologies that enable digital twin operations for optical networks, including a cloud-native architecture for coherent transceivers, remote transponder control, fast end-to-end optical path provisioning, transceiver-based physical-parameter estimation incorporating digital longitudinal monitoring, and optical line system calibration, demonstrating their feasibility through field validations.
Precise time synchronization at the single-photon level is a key requirement for quantum communication systems, including quantum key distribution (QKD) and entanglement-based networks. We demonstrate a single-shot time synchronization method based on a non-periodic weak coherent pulse pattern combined with a pilot sequence, transmitted only once on a single optical wavelength. Despite the probabilistic nature of photon detection and the presence of dark counts, the arrival time of the pulse pattern can be reliably identified. Using superconducting nanowire single-photon detectors and time-tagged detection, experiments demonstrate a synchronization precision of 24 ps (FWHM). The proposed method operates independently of the quantum encoding scheme and requires neither repeated transmissions nor wavelength-division multiplexing, making it broadly applicable to quantum communication architectures.
We experimentally verified an in-service frame-based delay measurement method using OpenZR+ transceivers, enabling latency-managed IP-over-DWDM for datacenter interconnects with precision comparable to OTN and OTDR.
Quantum conference key agreement (QCKA) enables multiple users to establish a common secret key with information-theoretic security and is regarded as a key primitive for secure communication in future quantum networks. However, practical implementations of QCKA typically suffer from higher noise levels than conventional bipartite quantum key distribution (QKD), making the improvement of the tolerable error threshold an important challenge. Gottesman and Lo proposed two preprocessing procedures for QKD with two-way classical communication, known as the B-step and the P-step, which enhance the tolerable error threshold. In this paper, we analyze the asymptotic security of QCKA with tripartite GHZ states and two measurement bases using two-way classical communication, including multiple B-steps and P-steps. We derive the corresponding secure key rate analytically and demonstrate that iterative B-steps can increase the tolerable error threshold beyond 20
Resilience in optical networks has traditionally relied on redundancy and pre-planned recovery strategies, both of which assume a certain level of disaster predictability. However, recent environmental changes such as climate shifts, the evolution of communication services, and rising geopolitical risks have increased the unpredictability of disasters, reducing the effectiveness of conventional resilience approaches. To address this unpredictability, this article introduces the concept of agile resilience, which emphasizes dynamic adaptability across multiple operators and layers. We identify key requirements and challenges, and present enabling technologies for the realization of agile resilience. Using a field-deployed transmission system, we demonstrate rapid system characterization, optical path provisioning, and database migration within six hours. These results validate the effectiveness of the proposed enabling technologies and confirm the feasibility of agile resilience.
Open optical networks have been considered to be important for cost-effectively building and operating the networks. Recently, the optical-circuit-switches (OCSes) have attracted industry and academia because of their cost efficiency and higher capacity than traditional electrical packet switches (EPSes) and reconfigurable optical add drop multiplexers (ROADMs). Though the open interfaces and control planes for traditional ROADMs and transponders have been defined by several standard-defining organizations (SDOs), those of OCSes have not. Considering that several OCSes have already been installed in production datacenter networks (DCNs) and several OCS products are on the market, bringing the openness and interoperability into the OCS-based networks has become important. Motivated by this fact, this paper investigates a software-defined networking (SDN) controller for open optical-circuit-switched networks. To this end, we identified the use cases of OCSes and derived the controller requirements for supporting them. We then proposed a multi-vendor (MV) OCS controller framework that satisfies the derived requirements; it was designed to quickly and consistently operate fiber paths upon receiving the operation requests. We validated our controller by implementing it and evaluating its performance on actual MV-OCS networks. It satisfied all the requirements, and fiber paths could be configured within 1.0 second by using our controller.
The Bennett-Brassard 1984 (BB84) protocol is one of the simplest protocols for implementing quantum key distribution (QKD). In the protocol, the sender and the receiver iteratively choose one of two complementary measurement bases. Regarding the basis choice by the receiver, a passive setup has been adopted in a number of its implementations, including satellite QKD and time-bin encoding. However, conventional theoretical techniques to prove the security of the BB84 protocol are not applicable if the receiver chooses their measurement basis passively, rather than actively, with a biased probability, followed by measurement with threshold detectors. Here we present a fully analytical security proof against coherent attacks for such a decoy-state BB84 protocol with the receiver's passive basis choice and measurement with threshold detectors. Numerical simulations under practical situations show that the difference in secure key rate between the active and passive implementations of the protocol is negligible except for long communication distances.
Optical link tomography (OLT) is a rapidly evolving field that allows the multi-span, end-to-end visualization of optical power along fiber links in multiple dimensions from network endpoints, solely by processing signals received at coherent receivers. This paper has two objectives: (1) to report the first field trial of OLT, using a commercial transponder under standard DWDM transmission, and (2) to extend its capability to visualize across 4D (distance, time, frequency, and polarization), allowing for locating and measuring multiple QoT degradation causes, including time-varying power anomalies, spectral anomalies, and excessive polarization dependent loss. We also address a critical aspect of OLT, i.e., its need for high fiber launch power, by improving power profile signal-to-noise ratio through averaging across all available dimensions. Consequently, multiple loss anomalies in a field-deployed link are observed even at launch power lower than the system-optimal level. The applications and use cases of OLT from network commissioning to provisioning and operation for current and near-term network scenarios are also discussed.
We review the needs and methods of automatic optimization in open optical networks, with a particular focus on digital coherent transmission systems for data center interconnections, including recent advancements and field experiment results.
As AI models grow in scale, the interconnect becomes a key bottleneck in large-scale GPU clusters. Conventional packet-switched networks face increasing challenges in power, cost, and scalability. This paper explores the use of optical circuit switching (OCS) as a spine-layer interconnect for AI training clusters. We analyze the traffic characteristics of AI workloads, particularly large language model (LLM) training, and argue that their structured, phase-based communication patterns align well with the slower reconfiguration speed of OCS. Our comparative evaluation shows that an OCS-based architecture can reduce spine-layer power consumption by nearly 99 % and 8 -year lifecycle costs by 76 % compared to electrical packet switching. We also discuss design extensions, such as supporting multi-tenant scheduling and integrating OCS into both spine and leaf layers. These results suggest that OCS offers a viable and energy-efficient alternative for future AI superclusters.
We report the first trial of network tomography over a live network in a multi-domain environment. We visualise end- to-end optical powers along multiple routes across multiple domains solely from a commercial 800G transponder, enabling performance bottleneck localisation, power and routing optimisation, and lightpath provisioning. (c) 2025 The Author(s)
There are increasing requirements for data center interconnection (DCI) services, which use fiber to connect any DC distributed in a metro area and quickly establish high-capacity optical paths between cloud services and mobile edge computing and the users. In such networks, coherent transceivers with various optical frequency ranges, modulators, and modulation formats installed at each connection point must be used to meet service requirements such as fast-varying traffic requests between user computing resources. This requires technology and architectures that enable users and DCI operators to cooperate to achieve fast provisioning of WDM links and flexible route switching in a short time, independent of the transceiver’s implementation and characteristics. We propose an approach to estimate the end-to-end (EtE) generalized signal-to-noise ratio (GSNR) accurately in a short time, not by measuring the GSNR at the operational route and wavelength for the EtE optical path but by simply applying a quality of transmission probe channel link by link, at a wavelength/modulation-format convenient for measurement. Assuming connections between transceivers of various frequency ranges, modulators, and modulation formats, we propose a device software architecture in which the DCI operator optimizes the transmission mode between user transceivers with high accuracy using only common parameters such as the bit error rate. In this paper, we first implement software libraries for fast WDM provisioning and experimentally build different routes to verify the accuracy of this approach. For the operational EtE GSNR measurements, the accuracy estimated from the sum of the measurements for each link was 0.6 dB, and the wavelength-dependent error was about 0.2 dB. Then, using field fibers deployed in the NSF COSMOS testbed, a Linux-based transmission device software architecture, and transceivers with different optical frequency ranges, modulators, and modulation formats, the fast WDM provisioning of an optical path was completed within 6 min.
We present a first-ever ultra-low-latency video-transmission system capable of transmitting and receiving uncompressed 8K120p video parallelizing SMPTE ST 2110. To reduce the delay and implementation difficulty, we propose a novel architecture based on SMPTE RP 2110-23 transmission. To enable the timestamp conformance without additional delay, timestamp addition/subtraction units are deployed. In the video transmission experiment, we confirmed that 8K120p video can be transmitted successfully within 1 millisecond.
Recent advances in wireless communication technology such as fifth-generation (5G) have enabled the creation of various novel applications. As a result, a large number of devices are now being connected to mobile networks, and mobile traffic is increasing year by year. Although the use of the millimeter-wave (mmWave) bands is a promising approach to increasing the capacity of mobile networks, there are many challenges to use mmWave bands. The link quality (LQ) of mmWave wireless links is impacted by the surrounding objects. Therefore, in order to stably utilize mmWave bands, we believe that it is necessary to predict future LQ predictions and adaptively control wireless communications. In this paper, we evaluated the throughput prediction methods using physical space information of the target UE and surroundings in a commercial 5G network. The evaluation entails measuring the throughput in an actual indoor environment where both the target UE and surrounding objects are moving. To create the huge dataset necessary to allow the moving terminal holder and surrounding pedestrian (objects) to be modelled, we develop two autonomous humanoid robots and make one move so as to block the LOS of the other robot, which is the UE holder. The experiments shows that our proposed method using physical space information yields a 57.5 % improvement in prediction accuracy at the 50th percentile absolute error value over a naive prediction model that uses past throughput information.
We propose a method to estimate the input power dependency of the transceiver BER-OSNR characteristic. Experiments using commercial transceivers show that estimation error in Q-factor is less than 0.2 dB.
We propose methods and an architecture to conduct measurements and optimize newly installed optical fiber line systems semi-automatically using integrated physics-aware technologies in a data center interconnection (DCI) transmission scenario. We demonstrate, for the first time, digital longitudinal monitoring (DLM) and optical line system (OLS) physical parameter calibration working together in real-time to extract physical link parameters for transmission performance optimization. Our methodology has the following advantages over traditional design: a minimized footprint at user sites, accurate estimation of the necessary optical network characteristics via complementary telemetry technologies, and the capability to conduct all operation work remotely. The last feature is crucial, as it enables remote operation to implement network design settings for immediate response to quality of transmission (QoT) degradation and reversion in the case of unforeseen problems. We successfully performed semi-automatic line system provisioning over field fiber networks facilities at Duke University, Durham, NC. The tasks of parameter retrieval, equipment setting optimization, and system setup/provisioning were completed within 1 hour. The field operation was supervised by on-duty personnel who could access the system remotely from different time zones. By comparing Q-factor estimates calculated from the extracted link parameters with measured results from 400G transceivers, we confirmed that our methodology has a reduction in the QoT prediction errors (+-0.3 dB) over existing design (+-10.6 dB).
To accommodate large-scale data processing requiring high bandwidth, low latency, and low power consumption, remote direct memory access (RDMA) is extensively utilized in today’s data centers (DCs). However, due to the limitations in constructing and operating large centralized DCs, major DC operators are gradually transitioning towards distributed DCs comprised of multiple smaller DCs. As a result of that transition, RDMA-based workload processing across widely distributed DCs will be required in the future. The most commonly used transport mode in RDMA, reliable connection (RC) mode, adopts an ACK scheme to prevent packet loss; however, that scheme exacerbates the "long fat pipe" problem as transmission distance increases, so throughput is degraded. In consideration of the above-described situation, an RDMA WAN accelerator—which improves throughput by sending pseudo-ACKs to the sender node and reduces the ACK waiting time on the sender side— is proposed. Additionally, to restore the reliability of packet delivery compromised by pseudo-ACKs, a mechanism called "segment-adaptive retransmission control" is also implemented in the accelerator. The performance (i.e., throughput relative to transmission distance) of the proposed accelerator was evaluated by using actual servers and RNICs and compared with that of standard RDMA. In the case of data transfers of 4096byte messages over a network spanning 100 km, the proposed RDMA WAN accelerator improves performance by about 20 times compared to that of standard RDMA.
Time-bin encoding is more favorable in fiber-based implementation of quantum key distribution (QKD) than polarization encoding as it avoids issues inherent for polarization encoding, such as birefringence, caused by optical fibers. QKD only with passive devices is desirable to prevent side-channel attacks possible in the case of use of active devices such as modulators. The Bennett-Brassard 1984 (BB84) protocol is a strong candidate for an implementation with satisfying these; it can be implemented using time bins with a passive delayed interferometer that inevitably generates "satellite time bins", two pulses outside the phase-interference timing. Although time-bin encoding BB84 has been frequently demonstrated, there is no consensus whether satellite time bins can be used to extract a key. Besides, there is no security proof for either case. Here, we prove the security of time-bin encoding BB84 protocol with a passive delayed interferometer and threshold detectors. If satellite time bins are used for key generation, we show that an additional operation is necessary for security. The result is not limited only to BB84 but can be applied to Bennett-Brassard-Mermin 1992 and quantum conference key agreement based on time bins.