The correct configurations of optical infrastructure are the cornerstone for telecoms to stably provision large-capacity and long-distance service pipelines. Although traditional configuration workflows can be automated using predefined message templates and scripting, the introduction of new multi-vendor devices or required updates to network configurations forces administrators to manually review extensive documentation to implement the necessary procedures. Fundamentally, these processes are a form of natural language understanding and transformation. To address the high effort-cost, time-investment, and cross-vendor interoperability challenges, we introduce large language models (LLMs) with additional reasoning enhancement to achieve Automation-from-Intentions . Specifically, we design a reinforcement-learning-based fine-tuning paradigm of an LLM and evaluate its performance on a deployed cross-vendor optical network. The results show perfect configuration accuracy—achieving a score of 100 without hallucinations—across a 12-vendor, 256-node network. By reducing reliance on vendor-provided professional services and lowering manual operational effort, our proposal can effectively decrease both CAPEX and OPEX for telecom operators.
We propose a federated learning-based inter-operator collaborative network modeling while preventing secured data breach. International trials have demonstrated that a globally optimized model accurately estimated QoT over cross-operator heterogeneous SDM optical networks.
We propose an OCS-based scale-across architecture with reach-and cluster-scalability, while reducing the number of deployed cables. Experimental results demonstrate that 30-km extending between clusters does not deteriorate AI job completion time.
We propose in-NIC AllReduce to lower communication times, adapting to optically switched GPU network for distributed training. The experimental results show a 1.89. acceleration aligning with a theoretical analysis of 1.90. acceleration on average.
We demonstrate the multivendor use case of LLM-based transport network assimilation, configuration and telemetry automations through parameter-efficiently fine tuning the LLM over an evolutionarily designed LLM-centric control and management framework.
Hybrid Point-to-Multipoint (PtMP) and Point-to-Point (PtP) metro-access converged optical network is introduced. We successfully demonstrate five scenarios of service provisioning and fault recovery in a proposed network with coordinated provisioning and control of both systems.
We propose the node structure, network service and strategy to enable re-grouping flexibility of DSCM P2MP. We demonstrate the re-grouping use cases, acheiving the fault recovery without service loss and the dynamic adaptation to the various day-night traffic over a metro-access integration ring network. (c) 2025 The Author(s)
We present an in-network optical AllReduce that eliminates in-cast problems in distributed training. Physical-layer experiments and large-scale evaluations verify the advantages of maximum 20× bandwidth compression, 50% equipment savings, and 22× faster communications in training.
We propose an integrated control mechanism of optical circuit switching for both general data center traffics and deep distributed learning applications. Semi-physical evaluations show a relative throughput of 1.27 and a 6.18x speedup in a 256-block network constructed by MEMS-based optical switches. (c) 2024 The Author(s)
Distributed deep training has become a significant consumer of bandwidth across datacenter-scale networks. The diverse parallel strategies employed in deep training require different communication patterns, necessitating the periodic adaptation of dynamic topologies. Since electrical switching approaches its capacity limit due to high bandwidths and has difficulties in regard to topology adaptation (i.e., logical and physical topologies are isomorphic), optical switching has become an attractive option to address these bottlenecks. In this paper, we propose Modoru, a wavelength- and datarate-agnostic Clos architecture with a switching speed of $O(10\;{\rm ns})$ . Modoru is a drop-in replacement solution that has no constraints on achieving a high radix. To verify its topological flexibility, we also develop topology-as-a-service, which provisions sequentially dynamic topologies for training jobs and guarantees high topology availability over the entire network. Large-scale simulations show a basic $7.9 \times$ acceleration in deep training jobs using Modoru. Additionally, experiments on the Modoru prototype demonstrate acceleration of deep training jobs through the provisioning of adaptive topologies.
Text-to-Table aims to generate structured tables to convey the key information from unstructured documents. Existing text-to-table datasets are typically oriented English, limiting the research in non-English languages. Meanwhile, the emergence of large language models (LLMs) has shown great success as general task solvers in multi-lingual settings (e.g., ChatGPT), theoretically enabling text-to-table in other languages. In this paper, we propose a Chinese text-to-table dataset, CT-Eval, to benchmark LLMs on this task. Our preliminary analysis of English text-to-table datasets highlights two key factors for dataset construction: data diversity and data hallucination. Inspired by this, the CT-Eval dataset selects a popular Chinese multidisciplinary online encyclopedia as the source and covers 28 domains to ensure data diversity. To minimize data hallucination, we first train an LLM to judge and filter out the task samples with hallucination, then employ human annotators to clean the hallucinations in the validation and testing sets. After this process, CT-Eval contains 88.6K task samples. Using CT-Eval, we evaluate the performance of open-source and closed-source LLMs. Our results reveal that zero-shot LLMs (including GPT-4) still have a significant performance gap compared with human judgment. Furthermore, after fine-tuning, open-source LLMs can significantly improve their text-to-table ability, outperforming GPT-4 by a large margin. In short, CT-Eval not only helps researchers evaluate and quickly understand the Chinese text-to-table ability of existing LLMs but also serves as a valuable resource to significantly improve the text-to-table performance of LLMs.
We train language models to automate the diagnosis of OTN configuration errors, and the diagnostic accuracy is up to 97.56%. We additionally demonstrate the effectiveness of the models on a real OTN system.
We propose O(10ns) optical switching structure with wavelength-elastic and topology-flexibility in support of distributed deep learning platform. Our experimental demonstration shows accelerations for three use cases.
Fitting stochastic input-process models to data and then sampling from them are key steps in a simulation study but highly challenging to non-experts. We present Neural Input Modeling (NIM), a Generative Neural Network (GNN) framework that exploits modern data-rich environments to automatically capture simulation input processes and then generate samples from them. The basic GNN that we develop, called NIM-VL, comprises (i) a variational autoencoder architecture that learns the probability distribution of the input data while avoiding overfitting and (ii) long short-term memory components that concisely capture statistical dependencies across time. We show how the basic GNN architecture can be modified to exploit known distributional properties-such as independent and identically distributed structure, nonnegativity, and multimodality-to increase accuracy and speed, as well as to handle multivariate processes, categorical-valued processes, and extrapolation beyond the training data for certain nonstationary processes. We also introduce an extension to NIM called Conditional Neural Input Modeling (CNIM), which can learn from training data obtained under various realizations of a (possibly time series valued) stochastic "condition," such as temperature or inflation rate, and then generate sample paths given a value of the condition not seen in the training data. This enables users to simulate a system under a specific working condition by customizing a pre-trained model; CNIM also facilitates what-if analysis. Extensive experiments show the efficacy of our approach. NIM can thus help overcome one of the key barriers to simulation for non-experts.
Simulation metamodeling is essential for speeding up optimization via simulation to support rapid decision making. During optimization, the metamodel, rather than expensive simulation, is used to compute objective values. We recently developed graphical neural metamodels (GMMs) that use graph neural networks to allow the graphical structure of a simulation model to be treated as a metamodel input parameter that can be varied along with scalar inputs. In this paper we provide novel methods for using GMMs to solve hybrid optimization problems where both real-valued input parameters and graphical structure are jointly optimized. The key ideas are to modify Monte Carlo tree search to incorporate both discrete and continuous optimization and to leverage the automatic differentiation infrastructure used for neural network training to quickly compute gradients of the objective function during stochastic gradient descent. Experiments on stochastic activity network and warehouse models demonstrate the potential of our method.
We demonstrate the fiber-to-application transport slicing architecture and mechanism. The experiment shows ultrahigh throughput (> 5Gbps per application) and significant acceleration for 100 applications in 4 categories. © 2022 The Author(s)
We designed the CoF protocol and demonstrated a slicing method based on this protocol. Experiment showed that the proposed method significantly increases the throughput and reduces the completion time of computing applications.
We experimentally demonstrate the dynamic reconfiguration of WDM VNTs in response to SDM spatial channel failures. We present an SDN control architecture with gRPC-telemetry and analytics to detect failures and restore failed virtual WDM links ©2022 The Authors.
Pre-trained Language Models (PLMs) are the cornerstone of the modern Natural Language Processing (NLP). However, as PLMs become heavier, fine tuning all their parameters loses their efficiency. Existing parameter-efficient methods generally focus on reducing the trainable parameters in PLMs but neglect the inference speed, which limits the ability to deploy PLMs. In this paper, we propose LayerConnect (hypernetwork-assisted inter-layer connectors) to enhance inference efficiency. Specifically, a light-weight connector with a linear structure is inserted between two Transformer layers, and the parameters inside each connector are tuned by a hypernetwork comprising an interpolator and a down-sampler. We perform extensive experiments on the widely used the GLUE benchmark. The experimental results verify the inference efficiency of our model. Compared to Adapter, our model parameters are reduced to approximately 11.75%, while the performance degradation is kept to less than 5% (2.5 points on average).