Network Function Virtualization (NFV) is now well known for making network services deployable in a virtual environment, for instance in data-centers. On another hand, the programability of functions has come down to the network forwarding equipment itself, such as programmable switches, using the Programming Protocol-independent Packet Processors (P4) solution. Each of the two programmable concepts has its own advantages and drawbacks and micro-services should be preferably developed in one or the other solution depending on their constraints and requirements. In this paper, our objective is to combine both approaches. We propose to leverage Segment Routing (SR) to define a solution allowing to chain the micro-services to be executed at both levels. This signaling protocol is integrated within network equipment but not within a NFV infrastructure. To overcome it, we design an intermediary proxy between the P4 nodes and the VNFs. This proxy is in charge of managing the SR labels and their association with the related VNF. The demonstrator we have developed proves the feasibility of the approach and opens the way towards a composition of network services taking the best of the two levels of programmable networks.
Some kinds of application traffic, such as Cloud Gaming (CG), are particularly demanding for a network to transport because they require at the same time a low-latency and a high-bitrate. Quality of Experience (QoE) can quickly deteriorate when the network Quality of Service (QoS) is not met regarding bandwidth and delay requirements. In particular, the competition with some capacity seeking flows may induce a high queuing latency on the bottleneck (buffer-bloat phenomenon). In this paper, we evaluate two network level solutions that allow CG traffic to be processed in specific queues, but exhibit different operational constraints. The first solution uses a class-based queuing policy (Hierarchical Token Buckets, HTB), which requires prior traffic classification and some traffic engineering. The second solution leverages the new Low Latency, Low Loss & Scalable Throughput (L4S) architecture and the DualPI2 Active Queue Management (AQM), but it needs the application support. We perform extensive measurements on an experimental CG platform that integrates the L4S-compliant SCReAM CCA and that we made to evaluate both approaches regarding their QoS enforcement capability and fairness against different competing flows that are driven by TCP CUBIC or BBRv2. We show that both solutions succeed to preserve the QoS of CG traffic.
Deep learning (DL) has been successfully applied to encrypted network traffic classification in experimental settings. However, in production use, it has been shown that a DL classifier's performance inevitably decays over time. Re-training the model on newer datasets has been shown to only partially improve its performance. Manually re-tuning the model architecture to meet the performance expectations on newer datasets is time-consuming and requires domain expertise. We propose AutoML4ETC, a novel tool to automatically design efficient and high-performing neural architectures for encrypted traffic classification. We define a novel, powerful search space tailored specifically for the early classification of encrypted traffic using packet header bytes. We show that with different search strategies over our search space, AutoML4ETC generates neural architectures that outperform the state-of-the-art encrypted traffic classifiers on several datasets, including public benchmark datasets and real-world TLS and QUIC traffic collected from the Orange mobile network. In addition to being more accurate, AutoML4ETC's architectures are significantly more efficient and lighter in terms of the number of parameters. Finally, we make AutoML4ETC publicly available for future research.
Anomaly detection (AD) plays a critical role in a wide variety of big data applications, including cybersecurity, monitoring, and network systems. It consists in finding patterns in time series data that indicate unexpected events such as faults or defects. Traditional AD approaches, predominantly based on reconstruction techniques, often yield suboptimal performance, particularly when anomalies are present in the training set. Conversely, contrastive learning (CL) has shown significant performance in image processing tasks and is increasingly applied in time series data classification and forecasting. However, traditional CL frameworks are not well-adapted for time series AD due to two key challenges. First, AD is typically performed only on normal instances, and thus CL does not benefit from knowledge about anomalous instances. Second, the temporal nature of time series data is often neglected when computing time series similarity, thereby hindering the effective learning of time series representation.To overcome these limitations, we propose CATS, a novel approach that leverages a temporal similarity measure to learn time series representations. Moreover, through negative data augmentation, CATS generates a more realistic distribution of anomalies, which enables anomaly-informed CL. Extensive experiments conducted on six real-world datasets demonstrate that CATS outperforms existing AD methods. Our results highlight the efficacy of CATS in enhancing time series AD performance in big data environment across various application domains.
New services with low-latency (LL) requirements are one of the major challenges for the envisioned Internet. Many optimizations targeting the latency reduction have been proposed, and among them, jointly re-architecting congestion control and active queue management (AQM) has been particularly considered. In this effort, the Low Latency, Low Loss and Scalable Throughput (L4S) proposal aims at allowing both Classic and LL traffic to cohabit within a single node architecture. Although this architecture sounds promising for latency improvement, it can be exploited by an attacker to perform malicious actions whose purposes are to defeat its LL feature and consequently make their supported applications unusable. In this paper, we exploit different vulnerabilities of L4S which are the root of possible attacks and we show that application-layer protocols such as QUIC can easily be hacked in order to exploit the over-sensitivity of those new services to network variations. By implementing such undesirable flows in a real testbed and characterizing how they impact the proper delivery of LL flows, we demonstrate their reality and give insights for research directions on their detection.
Detecting abnormal network events is an important activity of Internet Service Providers particularly when running critical applications (e.g., ultra low-latency applications in mobile wireless networks).Abnormal events can stress the infrastructure and lead to severe degradation of user experience.Machine Learning (ML) models have demonstrated their relevance in many tasks including Anomaly Detection (AD).While promising remarkable performance compared to manual or threshold-based detection, applying ML-based AD methods is challenging for operators due to the proliferation of ML models and the lack of well-established methodology and metrics to evaluate them and select the most appropriate one.This paper presents a comprehensive evaluation of eight unsupervised ML models selected from different classes of ML algorithms and applied to AD in the context of cloud gaming applications.We collect cloud gaming Key Performance Indicators (KPIs) time-series datasets in real-world network conditions, and we evaluate and compare the selected ML models using the same methodology, and assess their robustness to data contamination, their efficiency and computational complexity.In addition to the traditional F1-score performance metric used in anomaly detection, we use Matthews Coefficient Correlation (MCC) to better differentiate between models' efficiencies.Our proposed methodology relies on window-based anomaly detection techniques as they are more useful for network operators compared to single point detection approaches.However, we found most existing window-based approaches to lack in accuracy and may under or over-estimate a model's performance.Therefore, in this paper, we propose a novel Window Anomaly Decision (WAD) approach that overcomes these drawbacks.We leverage our experimental results to provide insights about the most relevant models for detecting QoE degradation and offer recommendations on their suitability for different application requirements.
With the recent technological evolutions in networks and increased deployment of multi-tier clouds, cloud gaming (CG) is gaining renewed interest and is expected to become a major Internet service in the upcoming years. Many companies have launched powerful platforms such as Google Stadia, Nvidia GeForce Now, Microsoft xCloud, Sony PlayStation Now, among others, to attract players. However, for all end-users to fully enjoy their gaming sessions over the wide range of network access qualities, CG platforms must adapt their traffic. In this paper, we present the outcomes of a comprehensive measurement study performed on the four aforementioned CG platforms, configuring different synthetic network constraints like packet loss, throughput decrease, latency increase and jitter variation to observe the traffic of these CG platforms under degraded network conditions and infer their adaptive behaviour. We also present how the four CG platforms behave when used under real cellular network conditions, captured on the Orange network in January 2022. Our findings show that the four platforms exhibit different adaptation behaviours. Moreover, many cases result in a degraded QoS, leaving room for further improvements at both application and/or network levels.
Over the past twenty years, a plethora of methods have been proposed for encrypted traffic classification (ETC), while the Server name indication (SNI) is deemed to solve the problem of classification for TLS traffic. However, SNI-based classification has its pitfalls and the SNI will likely be pushed into the encrypted tunnel in the future. In this work, we envision a futuristic scenario in which encrypted SNI is the norm and labeled traffic flows are scarce. In such settings, we tackle the problem of traffic classification at ISP level using few-shot learning. By means of six real-world ISP-level datasets collected between 2019 and 2021 and two publicly available client-side datasets, we study the performance of a few-shot learner on TLS data, including its cross-dataset generalizability. We further investigate the effect of the number of required labeled samples on the learner's performance. Our experiments show that the dataset-specificity of deep learners carries over to few-shot meta-learning, and calls for addressing the problem of generalizability for deep learning architectures.
The Low-Latency Low-Loss Scalable throughput (L4S) architecture has recently been proposed to reduce the network latency of low-latency services and to allow their flows to coexist with classic ones in the same domain. This coexistence implies monitoring and security challenges. However current monitoring methods, primarily based-on sampling and polling, exhibit performance and granularity limitations. This paper describes the challenges for monitoring LL services and details our solution when introducing a fine-grained and real-time monitoring capability in our P4-based L4S implementation using In-band Network Telemetry. The initial experimental evaluation shows that our solution is able to monitor the metrics of an L4S switch with very few networking and processing overhead and without disturbing the L4S behaviour.
Deep learning models have shown to achieve high performance in encrypted traffic classification. However, when it comes to production use, multiple factors challenge the performance of these models. The emergence of new protocols, especially at the application layer, as well as updates to previous protocols affect the patterns in input data, making the model's previously learned patterns obsolete. Furthermore, proposed model architectures for encrypted traffic classification are usually tested on datasets collected in controlled settings, which makes the reported performances unreliable for production use. In this paper, we study how the performances of two high-performing state-of-the-art encrypted traffic classifiers change on multiple real-world datasets collected over the course of two years from a major ISP's network. We investigate the changes in traffic data patterns highlighting the extent to which these changes, also known as data drift, impact the performance of the two models in service-level as well as application-level classification. We propose best practices for architecture adaptations to improve the accuracy of the model in the face of data drift. We show that our best practices are generalizable to other encryption protocols and different levels of labeling granularity.
Cloud Gaming (CG) has been gaining a lot of interest and major actors have entered this market such as Google, Nvidia, Sony or Microsoft. They operate CG platforms that attract an increasing number of players worldwide. This type of traffic is highly demanding for network infrastructures because it requests simultaneously high bandwidth, low delay and no traffic degradation (interruptions or jitter) to ensure a good end-user’s QoE. To improve the delivery of low-latency applications, new Active Queue Management architectures like L4S (Low Latency, Low Loss, Scalable Throughput) are proposed. Currently, traffic is routed to a low-latency queue only based on the presence of the Explicit Congestion Notification bit (ECN) in the IP header, but this is too restrictive and can be easily manipulated. Instead, we aim at analyzing and detecting CG traffic based on its inherent characteristics, to forward the packets in the low-latency queue. This paper presents our models to efficiently detect CG traffic based on flow-level features among other highbitrate applications transported over UDP. The evaluation proves that our model based on decision trees achieves very good results (98.5% accuracy) and can be realistically deployed as a Virtualized Network Function at the edge, handling more than 10Gb/s of medium-sized flows on a low-end server. Our network captures and source code are open to ensure reproducible results.
Low-Iatency (LL) applications, such as the increasingly popular cloud gaming (CG) services, have stringent latency requirements. Recent network technologies such as L4S (Low Latency Low Loss Scalable throughput) propose to optimize the transport of LL traffic and require efficient ways to identify it. A previous work proposed a supervised machine learning model to identify CG traffic but it suffers from limited processing rate due to a pure software approach and a lack of generalization. In this paper, we propose a hybrid P4/NFV architecture, where a hardware Tofino based P4 implementation of the feature extraction functionality is deployed in the data plane and a unsupervised model is used to improve classification results. Our solution has a better processing rate while maintaining an excellent identification accuracy thanks to model adaptations to cope with P4 limitations and can be deployed at ISP level to reliably identify the CG traffic at line rate.
Cloud gaming applications have gained great adoption on the Internet particularly benefiting from the wide availability of broadband access networks. However, they still fail to meet users' quality requirements when accessed using cellular networks due to common wireless channel degradations. Machine Learning (ML) techniques can be leveraged to detect such anomalies during users' cloud gaming sessions. In this respect, unsupervised ML approaches are particularly interesting since they do not require labeled datasets. In this work, we investigate these approaches to understand their performance and their robustness. Our dataset consists of game sessions played on the public Google Stadia Cloud Gaming servers. The game sessions are played using a 4G network emulation replicating the capacity variations sampled on a commercial 4G network. We compare different models ranging from traditional approaches to deep learning and we evaluate their default performance while varying the level of contamination in their training datasets. Our experiments show that Auto-Encoders models achieve the best performance without contamination while the OC-SVM and the Isolation Forest are the most robust to data contamination.
Traffic classification is essential in network management for a wide range of operations. Recently, it has become increasingly challenging with the widespread adoption of encryption in the Internet, for example, as a de facto in HTTP/2 and QUIC protocols. In the current state of encrypted traffic classification using deep learning (DL), we identify fundamental issues in the way it is typically approached. For instance, although complex DL models with millions of parameters are being used, these models implement a relatively simple logic based on certain header fields of the TLS handshake, limiting model robustness to future versions of encrypted protocols. Furthermore, encrypted traffic is often treated as any other raw input for DL, while crucial domain-specific considerations are commonly ignored. In this paper, we design a novel feature engineering approach used for encrypted Web protocols, and develop a neural network architecture based on stacked long short-term memory layers and convolutional neural networks. We evaluate our approach on a real-world Web traffic dataset from a major Internet service provider and mobile network operator. We achieve an accuracy of 95% in service classification with less raw traffic and a smaller number of parameters, outperforming a state-of-the-art method by nearly 50% fewer false classifications. We show that our DL model generalizes for different classification objectives and encrypted Web protocols. We also evaluate our approach on a public QUIC dataset with finer application-level granularity in labeling, achieving an overall accuracy of 99%.
Deep learning models have shown to achieve high performance in encrypted traffic classification. However, when it comes to production use, multiple factors challenge the performance of these models. The emergence of new network traffic protocols, especially at the application-layer, as well as updates to previous protocols affect the patterns in input data, making the model's previously learned patterns obsolete. Furthermore, proposed model architectures are usually tested on datasets collected in controlled settings, which makes the reported performances unreliable for production use. In this paper, we study how the performances of two high-performing traffic classifiers change on multiple real-world datasets collected over the course of two years. We investigate the changes in traffic data patterns showing the extent to which these changes reduce the performance of the two models. Furthermore, we propose architectural adaptations to a flow time-series based traffic classifier, showing that they improve accuracy by 4.8%.
Traffic classification is essential in network management for operations ranging from capacity planning, performance monitoring, volumetry, and resource provisioning, to anomaly detection and security. Recently, it has become increasingly challenging with the widespread adoption of encryption in the Internet, e.g., as a de-facto in HTTP/2 and QUIC protocols. In the current state of encrypted traffic classification using Deep Learning (DL), we identify fundamental issues in the way it is typically approached. For instance, although complex DL models with millions of parameters are being used, these models implement a relatively simple logic based on certain header fields of the TLS handshake, limiting model robustness to future versions of encrypted protocols. Furthermore, encrypted traffic is often treated as any other raw input for DL, while crucial domain-specific considerations exist that are commonly ignored. In this paper, we design a novel feature engineering approach that generalizes well for encrypted web protocols, and develop a neural network architecture based on Stacked Long Short-Term Memory (LSTM) layers and Convolutional Neural Networks (CNN) that works very well with our feature design. We evaluate our approach on a real-world traffic dataset from a major ISP and Mobile Network Operator. We achieve an accuracy of 95% in service classification with less raw traffic and smaller number of parameters, out-performing a state-of-the-art method by nearly 50% fewer false classifications. We show that our DL model generalizes for different classification objectives and encrypted web protocols. We also evaluate our approach on a public QUIC dataset with finer and application-level granularity in labeling, achieving an overall accuracy of 99%.
New types of services with low-latency requirements have become a major challenge for the future Internet. Many optimizations, all targeting the latency reduction have been proposed. Among them, jointly re-architecting congestion control and active queue management has been particularly considered. In this effort, the L4S (Low Latency, Low Loss and Scalable Throughput) proposal aims at allowing both classic and low-latency traffic to cohabit within a single node architecture. Although this architecture sounds promising for latency improvement, it can be exploited by an attacker to perform malicious actions whose purposes are to defeat its low-latency feature and consequently make their supported applications unusable. In this paper, we analyze a set of weaknesses of L4S architecture and show that application-layer protocols such as QUIC can easily be hacked in order to exploit the over-sensitivity of those new services to network variations. By implementing undesirable flows in a real testbed and evaluating how they impact the proper delivery of low-latency flows, we demonstrate their reality and relevance for future deployments.