Due to the proliferation of mobile applications, mobile traffic identification plays a crucial role in understanding the network traffic. However, the pervasive unconcerned apps and the emerging apps pose great challenges to the mobile traffic identification method based on supervised machine learning, since such method merely identifies and discriminates several apps of interest. In this paper we propose a three-layer classifier using machine learning to identify mobile traffic in open-world settings. The proposed method has the capability of identifying traffic generated by unconcerned apps and zero-day apps; thus it can be applied in the real world. A self-collected dataset that contains 160 apps is used to validate the proposed method. The experimental results show that our classifier achieves over 98% precision and produces a much smaller number of false positives than that of the state of the art.
With the coming of the big data age, data mining attracts more and more attention from all trades and professions. Due to the vast computation cost of data mining, the public service platform for big data mining has become the urgent needs, especially for the model training tasks. In this way, how to perform this kind of task scheduling becomes critical. This paper focuses on the assignment of tasks on multiple computing resources to optimize the total operation time. Firstly, a task scheduling algorithm based on the greedy and genetic algorithm is proposed to set the computation resource requirement for each task. Moreover, a greedy strategy is used to decide the task operation order and the assignment mapping between the tasks and the computation resources. Finally, the proposed algorithm proves to be efficient by several experiments.
Collaborative filtering is the mainstream approach to personalized recommendation. For the cold start problem faced in collaborative filtering, it is a hot research topic to introduce the user's side information into the recommendation model. Different from the matrix decomposition idea adopted in the existing methods, we propose a Top-N recommendation model using the side information of the user based on the reconstruction function of the stacked denoising auto-encoder. Experimental results show that the model outperforms the existing method in Recall. In addition, we explore the influence of missing ratings and user side information vector into the loss computation. The experimental results show that ignoring the missing ratings in the loss function is beneficial to improve the performance of the model.
为了解决传统滑模观测器方法应用在永磁同步电机无传感器矢量控制时所产生的抖振问题,使用RBF神经网络动态调节观测器的切换增益,即使其输入为传统滑模估计方案中的电流估计误差,输出为滑模增益;同时为了简化系统结构、提高方案可行性,将RBF神经网络设计为单输入单输出的结构,并将网络的学习和工作过程融合,使其在自身网络参数的不断优化中实时输出滑模增益,以增强系统鲁棒性.最后通过Matlab/Simulink软件对该系统进行建模仿真,并将该方法与传统滑模观测器方法进行对比.实验结果表明,该方案能够为矢量控制提供更加准确的转子位置及速度信息,提高了整个电机控制系统的稳定性.
Decision tree is one of the most popular supervised machine learning algorithms. Due to the rapid increase in the amount of data, many application scenarios require higher classification speed. Therefore, a variety of decision tree classification acceleration algorithms based on FPGA are proposed. These methods focus on improving the classification speed by designing effective pipeline architecture. However, the impact of floating-point numbers on storage and computing resources in the hardware implementation of the decision tree is ignored. In this paper, we present a discretization method for floating-point numbers in decision tree model by converting the floating-point numbers into integers. This method reduces the storage and computing resources required by the hardware implementation of decision tree, without affecting the classification performance of the classifier.
The upcoming Fifth Generation (5G) networks can provide ultra-reliable ultra-low latency vehicle-to-everything for vehicular ad hoc networks (VANET) to promote road safety, traffic management, information dissemination, and automatic driving for drivers and passengers. However, 5G-VANET also attracts tremendous security and privacy concerns. Although several pseudonymous authentication schemes have been proposed for VANET, the expensive cost for their initial authentication may cause serious denial of service (DoS) attacks, which furthermore enables to do great harm to real space via VANET. Motivated by this, a puzzle-based co-authentication (PCA) scheme is proposed here. In the PCA scheme, the Hash puzzle is carefully designed to mitigate DoS attacks against the pseudonymous authentication process, which is facilitated through collaborative verification. The effectiveness and efficiency of the proposed scheme is approved by performance analysis based on theory and experimental results.
In this paper, we consider the sparsely used dictionary learning problem and focus on the case that the dictionary is square with arbitrary entries and the coefficient matrix is a sparse with random entries, which is first proposed and studied by Wang and Spielman et al.[1], [2]. We improve the theoretical results with a tighter bounded uniqueness theorem, which says that O(n) samples are sufficient to uniquely determine the decomposition with probability one when given assumption is satisfied, and the proof is given.
为了提高实时性和准确性,提出一种改进的动态时间规整算法(Dynamic Time Warping-DTW),用于度量手势运动轨迹的相似性,实现了快速的精确动态手势识别.首先,通过Kinect2传感器实时地获取人体骨架的关节点坐标和手部的形状数据,然后构造矢量特征描述手的运动轨迹,运用动态时问规整方法进行模板匹配,并对特殊手势进行精确的二次分类,实现了基于轨迹匹配的快速动态手势识别.实验证明:该方法识别准确度高,实时性好,对光照强度和复杂背景干扰有很强的鲁棒性.
Sparse subspace clustering is a well-known algorithm, and it is widely used in many research field nowadays, and a lot effort has been contributed to improve it. In this paper, we propose a novel approach to obtain the coefficient matrix. Compared with traditional sparse subspace clustering (SSC) approaches, the key advantage of our approach is that it provides a new perspective of the self-expressive property. We call it rigidly self-expressive (RSE) property. This new formulation captures the rigidly self-expressive property of the data points in the same subspace, and provides a new formulation for sparse subspace clustering. Extensions to traditional SSC could also be cooperating with this new formulation. We present a first-order algorithm to solve the nonconvex optimization, and further prove that it converges to a KKT point of the nonconvex problem under certain standard assumptions. Extensive experiments on the Extended Yale B dataset, the USPS digital images dataset, and the Columbia Object Image Library shows that for images with up to 30 % missing pixels the clustering quality achieved by our approach outperforms the original SSC.
With the rapid advances in Internet technology, publishing real-time statistics data, in a privacy-preserving way, has led to a large body of research. The current state-of-the-art paradigm for privacy preserving with differential privacy on data stream is w-event privacy. But it neglects if only a few part of the elements of dataset change over time and others are substantially stabilize, then processing all the user data in specified timestamps will bring additional noise and reduce the utility of data. In this paper, a novel privacy preserving approach called G-event which follow the conventional use of w-event differential privacy is proposed. We group the statistics result at each timestamp based on difference calculation. Then the high difference group will publish more often than the similar group. We guarantee that all result with greater change will publish by adding noise, and the result with smaller change will be approximate with the corresponding lastly published statistics. Experiment using real-life dataset show that our approach improves the utility of data.
With the rapid increasing number of mobile devices being used as essential terminals or platforms for communication, security threats now target the whole telecommunication infrastructure. However, most of the software-based passive measurement system (PMS) could not achieve high performance to adapt the high-bandwidth mobile network. In this paper, we propose NTW, a real-time pre-processing system deploys between high-speed mobile core network and PMS to accomplish the large and repetitive work, such as decapsulation, decompression and PPP character unescape before the PMS receives the traffic flow. We evaluated the performance and accuracy of our NTW over a wide-area real network. Evaluation results indicate that NTW can achieve more than 15Gbps.
本文针对卫星网络资源有限的情况下,特别是当卫星网络带宽迅速降低至难以保障所有数据流的基本带宽时,基于服务满意度和中断服务不满意度的概念,设计并实现了一种面向服务满意度的数据流接纳控制决策机制.该决策机制给出了在无数据流中断模式下的最优求解算法和有数据流中断模式下的近似最优求解算法,对所有数据流进行接纳控制决策,决定各条数据流的通断情况及各自的带宽值,使得服务满意度达到最大.通过配置数据流对接纳控制决策机制进行仿真测试,实验结果验证了该决策机制在解决接纳控制决策优化问题上的优越性.
攻击源威胁行为评估是骨干网安全监测条件下海量报警信息处理的迫切需要.传统的安全评估方法研究侧重于信息系统的安全性评测,无法有效利用骨干网视窗优势评估攻击源威胁能力差异.本文在分析网络攻击源的行为特点的基础上,分类并量化多维度评估指标,并借助AHP层次分析法建立了基于“目标—准则—指标”三层评估体系的动态评估模型.实验结果表明,该方法能动态有效的评估网络攻击源在其所处监测环境下的威胁能力.
A promising approach to protect driver's location privacy in vehicular ad hoc network (VANET) suggests vehicle changing pseudonyms in regions called mix-zones, where the adversary cannot eavesdrop the vehicular communication. How to deploy mix-zones in a large city is a challenge problem and has not been well addressed in previously reported works. In this paper, we propose a statistics-based metric for evaluating the effectiveness of a mix-zone and selecting mix-zone candidates in term of privacy requirement. Furthermore, a cost-efficient mix-zones deployment scheme is presented to guarantee that vehicles at any place could pass through an effective mix-zone in certain driving time, and the extra overhead time of adjusting routes to across the mix-zone is small. Extensive simulations demonstrate that the proposed evaluation metric is viable under various traffic scenarios while the deployment plans generated by our scheme in a real-world map make drivers have more chances to pass through mix-zones on road.
Due to the limited communication range and the unbalance distribution of Wi-Fi access points, the disconnection gap during vehicular Internet access almost depends on the access point distribution along the driving path. In this way, with the vehicle mobility statistics, the service provider of account-based web service (i.e., Email and Weblog) could mine the current access point of an interested account owner by the sequence of disconnection gap. As the first paper to study this issue, we propose a probability algorithm for location mining that can be computed recursively with the length of communication feature sequence. Real map based simulations evaluate the location privacy threat and the relationship with the density of access points and communication feature classes.
In view of the present situation of large scale and high speed network.A method of worm detection was presented based on analysis of similarity of payload of connection which compute similarity of connection by using computing hanming distance of payload of connection.Comparing with arithmetic of longest common subsequence,this method can reduce computational resource consumption.And on this basis,present a detection system com bining with the coarse-grained anomaly detection and fine-grained analysis of behavior.Further exclude non worm traffic,focus on worm traffic and reduce the similarity calculation.The experiment proved this method can detect unknown worm.
Distributed denial of service(DDoS)attack is a serious threat to Internet security. Target networks and hosts will be overwhelmed by massive traffic when attack happens. It is important for the defense against DDoS attack to detect the attack quickly and accurately,discriminate the attack traffic from legitimate crowd traffic to eliminate attack traffic,and eliminate the attack traffic. The entropy is used to execute real-time statistics of some flow parameters for detecting the attack,and cumulative sum(CUSUM)algorithm is employed to track continuous changes of the entropy. According to the growth of destination IP quantity,victims can be discovered,and then the traffic swarming into the victims is observed emphatically. As the large-scale attack traffic and legitimate crowd traffic are very similar,it is difficult to recognize attack traffic. The correlation coefficient is used in this paper to check the similarity of the flow to discriminate the attack traffic from legitimate crowd traffic,which provides an evidence for subsequent elimination and filtering.
Successive interference cancellation (SIC) is an effective way of multipacket reception to combat interference. As conventional CSMA (Carrier Sense Multiple Access) is designed for single packet reception, it is unclear whether or not CSMA performs well to exploit the SIC capability. In this paper, we analyze the performance of a simple CSMA protocol in a network with SIC. For a given link, we derive the residing areas of an interfering node when simultaneous transmission is allowed and when the interference is harmful, respectively. We show that, though SIC provides many new transmission opportunities, CSMA cannot effectively exploit them. There is a fundamental tradeoff in a CSMA protocol between exploiting the transmission opportunities from SIC and capturing the harmful interference. In many cases, when CSMA achieves its best performance, almost all new transmission opportunities are not exploited. It is therefore very necessary to design a new distributed access protocol in wireless networks with SIC.
To deal with the rapid increment of network traffic,an Intrusion Detection System(IDS) based on commodity multi-core platform was proposed.This paper evaluated some critical factors for the system performance,such as hardware,resource-assigning and network traffic features.Extensive experiments demonstrate that number of traffic flow and pps index have larger impact on system performance.The ids performance can be improved obviously by activating the Hyper-Threading of multi-core processor and binding the detection engine with the CPU core.Our system is easy to be realized and has low price-performance ratio.