Archiving raw network traffic is fundamental for network forensics and troubleshooting. In order to efficiently retrieve specific entries from large-scale archives, multi-attribute queries are essential. However, most previous work mainly focus on single-attribute indexing, leading to limited performance for multi-attribute queries and index updates. To this end, we propose BK-Index, a novel multi-attribute index algorithm for network traffic. BK-Index combines multi-dimensional k-ary search trees and bitmap structures to efficiently handle both high-cardinality and low-cardinality attributes. This hybrid approach significantly reduces multi-attribute retrieval response time. Additionally, BK-Index employs a periodic batch construction scheme for real-time index updates in high-speed network links, eliminating the need for costly dynamic index maintenance. A comprehensive suite of experiments conducted on actual campus network traffic data has elucidated the superior performance of the BK-Index in terms of multi-attribute indexing and retrieval compared with other state-of-the-art index methods.
Socioeconomic Status (SES), an overall measure of a person's economic and social status relative to others combining factors such as economics and sociology, has received a lot of attention from researchers, as its assessment can help relevant orga-nizations to make various policies and decisions (governmental formulation of social policies, advertising personalized services, etc).In addition, with the development of big data technology and machine learning in recent years, assessing people's socioeconomic attributes (SEAs) and further obtaining the corresponding socioeconomic status with a data-driven approach can address the issue of extremely high cost of traditional methods.Therefore, this paper summarizes the research progresses of applying big data techniques to socioeconomic status analysis in recent years.It first introduces the basic concept of socioeconomic status and discusses the challenges posed by big data methods compared to traditional methods.After that, it systematically summarizes and classifies the state-of-the-art related methods based on the information in the learning process, and present them in detail, discusses the pros and cons of each type of method.Finally, it discusses the challenges and problems of inferring people's socioeconomic status and provides an outlook on future research directions.
Nowadays, research on social networks has attracted a large amount of attention from both academic and industrial societies. To understand the diffusion process and guide viral marketing, it is of importance to model and then estimate the influence of a seed user on a target user, which is defined as target influence in this paper. In famous diffusion models like independent cascade model and linear threshold model, tremendous computational costs are usually required in estimating influence probability through simulation. In this paper, we adopt duplicate forwarding model, and propose two measurements for the target influence, which can be analyzed theoretically. The former is the average number of duplicates the target user receives, and the latter is the probability of the target user receiving at least one duplicate. We further find the former will approach infinity if the spread intensity exceeds some threshold, but the latter can be adopted without this constraint. We also seek to use the latter to estimate the influence probability in the independent cascade model, and find it achieves much better accuracy than other heuristic metrics. All results are verified through simulations in real-world social networks, and we believe approach proposed here can provide insights to solve the problems like target influence maximization and influence maximization.
Predicting individual socioeconomic status (SES) from social media content benefits various applications in economic and social fields. Most previous works adopt machine learning methods with predefined features to infer SES. Nevertheless, they ignore some important information of social media content, such as order, structure and relation information, which leads to limited performance. In this paper, we propose a COupled social media content REpresentation model (CORE) for individual SES prediction, which efficiently exploits latent complex couplings of social media content. CORE devises a structure-aware social media text representation method to incorporate the order and the hierarchy of social media text, and leverages a coupled attribute representation method to take into account intra-coupled and inter-coupled interaction relationships among user level attributes. Our experiments on a real data set of a Chinese microblogging platform demonstrate that our approach significantly outperforms benchmark methods, which validates its efficiency and robustness. The proposed model could be applied to improve the SES prediction and other user profiling tasks.
With the explosive growth of users’ online information, user profiling, inferring users’ traits or interests, has attracted increasing attention due to its various applications in reliable personalized service and recommender system fields. Most existing studies regard user profiling as a node classification task and utilize graph-based methods to exploit users’ relations and affiliated information in heterogeneous graphs. However, they only consider single view of pair-wise relationships between users, ignoring the modeling of potential higher-order interactions among users. To this end, in this paper, we propose a novel Interaction-aware Hypergraph Neural Networks (IHNN) model for user profiling, which introduces hypergraphs to formulate high-order interactions among users from multiple views to achieve more accurate user profiles. Specifically, IHNN utilizes heterogeneous attention mechanism to exploit important interactive information from diversified users’ affiliated data and designs hypergraph convolutional operation to mine the interactions beyond pairwise among users. Besides, multiple hypergraphs are constructed from explicit and implicit interactions among users obtained from several types of data, aiming to leverage the complementarity of views. Extensive experiments on two real-world e-commerce datasets demonstrate our proposed model brings a significant performance boost compared with baseline models for user profiling.
Inferring users' occupational categories on the basis of user-generated content becomes an important issue in user profiling and applications such as personalised recommendation systems with the rapid explosion usage of online social me-dia. Although previous work has demonstrated that language features extracted from social media content can effectively predict users' occupations, work on overcoming the challenge of time-consuming, expensive expert knowledge and low prediction performance is fairly limited and mostly based on English social platforms. In this paper, we first investigate the relationship between users' language usage in their Chinese blogs and users' occupations, employing tools to extract quantitative features related to users' psychological states and social relationships. Additionally, We propose a novel content-aware hierarchical model called T-LSTM for the user occupation prediction, which is mainly divided into a word-level Transformer encoder layer overcoming the problem of neglecting mining the importance of words in users' texts and a blog-level bidirectional LSTM layer exploiting temporal information of blogs to obtain users' representations. Our experimental results on our collected real-world Chinese social media dataset shows that the proposed model greatly outperforms the baseline methods for occupation prediction and verifies the effectiveness of components as well as the robustness of the model.
Socioeconomic status (SES) is an important economic and social aspect widely concerned. Assessing individual SES can assist related organizations in making a variety of policy decisions. Traditional approach suffers from the extremely high cost in collecting large-scale SES-related survey data. With the ubiquity of smart phones, mobile phone data has become a novel data source for predicting individual SES with low cost. However, the task of predicting individual SES on mobile phone data also proposes some new challenges, including sparse individual records, scarce explicit relationships and limited labeled samples, unconcerned in prior work restricted to regional or household-oriented SES prediction. To address these issues, we propose a semi-supervised hypergraph-based factor graph model (HyperFGM) for individual SES prediction. HyperFGM is able to efficiently capture the associations between SES and individual mobile phone records to handle the individual record sparsity. For the scarce explicit relationships, HyperFGM models implicit high-order relationships among users on the hypergraph structure. Besides, HyperFGM explores the limited labeled data and unlabeled data in a semi-supervised way. Experimental results show that HyperFGM greatly outperforms the baseline methods on a set of anonymized real mobile phone data for individual SES prediction.
Accurate prediction is highly important for clinical decision making and early treatment. In this paper, we study the imbalanced data problem in prediction, a key challenge existing in the healthcare area. Imbalanced datasets bias classifiers towards the majority class, leading to an unsatisfied classification prediction performance on the minority class, which is known as imbalance problem. Existing imbalance learning methods may suffer from issues like information loss, overfitting, and high training time cost. To tackle these issues, we propose a novel ensemble learning method called Multiple bAlance Subsets Stacking (MASS) by exploiting a multiple balance subsets construction strategy. Furthermore, we improve MASS with introducing parallelism (Parallel MASS) to reduce the training time cost. We evaluate MASS on three real-world healthcare datasets, and experimental results demonstrate that its prediction performance outperforms the state-of-art methods in terms of AUC, F1-score and MCC. Through the speedup analysis, Parallel MASS reduces the training time cost greatly on large dataset, and its speedup increases as the data size grows.
The notion of socioeconomic status (SES) of a person or family reflects the corresponding entity's social and economic rank in society. Such information may help applications like bank loaning decisions and provide measurable inputs for related studies like social stratification, social welfare and business planning. Traditionally, estimating SES for a large population is performed by national statistical institutes through a large number of household interviews, which is highly expensive and time-consuming. Recently researchers try to estimate SES from data sources like mobile phone call records and online social network platforms, which is much cheaper and faster. Instead of relying on these data about users' cyberspace behaviors, various alternative data sources on real-world users' behavior such as mobility may offer new insights for SES estimation. In this paper, we leverage Smart Card Data (SCD) for public transport systems which records the temporal and spatial mobility behavior of a large population of users. More specifically, we develop S2S, a deep learning based approach for estimating people's SES based on their SCD. Essentially, S2S models two types of SES-related features, namely the temporal-sequential feature and general statistical feature, and leverages deep learning for SES estimation. We evaluate our approach in an actual dataset, Shanghai SCD, which involves millions of users. The proposed model clearly outperforms several state-of-art methods in terms of various evaluation metrics.
Social community question answering (SCQA) sites not only provide regular question answering (QA) service but also form a social network where users can follow each other. Identifying topical opinion leaders who are both expert and influential in SCQA becomes a hot research topic. However, existing works focus on either using knowledge expertise to find experts for improving the quality of answers, or measuring user influence to identify influential ones. In this paper, we propose QALeaderRank, a topical opinion leader identification framework, incorporating both the topic-sensitive influence and the topical knowledge expertise. To measure a user's topic-sensitive influence, we design a novel ranking algorithm that exploits both the social and QA features of SCQA, taking account of the network structure, topical similarity and knowledge authority. Besides, we incorporate three topic-relevant metrics to infer the topical expertise. Extensive experiments along with a user study demonstrate that QALeaderRank outperforms the compared state-of-the-art methods. QALeaderRank can also be used to identify multi-topic opinion leaders.
In Software-Defined Networking (SDN), central controllers can obtain global views of dynamic network statistics to manage their networks. In order to support SDN controllers to obtain global information of the networks, the data planes need to maintain a large number of counters, which are typically implemented in hardware such as ASIC. However, implementation of these counters in hardware faces critical challenges: high memory consumption, control inflexibility, and low statistical accuracy. In this paper, we present the concept of Software Defined Hardware Counters (SDHC) for SDN, which offloads the management of counter updating to software and still maintains practical execution efficiency in hardware. Therefore, SDHC can allocate counter memory on demand to enhance counter utilization, which greatly reduces on-chip memory consumption. It is able to allow controllers to flexibly control the counters through south-bound interface. Besides, with novel statistics feedback mechanisms, SDHC supports high-accuracy and active statistical requirement applications. Through a prototype implementation and performance evaluation based on FPGA and general processor, we reveal that the proposed SDHC is able to achieve high processing performance and high statistical accuracy, which incurs negligible updating delay to the switches.
Implementation of counters is a critical challenge for switches in today's Software-Defined Networking (SDN). In this paper, we address the current challenges in implementing SDN counters: high memory consumption, low utilization, and inflexibility. We introduce the concept of software defined hardware counters (SDHCs) for SDN. Our main idea is to make the switch-local CPU flexibly allocate memory space to each counter required by controllers. The ASIC of SDN switches transmits event records to the CPU, which contain updating information of the counters. Furthermore, the ASIC provides non-semantic counter memory space to be allocated by the CPU. Based on the proposed SDHC, an SDN controller can flexibly apply/release various counters for each counter category (e.g., each flow entry, each port) through the south-bound interface. It is shown that SDHC achieves high flexibility while reducing the memory space on ASIC. It also improves the update performance through alleviating the CPU overhead. Finally, we evaluate the performance of SDHC through comprehensive simulation study.
Software Defined Networking (SDN) provides efficient network and traffic management for data center network. As underlying devices in SDN, SDN switches must maintain a large number of hardware counters. Implementation of these counters faces serious challenges for SDN switches, i.e., High memory consumption and inflexibility. Thus, we previously proposed Software Defined Hardware Counters (SDHC), which decouples definition and implementation of counters to overcome these challenges. However, like traditional hardware counters, SDHC only supports passive statistical mode (i.e., The values of the counters can be only read passively by the controller). Based on the passive mode, most of applications need to send request messages at some frequency to obtain statistics, which causes some critical problems for SDN: i) low statistical accuracy, ii) high network bandwidth consumption. Hyper Software Defined Hardware Counters (Hyper SDHC) is thus proposed by extending SDHC. Through introducing the timer-triggering and updating-triggering statistics-reporting mechanisms, Hyper SDHC can naturally support active statistical mode, i.e., Counters actively report their values according to triggering condition. It can greatly enhance the statistical accuracy and reduce network bandwidth consumption between controller and switch. The demo of Hyper SDHC is implemented based on Net Magic platform. The demo will exhibit how Hyper SDHC works and how it supports a typical video quality monitor application.
The paper presents a remote-desktop based NetMagic remote debugging RNP (Remote-desktop based NetMagic Remote Debugging Platform) model for the situation that hardware programming experiment is not often conducted in the network innovation experiment and teaching. It allows researchers to operate NetMagic remotely, which is a kind of open source hardware network study experiment platform, using the remote-desktop key technology. Additionally, we preliminarily implemented the remote debugging platform based on this model and verified the key technology of the remote-desktop connection. The work of this paper has an important guiding significance for the network innovation experiment and teaching.