
Coordination and situation awareness are amongst the most important aspects of collaborative analysis in smart buildings. They are especially useful for emergency responses such as firefighters, police, and military soldiers. Currently, the communication between command centers and crews are often performed via voice, cameras, and possibly hand-held devices; which offer limited and in-efficient solutions. This work investigates a mixed-reality platform to improve the coordination and situation awareness for multiple users performing real-time operations in smart buildings. Our platform provides a flexible architecture to synchronize crew locations in buildings and vital information across the team and command center in real-time. We have also developed several immersive interaction functions to support efficient exchange of useful visual information. The case study and example results demonstrate the advantages of our immersive approach for on-site collaboration in real physical environment.
Spreading and transmitting pornographic images over the Internet in the form of either real or artificial images is illegal and harmful to teenagers. Because traditional methods are primarily designed to identify real pornographic images, they are less efficient in dealing with artificial images. Therefore, a novel feature selection and post-processing method for the recognition of artificial pornographic images in social networks was proposed in the work. Firstly, features related to image size, skin color region, gray histogram, image color, edge density and direction, Gray Level Co-occurrence Matrix (GLCM) and Local Binary Patterns (LBP) were selected. Secondly, a post-processing process for these multiple feature was proposed, which includes two steps. The first step is feature expansion, which is aimed at improving the generalization ability of the recognition model. The other step is rapid feature extraction, which is aimed at reducing the time required for image recognition in social networks. Finally, experimental results demonstrate that the proposed method is effective for the recognition of artificial pornographic images in social networks.
User-centric service is wildly adopted in the Cloud environment, but its comprehensible composition is still challenge. On the one hand, intuitive programming perspective is required to hide low-level details with the reasonable abstract; on the other hand, effective guarantee is necessary to verify the application’s legality. In this paper, we present a service composition method through multiple user-centric views in which different abstract is provided by orthogonal and orderly views. Domain applications can be synthesized and their consistency is guaranteed coordinately.
The value of GPS data has generated a group of location-based services. Pick-up points recommendation by mining taxis’ trajectories can effectively both improve drivers’ profits and reduce oil consumption. However, existing methods always ignore the spatial-temporal features and the drivers’ preferences. Therefore, we propose to recommend a personalized sequence of pick-up points taking the two preceding factors into account. Firstly, we extract historical pick-up points from taxis’ trajectories and use these points to generate candidate ones by a novel approach of spatial-temporal analysis. Secondly, we devise a collaborative filtering algorithm to choose candidate points again. According to the location and the time of historical pick-up points, our system can give taxi-drivers an optimal sequence of pick-up points. Experimental results show that our method can obviously improve both the accuracy and the preference of candidate pick-up points for taxi-drivers.
Facing massive services with different non-functional properties, obtaining the optimal composite service consumes considerable time. In order to reduce the search space and shorten the time of looking for the approximately optimal composite service, several methods based on skyline have been raised. However, these methods mainly focus on Quality of Service (QoS), which cannot describe non-functional properties adequately. Thus, service contracts are widely researched to make up for the deficiencies of QoS. Therefore, how to define the skyline services based on QoS and service contracts is becoming a critical challenge. To attach this issue, this paper proposes contract-oriented skyline services, including non-personalized skyline services and personalized skyline services. In addition, we discuss the natures of personalized skyline services. Besides, to deal with excessive skyline services, a method based on hierarchical clustering is presented, which contributes to the algorithm for contract-oriented service composition. Eventually, the verification experiments have been conducted, which illustrate the effectiveness of our methods.
The current Sensor Networks are generally domain-specific and task-oriented, tailored for particular applications with little possibility of sharing and reusing sensor data for different applications. The servitization of stream sensor data is an effective solution for sharing and reusing sensor data resources. Considering the limitations of existing methods on processing large-scale stream data and concurrent requests, this paper proposes a lightweight model for stream sensor data service, which processes sensor stream data with service modeling operations, and distributes data based on Pub/Sub mechanism. This paper fulfills the encapsulation by utilizing event-driven mechanism for stream data service processing, using SparkStreaming framework to process sensor events, and improving traditional matching-tree algorithm to distribute stream data efficiently. Finally, we evaluate our approach through experiments, and the stream data service can handle multiple requests and deliver data to corresponding applications in milliseconds level.
Running data-intensive applications across the geo-distributed data centers in cloud computing needs to address the problem of how to place the data items to the appropriate data centers. The general methods are mainly hash-based which could be understood as random placement intuitively when the query needs distributed data items. In this paper, We propose an genetic based data placement (GBDP) scheme in which a tripartite graph based model is constructed to formulate the data replica placement problem by leveraging the genetic algorithm, and decompose the original problem into two simplified subproblems, which are solved alternately. Through extensive experiments with synthesized and realistic data items, the performance of the proposed scheme is proved validated.
Entities play an important role in many natural language applications. Based on the Automatic content Extraction (ACE) conference, we study the extraction technologies of entity mentions in Chinese text. Compared to named entities, entity mentions have rich categories and complex structures, which bring great difficulty to the extraction task. To solve the above problems, we propose an unsupervised method to detect entity mentions and identify their categories in Chinese text, namely Un-MenEx. With the abundant data of Baidu Baike and Baidu search, Un-MenEx exploits a similarity calculation method to extract entity mentions in text, which solves the problem of identifying rare entity names difficultly and optimizes the mentions segmented wrongly. Moreover, Un-MenEx can meet the demand of processing massive data by reason of no manual annotation data. We conduct the experiments with the news text, and the experimental results show that this method has practical application value, and ensure the accuracy requirement.
Clustering is an important tool for data mining and analysis for massive data in big data. This paper proposes a clustering model of high-dimensional data based on the density peak cluster algorithm and accomplishes clustering for more than six-dimensional data with arbitrary shape simply and directly. This model achieves automatically pre-process and takes local points with larger density and far away from other local points as the clustering center followed by introducing the fine-tuning. Experimental results suggest that our model not only works for low-dimensional data, but also achieves promising performance for high-dimensional data.
The number of services is proliferating dramatically and the rate of services' evolution has also been increasingly fluctuating in recent years. The demands of service composition also show the characteristics of individuation and diversification at the same time. The traditional methods of service composition are difficult to meet the multiple granularity demands of users. This paper proposes a novel multiple granular service composition model based on services granular space. The model firstly constructs service granularity by service clustering. And then constructs the service granularity space according to the relationships between service granularities. So the process of getting appropriate service compositions can be transformed into getting service compositions from different granularity layers. Through experimental analysis, we can demonstrate that this model can provide users with different granularity service compositions which meet the multiple granularity demands of users. And can also decrease the response time of service composition at the same time.
In open conditions of Internet of Things, massive data would be rapidly accumulated from sensors in low quality. On huge size raw data, the correction for consistency is time-consuming and inaccurate to achieve, and the validation for legality is difficult to guarantee without prior knowledge. In this paper, time-based clustering and rule-based filtering for data cleaning is proposed on massive bus IC card data, which guarantees the consistency and legality among spatio-temporal attributes. Implemented through Hadoop MapReduce and evaluated on real data set, our method shows its efficiency and accuracy in extensive conditions.
With the rapid advances in Internet technology, publishing real-time statistics data, in a privacy-preserving way, has led to a large body of research. The current state-of-the-art paradigm for privacy preserving with differential privacy on data stream is w-event privacy. But it neglects if only a few part of the elements of dataset change over time and others are substantially stabilize, then processing all the user data in specified timestamps will bring additional noise and reduce the utility of data. In this paper, a novel privacy preserving approach called G-event which follow the conventional use of w-event differential privacy is proposed. We group the statistics result at each timestamp based on difference calculation. Then the high difference group will publish more often than the similar group. We guarantee that all result with greater change will publish by adding noise, and the result with smaller change will be approximate with the corresponding lastly published statistics. Experiment using real-life dataset show that our approach improves the utility of data.
In recent years, smartphone-based human activity recognition has become a promising research field of mobile computing, and is widely applied in inertial positioning, fall detection, and personalized recommendation. In practical scenario, smartphone can be placed at several body positions, such as trouser pocket, jacket pocket and so on. Since data is collected from the accelerometer embedded in smartphone, different body locations cannot generate consistent data for the same activity. As a result, the samples at a new position usually obtains low recognition rate from the classifier trained by the original data collected from other positions. In this paper, we propose a COntinuity-based POsition-adaptive recognition method, abbreviated COPO, for dealing with this problem. Considering the continuous results with high probability of correct recognition, we select them as the retraining data in COPO for updating the initial classifier. To prove the effectiveness of retraining data selecting method theoretically, we use Hidden Markov Model (HMM) to calculate the probability that the continuous recognition results are correctly recognized. Finally, a number of experiments are designed to verify our COPO, including data collection, performance comparison, and parameter analysis. The results show that the recognition rate of COPO is 2.62 % higher than other common methods.
Performance modeling for MapReduce applications with large-scale data is a very important issue in the study of optimization, evaluation, prediction and resource scheduling of the jobs over big data and cloud computing platforms. In this paper, we study the Hadoop distributed computing framework, which is the current trend of Big Data solutions. We use the locally weighted linear regression (LWLR) algorithm and linear regression (LR) algorithm to establish three kinds of computing models based on different characteristics to estimate the execution time of the applications that have large-scale data and run on the Hadoop framework, and at the same time we make comparison and improvement to the three models. By building different types of experimental environments, and running different types of jobs, we can draw a conclusion that all the three models have very good results in predicting the execution time and evaluating the performance of large-scale data applications with small-scale data.
In this paper, we present a mixed reality environment (MIXER) for immersive interactions. MIXER is an agent based collaborative information system displaying hybrid reality merging interactive computer graphics and real objects. MIXER is an agent based collaborative information system displaying hybrid reality merging interactive computer graphics and real objects. The system comprises a sensor subsystem, a network subsystem and an interaction subsystem. Related issues to the concept of mixed interaction, including human aware computing, mixed reality fusion, agent based systems, collaborative scalable learning in distributed systems, QoE-QoS balanced management and information security, are discussed. We propose a system architecture to perform networked mixed reality fusion for Ambient Interaction. The components of the mixed reality suit to perform human aware interaction are Interaction Space, Motion Monitoring, Action and Scenario Synthesisers, Script Generator, Knowledge Assistant Systems, Scenario Display, and a Mixed Reality Module. Thus, MIXER as an integrated system can provide a comprehensive human-centered mixed reality suite for advanced Virtual Reality and Augmented Reality applications such as therapy, training, and driving simulations.
In the era of big data, the cloud infrastructure needs to strongly support big data. As a distributed computational framework, Hadoop is one of the de facto leading software tools for solving big data problems. The cloud infrastructure has been proven to be a good support for three-tier architecture applications. In this paper, we construct a Hadoop big data platform based on OpenStack cloud. At the same time, we design three experimental scenarios, carry out a set of experiments using the standard Hadoop benchmarks TestDFSIO, TeraSort and PI, and examine the performance. Our experiments reveal that the disk read operation of physical servers can be a bottleneck for TestDFSIO and TeraSort. Wider allocation of VMs over physical servers achieves better performance for read jobs of TestDFSIO and TeraSort. For CPU-intensive job PI, the best practice is to centralize the allocation of VMs over physical machines.
With the ever-increasing number of web services registered in service communities, many users are apt to find their interested web services, through various recommendation techniques, e.g., Collaborative Filtering (i.e., CF)-based recommendation. Generally, the CF-based recommendation approaches can work well, when the target user has similar friends or the target services (i.e., the services preferred by target user) have similar services. However, in certain situations when user-service rating data is sparse, it is possible that target user has no similar friends and target services have no similar services; in this situation, traditional CF-based recommendation approaches fail to generate a satisfying recommendation result, which brings a great challenge for accurate service recommendation. In view of this challenge, we combine Social Balance Theory (i.e., SBT) and CF to put forward a novel recommendation approach Rec SBT+CF . Finally, the feasibility of our proposal is validated, through a set of simulation experiments deployed on MovieLens-1M dataset.
Health states, which are the abstract concepts for body recognition, represent a set of body conditions. Forming the accurate concepts, which are constantly tuning as the growth of human experiences, may be a long-term process in the human history. Nowadays, advances in technology have made monitoring the various data of body condition available. As the explosion of the data, it goes far beyond the ability of our brain to handle. But the computer can help us discover the patterns from the big data. In order to discover new representative health states, we propose forming them based on clustering. We use K-means clustering algorithm to discover the nine body constitution (BC) types in traditional Chinese medicine (TCM) according to the items in Constitution in Chinese Medicine Questionnaire (CCMQ). The results illustrate the ability of computer system to discover human health states.
Considering the high-speed and low-latency communication requirements of future 5G networks, a Stackelberg game based interference suppression approach is proposed. We analyze the uplink interference of macrocell, which is located in ultra-dense heterogeneous cloud access networks. Dense deployment brings relief of traffic, but leads to new interference problems. A power pricing game model between macrocell user end (MUE) and RRH user ends (RUEs) is formulated and Nash equilibrium is analyzed. Different from traditional methods concentrating on power, our proposed approach can maintain the power and energy efficiency of different kinds of user ends, so as to increase the spectrum efficiency of the whole heterogeneous networks. Simulations validate the results and demonstrate the superiority of the approach.
Many indexes have been designed to solve the problem of the point-to-point distance query on big graphs. In this paper, we design an incremental updating method for the index construction on dynamic graphs. The results show that the method is much time-saving compared with the way of index reconstruction. We also propose an exact method for distance labeling focused on undirected unweighted dense graphs. A kind of tight substructure named clique commonly exists in some dense graphs, such as social networks and communication networks. We take advantage of the cliques to compress the index. The experiments show that the technique can save index space and bring about comparable query time.