
The ability to detect when a person change their place of residence in a city or country is vitally important not just for urban planning but also for business intelligence. Although there are traditional approaches such as population census to collect this type of data, they have serious drawbacks. Thanks to the ubiquity of mobile phones, researchers have demonstrated that data generated from cellular network such as Call Detailed Records(CDR) can provide similar information at a relatively lower cost and higher temporal resolution. In this paper, we investigate two research questions: first, whether we can reliably discover a person's residence change from unlabeled CDR data. Second, if we can develop an algorithm that can autamatically carry out this task. To this end, we first formulate the residence change discovery problem by learning from population census approach and then propose a sequential spatio-temporal clustering technique-MoveSense to solve this problem. We use a large scale CDR dataset with over 3.5 billion call records and 16 million unique users to conduct experiments to validate our technique. We find that across the three categories of test datasets, the technique performed well with average detection rate of 71 percent, 68 percent and 72 percent.
In the field of Artificial Intelligence the task of spatial language understanding is a particularly complex one. Textual spatial information is frequently represented by so-called locative expressions, incorporating spatial prepositions. However, apart from the spatial domain, these prepositions can occur in a wide range of senses (e.g., temporal, manner, cause, instrument) as well as in semantically transformed senses (e.g., metaphors and metonymies). Existing practical approaches usually disregard semantic transformations or falsely classify them as spatial, although they represent the majority of cases. For the efficient extraction of locative expressions from data streams (e.g. from social media sources), a fast filter mechanism for this non-spatial information is needed. Hence, we present a classification schema to quickly and robustly disambiguate spatial from non-spatial uses of prepositions. We conduct an inter-annotator agreement test to highlight the feasibility and comprehensibility of our schema based on examples sourced from a large social media corpus. We further identify the most promising existing natural language processing tools in order to combine machine learning features with fixed rules.
We address the problem of extending the querying capabilities of Trajectories Data Warehouses (TDW) for symbolic trajectories, by introducing Semantic Relatedness (SR) as part of the formal model. This enables capturing the similarity between different annotations describing Points of Interest (POI), locations and activities. We formally define the inclusion of the relationship between different terms used as descriptors in symbolic trajectories and present the Semantic Relatedness in Trajectories Data Warehouse (SR-TDW) model. We introduce newly enabled queries in the SR-TDW model and illustrate the impacts of the added functionality. Our experiments demonstrate the benefits of the proposed approaches in terms of enriching the answer-sets for the common OLAP-based queries, and the sensitivity in terms of the various measures of semantic similarity.
Many natural phenomena are intuitively represented as spatiotemporal data objects, or moving objects. For example, vehicles, rivers, hurricanes, low pressure systems, areas of high density of foliage, etc align well with a geometric representation, and all change position or shape over time. Moving object models exist that represent real world objects as point, line, and region geometries that change continuously over time, leading to research into spatiotemporal analysis functionality over these objects. Models of moving objects are ideal for representing data streams that record the motion of spatial data over time. However, the implementation of operations to support spatiotemporal analysis over moving objects, particularly over moving regions, has proven difficult. In this paper, we develop a mechanism to support the implementation of the set operations of intersection, union, and difference between pairs of moving regions. The mechanism builds on the Component Model of Moving Regions and the semantic specifications of its operations. Specifically, we develop a generalized method of computing an intermediate data structure from which the results of various operations are then derived. The mechanism utilizes well-known 2D and 3D operational primitives and achieves O(n lg n) time complexity using appropriate data structures.
When a large-scale natural disaster occurs, it is necessary to collect damage information within about 10 minutes so that disaster-relief operations and wide-area support (depending on the the scale of the natural disaster) can be initiated. A high-performance method for "spatio-temporal join" which joins time-series grid data (such as results of simulations of natural disasters like tsunamis and fire spreading after a large-scale earthquake) and time-series point data representing people flows is proposed and applied to estimate damage situations following a natural disaster. The results of a performance evaluation of the method show that the response time for joining 100,000 point data and 250,000 grid data is about 50 seconds. They also show that it is possible to apply the proposed method to a real environment in which it is necessary to join one-million point data and hundreds of thousands of grid data within 10 minutes.
Social media data provide insight into people's opinions, thoughts, and reactions about real-world events such as hurricanes, infectious diseases, or urban crimes. In particular, the role of location-embedded social media is being emphasized to monitor surrounding situations and predict future effects by the geography of data shadows. However, it brings big challenges to find meaningful information about dynamic social phenomena from the mountains of fragmented, noisy data flooding. This paper proposes a data model to represent local flock phenomena as collective interests in geosocial streams and presents an interactive visual analysis process. In particular, we show a new visualization tool, called RendezView, composed of a three-dimensional map, word cloud, and Sankey flow diagram. RendezView allows a user to discern spatio-temporal and semantic contexts of local social flock phenomena and their co-occurrence relationships. An explanatory visual analysis of the proposed model is simulated by the experiments on a set of daily Twitter streams and shows the local patterns of social flocks with several visual results.
Mobility analysis is involved in many areas such as urban planning, traffic monitoring, climatology, study of social and animal phenomena to mention a few examples. The emergence and proliferation of mobile and sensor-based systems generate a significant increase of spatial and temporal data in terms of volume and frequency of update. In particular, the storage, management and analysis of the large data sets generated become a non straightforward task. Current works related to the manipulation of mobility data have been directed towards either mining archived historical data or continuous processing of incoming data streams. Our research introduces a hybrid approach whose objective is to provide a combined processing of real-time data streams and archived data. The principles of our approach is to promote the distributed and parallelized processing of mobility data. The whole framework is currently applied to the real-time monitoring of maritime traffic.
In this paper, we address the problem of recommending new locations to the users of a Location Based Social Network (LBSN). LBSNs are social and physical information-rich networks that incorporate mobility patterns and social ties of humans. Most of the existing recommender systems are build on variants of graph-based techniques that utilize complete knowledge of location history and social ties of all users. Therefore, these recommender systems are computationally expensive for large scale LBSNs. Further, these systems do not take into account the mobility habits of humans. Recent studies on human mobility patterns have highlighted that people frequently visit a set of locations and go to places closer to them. In this paper, we validate the existence of these human mobility aspects in LBSN through the analysis of user check-in behavior and derive a set of observations. Further, we propose REGULA-- A location recommendation algorithm that exploits three behavior patterns of humans: 1) People regularly (or habitually) visit a set of locations 2) People go to places close to these regularly visited locations and 3) People are more likely to visit places that were recently visited by others like friends. Using these behavior patterns, REGULA minimizes the computational complexity by reducing the set of candidate locations to recommend. We evaluate the performance of REGULA by employing two large scale LBSN datasets: Gowalla and Brightkite. Based on our results, we show that REGULA outperforms existing state of the art recommendation algorithms for LBSNs while reducing the complexity.
Dense line graphs and polyline maps are challenging for interactive visualization in geographic information systems (GIS). Bundling techniques are a common approach to reduce clutter and have successfully been demonstrated for the display of complex planar graphs. Previous techniques typically employed some form of attraction or repulsion forces to bundle edges in two dimensions, and while in principle extensible to 3D they do not directly support hard intersection constraints in a 3D environment. In geographic visualization systems, e.g. such as interactive virtual globes or 3D GIS viewers, it is often necessary to take the 3D environment into account and to: (1) bundle lines and paths in 3D, (2) constrain path bundles to follow some reference network vector map, as well as (3) avoid intersections with the digital elevation model (DEM). In this paper we introduce a novel method which uses geographic vector map reference information to route, visualize and simplify path bundles along their network paths in a constrained 3D environment using adaptive B-splines. Moreover, we describe an efficient rendering architecture to flexibly display bundled paths within a 3D rendering pipeline at varying level of detail (LOD).
We present a new algorithm for measuring the similarity between trajectories, and in particular between GPS traces. We call this new similarity measure the Merge Distance (MD). Our approach is robust against subsampling and supersampling. We perform experiments to compare this new similarity measure with the two main approaches that have been used so far: Dynamic Time Warping (DTW) and the Euclidean distance.
Integrating a raw GPS trajectory with spatial road networks is often referred to as the Map Matching problem. It's a fundamental component to support further analysis of intelligent transport systems. However, currently the occurrence of low-frequency trajectories (e.g. one point every 1-2 minutes) has brought lots of challenges to existing map matching algorithms. In this paper, we propose a novel global map-matching algorithm called ST-CRF based on the following insights: 1) the spatial positioning accuracy of GPS points as well as the topological information of the underlying road networks; 2) the spatial-temporal accessibility of a floating car; 3) the spatial distribution of the middle point between two consecutive GPS points; 4) the directional consistency of a GPS trajectory. Based on the spatial-temporal analysis, we construct a conditional random field (CRF) model and identify the best matching path sequence from all the candidate points. ST-CRF algorithm not only overcomes the long-existing "label-bias" problem of HMM-based models (e.g. ST-Matching, IVMM), but also performs more effective and robust based on a real trajectory dataset from Beijing. As a result, the ST-CRF algorithm outperforms the related models (Point-Line, ST-Matching, and IVMM).
This paper introduces a qualitative reasoning model for the representation of the trajectory of a moving point with respect to a region. The approach is based on a formal model of topological relations between a directed line and a region in a two-dimensional space. The approach is flexible enough to qualify possible movements according to several topological properties such as the dimension and cardinality of the intersections between a directed line and a region. We introduce the notion of conceptual transition that favors the exploration of possible trajectories in the case of incomplete knowledge configurations. A composition of DL-RE topological relations supports the derivation of complex movement patterns. The whole approach is experimented by a prototype development and applied to a large maritime trajectory database.
GPS-based navigation systems widely available on automobiles and smartphones nowadays are essential to find the best routes in the complicated urban space. However, it is still difficult for bikers to take full advantages of such navigation systems due to the lack of consideration on the different driving conditions. Generally, motorcyclists and cyclists take rides on narrow alleys and sidewalks which have a high risk of bumping against pedestrians. Therefore, it is necessary to find comfortable driving routes, also possibly avoiding areas congested by crowds. However, it is impractical to monitor crowd's existence everywhere at all times for such crowd-aware navigation. To overcome this limitation, we attempt to utilize location-based social network services where geo-tagged microblogs from massive crowd can be a good alternative source to measure pedestrian congestion in urban areas. In this paper, we introduce a route search method for bikers particularly to exploit crowd's volunteering reports being streamed via microblogs. In order to estimate human traffic from microblogs, we develop a crowd flow network which captures probable crowd movement on an urban network. We also examine the possible intersections which are expected to be highly congested based on the model. On the crowd flow network, we will find the best routes consisting of comfortable intersections and streets for the bike navigation systems.
A temporal coverage operation computes the duration that a moving object covers a spatial area. We extend this notion into temporal coverage aggregates, in which the spatial area covered for a maximum or minimum amount of time by a moving region, or set of moving regions, is discovered.
Due to the booming industry of location-based services, the analysis of human location histories is increasingly important. Next location prediction is essential to many location-based services. Predicting user's next location usually involves obtaining significant places from the history trajectories and predicting location with a certain statistic model. This paper presents new approaches to deal with both of above problems. For the former problem, a hierarchical clustering algorithm is proposed. We first identify specific features of stay points and then group the GPS points satisfying the identified features to form stay points by a new algorithm which is a variant of DBSCAN clustering algorithm. After that these stay points can be clustered to form significant places. For the later problem, taking the drawbacks like high space complexity and zero frequency problem in N-order Markov Model into consideration, we train a variable order Markov Model to predict next location. The variable order Markov Model uses escape mechanism to address the zero frequency problem and uses a tree structure to decrease the amount of memory needed in N-order Markov Model. An extensive set of experiments have been conducted to demonstrate the performance of proposed methods based on a real-world dataset, GeoLife.
Many natural phenomena can be nicely represented by concepts of moving regions. For example, hurricanes, rain clouds, pollution zones, etc., change shape and position over time. Current models of moving regions have proven to be difficult to translate effectively to implementation for two reasons: i) algorithms for operations, such as intersection, are difficult to implement, and ii) creating instances of moving regions from data sources is difficult. In this paper, we create a new model of moving regions at the abstract, discrete, and implementation levels that overcome the difficulties of previous models. The CMR Model aligns well with data collection techniques, can be implemented easily, and allows complex movement patterns to be easily depicted.
Technological advances have created an unprecedented availability of inexpensive sensors able to stream environmental data in real-time. However, we still seek appropriate data management technology capable of handling this onslaught of sampling in previously unavailable spatial and temporal density. Data stream engines (DSEs) are state of the art data management tools that have update throughput rates of up to 500k tuples/s. In previous work we have shown that DSEs can be extended to generate smooth representations of continuous spatio-temporal fields sampled by up to 250K sensors on-the-fly in near real-time, creating a new representation every second. In this paper we investigate a spatio-temporal stream operator framework that can efficiently execute predicate operators over such spatio-temporal fields. Typical predicates are e.g. "find all sub-areas in a field that are below or above a certain threshold value". We present the requirements, the approach taken, and our results along with a performance evaluation.
The continuously increasing popularity of social media sites such as Twitter and Facebook has recently led to a number of approaches to detect and extract event information from social media streams. Such events play an important role, e.g., in supporting location-based services and improving situational awareness. Moreover, the introduction of GPS-equipped communication devises has led to an increase in the percentage of geo-tagged messages. These help to detect localized events, i.e., events occurring at a certain location, such as sport events or accidents. The main entities that indicate a localized event are local keywords that exhibit a surge in usage at the event location. In this paper, we propose an approach to extract local keywords from a Twitter stream by (1) identifying local keywords, and (2) estimating the central location of each keyword. This extraction process is performed in an online fashion using a sliding window on the Twitter stream. In addition, we address the problem of spatial outliers that adversely affect a proper identification of local keywords. Outliers occur when people far away from an event location use related keywords in their Tweets. We handle this problem by adjusting the spatial distribution of keywords based on their co-occurrence with place names that may refer to the location of an event. We evaluate the performance of our framework to reliably and efficiently extracting local keywords and estimating their central locations using a Twitter dataset.
The emergence of internet advertising, email marketing and social networking has given rise to a new world of digital advertising used by stores and consumers alike. While retailers aim to promote all types of products, consumers also want to share this information via social media. This paper presents Shopaholic, a system that leverages social media to provide information on trending deals and store sales in any given location. It is intended to help shoppers identify great deals from the vast amounts of data scattered among social networks. Personalized search results, visualization of trends and sentiment analysis provided by Shopaholic allow the user to identify optimal deals. The application accounts for spatial and temporal data via a customized ranking algorithm and features integration with Twitter so that the user can share his or her actual experience using a deal. Ultimately, the system gives back to the shopping community by allowing users to share their experiences and evaluations of deals. A recommendation algorithm uniquely identifies the user's tastes, shopping history and current location to provide deal suggestions, thereby integrating temporal and spatial entities in recommendations.