
In this talk we present two historical models of human migration from the 19th and 20th centuries, and discuss how they apply to location data on the Web, in the 21st Century.
Record linkage is the task of identifying which records in one or more data collections refer to the same entity, and address is one of the most commonly used fields in databases. Hence, segmentation of the raw addresses into a set of semantic fields is the primary step in this task. In this paper, we present a probabilistic address parsing system based on the Hidden Markov Model. We also introduce several novel approaches of synthetic training data generation to build robust models for noisy real-world addresses, obtaining 95.6% F-measure. Furthermore, we demonstrate the viability and efficiency of this system for large-scale data by scaling it up to parse billions of addresses.
This paper is concerned with the automatic prediction of the zoom level at which to present a map with results for an informal location description. We propose the use of identifiability (relative uniqueness) and zoom level of each component geospatial expression (GE) in the location description, as a means of predicting the appropriate zoom level for the overall description. We apply a simple classification approach to zoom level prediction, and compare results using gold-standard and automatically-inferred GE information. We find the approach to have strong promise, including relative to the zoom level used in results from Google Maps for our location descriptions dataset.
Location nowadays is an important aspect of the Web. One scenario in this respect are archives or collections of geo-tagged media items. More concretely, we can think of collections in the arts and humanities available via OAI-PMH (a protocol for metadata harvesting) on the web or web accessible personal media archives maintained in a peer-to-peer manner. In such scenarios for search problems, source selection becomes an important aspect. For example, we would like to access only those collections containing media items in a certain geospatial region (maybe we are interested in images from Shanghai only). Here, the geospatial search criterion allows for a high selectivity. What is needed in such a scenario are expressive and nevertheless compact representations or descriptions of the ``geospatial footprint'' of each collection. A minimum bounding rectangle would be a trivial but not very accurate option. Generally, summarization techniques for this purpose can be distinguished into three categories, geometric approaches, space partitioning approaches and hybrid approaches. In this work, we present novel hybrid techniques, which mostly apply a set of approximating minimum area rectangles for subspace description together with quantization techniques in order to increase the selectivity of the summaries and, at the same time, keep the storage requirements small.
Indoor positioning system (IPS) identifies positions of various indoor objects, and is a key technology to achieve sophisticated Indoor Location-Aware Services (InLAS). In most conventional systems, InLAS and IPS are tightly coupled. That is, one system does not supposed to reuse indoor location data and program of another system. This makes individual systems complex and difficult to manage. To cope with the problem, we propose Data Model for Indoor Location (DM4InL), which prescribes a common data schema, independent of implementation of IPS or the usage of InLAS. The proposed DM4InL represents the location of every indoor object in a standard way, by using three kinds of models: location, building and object models. We also design the fundamental API, which implements typical queries to the indoor location data from external applications. The proposed method achieves loose-coupling of InLAS and IPS, which significantly improves the efficiency and reusability in the InLAS development.
Over the recent years smart devices have become a ubiquitous medium supporting various forms of functionality and are widely accepted for common users. One distinguishing feature for smart devices is the ability of positioning the physical location of a device, and numerous applications based on user location information have been proposed. While the potentials have been foreseen, location based services fundamentally suffer from the problem of lacking an effective and scalable mechanism to bridge the gap between the machine-observed locations and the human understandable places. In this study, we contribute on this fundamental problem. Differing from the existing solutions on this subject, we start from a novel perspective; we propose to address the place semantic understanding problem by casting it as a classification problem and employ machine learning techniques to automatically infer the types of the places. The key observation is that human behaviors are not random, e.g., people visit restaurants around noon, go for work in the daytime, and stay at home at night. Namely, by properly selecting features, a mechanism for automatically inferring place type semantics can be achieved. This paper summarizes our treatment and findings of leveraging the human behaviors patterns to infer the type of a place. Experiments using month-long trace logs from the recruited participants are conducted, and the experiment results demonstrate the effectiveness of the proposed method.
With the proliferation of smartphones and the increasing popularity of social media, people have developed habits of posting not only their thoughts and opinions, but also content concerning their whereabouts. On such highly-interactive yet informal social media platforms, people make heavy use of informal language, including when it comes to locative expressions. Such usage inhibits the ability of traditional Natural Language Processing approaches to retrieve geospatial information from social media text. In this research, we: (1) develop a medium-scale corpus of "locative expressions" derived from a variety of social media sources; (2) benchmark the performance of a range of geoparsers over the corpus, with the finding that even the best-performing systems are substantially lacking; and (3) carry out extensive error analysis to suggest ways of improving the accuracy and robustness of geoparsers.
This contribution investigates how accurate location information is on smartphones. Our research is based on a data set consisting of 2289 locations gathered from a marketed iPhone application. In a first analysis, it became evident that the accuracy information differs significantly among iPhone and iPod and iPad devices. A second analysis of the accuracy values reveals clusters of accuracy values at above 1 km, at 500 meters, and at an accuracy of below 300 meters. Information with an accuracy of above 500 meters originated from Cell-ID based positioning. Finally, an analysis revealed that the accuracy is significantly reduced for locations with more than 500 meters away from the next populated area. The overall results suggest that additional Cell-ID based positioning technology allows for higher coverage at the costs of a significantly reduced accuracy. If location information is required, with an accuracy of below 300 meters, the technology should be limited to GPS and WLAN based positioning. Adding Cell-ID based positioning increases the coverage while accuracy is reduced.
This paper describes an experimental setup for the analysis of e-bike usage characteristics based on GPS data. Usage characteristics include parameters such as average and maximum velocity, trip lengths and distribution over daytime. Based on high resolution position measurement these parameters are extracted and compared to other studies on both e-bikes and conventional bicycles. We show that applying location technology to concurrent monitoring of a fleet of e-bikes yields higher quality in terms of resolution and accuracy (1), and is less intrusive (2) than obtaining these data by conventional user surveys. The findings form a proof-of-concept for the adoption of location technology to transportation and behavioral sciences and suggest further interdisciplinary collaboration in these fields.
Following paradigms like Ubiquitous Computing or the Internet of Things, modern factories are developing to intelligent environments in which wireless technology, sensor networks and mobile information access close the gap between the physical and the digital world. This article assumes that in such a versatile Factory of Things location information will play an important role for the transparent and efficient design of mobile and adaptive processes. Based upon a maintenance use case in the SmartFactoryKL it will be examined how location information can contribute to the optimization of maintenance processes. As research questions regarding the development of an appropriate system architecture, the definition of a consistent data representation for location information as well as mechanisms for its semantic interpretation are focused. Finally, the desired architecture is evaluated regarding its benefits, limitations and role as an enabler for a lean information management suitable to apply in future intelligent factories.
Despite much advances in both general and targeted Social Network Services (SNS) and Location-Based Social Networks (LBSN), there is currently a void in literatures on SNS that form temporary social networks to address specific problems and employ intelligent classification of members and coordination of tasks toward goal oriented action. In this position paper, we present Genetic Location-Based Social Networks (G-LBSN) which is a new concept in social networking where temporary social networks, to address specific problems, are formed. Unique characteristics of G-LBSN include: (a) formation of temporary social networks (each network has a life time which spans from its inception until the problem is solved); (b) classification of members based on their proximity to given locations, their contribution to the solution, and their availability; (c) assignment of tasks to selected members; and (d) coordination and supervision of members activities and the progress toward a solution.
In providing a service to mobile users, it would be critical to know what types of information they would look for in association with geo-referenced entities that may be extractable from queries or contexts. While understanding high-level user intentions in accessing the Web, such as informational, navigational, and transactional, is useful, a finer-level classification of user interests would further help adapting mobile search results to user intensions. Our research focuses on understanding what aspects of geo-referenced entities are mentioned in user queries in an attempt to create a model for user intents in geo-referenced Web searching. By collecting and analyzing geo-referenced questions posed to operational question answering systems, we delineated major aspects of non-topical information that people would seek in association with geographic information. The identified aspects were further conceptualized to develop a user interest model with three dimensions, which was validated with two sets of data. The model can be a basis for identifying user's intent in a mobile search context as well as classifying geo-related text to be retrieved for its aspectual category.
In this paper we propose a new algorithm for finding the frequent routes that a user has in his daily routine, in our method we build a grid in which we map each of the GPS data points that belong to a certain sequence. (We consider that each sequence conforms a route) we then carry out an interpolation procedure that has a probabilistic basis and find a more precise description of the user's trajectory. For each trajectory we find the edges that were crossed, with the crossed edges we create a histogram in which the bins denote the crossed edges and the frequency value the number of times that edge was crossed for a certain user. We then select the K most frequent edges and combine them to create a list of the most frequent paths that a user has. We compared our results with the algorithm that was proposed in Adaptive learning of semantic locations and routes [6] to find frequent routes of a user, and found that our implementation on the contrary of [6] can discriminate directions, ie routes that go from A to B and routes that go from B to A are taken as different. Furthermore our implementation also permits the analysis of subsections of the routes, something that to our knowledge had not been carried out in previous related work.
During the last few years, the amount of online descriptive information about places and their dynamics has reached reasonable dimension for many cities in the world. Such enriched information can now support semantic analysis of space, particularly in which respects to what exists there and what happens there. We present a methodology to automatically label places according to events that happen there. To achieve this we use Information Extraction techniques applied to online Web 2.0 resources such as Zvents and Boston Calendar. Wikipedia is also used as a resource to semantically enrich the tag vectors initially extracted. We describe the process by which these semantic vectors are obtained, present results of experimental analysis, and validated these with Amazon Mechanical Turk and a set of algorithms. To conclude, we discuss the strengths and weaknesses of the methodology.
We introduce an iPhone / iPod touch timetable application using WiFi location system named "Eki.Locky". This application adopts UGC (User Generated Content) approach to collect TimeTable information and WiFi access point (AP) information from public users. Since the service started in October 2009, Eki. Locky has been used by over 440,000 people, posted timetable information covers 98% of all stations in Japan and 350,000 WiFi AP information were collected. In addition, since June 2010, we started a new version of this application named "TimeTable.Locky" which supports any timetable such as buses and airplanes. TimeTable.Locky also has been used large number of people, and collected over 23,000 of timetable information.
The Third International Workshop on Location and the Web (LocWeb 2010) focuses on research and development that targets the intersection of Internet-enabled location-aware and/or located devices, and services based on Web technologies and Web architecture. The rapid rise of multi-sensory mobile devices and Internet-enabled "things" equipped with sensors and ubiquitous connectivity opens new possibilities and provides the foundations to capture, share and use Web services and applications in ways which go beyond the traditional scenarios of stationary or even mobile computer-like devices. Increasingly, applications will have to bridge the physical world and the Web space, and location is one of the major connecting links. When Web services will "surround" users, designers have to address the challenges of scalability and interoperability on the Web, and designers also have to look at policy, regulatory, and legislative responses to the privacy and security challenges created by something as sensitive as location information.
In this paper we outline a unified architecture for representing locations of people, places and things in real or virtual worlds, called realms, on the web. Our architecture is based on the location graph that encodes web-level containment and connectedness relationships between locations. The architecture provides an information processing model that allows realm independent queries such as position, range and path, and realm specific queries, such as distance. We present existing systems that are enablers for the proposed architecture. With this architecture we enable a common way to develop location based services and applications across real or virtual realms, avoiding fragmentation.
While the demand for Location-based Services (LBS) is strongly increasing, technical laymen are not yet able to build and provide location-aware applications. This paper presents a radical simplification of the lifecycle of LBSs. An authoring toolkit enables non-technicians to easily develop context-aware mobile applications. In addition, an adaptive user interface makes the consumption of LBSs easier. The platform we present, in covering the whole LBS supply chain, is hiding the complexity of providing and consuming LBSs from the end-users.
In this position paper the question of how location is employed in location based services (LBS) is considered. The importance of the notion of location is highlighted as a means of blurring the boundary between forms of experiences that are direct, and sensed in the environment, and those that are indirect, and learned from information. It is suggested that current methods for modeling location are limited by their lack of strong theoretical underpinning. To help bridge this gap the notions of Space, Place, and Region, from geographical theory, are proposed and implications of these for considering location in LBS outlined.
We propose a novel method to detect cultural differences over the world automatically by using a large amount of geotagged images on the photo sharing Web sites such as Flickr. We employ the state-of-the-art object recognition technique developed in the research community of computer vision to mine representative photos of the given concept for representative local regions from a large-scale unorganized collection of consumer-generated geotagged photos. The results help us understand how objects, scenes or events corresponding to the same given concept are visually different depending on local regions over the world.