This article presents an approach to place reference corpus building and application of the approach to a Geo-Microblog Corpus that will foster research and development in the areas of microblog/twitter geoparsing and geographic information retrieval. Our corpus currently consists of 6000 tweets with identified and georeferenced place names. 30% of the tweets contain at least one place name. The corpus is intended to support the evaluation, comparison, and training of geoparsers. We introduce our corpus building framework, which is developed to be generally applicable beyond microblogs, and explain how we use crowdsourcing and geovisual analytics technology to support the construction of relatively large corpora. We then report on the corpus building work and present an analysis of causes of disagreement between the lay persons performing place identification in our crowdsourcing approach.
The world has become a complex set of geo-social systems interconnected by networks, including transportation networks, telecommunications, and the internet. Understanding the interactions between spatial and social relationships within such geo-social systems is a challenge. This research aims to address this challenge through the framework of geovisual analytics. We present the GeoSocialApp which implements traditional network analysis methods in the context of explicitly spatial and social representations. We then apply it to an exploration of international trade networks in terms of the complex interactions between spatial and social relationships. This exploration using the GeoSocialApp helps us develop a two-part hypothesis: international trade network clusters with structural equivalence are strongly 'balkanized' (fragmented) according to the geography of trading partners, and the geographical distance weighted by population within each network cluster has a positive relationship with the development level of countries. In addition to demonstrating the potential of visual analytics to provide insight concerning complex geo-social relationships at a global scale, the research also addresses the challenge of validating insights derived through interactive geovisual analytics. We develop two indicators to quantify the observed patterns, and then use a Monte-Carlo approach to support the hypothesis developed above.
Associating place name mentions in unstructured text with their actual references in geographic space is vital to enable spatial queries and analysis. In this paper, we introduce GeoTxt, a web API plus human-usable web tool designed and implemented to tackle three components of place-reference processing from text, namely: extraction, disambiguation, and geolocation of place names mentioned in unstructured text. Current GeoTxt development is focused particularly on support for processing short microblog posts.
The utilization of mathematical and computational tools for pollutant assessment frameworks has become increasingly valuable due to the capability to interpret integrated variable measurements. Artificial neural networks (ANNs) are considered as dependable and inexpensive techniques for data interpretation and prediction. The self-organizing map (SOM) is an unsupervised ANN used for data training to classify and effectively recognize patterns embedded in the input data space. Application of SOM-ANN is useful for recognizing spatial patterns in contaminated zones by integrating chemical, physical, ecotoxicological and toxicokinetic variables in the identification of pollution sources and similarities in the quality of the samples. Water (n = 11), soil (n = 38) and sediment (n = 54) samples from four areas in the Niger Delta (Nigeria) were classified based on their chemical, toxicological and physical variables applying the SOM. The results obtained in this study provided valuable assessment using the SOM visualization capabilities and highlighted zones of priority that might require additional investigations and also provide productive pathway for effective decision making and remedial actions. (c) 2012 Elsevier Ltd. All rights reserved.
Scalable Vector Graphics (SVG) is an effective web map technology for client-side vector mapping. We demonstrate this through a web map application developed for the Trans-Border Institute of the University of San Diego. The Trans-Border Institute currently releases reports summarizing data on the number of drug-related homicides in Mexico and produces maps to show drug-related killings in each Mexican state. Together, the reports and maps are used to inform U. S. audiences about the public security situation in Mexico and the effects of the war on drugs. The deadly consequences of the drug-war in Mexico become clear in the maps produced by the Trans-Border Institute. The maps, however, are not as effective as they could be because they are released as static images. These static images do not allow for the exploration of the data themselves and make it difficult to view changes over time. To facilitate data exploration and thereby assist the Trans-Border Institute in more effectively disseminating information on the drug-war in Mexico, an interactive web map was developed using SVG. For client side vector graphics, SVG provides clear advantages in that it is open, interoperable, and extensible and is resolution independent. In addition, behaviors, including animation, can be included in the markup file itself or added through scripting. These advantages make SVG optimal for developing high quality interactive (and non-interactive) vector maps. Despite these advantages, SVG has not been widely adopted. However, recent technological changes and trends have made competing systems, like Flash, less optimal and have heightened the awareness and ease of using SVG. Combined, these changes have paved the way for SVG becoming the most widely used client-side vector standard. Its effectiveness is demonstrated in an SVG-based, interactive web map developed for the Trans-Border Institute and shown in this presentation.
In this paper, we present the GeoViz Toolkit (GVT), an open-source, Internet-delivered program for geographic visualization and analysis that features a diverse set of software components which can be flexibly combined by users who do not have programming expertise. The design and architecture of the GVT allows us to address three key research challenges in geovisualization: allowing end users to create their own geovisualization and analysis component set on the fly, integrating geovisualization methods with spatial analysis methods, and making geovisualization applications sharable between users. Each of these tasks necessitates a robust yet flexible approach to intertool coordination. The coordination strategy developed for the GVT, called Introspective Observer Coordination, leverages and combines key advances in software engineering from the last decade, such as automatic introspection of objects, software design patterns, and reflective invocation of methods.
The characterization, identification, and understanding of spatial patterns are central concerns of geography. Deeply rooted in the notion that geographic location matters, one testable assumption is that near things are more related than distant things—a concept often referred to as Tobler's first law of geography. One means of quantifying this assumption is using measures of spatial autocorrelation. Several such measures have been developed to test whether a pattern is indeed clustered, or dispersed, or whether it is, from a statistical perspective, random. To shed light on how spatial patterns are understood from a cognitive perspective, this article reports results from studies of spatial pattern interpretation represented in maps. For the purpose of experimental validation, we used a two-color map. We systematically varied the ratio of the colors as well as the level of significance of clustering and dispersion; we targeted two groups: experts and nonexperts. The task for both experts and nonexperts was to sort patterns according to five specified categories of spatial autocorrelation structures. The results show clearly that patterns are understood on the basis of the dominant color, by both experts and nonexperts. A third experiment, using a free classification paradigm, confirmed the dominance of the color effect. These results are important, as they point to critical aspects of pattern perception and understanding that need to be addressed from the perspective of spatial thinking, especially how people relate concepts of randomness with spatial patterns (represented in maps).
There has been considerable interest in applying social network analysis methods to geographically embedded networks such as population migration and international trade. However, research is hampered by a lack of support for exploratory spatial-social network analysis in integrated tools. To bridge the gap, this research introduces a spatial-social network visualization tool, the GeoSocialApp, that supports the exploration of spatial-social networks among network, geographical, and attribute spaces. It also supports exploration of network attributes from community-level (clustering) to individual-level (network node measures). Using an international trade case study, this research shows that mixed methods --- computational and visual --- can enable discovery of complex patterns in large spatial-social network datasets in an effective and efficient way.
Penn State Public Broadcasting has produced the Geospatial Revolution Project, an integrated public media and outreach initiative about the world of digital mapping and how it is changing the way we think, behave and interact. With the goal of increasing public awareness of geospatial technologies, the project offers four 15-minute online mini-documentary episodes, 3-minute shorter chapters, as well as K-16 educational materials. The episodes share compelling human stories that clarify the complex and decode the mysterious, explain the virtues and explore the potential dangers of these emerging technologies. The Geospatial Revolution Project explores the seamless layers of satellites, surveillance, and location-based technologies creating a worldwide geographic knowledge base vital to solving myriad social and environmental problems in the interconnected global community. Faculty from the Dutton e-Education Institute at The Pennsylvania State University served on the project advisory board, and Frank Hardisty from the Institute will play a selection from one of the following episodes during the demo presentation: Episode One -- defining the geospatial revolution and its historical origins; includes a story on the Haitian earthquake Episode Two -- geospatial technology in interactive city and business management Episode Three -- mapping in war and peace, police protection, and personal privacy and safety Episode Four -- agriculture and the environment, mapping disease, and human rights and aid www.geospatialrevolution.psu.edu; twitter.com/geospatialrev; Facebook.com/geospatialrev
Analysts are faced with increasing volume and complexity of spatially and spatio-temporally referenced events to analyze. One means of taming this volume and complexity is to develop methods and tools that can identify patterns, including spatio-temporal structure like clusters, in event data. To understand these methods and tools, we first present some of the motivation for the work, and then we detail the software architecture that we will use to support categorical analysis of spatio-temporal events. Methods for analyzing spatial events, and spatio-temporal events, have experienced a recent renaissance. This upsurge in interest has occurred in part because of novel high-quality event sources which can provide complex data with geographic and temporal referents. Examples of such event sources include geographically located Twitter postings linked to documents, or photographs and video taken with GPS-enabled mobile phones. The products of these event sources are a vast stream of events that are linked with heterogeneous and voluminous data, including textual, imagery data as well as numerical data. One of the ways of making numerical sense of such heterogeneous data is to consider the text or media as a set of tagged categories. We can then apply methods for detecting structure in spatio-temporal events, including recently developed methods for disease outbreak detection, to these categories. We are developing methods and associated software that will allow users to tag or label events, analyze them, and interact with visual representations of the event structure detected by the analysis. An integrated software system, called STempo, will provide the user with a fixed set of analysis tools and coordination topology to work from. The category tagging and structure detection tools will also be worked into the larger set of tools available in the GeoViz Toolkit, an interactive system for geographic visualization and analysis.
Traditional models of migration assume that migrants move to places of greatest economic incentive, and are more likely to move when current economic conditions `push' migrants from their origin. Although prospective income at a destination has been a major determining factor for migration in preexisting migration models, and distance between origin and destination is also a major consideration, we take a new approach with a model that reflects migration `chaining', where migrants to a city B send information back to their origin city A, and interest other members of A to migrate to B. We isolate the social factors of place-pair synergies through components from Bayes' Law: conditional probability and posterior probability of unique origin/destination migrant volume, and a system-wide probability of unique O/D transfer. These allow us to model social space as well as physical space, rather than physical space alone. We test these variables' power for predicting future migration against four other predictive models: the traditional gravity model, transit data, airline and trip data, and linear trends. We use a case study of U.S. Migration flows in a system of major cities, given annual data from 1996-2004 to predict city-to-city flows annually for 2005-2008, and find that conditional and posterior probabilities outperform system-wide probabilities, gravity, transit and linear forecast models. These probabilities also exhibit a surprising level of steady-state stationarity, and therefore are a promising avenue for more accurately modelling future migration flows.
Behavioral experiments addressing the conceptualization of geographic events are few and far between. Our research seeks to address this deficiency by developing an experimental framework on the conceptualization of movement patterns. In this paper, we report on a critical experiment that is designed to shed light on the question of cognitively salient invariants in such conceptualization. Invariants have been identified as being critical to human information processing, particularly for the processing of dynamic information. In our experiment, we systematically address cognitive invariants of one class of geographic events: single entity movement patterns. To this end, we designed 72 animated icons that depict the movement patterns of hurricanes around two invariants: size difference and topological equivalence class movement patterns endpoints. While the endpoint hypothesis, put forth by Regier (2007), claims a particular focus of human cognition to ending relations of events, other research suggests that simplicity principles guide categorization and, additionally, that static information is easier to process than dynamic information. Our experiments show a clear picture: Size matters. Nonetheless, we also find categorization behaviors consistent with experiments in both the spatial and temporal domain, namely that topology refines these behaviors and that topological equivalence classes are categorized consistently. These results are critical steppingstones in validating spatial formalism from a cognitive perspective and cognitively grounding work on ontologies.
Many interesting analysis problems (e.g. disease surveillance) would become more tractable if their spatio-temporal structure was better understood. Specifically, it would be helpful to be able to identify autocorrelation in space and time simultaneously. Some of the most commonly used measures of spatial association are LISA statistics, such as the Local Moran's I or the Getis-Ord Gi*; however, these have not been applied to the spatio-temporal case (including many time steps) because of computational limitations. We have implemented a spatio-temporal version of the Local Moran's I and claimed two advances: first, we exploit the fact that there are a limited number of topological relationships present in the data to make Monte Carlo's estimation of probability densities computationally practical, and thereby bypass the 'curse of dimensionality'. We term this approach 'spatial memoization'. Second, we developed a tool (LISTA-Viz) for interacting with the spatio-temporal structure uncovered by the statistics that contains a novel coordination strategy. The potential usefulness of the method and the associated tool are illustrated by an analysis of the 2009 H1N1 pandemic, with the finding that there was a critical spatio-temporal 'inflection point' at which the pandemic changed its character in the United States.
Directional flows created from an origin/destination matrix have been traditionally difficult to visualize because of the number of flows to be rendered in a small cartographic space. Because visualizing geographic flow dynamics are useful for understanding the complex dynamics of human and information flow that connect non-adjacent space, techniques that allow for visual data mining or static representations of system dynamics are a growing field of research. Here, we use a Weighted Radial Variation (WRV) technique to classify places based on their group’s radially-emanating vector flows. Each entity’s vector are syncopated in terms of cardinality, direction, length, and flow magnitude. The WRV process unravels each star-like entity’s individual flow vectors on a 0-360° spectrum, to form a unique signal whose distribution depends on the flow presence at each step around the entity, and is further characterized by flow distance and magnitude. The signals are processed with a supervised classification method that clusters entities with similar signatures or trajectories in order to learn about types and geographic distribution of flow dynamics. We use U.S. county-to-county human incoming and outgoing migration data to test our method.
The GeoJabber concept, protocol, and working prototype software introduced here enable same-time, different-place collaborative geovisualization. The key problem this work addresses is how to turn geovisual software states into persistent textual representations that can be shared between users. In the current implementation, GeoJabber leverages three key Open Source technologies: the GeoViz Toolkit, the Jabber protocol, and XStream. GeoJabber is the first project to Support same-time different-place geovisualization tool state sharing. As part of this effort, this paper presents a typology of sharable geovisualization software states, rooted in the concepts of data, display, and category.
Our research addresses the question of how to design interfaces for spatial analysis such that they support cognitive processes. In this paper we specifically target the question of map symbol design for the analysis of multivariate data, which is a common problem in cartography and related fields. We focus on star plots and the largely unaddressed question of how to assign variables to rays in a star plot and which consequences specific shapes have-as the result of data characteristics and the assignment of variables to rays-on interpretation and classification. We conducted an experiment with two conditions that were designed to shed light on the question: Does the shape of a star plot influence the interpretation (meaning) of the data it represents in a classification task? While previous research on multivariate point symbols has addressed this question for Chernoff faces, for example, few connections have been made to the shape of a star plot and its potential influence on meaning. We found that certain salient shape characteristics induced by variations along the horizontal and vertical axis increase the classification speed. However, we also found that salient shapes, such as has one spike, introduce a perceptual similarity that overrides the assumed similarities in the meaning of the represented data.
Prasenjit Mitra合作论文数College of Information Sciences and Technology, Penn State University1