Over the last decade, event prediction has drawn attention from both academic and industry communities, resulting in a substantial volume of scientific papers published in a wide range of journals by scholars from different countries and disciplines. However, thus far, a comprehensive and systematic survey of recent literature has been lacking to quantitatively capture the research progress as well as emerging trends in the event prediction field. Aiming at addressing this gap, we employed CiteSpace software to analyze and visualize data retrieved from the Web of Science (WoS) database, including authors, documents, research institutions, and keywords, based on which the author co-citation network, document co-citation network, collaborative institution network, and keyword co-occurrence network were constructed. Through analyzing the aforementioned networks, we identified areas of active research, influential literature, collaborations at the national level, interdisciplinary patterns, and emerging trends by identifying the central nodes and the nodes with strong citation bursts. It reveals that sensor data has been widely used for predicting weather events and meteorological events (e.g., monitoring sea surface temperature and weather sensor data for predicting El Nino). The real-time and multivariable monitoring features of sensor data enable it to be a reliable source for predicting multiple types of events. Our work offers not only a comprehensive survey of the existing studies but also insights into the development trends within the event prediction field. These findings will assist researchers in conducting further research in this area and draw a large readership among academia and industrial communities who are engaged in event prediction research.
Small-scale events involve interactive human movement in limited space and time. Social media platforms possibly generate large amount of geospatially-referenced information related to small-scale events. It benefits individuals, management departments, and urban systems if small-scale events can be timely detected from social media platforms, where measuring the abnormal patterns of human movement to discover events and analyzing associated texts to interpret the reasons behind abnormal movement are two keys. Through investigating how people move as different events occur and measuring the patterns on social media platforms, small-scale events can be generally classified into two types, namely type I events with abrupt patterns and type II events with random occurrence of key factors, where social events and traffic events are representative correspondingly.Despite many studies have been conducted to detect social events and traffic events usinggeosocial media data, there still are some un-answered questions requiring further research. Mostexisting studies did not identify occurring events from a full coverage of spatial, temporal, andsemantic perspectives. Studies concerning social event detection lack efficient semantic analysis summarizing event content to infer the reasons driving the abnormal movement. The typicalclassification-based method regarding traffic event detection lacks investigation on how the spatiotemporal distribution of traffic relevant posts associate with the occurring traffic events, andsimply assigns the detected events with predefined categories, missing events that indicate trafficanomalies but go beyond the predetermined categories.In this thesis, spatial-temporal-semantic approaches are proposed to measure spatiotemporalpatterns of posts and users of social media platforms to capture abnormal human movement, andanalyze the content of associated posts to mine the reasons driving the movement. A variety oftechniques including machine learning, natural language processing, and spatiotemporal analysisare adopted to realize effective detection. Based on one-year Twitter data collected in Toronto,2014 Toronto International Film Festival and traffic anomaly detection are selected as two casestudies to evaluate the performance of proposed approaches. Through comparing with the groundtruth data, the result reveals that more than 80% of the detected events do refer to real-world events,which illustrates the feasibility and efficiency of proposed approaches.Keywords: Small-scale event, Event detection, Geosocial media data, Traffic event, Social event,Twitter, Spatiotemporal clustering
The emergence of micromobility, exemplified by bike sharing and e-bike sharing systems, has ushered in a low-carbon, environmentally friendly, and sustainable revolution in urban transportation. This transformative shift addresses the "first and last mile" challenge and holds immense potential for urban mobility enhancement. Nevertheless, the existing literature predominantly investigates the spatiotemporal travel patterns and influencing factors of bike sharing and e-bike sharing systems in isolation, overlooking comparative analyses grounded in quantitative methodologies. In order to fill this gap, this study first compares and analysis their spatiotemporal travel patterns, which are measured by travel distance, travel time, and travel volume. A Multiscale Geoweighted Regression (MGWR) model was constructed using various data sources, such as Point of Interest (POI) data, metro station data, and bus stop data, to conduct a spatiotemporal correlation analysis of land use and public transport factors with the travel volume of the shared system. Our study centers on Manhattan, New York City, utilizing data from May 2022 for both bike sharing and e-bike sharing systems. The study analysis reveals that hourly trip volumes are higher for bike sharing than for e-bike sharing, exhibiting substantial spatial variation across different regions within the city. The MGWR model's findings suggest that educational facilities exert a negative influence on bike sharing in the northeast and on e-bike sharing trips in the central region, with this impact being more pronounced on weekends. Similarly, cultural facilities negatively affect the Central region's bike sharing system and the citywide e-bike sharing system, with a milder effect during weekends. Moreover, bus stops exhibit a significant negative impact on bike sharing and e-bike sharing at Chelsea Waterside Park (only weekdays), while displaying a positive influence on both systems during weekends. To validate the MGWR model's efficacy, we conducted a comparative analysis with a Geographically Weighted Regression (GWR) model. The results demonstrate that MGWR can be more effective in correlating and quantitatively explaining the effects of different factors on spatiotemporal travel patterns. In conclusion, this study furnishes valuable insights for optimizing urban infrastructure rebalancing strategies and advancing sustainable urban infrastructure development.
In the last decade, the event knowledge graph field has received significant attention from both academic and industry communities, leading to the proliferated publication of numerous scientific papers in diverse journals, countries, and disciplines. However, a comprehensive and systematic survey of the recent literature in this area to obtain how the development of event knowledge graph evolves over time is lacking. To address this gap, we performed scientometric analyses utilizing the CiteSpace software of version 6.2.R4 package to extract and analyze data from the Web of Science database, including information about authors, journals, countries, and keywords. We then constructed four networks, including the author co-citation network, journal co-citation network, collaborative country network, and keyword co-occurrence network. Analyzing these networks allowed us to identify core authors, research hotspots, landmark journals, and national collaborations, as well as emerging trends by assessing the central nodes and nodes with strong citation bursts. Our contribution mainly lies in providing a scientometric way to quantitatively capture the research patterns in the last decade in the event knowledge graph field. Our work provides not only a structured view of the state-of-the-art literature but also insights into future trends in the event knowledge graph field, aiding researchers in conducting further research in this area.
Air quality acts as an important factor that human may consider as they make decisions on when and where they would go. In order to access how much the air quality affects human mobility patterns, the air quality was measured using air quality index (AQI) and human mobility patterns were measured by travel volume and travel distance of shared bikes. Their correlation that presents on weekdays and weekends as well as in different administrative districts were investigated using Spearman correlation analysis method. A case study was conducted in Beijing, China using bike sharing data and air quality data ranging from May 10 to 16, 2017. The results show that travel distance is more sensitive to air quality on weekdays such as Changping District (−0.20), Haidian District (−0.13), Shunyi District (−0.12). The travel volume on weekdays is less sensitive to air quality due to commuting. The travel volume has a negative relationship with AQI on weekends. Fengtai District, Huairou District, Pinggu District are more susceptible to severe air quality, leading to a reduction in bike traveling distance. This work sheds light on understanding human-environment coupling mechanism and promoting urban sustainable development.
Financial outlets are branches or business outlets of financial institutions such as banks, investment and finance companies and insurance companies in cities, which are used to provide a variety of financial services and businesses. The location of financial outlets is usually determined on the basis of factors such as population density, commercial areas, residential areas and transportation convenience. As an important carrier of urban financial services, the rationality of the layout of financial outlets can reflect the financial needs and behavioral patterns of people in the city. Urban functional area is a product after urbanization, different from urban land use type which focuses on the classification of urban land use through natural conditions, urban functional area mainly analyzes the internal spatial structure of the city, focusing on the discernment of the dominant functions within the built-up area of the city. By categorizing urban functions such as residence, commercial service and industry, the spatial distribution of urban functions can be grasped to provide support for urban planning and development. Urban functional areas can reveal human behavior patterns, needs and preferences in different areas. The current research on financial outlets mainly focuses on the accessibility and coverage of financial outlets, and researchers treat financial outlets as spatially undifferentiated points, ignoring the problem of spatial heterogeneity of financial outlets within different functional zones. When studying the rationality of the layout of urban financial outlets, considering different functional zones within the city can better meet the financial needs of residents, enterprises and institutions in the region and improve the accessibility and convenience of financial services. Analyzing the comprehensive service capacity of financial outlets within different functional areas is obviously insufficient in the current study, which is not conducive to the configuration planning of financial outlets and the promotion of the accessibility of urban financial services. The wide application of Point of Interest (POI) data in spatial analysis provides a new perspective for the study of urban financial outlets, for this reason, this paper analyzes and researches the current situation and rationality of the spatial layout of financial outlets based on POI data, under the perspective of the functional areas of the city, and combined with the population travel vitality index. The Term Frequency-Inverse Document Frequency (TF-IDF) algorithm is used to process the POI data to generate the urban functional area, on the basis of which the travel vitality index in the city is calculated by combining the population coverage and road network density, so as to judge the reasonableness of the layout of urban financial outlets. This helps us to better understand the behavioral decision-making process and social interaction of human beings about financial life in the city, and provides guidance and decision-making support for urban financial layout planning, financial public services and financial policies.
Social media platforms enable efficient traffic event detection by allowing users to produce geo-tagged content (e.g., tweets) known as geosocial media data. Geosocial media data improve road safety by providing timely updates for traffic flow and traffic control. Recent studies on traffic event detection with geosocial media data have been focused around keyword-based query approaches, where the event content was inferred by predetermined categories, to retrieve relevant traffic events. Spatiotemporal features associated with traffic-related posts have not been fully investigated. In this study, we filtered irrelevant posts with association rules. A spatiotemporal clustering-based method was then used to retrieve traffic events from these filtered posts, where the content of detected events was automatically inferred with a set of representative terms. For comparison, a typical text classification-based method was also used by classifying the posts filtered from association rules into different categories. By validating the detection results with vehicle travel speed data, we demonstrate that the former outperforms the latter in terms of the number of correctly detected traffic events from one-year of Twitter data in Toronto, Canada. Our proposed approach helps organizations and governments to be aware of when and where traffic events occur by identifying event hotspots and peak periods, which improves both traffic management and urban planning.
Urban Functional Zone (UFZ) identification facilitates the understanding of urban systems, which are complex and huge, and helps promote sustainable urban development. Existing studies on UFZ identification with Points of Interests (POIs) have focused much on more accurately extracting functional semantics, but ignored the fine delineation of UFZs in the spatial domain. The fine delineation of the spatial units of UFZs is also a key issue in UFZ identification. Since the sizes of UFZs can be different in practice, it is difficult to delineate spatially heterogeneous UFZs on a fixed scale. To solve the issue, a novel multi-scale spatial segmentation method was proposed in this study. Through taking the homogeneous socio-economic attributes of UFZs into account, we firstly generated a number of multi-scale spatial units by computing the mixed degree of POIs types, which reflects the mixed functions of each UFZs, using information entropy. Subsequently, we constructed the urban functional corpus of each spatial unit by measuring the spatial distribution pattern of POIs. The Word2Vec model was employed to obtain the semantic embedding vectors of UFZs, following which we adopted cosine distance-based K-means clustering method to group similar UFZs into one cluster. Finally, the enrichment factor was used to help annotate each functional cluster with a specific label. The UFZ identification results were compared with the Baidu e-maps and Baidu street view images for evaluation, and an accuracy of 82.7% was obtained. This study considering the heterogeneous distribution of POIs supports the fine-grained identification of UFZs, providing reference for urban planning.
The identification of urban functional regions (UFRs) is important for urban planning and sustainable development. Because this involves a set of interrelated processes, it is difficult to identify UFRs using only single data sources. Data fusion methods have the potential to improve the identification accuracy. However, the use of existing fusion methods remains challenging when mining shared semantic information among multiple data sources. In order to address this issue, we propose a context-coupling matrix factorization (CCMF) method which considers contextual relationships. This method was designed based on the fact that the contextual relationships embedded in all of the data are shared and complementary to one another. An empirical study was carried out by fusing point-of-interest (POI) data and taxi origin–destination (OD) data in Beijing, China. There are three steps in CCMF. First, contextual information is extracted from POI and taxi OD trajectory data. Second, fusion is performed using contextual information. Finally, spectral clustering is used to identify the functional regions. The results show that the proposed method achieved an overall accuracy (OA) of 90% and a kappa of 0.88 in the study area. The results were compared with the results obtained using single sources of non-fused data and other fusion methods in order to validate the effectiveness of our method. The results demonstrate that an improvement in the OA of about 5% in comparison to a similar method in the literature could be achieved using this method.
Three-dimensional point cloud has been widely used in the cultural heritage field in the last two decades, gaining attention from both academic and industry communities. A large number of scientific papers have been published concerning this topic, which covers a wide range of journals, countries, and disciplines. There has been no comprehensive and systematic survey of recent literature performed in a scientometric way based on the complex network analysis methods. In this work, we extracted the terms (i.e., noun phrases included in the title, abstract and keywords), the documents, the countries that the research institutions are located in, and the categories that the literature belongs to from the Web of Science database to compose a term co-occurrence network, document co-citation network, collaborative country network and category co-occurrence network using CiteSpace software. Through visualizing and analyzing those networks, we identified the research hotspots, landmark literature, national collaboration, interdisciplinary patterns as well as the emerging trends through assessing the central nodes and the nodes with strong citation bursts. This work not only provides a structured view on state-of-art literature, but also reveals the future trends of employing 3D point cloud data for cultural heritage, aiding researchers carry out further research in this area.
Traffic flow forecasting is a crucial task in transportation and necessary for congestion mitigation, traffic control, and intelligent traffic management. Deep learning models can aid in high-accuracy traffic flow forecasting; however, the current research focuses only the ability of the model to capture dynamic spatiotemporal features, and studies on the effect of deeper network layers on spatiotemporal features—a critical factor affecting traffic flow forecasting accuracy—are limited. In this paper, we propose an attention-based spatiotemporal graph attention network (ASTGAT) model designed for network degradation and over-smoothing problems to investigate in-depth spatiotemporal information. Compared to other networks, ASTGAT can capture dynamic spatiotemporal correlations in data and deepen the network to improve prediction accuracy through multiple residual convolution and high-low feature concat. ASTGAT comprises three components that separately model the temporal relationships of the recent, daily, and weekly periods. Each component stacks multiple spatiotemporal blocks constructed using the attention mechanism, dilated gated convolution, and graph attention network. The graph and temporal attention layers capture spatiotemporal information dynamically, and the graph attention layer alleviates the over-smoothing phenomenon to deepen the network. The combined utilization of the attention mechanism and dilated gated convolution layer improves the medium and long temporal span prediction ability. We validated ASTGAT using two open highway datasets, and the results demonstrated that our ASTGAT model effectively extracts in-depth spatiotemporal information and the prediction results outperform those predicted by the current eight baselines. Our research is dedicated to establishing a better scientific basis for intelligent traffic management that can assist in decision making.
Urban functional regions (UFRs) are formed and developed with human social actions and can reflect urban land use types. Appropriately identifying UFRs helps solve existing urban problems, optimize the spatial structure of cities, and provide a database for sustainable urban development. Most existing studies focus on developing novel methods and fusing multiple data sources, but neglect the impact of heterogeneous spatial units on UFR identification results. In this work, a hierarchical spatial unit partition approach was proposed to split the research area into many hierarchical units whilst considering the mixed degree of each unit. We further explored its capacity for identifying UFRs by integrating it with three widely used UFR identification methods. The results reveal that our proposed approach can correctly identify 10% more UFRs compared with the traditional grid‐based method, showing the efficiency of our proposed approach for identifying UFRs at a finer scale. This work provides support for those who are engaged in urban planning and urban policymaking, promoting urban sustainable development.
Urban Functional Zones (UFZs) can be identified by measuring the spatiotemporal patterns of activities that occur within them. Geosocial media data possesses abundant spatial and temporal information for activity mining. Identifying UFZs from geosocial media data aids urban planning, infrastructure, resource allocation, and transportation modernization in the complex urban system. In this work, we proposed an integrated approach by combining the spatiotemporal clustering method with a machine learning classifier. The spatiotemporal clustering method was used to mine the spatiotemporal patterns of activities, of which the distinctive features were extracted as inputs into a machine learning classifier for UFZ identification. The results show that more than 80% of the UFZs can be correctly identified by our proposed method. It reveals that this work serves as a functional groundwork for future studies, facilitating the understanding of urban systems as well as promoting sustainable urban development.
Urban Functional Zone (UFZ) identification is vital for urban planning, renewal, and development. Point of Interest (POI), as one of the most popular data in UFZ studies, is transformed into a geo-corpus under specific sampling strategies, which can be used with Natural Language Processing (NLP) technology to extract geo-semantic features and identify UFZs. However, existing studies only capture a single spatial distribution pattern of POIs, while ignoring the other spatial distribution information. In this paper, we developed an integrated geo-corpus construction approach to capture multi-spatial distribution patterns of POIs that were represented by different modal POI embeddings. Subsequently, random forest model was leveraged to classify UFZs based on those embeddings. A set of combination experiments were designed for performance validation. The results show that our proposed method can effectively identify UFZs with an accuracy of 72.9%, with an improvement of 8.5% compared to the baseline methods. The outcome of this study will help urban planners to better understand UFZs through investigating the integrated spatial distribution patterns of POIs embedded in UFZs.
At present, China's tourism industry has entered an era of rapid development. It is of great significance to predict the precise and real‐time tourist flow of the popular tourist areas with dense crowds, providing decision support for rational evacuation of tourist flow, activation of emergency plans, and prevention of security accidents. This article presents a long short‐term memory (LSTM) neural network model based on the global attention mechanism. The model first uses two LSTM layers to calculate attention weights using multi‐source data at different time steps to improve the model’s learning ability. Then the model predicts the tourist flow of the scenic spots in the next time period based on the weights and outputs previously obtained. At the final stage, the pre‐trained parameters of the network are used to initialize the model. To verify the validity of the model, we compared it with the LSTM model, back propagation neural network model, and autoregressive integrated moving average model based on the data of Beijing’s South Luogu Lane scenic spot. It turned out that the results of our solution were more in line with the true value, which proves that it is feasible for real‐time prediction of tourist flow in popular scenic spots.
Grottoes, with caves and statues, are an important part of immovable heritage. Statues in a particular grotto setting are often similar in geometric form and artistic style, and identifying the similarity between these statues can help provide important references for value recognition, condition assessment, repair, and the virtual restoration of statues. Traditionally, such reference information mainly depended on expert empirical judgment, which is highly subjective, lacks quantitative analysis, and cannot provide effective scientific support for the virtual restoration of grotto statues. This paper presents a similarity index based approach for identifying similarities between grotto statues by studying 11 small Buddhist statues carved on the 18th cave in the Yungang Grottoes, located in Datong, China. The similarity index is determined according to the hash values calculated based on the pHash method using the orthophoto images of Buddhist statues to identify similar statues. Similar feature points between the identified statues are then matched using the Scale Invariant Feature Transform (SIFT) operator to support the repair and reconstruction of damaged statues. The experimental results show that the variation of similarity index values confirms the visual inspection of the statues’ appearance in the orthophotos. The additional analysis of three-dimensional (3D) point clouds also confirms that the similarity index based approach is accurate in the initial screening of similar grotto statues.
Social media networks allow users to post what they are involved in with location information in a real‐time manner. It is therefore possible to collect large amounts of information related to local events from existing social networks. Mining this abundant information can feed users and organizations with situational awareness to make responsive plans for ongoing events. Despite the fact that a number of studies have been conducted to detect local events using social media data, the event content is not efficiently summarized and/or the correlation between abnormal neighboring regions is not investigated. This article presents a spatial‐temporal‐semantic approach to local event detection using geo‐social media data. Geographical regularities are first measured to extract spatio‐temporal outliers, of which the corresponding tweet content is automatically summarized using the topic modeling method. The correlation between outliers is subsequently examined by investigating their spatial adjacency and semantic similarity. A case study on the 2014 Toronto International Film Festival (TIFF) is conducted using Twitter data to evaluate our approach. This reveals that up to 87% of the events detected are correctly identified compared with the official TIFF schedule. This work is beneficial for authorities to keep track of urban dynamics and helps build smart cities by providing new ways of detecting what is happening in them.
Social media platforms allow millions of people worldwide to instantly share their thoughts online. Many people use social media to share traffic related experiences and events with online posts. A large amount of traffic related data can be obtained from these online posts – especially geosocial media data, where posts are tagged with geolocation information such as coordinates or place names. By extracting traffic events from geosocial media data, drivers can adapt to changing traffic conditions, while traffic management departments can propose timely and effective plans to improve traffic conditions. Most of the existing studies query traffic-related information based on a list of single keywords, which result in large amounts of noisy data – negative data containing one or more traffic-related keywords, but do not actually represent real-world traffic events. This paper aims to filter noisy data by mining association rules among words in positive data containing messages representing traffic events. Messages are more likely to be true traffic events if they follow the co-occurrence pattern of words mined from positive samples. A case study was conducted in Toronto, Canada using Twitter data. The tweets queried by the association rules were classified into non-traffic event, traffic accidents, roadwork, severe weather conditions, and special events with an 85% accuracy based on supervised machine learning methods. Compared with hourly average travel speed data, 81% of detected events were identified as real-world traffic events. This research sheds light on traffic condition monitoring in smart transportation platforms, which plays an important role for smart cities.
Social media platforms, or social networks, have allowed millions of users to post online content about topics related to our daily lives. Traffic is one of the many topics for which users generate content. People tend to post traffic related messages through the ever-expanding geosocial media platforms. Monitoring and analyzing this rich and continuous user-generated content can yield unprecedentedly valuable traffic related information, which can be mined to extract traffic events to enable users and organizations to acquire actionable knowledge. A great number of literature has reported on the methods developed for detecting traffic information from social media data, especially geosocial media data when geo-tagged. However, a systematic review to synthesize the state-of-the-art developments is missing. This paper presents a systematic review of a wide variety of techniques applied in detecting traffic events from geosocial media data, arranged based on their adoption in each stage of an event detection framework developed from the literature review. The paper also highlights some challenges and potential solutions. The aim of the paper is to provide a structured view on current state-of-art of the geosocial media based traffic event detection techniques, which can help researchers carry out further research in this area.