
Based on spatial object and relation model of spatial database management system, this paper propose an “integrated spatial data-attribute data” standardized data structure of multisource heterogeneous satellite data. The structure is described by spatial table, and provides a unified access and processing interface of multisource heterogeneous satellite data. Using this standardized data structure, we can provide protection for satellite data management application system of "Single-requirement and multi-satellite".
The amount of digital data is increasing beyond any previous estimation and data stores and sources are more and more pervasive and distributed. Professionals and scientists need advanced data analysis tools and services coupled with scalable architectures to support the extraction of useful information from big data repositories. Cloud computing systems offer an effective support for addressing both the computational and data storage needs of big data mining and parallel knowledge discovery applications. In fact, complex data mining tasks involve data- and compute-intensive algorithms that require large and efficient storage facilities together with high performance processors to get results in acceptable times. In this paper we introduce the topic and the main research issues. We discuss how to make knowledge discovery services scalable and present the Data Mining Cloud Framework designed for developing and executing distributed data analytics applications as workflows of services. In this environment we use data sets, analysis tools, data mining algorithms and knowledge models that are implemented as single services that can be combined through a visual programming interface in distributed workflows to be executed on Clouds. The main features of the programming interface are described and performance evaluation of knowledge discovery applications are reported.
Lines and Polygons Overlay Analysis (LP-OA) is an important analysis method that has been widely used in geographic information systems (GIS). In this study, a new algorithm, Uniform Spatial Grid Indexing (USGI), is proposed to study LP-OA. By applying this method, the theoretical position is evaluated and the theoretical results are verified, suggesting that this algorithm can contribute to the software development in commercial GIS.
Laser scanning (also known as Light Detection And Ranging) has been widely applied in various application. As part of that, aerial laser scanning (ALS) has been used to collect topographic data points for a large area, which triggers to million points to be acquired. Furthermore, today, with integrating full wareform (FWF) technology during ALS data acquisition, all return information of laser pulse is stored. Thus, ALS data are to be massive and complexity since the FWF of each laser pulse can be stored up to 256 samples and density of ALS data is also increasing significantly. Processing LiDAR data demands heavy operations and the traditional approaches require significant hardware and running time. On the other hand, researchers have recently proposed parallel approaches for analysing LiDAR data. These approaches are normally based on parallel architecture of target systems such as multi-core processors, GPU, etc. However, there is still missing efficient approaches/tools supporting the analysis of LiDAR data due to the lack of a deep study on both library tools and algorithms used in processing this data. In this paper, we present a comparative study of software libraries and new algorithms to optimise the processing of LiDAR data. We also propose new method to improve this process with experiments on large LiDAR data. Finally, we discuss on a parallel solution of our approach where we integrate parallel computing in processing LiDAR data.
Distributed data mining techniques and mainly distributed clustering are widely used in the last decade because they deal with very large and heterogeneous datasets which cannot be gathered centrally. Current distributed clustering approaches are normally generating global models by aggregating local results that are obtained on each site. While this approach mines the datasets on their locations the aggregation phase is complex, which may produce incorrect and ambiguous global clusters and therefore incorrect knowledge. In this paper we propose a new clustering approach for very large spatial datasets that are heterogeneous and distributed. The approach is based on K-means Algorithm but it generates the number of global clusters dynamically. Moreover, this approach uses an elaborated aggregation phase. The aggregation phase is designed in such a way that the overall process is efficient in time and memory allocation. Preliminary results show that the proposed approach produces high quality results and scales up well. We also compared it to two popular clustering algorithms and show that this approach is much more efficient.
Multi-temporal Landsat images in 1999, 2006 and 2013 were used to reveal the spatial-temporal changes of the coastal locations and types of Xiamen Island. The high tide water levels were extracted as coastline approximations and the coastal types were made by manual interpretation based on Google earth high resolution images in this study. The results indicated that the coast changes mainly happened at the northern and eastern parts of the island. The artificial coast came up to 46.4 km with an increment about 8.8 km due to sea reclamation while the mudflat coast decreased about 18.1 km and faded away from 1999 to 2006. Then there was another increase of artificial coast due to returning ponds back to sea, and the sand coast also increased due to sand supply project from 2006 to 2013. Thus the artificial coast increased to 48.0 km and the sand coast increased from 10.1 km to 16.9 km. The total coastline decreased from 66.4 km to 57.0 km first but then increased back to 65.5 km during the whole period while the island area had a persistent growth from 128.5 km2 to 140.0 km2. The coastlines moved northward and seaward with less curves due to local constructions except the parts returning ponds back to sea. The rapid social and economic development raised great demand of land use as well as the environmental quality, which made sea reclamation and returning ponds back to sea the main causes of coast changes of Xiamen Island.
Terrestrial Laser Scanner (TLS) has been used to record three-dimensional (3D) point cloud data of trees, and this data is being further processed to extract morphological parameters of standing trees and reconstruct 3D geometric model. We present an efficient, high-accuracy and photorealistic 3D single tree reconstruction technique based on the point cloud of individual standing tree acquired mainly from TLS. Our method can process point cloud data of the whole tree directly without the need to first segment data into leaves and branches. It also allows user to interactively edit and adjust model parameters to better fie the data.
The context for geographic research has shifted from a data-scarce to a data-rich environment, in which the most fundamental changes are not just the volume of data, but the variety and the velocity at which we can capture geo-referenced data; trends often associated with the concept of Big Data. A data-driven geography may be emerging in response to the wealth of geo-referenced data flowing from sensors and people in the environment. Although this may seem revolutionary, in fact it may be better described as evolutionary. Some of the issues raised by data-driven geography have in fact been longstanding issues in geographic research, namely, large data volumes, dealing with populations and messy data, and tensions between idiographic versus nomothetic knowledge. The belief that spatial context matters is a major theme in geographic thought and a major motivation behind approaches such as time geography, disaggregate spatial statistics and GIScience. There is potential to use Big Data to inform both geographic knowledge discovery and spatial modeling. However, there are challenges, such as how to formalize geographic knowledge to clean data and to ignore spurious patterns, and how to build data-driven models that are both true and understandable.
Rapid growth of urban public facilities increased the important impact on urban component management in case number and hotspot space distribution. Many urban component management events such as huddle of wastes and illegal ads have been frequently hot issues. To support the government decision, methods and flows based on spatial analyst are put forward to explore the hot spot of urban component management event. Illegal ads event is taken as case study. The event's spatial correlation is decided by using exploratory spatial data analysis and Global Moran's I. Ripley's K is used to explore its spatial distribution pattern, which provides the basis for urban management events data mining. To identify the spatial border of hot spots, the spatial hot spot analysis is employed. Finally, the overlay spatial analyst between hot spot and business area layer and/or residential building layer is used to validate the feasibility and rationality of the hot spot analysis. It is concluded that illegal ads event has obvious spatial clustering and coincides exactly with the existing business area and places where population flows more frequently. According to the result, the urban management administrators can make decision more conveniently and reasonably. Comparing with time series analysis methods, this method has advantage in space information, good visualization and dynamic monitor which are useful for government decision.
Moving characteristics of ocean eddy have become one of the research foci, based on more advanced detection and tracking method. This study applied trajectory data mining to the analysis of eddies trajectories in the South China Sea (SCS), and experimented with two methods. One of them identified moving patterns of eddies by trajectory clustering and obtained representative paths. The other one discovered flow patterns using regionalization method with contiguous constrained, which partitioned of SCS into 4 regions. In the SCS, eddies mainly propagate in three parts with various moving characteristics and present regional flow patterns. Both of these methods validate the effectiveness of the application of trajectory data mining to the ocean eddies study.
Two advanced modelling approaches, Multi-Level Models and Artificial Neural Networks are employed to model house prices. These approaches and the standard Hedonic Price Model are compared in terms of predictive accuracy, capability to capture location information, and their explanatory power. These models are applied to 2001-2013 house prices in the Greater Bristol area, using secondary data from the Land Registry, the Population Census and Neighbourhood Statistics so that these models could be applied nationally. The results indicate that MLM offers good predictive accuracy with high explanatory power, especially if neighbourhood effects are explored at multiple spatial scales.
Online social networks have played an important role in people's common life. Most existing social network platforms, however, face the challenges of dealing with undesirable users and their malicious spam activities that disseminate content, malware, viruses, etc. to the legitimate users of the service. In this paper, an Extreme Learning Machine based supervised machine is proposed for effective spammer detection. The experiment and evaluation show that the proposed solution provides excellent performance with a true positive rate of spammers and non-spammers reaching 99% and 99.95%, respectively. As the results suggest, the proposed solution could achieve better reliability and feasibility compared with existing SVM based approaches.
This paper introduces the binary granule into the algorithm in computing and mining meteorological data. By redefining the binary algorithm, matching operators, convergence operators, disjunction operators, shift algorithm, and the shielding granule are given based on the binary granule. The function of operators is also explained, then different operators are applied on specific computing methods according to various requirements. Based on the binary granule, the whole sequence match algorithm is put forward, thus space-time efficiency of computing methods can be increased. Based on the knowledge of drought and flood distributions in China, this paper transforms meteorological drought and flood data into event set in the corresponding space and time after they are pre-processed using Standardized Precipitation Index (SPI) algorithm. Made by binary granulation, this event set is changed into the binary granule drought and flood event sets in different spatiotemporal scales. Granule operators are then applied to the whole sequence match algorithm, which describes the measuring module and definition of the whole sequence similarity matching, lastly the match mining is carried out. The result shows that research and application of the algorithm can provide a new method for monitoring and predicting drought and flood events in studying regional and local climate similarity further.
This paper aims to improve our knowledge of the complex vegetation-climate relationship in subtropical humid region in the context of global warming, by taking into considerations of spatio-temporal variation of both vegetation and climate change. A multi-resolution analysis (MRA) based on the wavelet transform (WT) is applied to examine the vegetation growth and its relationship with climate factors based on 250m 16-day composites MODIS vegetation EVI dataset in subtropical humid region of China over the period 2001-2010. A general greening up (68%) was observed over the period 2001-2010, as well as rather local negative trends. A trend toward global warming was also observed for the whole study region, whereas no obvious trend of precipitation is examined in most areas. Temperature generally has a positive influence on vegetation; with only very few negative EVI-temperature coefficients observed on the south portion principally due to changes in land use, land degradation, and cloud noise. However, nearly equally positive and negative EVI-rainfall relationship is observed on the inter-annual level, with negative coefficients principally observed in the northwest portion with abundant precipitation. Very strong positive relationship is observed between both EVI-temperature and EVI-precipitation at seasonal level. It is revealed that spatio-temporal variation of both vegetation and climate should be taken into considerations when analyzing the long-term effects of global climate change. Interactions between vegetation dynamics and climate variability must be studied through spatially and temporally explicit multi-scale analysis to investigate the influence of long-term climate change on the vegetation growth.
China has been the study area for this paper. In this paper, we took series of MODIS Enhanced Vegetation Index (EVI) products from 2001 to 2012 to create the temporal vegetation coverage (TVC) as indication factor of vegetation cover. Then we selected a latitudinal transect along the latitude 30°N and used the continuous wavelet variance to test the characteristic scale of vegetation distribution with the wavelet coherency to detect the multi-scale relationship between vegetation cover with the effect factors through the continuous wavelet transform. It was shown that: (1) The study area is divided into multiple nested pattern levels which are larger river effect pattern, physiognomy effect pattern and the macro structure of Chinese Three Gradient Terrains. (2) In the scales of 0-300km, vegetation cover were greatly influenced by factitious factors; In the 300-600km scales, a major scale characteristic, vegetation cover has been influenced by elevation, temperature, land use/cover and factitious factors, and vegetation heterogeneity is greatly strong on this scale; In the scales of 600-1,200km, vegetation cover is intensity influenced by climate factors. (3) Elevation plays an important role to vegetation cover. At low elevation area, vegetation cover increases with elevation growth, while it appears the reverse relation in high elevation area. In the west area, the correlation between vegetation cover and precipitation is negative, while vegetation cover and temperature presents a strong positive correlation. In the whole eastern area, this relationship is not as strong as in western area, and in the mountains or hilly area, temperature and vegetation cover has a positive correlation. (4) Artificial factors show some inhibitory effects of vegetation cover that the vegetation coverage is relatively poor in the area with intensive human activities.
The airborne laser scanning (ALS) has become the most effective way to acquire 3D information of the earth's surface. However, data post-processing lags far behind hardware especially the filtering algorithm. At present, most algorithms for DEM generation are based on the mutation of the elevation, and every method has its own defects. All of the existing filter algorithms are based on a priori information, so that the accuracy of the filtering is affected by the terrain fluctuation. In this paper, a new algorithm called Hybrid Filtering of Lidar Data based on the Echoes was proposed. First, we make an analysis of echo characteristic of airborne Lidar data, and propose classification schemas based on the characteristic of echo characteristic. Then we use the Pseudo Grid for segmentation and index. Furthermore, using the existing filtering algorithm to distinguish the seed point is terrain point or not. At last, we filter the cooked data. In theory, because of removing part of object point the accuracy and the efficiency must be significantly increased. Experiments show that it is effectively to eliminate most of vegetation and building points, that means, the method proposed in this paper not only can reduce the amount of computing data, but also improve the effect of filtering algorithm for eliminating the building and vegetation.
Forest fire in Indonesia is considered as an annual event that causes serious problems in health and environment especially in Sumatera and Kalimantan Islands. Studies on analyzing hotspot data as forest fire indicators are required for predicting hotspot occurrences. The objective of this work is to detect global and collective outliers on hotspot data in Riau Province in Sumatera Island for the period 2001-2012. The data used in this work are 4383 daily hotspots and 144 monthly hotspots. The method applied to discover outliers is the k-means clustering algorithm. The best clustering results are obtained on the number of clusters of 10 and the sum of squared error value is 18526.14. Based on the clustering results, we obtain 59 collective outliers and 30 global outliers on the hotspot dataset. The outliers on the hotspot data mostly occur in February, March, June, July, and August. The average frequency of outliers is 482.22 and the highest frequency of outliers is occurred in 2005. As many 1118 hotspots were found in the northern part of the Riau province on 21 June 2005. In August 2005 outliers spread on the whole area of Riau Province. For the period 2001-2012 there are no outliers occurred in April, November and December. This information is essential for an early warning system in forest fires prevention.
Identifying spatial patterns of geographic entities such as retail stores is important in city for understanding how they behave. The pattern formed by the distribution of points can be measured by some quantitative methods. In Big Data era, the data sets for spatial patterns analysis are various including traditional street network data and points of interest (POIs) data in LBS (Location based services) application. This paper analyzed the spatial pattern of retail stores and its correlations with street centrality using POIs data in Zhengzhou, China. Firstly, the paper provided an exploratory analysis of spatial patterns using the centrographic methods including Standard Deviational Ellipse and Average Nearest Neighbor. Secondly, the paper uncovered the spatial distribution of retail stores using the kernel density estimation (KDE). Finally, the paper calculated the street centrality of Zhengzhou using three centrality assessment indexes and converted all nodes centrality index values to raster pixel using KDE for correlation analysis. Results show that the retail stores are clustering pattern and mainly elongated along the west-east direction. The street centralities are correlated with the retail store location in Zhengzhou, and there is a different level of correlation between them. The paper reveals that the spatial pattern analysis and street centralities index are valuable in location analysis or urban planning.
Recovering 3D depth from a single outdoor image is a basic problem in Computer Vision and Close-Range Photogrammetry. In this paper, an efficient depth estimation approach from a single outdoor image is presented. According to scene classification, depth of regions marked as sky, ground and vertical labels is respectively predicated. Firstly, a more accurate depth calculation model for ground regions is deduced from the camera imaging model, which takes camera pitch into account. Then the depth of vertical regions is estimated based on the ground regions depth. To consider the depth change of the regions, the vertical regions with visible ground-contain points are regarded as composed of one or multi-planes, and the vanishing points are incorporated into those points to further improve the position accuracy of intersecting lines. The contrastive experimental results are shown to verify the effectiveness and validity of the proposed approach.
In view of the problems existing in traditional recommendation algorithm of low accuracy and low efficiency, this paper presents a machine learning based social media recommendation algorithm. The algorithm is based on the traditional personalized collaborative filtering algorithm, and combines with the correlation characteristics among users in a social network. Besides, the algorithm also considers the network rating factors and upgrade its efficiency by using clustering algorithm. At last, the algorithm is realized on the Hadoop cloud platform.