In light of the rapidly growing passenger and flight volumes, airports seek for sustainable solutions to improve passengers’ experience and comfort, while maximizing their profits. A major technological solution towards improving service quality and management processes in airports comprises Internet of Things (IoT) systems that realize the concept of smart airports and offer interconnection potential with other public infrastructures and utilities of smart cities. In order to deliver smart airport services, real-time flight delay data and forecasts are a critical source of information. This paper introduces an essential methodology using machine learning techniques on Apache Spark, a cloud computing framework, with Apache MLlib, a machine learning library to develop and implement prediction models for air flight delays that could be integrated with information systems in order to provide up-to-date analytics. The experimental results have been implemented with various algorithms in terms of classification as well as regression, thus manifesting the potential of the proposed framework.
We investigate the problem of finding the visible pieces of a scene of objects from a specified viewpoint. In particular, we are interested in the design of an efficient hidden surface removal algorithm for a scene comprised of iso-oriented rectangles. We propose an algorithm where given a set of $n$ iso-oriented rectangles we report all visible surfaces in $O((n+k)\log n)$ time and linear space, where $k$ is the number of surfaces reported. The previous best result by Bern, has the same time complexity but uses $O(n\log n)$ space.
Privacy Preserving and Anonymity have gained significant concern from the big data perspective. We have the view that the forthcoming frameworks and theories will establish several solutions for privacy protection. The k-anonymity is considered a key solution that has been widely employed to prevent data re-identifcation and concerns us in the context of this work. Data modeling has also gained significant attention from the big data perspective. It is believed that the advancing distributed environments will provide users with several solutions for efficient spatio-temporal data management. GeoSpark will be utilized in the current work as it is a key solution that has been widely employed for spatial data. Specifically, it works on the top of Apache Spark, the main framework leveraged from the research community and organizations for big data transformation, processing and visualization. To this end, we focused on trajectory data representation so as to be applicable to the GeoSpark environment, and a GeoSpark-based approach is designed for the efficient management of real spatio-temporal data. Th next step is to gain deeper understanding of the data through the application of k nearest neighbor (k-NN) queries either using indexing methods or otherwise. The k-anonymity set computation, which is the main component for privacy preservation evaluation and the main issue of our previous works, is evaluated in the GeoSpark environment. More to the point, the focus here is on the time cost of k-anonymity set computation along with vulnerability measurement. The extracted results are presented into tables and figures for visual inspection.
Identifying differentially expressed subpathways connected to the emergence of a disease that can be considered as candidates for pharmacological intervention, with minimal off-target effects, is a daunting task. In this direction, we present a bilevel subpathway analysis method to identify differentially expressed subpathways that are connected with an experimental condition, while taking into account potential crosstalks between subpathways which arise due to their connectivity in a combined multi-pathway network. The efficacy of the method is demonstrated on a hematopoietic stem cell aging dataset, with findings corroborated using recent literature.
The vast accumulation of electronically available literature in the fields of Biology and Medicine has raised new challenges in Knowledge Discovery technology and provides increasingly attractive opportunities for Text Mining. In this work we present a methodology for concept discovery from the Molecular Biology Literature. Our approach combines Natural Language Processing Techniques and Clustering Methods in order to produce clusters of biological abstracts based on term co-occurence. Experiments show that the resulting document clusters are meaningful as assesed by cluster-specific terms. The application of this method to a collection of abstracts relevant to transcription factors provided a shallow description of the document corpus and supported classification of cancer specific terms.
Twitter Sentiment Classification is undergoing great appeal from the research community; also, user posts and opinions are producing very interesting conclusions and information. In the context of this paper, a pre-processing tool was developed in Python language. This tool processes text and natural language data intending to remove wrong values and noise. The main reason for developing such a tool is to achieve sentiment analysis in an optimum and efficient way. The most remarkable characteristic is considered the use of emojis and emoticons in the sentiment analysis field. Moreover, supervised machine learning techniques were utilized for the analysis of users' posts. Through our experiments, the performance of the involved classifiers, namely Naive Bayes and SVM, under specific parameters such as the size of the training data, the employed methods for feature selection (unigrams, bigrams and trigrams) are evaluated. Finally, the performance was assessed based on independent datasets through the application of k-fold cross validation.
The emergence of diseases and drug-induced perturbations are oftentimes the cause of biological pathway deregulations. Identifying differentially expressed subpathways in organism-level networks of signaling pathways can be a computationally intensive undertaking, due to their complexity. In this direction, we present a subpathway analysis method which refines organism-level networks via a two step-approach. The method first constructs a core-network of differentially expressed genes and subsequently includes a set of topologically significant non-differentially expressed genes, that both exhibit correlated expression levels with their neighbors and facilitate the signal propagation within the core-network. The refined network is then searched for differentially expressed subpathways using a plethora of subpathway identification methods and checked for enrichment in functional terms such as drugs and diseases. The approach assesses the differential expression of the subpathways using a consensus approach, by detecting even weak signals of differential expression, while accounting for correlations which arise in gene expression data.
In the context of this research work, we studied the problem of privacy preserving on spatiotemporal databases. In particular, we investigated the k-anonymity of mobile users based on real trajectory data. The k-anonymity set consists of the k nearest neighbors. We constructed a motion vector of the form (x,y,g,v) where x and y are the spatial coordinates, g is the angle direction, and v is the velocity of mobile users, and studied the problem in four-dimensional space. We followed two approaches. The former applied only k-Nearest Neighbor (k-NN) algorithm on the whole dataset, while the latter combined trajectory clustering, based on K-means, with k-NN. Actually, it applied k-NN inside a cluster of mobile users with similar motion pattern (g,v). We defined a metric, called vulnerability, that measures the rate at which k-NNs are varying. This metric varies from 1/k (high robustness) to 1 (low robustness) and represents the probability the real identity of a mobile user being discovered from a potential attacker. The aim of this work was to prove that, with high probability, the above rate tends to a number very close to 1/k in clustering method, which means that the k-anonymity is highly preserved. Through experiments on real spatial datasets, we evaluated the anonymity robustness, the so-called vulnerability, of the proposed method.
One of the main characteristics of our time is the growth of the data volumes. We collect data literally from everywhere; smart phones, smart devices, social media and the health care system, which defines a small portion of the sources of the big data. The big data growth poses two main difficulties, storing and processing them. For the former, there are certain new technologies that enable us to store large amounts of data in a fast and reliable way. For the latter, new application frameworks have been developed. In this paper, we perform classification analysis using Apache Spark in one real dataset. The classification algorithms that we have used are multiclass, and we are going to examine the effect of the dataset size and input features on the classification results.
Nowadays, digital data are the most valuable asset of almost every organization. Database management systems are considered as storing systems for efficient retrieval and processing of digital data. However, effective operation, in terms of data access speed and relational database is limited, as its size increases significantly [6]. Bloom filter is a special data structure with finite storage requirements and rapid control of an object membership to a dataset. It is worth mentioning that the Bloom filter structure has been proposed with a view to constructively increase data access in relational databases. Since the characteristics of a Bloom filter are consistent with the requirements of a fast data access structure, we examine the possibility of using it in order to increase the SQL query execution speed in a database. In the context of this research, a database in a RDBMS SQL Server that includes big data tables is implemented and in following the performance enhancement, using Bloom filters, in terms of execution time on different categories of SQL queries, is examined. We experimentally proved the time effectiveness of Bloom filter structure in relational databases when dealing with large scale data.
In the era of Systems Biology and growing flow of omics experimental data from high throughput techniques, experimentalists are in need of more precise pathway-based tools to unravel the inherent complexity of diseases and biological processes. Subpathway-based approaches are the emerging generation of pathway-based analysis elucidating the biological mechanisms under the perspective of local topologies onto a complex pathway network. Towards this orientation, we developed PerSub, a graph-based algorithm which detects subpathways perturbed by a complex disease. The perturbations are imprinted through differentially expressed and co-expressed subpathways as recorded by RNA-seq experiments. Our novel algorithm is applied on data obtained from a real experimental study and the identified subpathways provide biological evidence for the brain aging.
During the last years, there is a huge proliferation in the usage of location-based services (LBSs), mostly through a multitude of mobile devices (GPS, smartphones, mapping devices, etc.). The volume of the data derived by such services, grows exponentially and conventional databases tend to be ineffective in storing and indexing them efficiently. Ultimately, we need to turn to scalable solutions and methods using the NoSQL database model. Quite a few indexing methods exist in literature that work on top of NoSQL database. In this spirit, we deploy a new distributed indexing structure based on M-tree and perform a thorough experimental analysis to display its benefits.
Community discovery is central to social network analysis as it provides a natural way for decomposing a social graph to smaller ones based on the interactions among individuals. Communities do not need to be disjoint and often exhibit recursive structure. The latter has been established as a distinctive characteristic of large social graphs, indicating a modularity in the way humans build societies. This paper presents the implementation of four established community discovery algorithms in the form of Neo4j higher order analytics with the Twitter4j Java API and their application to two real Twitter graphs with diverse structural properties. In order to evaluate the results obtained from each algorithm a regularization-like metric, balancing the global and local graph self-similarity akin to the way it is done in signal processing, is proposed.
It is vastly acknowledged that analyzing social networks is a very challenging research area. Take as a striking example the organization of vertices in clusters, with many edges joining vertices of the same cluster and comparatively few edges joining vertices of different clusters. This comprises a fundamental aspect, which concerns the detection of user communities. In certain fields such as sociology and computer science where interactions and associations are often represented in the form of graphs, detecting communities is of vital importance. This paper addresses the need for an efficient and innovative methodology for community detection that will also leverage users’ behavior on emotional level. Ekman emotional scale is the key point with which the methodology analyzes user’s tweets in order to determine their emotional behavior. Consequently, the derived communities are estimated with the use of three different metrics, while the weighted version of a modularity community detection algorithm is utilized. There is substantial evidence indicating that our proposed methodology creates influential enough communities.
Sentiment Analysis on Twitter Data is a challenging problem due to the nature, diversity and volume of the data. In this work, we implement a system on Apache Spark, an open-source framework for programming with Big Data. The sentiment analysis tool is based on Machine Learning methodologies alongside with Natural Language Processing techniques and utilizes Apache Spark's Machine learning library, MLlib. In order to address the nature of Big Data, we introduce some pre-processing steps for achieving better results in Sentiment Analysis. The classification algorithms are used for both binary and ternary classification, and we examine the effect of the dataset size as well as the features of the input on the quality of results. Finally, the proposed system was trained and validated with real data crawled by Twitter and in following results are compared with the ones from real users.
Complex networks can be considered as a new field of scientific research inspired by the empirical study of real-world networks such as computer, social as well as biological ones. More to this point, the study of complex networks has expanded in many disciplines including mathematics, physics, biology, telecommunications, computer science, sociology, epidemiology and others. An important type of complex networks are called biological dealing with the mathematical analysis of connections - interfaces that are ecological, evolutionary and physiological studies, such as neural networks or network epidemic models. The analysis of biological networks in connection with human diseases has led to expand science and examine medical supplies networks for their deeper understanding. In this paper, an implementation of epidemic/networks models is introduced concerning the HIV spreading in a sample of people who are needle drug users.
Sentiment Analysis on Twitter Data is indeed a challenging problem due to the nature, diversity and volume of the data. People tend to express their feelings freely, which makes Twitter an ideal source for accumulating a vast amount of opinions towards a wide spectrum of topics. This amount of information offers huge potential and can be harnessed to receive the sentiment tendency towards these topics. However, since no one can invest an infinite amount of time to read through these tweets, an automated decision making approach is necessary. Nevertheless, most existing solutions are limited in centralized environments only. Thus, they can only process at most a few thousand tweets. Such a sample is not representative in order to define the sentiment polarity towards a topic due to the massive number of tweets published daily. In this work, we develop two systems: the first in the MapReduce and the second in the Apache Spark framework for programming with Big Data. The algorithm exploits all hashtags and emoticons inside a tweet, as sentiment labels, and proceeds to a classification method of diverse sentiment types in a parallel and distributed manner. Moreover, the sentiment analysis tool is based on Machine Learning methodologies alongside Natural Language Processing techniques and utilizes Apache Spark's Machine learning library, MLlib. In order to address the nature of Big Data, we introduce some pre-processing steps for achieving better results in Sentiment Analysis as well as Bloom filters to compact the storage size of intermediate data and boost the performance of our algorithm. Finally, the proposed system was trained and validated with real data crawled by Twitter, and, through an extensive experimental evaluation, we prove that our solution is efficient, robust and scalable while confirming the quality of our sentiment identification.
In the recent decades, changes regarding aspects of human's daily needs are escalating. Technology has had a huge impact on users' everyday life and thus, everything has been modified to match the new data. Achieved advances in information, communication and network technology field have significantly influenced the way the health sector operates. Different healthcare systems have been developed in order to provide solutions. There are, however, many health issues that are yet to be solved, being a significant fact that leads to an increasing demand for more efficient and advanced healthcare services. The main objective of this paper is to pave the way for the Internet of Things to apply technological breakthroughs in the health sector so as to efficiently deal with major health problems and also contribute in the decreasing of healthcare costs.