
The efficiency of a road network may be improved by making the traffic intersections more efficient. A smart traffic intersection is equipped with different sensors from which it is possible today ...
Children's privacy compliance assessment in the area of smart connected toy (SCT) or play robot is a challenge, considering the plethora of device manufacturers using varied controls in an attempt ...
Network traffic has been a critical issue that has attracted massive attention in network operations research and the industry. This paper tackles the need to understand traffic patterns across a h...
Since the complete digitisation of civil processes that took place in Italy in 2008, a lot of data regarding the life cycle of thousands of civil proceedings has been collected. However, despite th...
Effective monitoring and analysis of network traffic are vital for scientific computing, since scientific applications often require moving massive data from one site to another. A body of statisti...
This paper inquires on the options pricing modeling using Artificial Neural Networks to price Apple's European Call Options. The model is based on the premise that ANNs can be used as functional approximators and used as an alternative to the numerical methods to some extent, for a viable and faster solution. We evaluate our predictions using the existing numerical solutions for the same, the analytic solution for the Black-Scholes equation, COS-Model for Heston's Stochastic Volatility Model and Standard Heston-Quasi analytic formula. The aim of this study is to find a viable time-efficient alternative to existing quantitative models for option pricing.
The online world has deeply changed the rules of information: a few selected systems have emerged as centralisers, providing simplified access. On the one side, search engines compress the number of pages to interact with, and on the other side Wikipedia tries to compress information itself. These systems have had an enormous success, but success also brings problems. In the case of Wikipedia, these problems are due to its distributed nature: everybody can contribute and so also manipulate information in a way that is practically invisible to the general public. We describe the Negapedia system, an online public service providing a more complete picture of this underlying layer. We explain the challenges and choices that had to be made: big data analysis, potential information overload, and novel insights on the important issue of Wikipedia categorisation, analysing the problem of presenting general users with easy and meaningful category information.
Massive volumes of data streams are expected to be generated by the internet of things (IoT). Due to their dispersed and mobile nature, they need to be processed using automated analytical tasks. The research challenge is to uncover whether the data streams, which are being generated by billions of IoT devices, actually conform to a data flow that is required to perform streaming analytics. In this paper, we propose process discovery and conformance checking techniques of process mining in order to expose the flow dependency of IoT data streams between automated analytical tasks running at the edge of a network. Towards this end, we have developed a Petri Net model to ensure the optimal execution of analytical tasks by finding path deviations, bottlenecks, and parallelism. A real-world scenario in smart transit is used to evaluate the full advantage of our proposed model. Uncovering the actual behaviour of data flows from IoT devices to edge nodes has allowed us to detect discrepancies that have a negative impact on the performance of automated analytical tasks.
Data security is an important issue in big data applications. The sheer data volume provides way more opportunities for a potential attacker to observe and identify patterns in computation and data. In this paper, we reveal that the data/computation patterns derived from the observation of large volume of data can be associated with the key used in the AES-GCM algorithm, one of the foundation algorithms in data security. The paper presents a software-based cache-collision timing attack against the well known authenticated encryption scheme AES-GCM. The attack can be successful if enough data (plaintext-ciphertext pairs) are processed and the hash key H used for generating look-up tables in software implementation. We present an attack model and an implementation of the attack based on OpenSSL, a widely used library that provides security-related functions for many applications. In most cases, our attack methodology is able to converge and extract the hidden key.
In the information age, data integration has become easier than ever. Enterprises integrate a wide range of data sources to enrich big data lakes. Enterprise big data lake made data consumption simpler and faster for all stakeholders. Often, stakeholders face challenges to limit data that they need for analysis and making effective decisions. As more data from ever-growing data sources is coming in, users are flooded with a variety of data. Data models alleviated the pain to serve insights to enterprise users. Data models provided insights after data cleansing, aggregating, and applying business rules. As data models in big data grow, queries and analysis require processing the large volume of data and big joins. It leads to long response and processing times. Data modelling in big data platforms needs attention to effectively cleanse, organise, and store big data to ensure timely availability of enterprise insights. As the scale is a critical aspect of the big data platform, big data should be modelled in a way that accessibility and delivery of insights should not be affected when the scale goes up. This paper presents best practices to model structured and semi-structured data in the big data platform.
Stock index prediction has been a challenging problem due to difficult to model complexities of the stock market. More recently deep learning approaches have become an important method in modelling complex relationships in time-series data. In this paper: we propose novel deep learning models that combine multiple pipelines of convolutional neural network and uni-directional or bi-directional gated recurrent units. Proposed models improve prediction performance and execution time upon previously published models on large scale S&P 500 dataset. We present several variations of multiple and single pipeline deep learning models based on different CNN kernel sizes and number of GRU units.
The value of conversation intelligence in deepening the insights of authentic conversations is a common ground nowadays between researchers and the business community. The rapid development of big data algorithms and technology enables massive amounts of data and meta-data processing, including content, vocal features and body gestures. This study is based on 358 business-to-business (B2B) sales calls at the discovery stage. We propose a model to capture the dynamics of acoustic gaps between the sales representatives and customers by relying solely on the acoustic signal. We extract basic features from the acoustic signal: speech proportion, fundamental frequency (F0), intensity, harmonics-to-noise ratio (HNR), jitter and shimmer. We focus on the differences between the four speakers' role-gender groups (e.g., female-representative with female-customer). We found significant differences in the behavioural patterns of the dynamics between these four groups. The study demonstrates that using delta metrics to assess the interactions leads to new insights.
Nature of data stream is determined after complete scanning of whole data sets during real-time data processing. However, it becomes inconvenient to process entire data stream at once in real-time data stream processing. Thus, a sheer sized fixed window of data streams is processed at a particular time. The intensification of sheer sized fixed window at processing node is mitigated by reducing the flowing rate of data stream. Heuristic clustering windowing (HCW) approach and partial blind window (PBW) algorithms are proposed for reducing the flow of data stream with least sampling error. These approaches consist of the combination of systematic sampling and clustering mechanism. A clustering approach is applied on one fraction of data streams whereas systematic sampling handles other portion of streams. These approaches are helpful in reducing flow of data streams in minimum latency.
This paper presents various transaction sampling algorithms for the proposed real-time crypto computing, and analytical model to assure their dependability under stringent real-time requirement. Efficacy of the algorithms is assessed in terms of the block dependability that expresses the probability for the pending transactions to be posted within the current or the target block delay. Algorithms on prioritising and sampling transactions from pool, to facilitate execution of those transactions within their deadline requirements, such as normal, random, sorted, and stratified, are proposed and simulated. Performance variables such as the number of pending transactions, average speed, gas fees, deadlines, number of miners, are identified and taken into the block dependability in order to reveal the influence of those variables. Extensive parametric simulation results are presented and discussed in the cases of the random and sorted transaction sampling algorithms along with a prototype built based on the Ethereum open source.
The cloud logging service is a core component for the operation and management of the production system. The service is usually a central server deployment whereby the dedicated central servers accept all log messages from leaf computing nodes. As the number of applications and solutions on the cloud changes dynamically, the amount of log messages that are forwarded to the logging service is also changed. The paper proposes a distributed logging service (DLS) that distributes log messages to multiple leaf computing nodes. No central server is required to manage the logging service. DLS also provides alert notifications, authentication, lifetime management and resilience, which are required for the production system. Results of an evaluation of the emulated environment show that DLS is suitable for use with applications and solutions which are used for production usages.