We demonstrate our topic extraction method in which topics are treated as clusters of word embeddings. The OPTICS algorithm is used to find small and arbitrarily-shaped clusters of embeddings, produced by a fastText model. The result is a set of dominant and non-dominant domain-specific topics. The focus of the method is on short online posts which are difficult to analyze with traditional topic extraction approaches because of the word collocation scarcity. The method is tested on dataset of posts from Twitter, LinkedIn and company blogs related to industrial automation. The method significantly outperforms traditional topic extraction approaches by finding relevant and understandable topics with related tokens.
Context: Big Data systems are a class of software systems that ingest, store, process and serve massive amounts of heterogeneous data, from multiple sources. Despite their undisputed impact in current society, their engineering is still in its infancy and companies find it difficult to adopt them due to their inherent complexity. Existing attempts to provide architectural guidelines for their engineering fail to take into account important Big Data characteristics, such as the management, evolution and quality of the data.Objective: In this paper, we follow software engineering principles to refine the lambda-architecture, a reference model for Big Data systems, and use it as seed to create Bolster, a software reference architecture (SRA) for semantic-aware Big Data systems.Method: By including a new layer into the lambda-architecture, the Semantic Layer, Bolster is capable of handling the most representative Big Data characteristics (i.e., Volume, Velocity, Variety, Variability and Veracity).Results: We present the successful implementation of Bolster in three industrial projects, involving five organizations. The validation results show high level of agreement among practitioners from all organizations with respect to standard quality factors.Conclusion: As an SRA, Bolster allows organizations to design concrete architectures tailored to their specific needs. A distinguishing feature is that it provides semantic-awareness in Big Data Systems. These are Big Data system implementations that have components to simplify data definition and exploitation. In particular, they leverage metadata (i.e., data describing data) to enable (partial) automation of data exploitation and to aid the user in their decision making processes. This simplification supports the differentiation of responsibilities into cohesive roles enhancing data governance. (C) 2017 Elsevier B.V. All rights reserved.
Gamification has been applied in software engineering contexts, and more recently in requirements engineering with the purpose of improving the motivation and engagement of people performing specific engineering tasks. But often an objective evaluation that the resulting gamified tasks successfully meet the intended goal is missing. On the other hand, current practices in designing gamified processes seem to rest on a try, test and learn approach, rather than on first principles design methods. Thus empirical evaluation should play an even more important role.We combined gamification and automated reasoning techniques to support collaborative requirements prioritization in software evolution. A first prototype has been evaluated in the context of three industrial use cases. To further investigate the impact of specific game elements, namely point-based elements, we performed a quasi-experiment comparing two versions of the tool, with and without pointsification. We present the results from these two empirical evaluations, and discuss lessons learned.
Continuous software engineering is a new trend that is gaining increasing attention of the research community in the last years. The main idea behind this trend is to tighten the connection between the software engineering lifecycle activities (e.g., development, planning, integration, testing, etc.). While the connection between development and integration (i.e., continuous integration) has been subject of research and is applied in industrial settings, the connection between other activities is still in a very early stage. We are contributing to this research topic by proposing our ideas towards connecting the software development and software release planning activities (i.e., continuous software release planning). In this paper we present our initial findings on this topic, how we envision to address the continuous software release planning, and a research agenda to fulfil our objectives.
Software release planning is the activity of deciding what is to be implemented, when and by who. It can be divided into two tasks: strategic planning (i.e., the what) and operational (i.e., the when and the who). Replan, the tool that we present in this demo, handles both tasks in an integrated and flexible way, allowing its users (typically software product managers and developer team leaders) to (re)plan the releases dynamically by assigning new features and/or modifying the available resources allocated at each release. A recorded video demo of Replan is available at https://youtu.be/PNK5EUTdqEg.
The new data collection opportunities offered by IoT and the recent advancements in data processing infrastructure have provided a significant impulse to the development of smart cities. Although reliable data collection and analysis is a fundamental prerequisite for urban intelligence, the smartness of a smart city also depends on the ability to translate IoT data into useful and innovative services: it is important to conceive new design principles for the development of applications for smart cities and IoT. To this extent, this paper presents a simplified conceptual framework for describing smart city ecosystems based on concepts of natural ecosystems. The goal of the framework is to derive new metaphors for designing IoT applications based on the observation of communication and interaction patterns in natural ecosystems. The proposed framework is used to describe a real smart city project in Vienna, Austria.
The abundance of data in the context of smart cities yields huge potential for data-driven businesses but raises unprecedented challenges on data privacy and security. Some of these challenges can be addressed merely through appropriate technical measures, while other issues can only be solved through strategic organizational decisions. In this paper, we present few cases from a real smart city project. We outline some exemplary data analytics scenarios and describe the measures that we adopt for a secure handling of data. Finally, we show how the chosen solutions impact the awareness of the public and acceptability of the project.
Mobile cellular networks can serve as ubiquitous sensors for physical mobility. We propose a method to infer vehicle travel times on highways and to detect road congestion in real-time, based solely on anonymized signaling data collected from a mobile cellular network. Most previous studies have considered data generated from mobile devices active in calls, namely Call Detail Records (CDR), an approach that limits the number of observable devices to a small fraction of the whole population. Our approach overcomes this drawback by exploiting the whole set of signaling events generated by both idle and active devices. While idle devices contribute with a large volume of spatially coarse-grained mobility data, active devices provide finer-grained spatial accuracy for a limited subset of devices. The combined use of data from idle and active devices improves congestion detection performance in terms of coverage, accuracy, and timeliness. We apply our method to real mobile signaling data obtained from an operational network during a one-month period on a sample highway segment in the proximity of a European city, and present an extensive validation study based on ground-truth obtained from a rich set of reference datasources - road sensor data, toll data, taxi floating car data, and radio broadcast messages.
In dieser Dissertation wird die Verwendung von Mobilfunknetzen als universeller Sensor fur menschliche Mobilitat vorgeschlagen. Der Grundgedanke besteht darin, die Erhebung, Verarbeitung und Analyse von Signalisierungsdaten aus einem Mobilfunknetz durchzufuhren, um Mobilitatsmuster von mobilen Endgeraten zu beobachten. Diese Muster konnen wiederum verwendet werden, um Informationen uber Fahrzeugstrome und Strasenverkehr zu extrahieren. Der entscheidende Beitrag dieser Arbeit ist die Konzeption und Realisierung eines Frameworks, das solche Funktionalitaten bietet und alle Stufen der Verarbeitungskette abdeckt, ausgehend vom passiven Monitoring der Signalisierung im Mobilfunknetz bis zur Extraktion von Verkehrsinformationen und ihrer Auslieferung an die VerkehrsteilnehmerInnen. Dazu werden umfangreiche Validierungsstudien entlang der landlichen und stadtischen Autobahnen vorgestellt. Die entwickelten Algorithmen und Techniken werden auf Daten eines grosen nationalen Mobilfunkanbieters angewendet und deren Ergebnisse zudem mit mehreren bereits vorhandenen Quellen von Verkehrsdaten (z.B. Verkehrsdetektoren des Autobahnbetreibers) verglichen. Diese Arbeit zeigt, dass Mobilfunknetze ein integraler Bestandteil von Intelligenten Verkehrssystemen (IVS) sein konnen, in dem diese einerseits Mobilitats- und Strasenverkehrsdaten zur Verfugung stellen und andererseits als effizientes Transportmedium fur die Kommunikation zwischen Infrastruktur und Fahrzeugen fungieren.
Road traffic can be monitored by means of static sensors and derived from floating car data, i.e., reports from a sub-set of vehicles. These approaches suffer from a number of technical and economical limitations. Alternatively, we propose to leverage the mobile cellular network as a ubiquitous mobility sensor. We show how vehicle travel times and road congestion can be inferred from anonymized signaling data collected from a cellular mobile network. While other previous studies have considered data only from active devices, e.g., engaged in voice calls, our approach exploits also data from idle users resulting in an enormous gain in coverage and estimation accuracy. By validating our approach against four different traffic monitoring datasets collected on a sample highway over one month, we show that our method can detect congestions very accurately and in a timely manner.
This paper describes various techniques that are developed for traffic analysis purposes based on cellular network signaling events. Starting from the mobility reported by the mobile devices of a cellular network, traffic volumes and origin-destination matrices are calculated, zones of interests are analyzed and users on different modes of transport are identified for rural and urban areas and different times of the day (e.g. peak hours and night traffic). In addition a tool is described that can process that cellular data for import into expert traffic modeling software which assists during traffic planning and civil engineering exercises. The new extraction methods increase the accessibility of traffic data and reduce the costs of data collection in comparison to conventional collection and analysis methods. Furthermore the results are checked against reference traffic counts and additional traffic engineering statistics for validation and quality assurance purposes.
The signaling traffic of a cellular network is rich of information related to the movement of devices across cell boundaries. Thus, passive monitoring of anonymized signaling traffic enables the observation of the devices' mobility patterns. This approach is intrinsically more powerful and accurate than previous studies based exclusively on Call Data Records as significantly more devices can be included for investigation, but it is also more challenging to implement due to a number of artifacts implicitly present in the network signaling. In this study we tackle the problem of estimating vehicular trajectories from 3G signaling traffic with particular focus on crucial elements of the data processing chain. The work is based on a sample set of anonymous traces from a large operational 3G network, including both the circuit-switched and packet-switched domains. We first investigate algorithms and procedures for preprocessing the raw dataset to make it suitable for mobility studies. Second, we present a preliminary analysis and characterization of the mobility signaling traffic. Finally, we present an algorithm for exploiting the refined data for road traffic monitoring, i.e., route detection. The work shows the potential of leveraging the 3G cellular network as a complementary "sensor" to existing solutions for road traffic monitoring.
This paper demonstrates the potential of using anonymized cellular network signaling data to extract and analyse macro-mobility patterns. The authors show that, by properly processing signaling data passively collected from a cellular core network, the approach is able to catch crucial characteristics of wide-area mobility patterns. By means of simple illustrative examples, the authors present an analysis of country-wide travel relations, the dynamics of travel times between cities, and commute behaviour.
Future intelligent transportation systems (ITS) will necessitate wireless vehicle-to-infrastructure (V2I) communications. This wireless link can be implemented by several technologies, such as digital broadcasting, cellular communication, or dedicated short-range communication (DSRC) systems. Analyses of the coverage and capacity requirements are presented when each of the three systems is used to implement the V2I link. We show that digital broadcasting systems are inherently capacity limited and do not appropriately scale. Furthermore, we show that the Universal Mobile Telecommunications System (UMTS) can implement the V2I link using either a dedicated channel (DCH) or a multimedia broadcast/multicast service (MBMS), as well as a hybrid approach. In every case, such V2I systems scale well and are capacity limited. We also show that wireless access in vehicular environment (WAVE) systems scale well, provide ample capacity, and are coverage limited. Finally, a direct quantitative comparison of the presented systems is given to show their scaling behavior with the number of users and the geographical coverage.
In this work we present an implementation of a fully functional IEEE 802.11p transmitter in software-defined radio. We describe the rapid-prototyping methodology that was used to implement the frame-encoder within the open-source GNU Software Radio (GNURadio) platform [1]. The encoder generates OFDM frames in digital complex base-band representation and uses the USRP2 [2] as digital-to-analog front-end for up-conversion and final transmission. Since the actual encoding process involves a large number of complex steps we split the development approach into three sequential stages. First, a reference-encoder in a high-level language (MATLAB) is derived from the IEEE standard documents. Second, the individual blocks of the MATLAB encoding chain are progressively ported to GNURadio, cross-checking with the reference after each step. Finally, standard compliance is verified by conducting comparative over-the-air measurements with an early prototype of a commercial 11p transceiver. Initial measurement results indicate that the fidelity of the resulting GNURadio implementation is on par with non-software-defined radio industry solutions and capable of generating truly standard-compliant OFDM frames. The encoder presented here has been released under GPLv3 and is also capable of encoding frames according to the 11a and 11g amendments, thus making it a valuable building block for upcoming software-defined radio projects.
In this paper we present a road traffic estimation system built on top of the cellular network infrastructure. Based on the concept that many road users are also customers of a cellular operator, we show that specific road conditions map to certain signaling patterns in the cellular core network. In order to estimate the road traffic, signaling is collected from the core network of an operational mobile network. We explore the feasibility of using mobility-related signaling for pantomiming local-loop sensors, i.e. counting the number of vehicles crossing a specific road section. In this work we present the main system component and discuss a number of practical issues to be considered in the deployment of such system. Based on the explorative analysis of real signaling data, we show how normal and abnormal road conditions (e.g. accidents) map into mobility signaling in a real cellular network.
In this contribution we address the problem of using cellular network signaling for inferring real-time road traffic information. We survey and categorize the approaches that have been proposed in the literature for a cellular-based road monitoring system and identify advantages and limitations. We outline a unified framework that encompasses UMTS and GPRS data collection in addition to GSM, and prospectively combines passive and active monitoring techniques. We identify the main research challenges that must be faced in designing and implementing such an intelligent road traffic estimation system via third-generation cellular networks.