Monitoring of water quality should not be solely based on laboratory samples.Such activity, although producing reliable results, cannot provide an accurate enough temporal coverage for water quality monitoring.The Finnish Environment Institute, SYKE, has therefore established numerous online water monitoring stations that continuously monitor water quality.The problem with the automatic monitoring, however, is that the recorded values are not reliable as such and need to be subject to quality control and uncertainty estimation.Here, as the main contribution, we present a computational service that we have implemented to automate and integrate the water quality monitoring process.We also present a case study regarding the river Väänteenjoki and discuss the obtained uncertainty results and their implication.
ions from Sensor Data with Complex Event Processing and Machine Learning Markus Stocker University of Eastern Finland, markus.stocker@uef.fi Mauno Rönkkö University of Eastern Finland, mauno.ronkko@uef.fi Mikko Kolehmainen University of Eastern Finland, mikko.kolehmainen@uef.fi Follow this and additional works at: https://scholarsarchive.byu.edu/iemssconference Part of the Civil Engineering Commons, Data Storage Systems Commons, Environmental Engineering Commons, Hydraulic Engineering Commons, and the Other Civil and Environmental Engineering Commons Stocker, Markus; Rönkkö, Mauno; and Kolehmainen, Mikko, "Abstractions from Sensor Data with Complex Event Processing and Machine Learning" (2014). International Congress on Environmental Modelling and Software. 8. https://scholarsarchive.byu.edu/iemssconference/2014/Stream-F/8 This Event is brought to you for free and open access by the Civil and Environmental Engineering at BYU ScholarsArchive. It has been accepted for inclusion in International Congress on Environmental Modelling and Software by an authorized administrator of BYU ScholarsArchive. For more information, please contact scholarsarchive@byu.edu, ellen_amatangelo@byu.edu. International Environmental Modelling and Software Society (iEMSs) 7th Intl. Congress on Env. Modelling and Software, San Diego, CA, USA, Daniel P. Ames, Nigel W.T. Quinn and Andrea E. Rizzoli (Eds.) http://www.iemss.org/society/index.php/iemss-2014-proceedings Abstractions from Sensor Data with Complex Event Processing and Machine Learningions from Sensor Data with Complex Event Processing and Machine Learning Markus Stocker, Mauno Rönkkö, Mikko Kolehmainen University of Eastern Finland, P.O. Box 1627, 70211 Kuopio, Finland markus.stocker@uef.fi, mauno.ronkko@uef.fi, mikko.kolehmainen@uef.fi Abstract: Environmental knowledge systems that build on sensor-based environmental monitoring rely on techniques in knowledge acquisition and representation to interpret the numbers obtained in measurement for what they tell about the monitored environment. Languages and systems in knowledge representation and reasoning, specifically Semantic Web technologies, support the formulation and execution of rules, a technique that enables deductive inference in a knowledge base. This technique has been used to demonstrate inference on sensor data. While the approach certainly has its merits, it is often demonstrated for numerical thresholds and, thus, for relatively trivial “semantic enrichment.” In reality, knowledge acquisition tasks of interest to environmental knowledge systems that build on sensor-based environmental monitoring are often more challenging. They rely on advanced computational techniques and models, e.g. in machine learning or complex event processing. In order to ease the formulation and execution of such tasks, systems need to integrate such techniques. Towards this end, we present the integration of machine learning with WEKA and complex event processing with Esper in Wavellite. Environmental knowledge systems that build on sensor-based environmental monitoring rely on techniques in knowledge acquisition and representation to interpret the numbers obtained in measurement for what they tell about the monitored environment. Languages and systems in knowledge representation and reasoning, specifically Semantic Web technologies, support the formulation and execution of rules, a technique that enables deductive inference in a knowledge base. This technique has been used to demonstrate inference on sensor data. While the approach certainly has its merits, it is often demonstrated for numerical thresholds and, thus, for relatively trivial “semantic enrichment.” In reality, knowledge acquisition tasks of interest to environmental knowledge systems that build on sensor-based environmental monitoring are often more challenging. They rely on advanced computational techniques and models, e.g. in machine learning or complex event processing. In order to ease the formulation and execution of such tasks, systems need to integrate such techniques. Towards this end, we present the integration of machine learning with WEKA and complex event processing with Esper in Wavellite.
BACKGROUND:The redundancy of information is becoming a critical issue for epidemiologists. High-dimensional datasets require new effective variable selection methods to be developed. This study implements an advanced evolutionary variable selection method which is applied for cardiovascular predictive modeling. The epidemiological follow-up study KIHD (Kuopio Ischemic Heart Disease Risk Factor Study) was used to compare the designed variable selection method based on an evolutionary search with conventional stepwise selection. The sample contains in total 433 predictor variables and a response variable indicating incidents of cardiovascular diseases for 1465 study subjects.RESULTS:The effectiveness of variable selection methods was investigated in combination with two models: Generalized Linear Logistic Regression and Support Vector Machine. We managed to decrease the number of variables from 433 to 38 and save the predictive ability of the models used. Their performance was evaluated with an F-score metric. At most, we gained 65.6% and 67.4% of the F-score before and after variable selection respectively. All the results were averaged over 5-folds of a cross-validation procedure.CONCLUSIONS:The presented evolutionary variable selection method allows a reduced set of variables to be chosen which are relevant to predicting cardiovascular diseases. A reference list of the most meaningful variables is introduced to be used as a basis for new epidemiological studies. In general, the multicollinearity of variables enables different combinations of predictors to be used and the same performance of models to be attained.
Over the past decades, sensor networks have been deployed around the world to monitor over time and space a large number of properties appertaining to various environmental phenomena. A popular example is the monitoring of particulate matter and gases in ambient air undertaken, for instance, to assess air quality and inform decision makers and the public. Such infrastructure can generate large amounts of data, which must be processed to derive useful information. Infrastructure may be for environmental research, specifically. In order to reduce duplication and improve interoperability, efforts have been initiated more recently that aim at abstract architectural descriptions of infrastructure that supports the acquisition, curation, access, and processing of measurement and observation data. The ENVRI Reference Model (ENVRI-RM) is an example for an abstract architectural description of infrastructure tailored for environmental research. We briefly summarize ENVRI-RM and provide an overview of its subsystems, functionality, and viewpoints. We highlight that its primary focus is on the data life-cycle in environmental research infrastructure. As our contribution, weextend ENVRI-RM with functionality for the acquisition of knowledge from data, and the curation, access, and processing of knowledge. The extension, which we name +K, aims at addressing the knowledge life-cycle in environmental research infrastructure. We present the +K subsystems and functionality, and discuss the extension from ENVRI-RM viewpoints. We argue that the +K extension can be superimposed on ENVRI-RM to form the ENVRI-RM+K model for the ‘archetypical’ knowledge-based environmental research infrastructure that addresses both data and knowledge life-cycles. We demonstrate the application of the extension to a concrete use case in aerosol science.
We present a software system for automated projection of situational knowledge for disease outbreak in agriculture. The system supports farmers and agricultural advisers in obtaining and maintaining awareness of present and future disease outbreaks in crops grown at agricultural parcels. It models objects such as plant pathogens and agricultural parcel crops, and their relations, as entities in situations observed by an environmental monitoring system. It utilizes a mechanistic disease pressure model to obtain knowledge about observed situations from forecast data for various weather parameters. It represents obtained situational knowledge explicitly and manages represented knowledge in a knowledge base. We evaluate the system for 3 fungal plant pathogens, 2 cereal crops, and 17 agricultural parcels located in Finland, for a growing season. We underscore how the explicit representation of situational knowledge is useful toward various purposes, including reasoning, query and visualization, and is, thus, vastly superior to having situational knowledge only implicit in high-level data products such as maps.
Road vehicle detection and, to a lesser extent, classification have received considerable attention, in particular for the purpose of traffic monitoring by transportation authorities. A multitude of sensors and systems have been developed to assist people in traffic monitoring. Camera-based systems have enjoyed wide adoption over the last decade, partially substituting for more traditional techniques. Methods based on road-pavement vibration are not as common as camera-based systems. However, vibration sensors may be of interest when sensors must be out of sight and insensitive to environmental conditions, such as fog. We present and discuss our work on detection and classification of vehicles by measurement of road-pavement vibration and by means of supervised machine learning. We describe the entire processing chain from sensor data acquisition to vehicle classification and discuss our results for the task of vehicle detection and the task of vehicle classification separately. Using data for a single vibration sensor, our results show a performance ranging between 94% and near 100% for the detection task (1340 samples) and between 43% and 86% for the classification task (experiment specific, between 454 and 1243 samples).
The design of ontologies for sensor data and metadata has received considerable attention. The most prominent is arguably the Semantic Sensor Network (SSN) ontology. For persistence and retrieval of sensor observations, systems that adopt the SSN ontology most obviously build on an RDF database (triple store). However, large volumes of collected sensor data can be challenging for RDF databases, as the evaluation of SPARQL queries for SSN observations quickly becomes prohibitively expensive. This is arguably due to the fact that triple stores are optimized to efficiently evaluate graph pattern queries, not time series interval queries. As our main contribution, we present Emrooz, a scalable database capable of consuming SSN observations represented in RDF and evaluating queries for SSN observations formulated in SPARQL. We present the Emrooz implementation on Apache Cassandra and Sesame and its performance compared to two state-of-the-art RDF databases. The results show that Emrooz query performance outperforms the two RDF databases by orders of magnitude with increasingly large datasets. We motivate the need for scalable databases for SSN observations on a case study in micrometeorology.
We present an environmental software system that obtains, integrates, and reasons over situational knowledge about natural phenomena and human activity. We focus on storms and driver directions. Radar data for rainfall intensity and Google Directions are used to extract situational knowledge about storms and driver locations along directions, respectively. Situational knowledge about the environment and about human activity is integrated in order to infer situations in which drivers are potentially at higher risk. Awareness of such situations is of obvious interest. We present a prototype user interface that supports adding scheduled driver directions and the visualization of situations in space-time, in particular also those in which drivers are potentially at higher risk. We think that the system supports the claim that the concept of situation is useful for the modelling of information about the environment, including human activity, obtained in environmental monitoring systems. Furthermore, the presented work shows that situational knowledge, represented by heterogeneous systems that share the concept of situation, is relatively straightforward to integrate.
The emergence of Web 2.0 technologies has changed dramatically not only the way users perceive the Internet and interact on it but also the way they influence a community and act in real life aspects. With the rapid rise in use and popularity of social media, people tend to share opinions and observations for almost any subject or event in their everyday life. Consequently, microblogging websites have become a rich data source for user-generated information. The leading opportunity is to take advantage of the wisdom of the crowd and to benefit from collective intelligence in any applicable domain. Towards this direction, we focus on the problem of mining and extracting knowledge from unstructured textual content, for the atmospheric environment domain and its effect to quality of life. As the main contribution, we propose a combined methodology of unsupervised learning methods for analyzing posts from Twitter and clustering textual data into concepts with semantically similar context. By applying Self-Organizing Maps and k-means clustering, we identify possible inter-relationships and patterns of words used in tweets that can form upper concepts of atmospheric and health related topics of discussion. We achieve to group together tweets, from more generic to more specific description levels of their content, according to the selected number of clusters. Strong clusters with significant semantic relatedness among their content are revealed, and hidden relations between concepts and their related semantics are acquired. The results highlight the potential use of social media text streams as a highly-valued supplement source of environmental information and situation awareness.
In this article we discuss automated preprocessing of environmental data for further use. Environmental data is by default heterogeneous, as it may consist of data from sources such as weather stations, weather radars, chemical sensors, acoustic sensors, and off-line laboratory analysis. When integrating data from such heterogeneous sources, it needs to be processed in a context dependent manner. In addition, there is no single generic processing method; rather, several atomic methods need to be applied and in an appropriate sequence. Furthermore, the problem is complicated by the requirements set by the intended use of the data. The requirements influence not only the set of applicable methods but also the application sequence. In this article, we study automation of the selection and sequencing of preprocessing methods based on the user requirements. As the main contribution, we propose here the use of characterizations and a reachability algorithm to solve the selection and sequencing problem. In this article, we present the algorithm and argue for its correctness. We also discuss, how the algorithm is implemented as a cloud service, and illustrate the use of the service with simple case studies.
We discuss quality control of environmental measurement data. Typically, environmental data is used to compute some specific indicators based on models, historical data, and the most recent measurement data. For such a computation to produce reliable results, the data must be of sufficient quality. The reality is, however, that environmental measurement data has a huge variation in quality. Therefore, we study the use of quality flagging as a means to perform both real-time and off-line quality control of environmental measurement data. We propose the adoption of the quality flagging scheme introduced by the Nordic meteorological institutes. As the main contribution, we present both a uniform interpretation for the quality flag values and a scalable Enterprise Service Bus based architecture for implementing the quality flagging. We exemplify the use of the quality flagging and the architecture with a case study for monitoring of built environment.
As environmental monitoring systems increasingly automate the collection and processing of environmental sensor network data, the technical components of such systems can automatically obtain and maintain higher levels of situation awareness—awareness of the monitored part of reality. In order to increase confidence in the correctness of situation awareness maintained by such systems it is important to explicitly model provenance. We present an alignment of the PROV ontology with ontologies used in a software framework for situation awareness in environmental monitoring, called Wavellite. The extended vocabulary enables the explicit representation of provenance in Wavellite applications. We demonstrate the implementation for a concrete scenario.
Smart electrical grids refer to networked systems for distributing and transporting electricity from producers to consumers, by dynamically configuring the network through remotely controlled (dis)connectors. The consumers of the grid have typically distinct priorities, e.g., a hospital and an airport have the highest priority and the street lighting has a lower priority. This means that when electricity supply is compromised, e.g., during a storm, then the highest priority consumers should either not be affected or should be the first for whom electricity provision is recovered. In this paper, we propose a general formal model to study the provability of such a property. We have chosen Event-B as our formal framework due to its abstraction and refinement capabilities that support correct-by-construction stepwise development of models; also, Event-B is tool supported. Being able to prove various properties for such critical systems is fundamental nowadays, as our society is increasingly powered by dynamic digital solutions to traditional problems.
A recurrent problem in applications that build on environmental sensor networks is that of sensor data organization and interpretation. Organization focuses on, for instance, resolving the syntactic and semantic heterogeneity of sensor data. The distinguishing factor between organization and interpretation is the abstraction from sensor data with information acquired from sensor data. Such information may be situational knowledge for environmental phenomena. We discuss a generic software framework for the organization and interpretation of sensor data and demonstrate its application to data of a large scale sensor network for the monitoring of atmospheric phenomena. The results show that software support for the organization and interpretation of sensor data is valuable to scientists in scientific computing workflows. Explicitly represented situational knowledge is also useful to client software systems as it can be queried, integrated, reasoned, visualized, or annotated.
We regard the basic unit of the organism, the cell, as a complex dissipative natural process functioning under the second law of thermodynamics and the principle of least action. Organisms are conglomerates of information bearing cells that optimise the efficiency of energy (nutrient) extraction from its ecosystem. Dissipative processes, such as peptide folding and protein interaction, yield phenotypic information from which form and function emerge from cell to cell interactions within the organism. Organisms, in Darwin's proportional numbers', in turn interact to minimise the free energy of their ecosystems. Genetic variation plays no role in this holistic conceptualisation of the life process.
Information systems that build on sensor networks often process data produced by measuring physical properties. These data can serve in the acquisition of knowledge for real-world situations that are of interest to information services and, ultimately, to people. Such systems face a common challenge, namely the considerable gap between the data produced by measurement and the abstract terminology used to describe real-world situations. We present and discuss the architecture of a software system that utilizes sensor data, digital signal processing, machine learning, and knowledge representation and reasoning to acquire, represent, and infer knowledge about real-world situations observable by a sensor network. We demonstrate the application of the system to vehicle detection and classification by measurement of road pavement vibration. Thus, real-world situations involve vehicles and information for their type, speed, and driving direction.
ions from Sensor Data with Complex Event Processing and Machine Learning Markus Stocker, Mauno Rönkkö, Mikko Kolehmainen University of Eastern Finland, P.O. Box 1627, 70211 Kuopio, Finland markus.stocker@uef.fi, mauno.ronkko@uef.fi, mikko.kolehmainen@uef.fi Abstract: Environmental knowledge systems that build on sensor-based environmental monitoring rely on techniques in knowledge acquisition and representation to interpret the numbers obtained in measurement for what they tell about the monitored environment. Languages and systems in knowledge representation and reasoning, specifically Semantic Web technologies, support the formulation and execution of rules, a technique that enables deductive inference in a knowledge base. This technique has been used to demonstrate inference on sensor data. While the approach certainly has its merits, it is often demonstrated for numerical thresholds and, thus, for relatively trivial “semantic enrichment.” In reality, knowledge acquisition tasks of interest to environmental knowledge systems that build on sensor-based environmental monitoring are often more challenging. They rely on advanced computational techniques and models, e.g. in machine learning or complex event processing. In order to ease the formulation and execution of such tasks, systems need to integrate such techniques. Towards this end, we present the integration of machine learning with WEKA and complex event processing with Esper in Wavellite. Environmental knowledge systems that build on sensor-based environmental monitoring rely on techniques in knowledge acquisition and representation to interpret the numbers obtained in measurement for what they tell about the monitored environment. Languages and systems in knowledge representation and reasoning, specifically Semantic Web technologies, support the formulation and execution of rules, a technique that enables deductive inference in a knowledge base. This technique has been used to demonstrate inference on sensor data. While the approach certainly has its merits, it is often demonstrated for numerical thresholds and, thus, for relatively trivial “semantic enrichment.” In reality, knowledge acquisition tasks of interest to environmental knowledge systems that build on sensor-based environmental monitoring are often more challenging. They rely on advanced computational techniques and models, e.g. in machine learning or complex event processing. In order to ease the formulation and execution of such tasks, systems need to integrate such techniques. Towards this end, we present the integration of machine learning with WEKA and complex event processing with Esper in Wavellite.
In this article we discuss how to improve the resilience of an existing control system. In recent years, our environment has become populated with numerous control systems due the to availability of low-cost technologies. For instance, modern home automation has become a cooperative network of multiple control systems, many of which communicate over the Internet. Many of these systems hardly address resilience, and improving them is hard, as many of them are provided as “black boxes”. Consequently, as the main contribution, we propose a method for introducing resilience to an existing control system. The method is based on designing and adding resilience mediators that act in between the components to correct faulty communication and to mediate state awareness. In the method, behavioral analysis and HAZOP tables are used as tools to identify and design the relevant resilience mediators. We illustrate the use of the proposed method with an ongoing, real-life case study involving the control of a residential building that adapts to occupant’s behavior.
Anders P Ravn合作论文数Department of Computer Science;Aalborg University2