The aim of this study is to collect information about events in the city of Oslo, Norway, that produce a seismic signature. In particular, we focus on blasts from the ongoing construction of tunnels and under-ground water storage facilities under populated areas in Oslo. We use seismic data recorded simultaneously on up to 11 Raspberry Shake sensors deployed between 2021 and 2023 to quickly detect, locate, and classify urban seismic events. We present a deep learning approach to first identify rare events and then to build an automatic classifier from those templates. For the first step, we employ an outlier detection method using auto-encoders trained on continuous background noise. We detect events using an STA/LTA trigger and apply the auto-encoder to those. Badly reconstructed signals are identified as outliers and subsequently located using their surface wave (Rg) signatures on the seismic network. In a second step, we train a supervised classifier using a Convolutional Neural Network to detect events similar to the identified blast signals. Our results show that up to 87% of about 1,900 confirmed blasts are detected and locatable in challenging background noise conditions. We demonstrate that a city can be monitored automatically and continuously for explosion events, which allows implementing an alert system for future smart city solutions.
Seismic phase detection and classification using deep learning is so far poorly investigated for regional events since most studies focus on local events and short time windows as the input to the detection models. To evaluate deep learning on regional seismic records, we create a data set of events in Northern Europe and the European Arctic. This data set consists of about 151 000 three component event waveforms and corresponding phase arrival picks at stations in mainland Norway, Finland and Svalbard. We train several state-of-the-art and one newly developed deep learning model on this data set to pick P- and S-wave arrivals. The new method modifies the popular PhaseNet model with new convolutional blocks including transformers. This yields more accurate predictions on the long input time windows associated with regional events. Evaluated on event records not used for training, our new method improves the performance of the current state-of-the-art methods when it comes to recall, precision and pick time residuals. Finally, we test our new model for continuous mode processing on 4 d of single-station data from the ARCES array. Results show that our new method outperforms the existing array detector at ARCES. This opens up new opportunities to improve automatic array processing with deep learning detectors.
Various knowledge bases (KBs) have been constructed via information extraction from encyclopedias, text and tables, as well as alignment of multiple sources. Their usefulness and usability is often limited by quality issues. One common issue is the presence of erroneous assertions and alignments, often caused by lexical or semantic confusion. We study the problem of correcting such assertions and alignments, and present a general correction framework which combines lexical matching, context-aware sub-KB extraction, semantic embedding, soft constraint mining and semantic consistency checking. The framework is evaluated with one set of literal assertions from DBpedia, one set of entity assertions from an enterprise medical KB, and one set of mapping assertions from a music KB constructed by integrating Wikidata, Discogs and MusicBrainz. It has achieved promising results, with a correction rate (i.e., the ratio of the target assertions/alignments that are corrected with right substitutes) of 70.1 %, 60.9 % and 71.8 %, respectively.
Automatic detection of seismic events in processing pipelines at the IDC and many NDCs is mostly done using beamforming on arrays; however, extensive use of single stations can improve the detection capability and accuracy of event location. Advances in deep learning methods enable faster and more accurate processing of large quantities of single station data not seen previously. We use event catalogues including phase picks on a range of arrays in Scandinavia at regional distances (200-2000km), i.e., up to 3min separation between P and S arrivals, to train several deep learning models (PhaseNet and EQTransformer variants) using single stations within the arrays. The models are trained on clips of 324s to capture the multiple arrivals. The models are then applied to various single stations in Norway to assess their generalization. We can detect events at a variety of back-azimuths and distances. Furthermore, we expand the existing deep learning models to provide predictions for back-azimuth and distance. This imposes physical restrictions on the models, leading to increased picking accuracy for the predicted phase arrivals. Moreover, this enables us to use the vast number of single stations available to efficiently detect and locate distant events.
<p>The real-time seismo-acoustic monitoring of military conflicts can provide a unique alternative to conventional ground reports and sparse satellite coverage. The pressure waves generated by an explosion travel through the atmosphere and subsurface as sound and seismic waves, and their signature can be recorded by arrays of seismometers for ground motion or microbarometers for sound propagation. However, standard monitoring techniques can be both computationally expensive when localizing signals over large regions and/or prone to false detections when signals have low amplitudes. In this contribution we propose a Machine-Learning (ML) based solution to detect seismic and infrasound arrivals and locate sources close to real time. To validate our model we leverage the seismic data collected during the Russia-Ukraine conflict started in February 2022 using the Ukrainian primary station of the International Monitoring System (IMS), the Malin array (AKSAG). We test both the accuracy and computational efficiency of our approach against a threshold-based migration stacking model developed for near-real time monitoring in Ukraine. We hope that this first-ever ML detector of both seismic and acoustic phases could be employed for real-time monitoring of conflicts around the world across different network geometries and noise conditions.</p>
Array processing is routinely used to measure apparent velocity and back-azimuth of seismic arrivals. Being an integral part of automatic processing pipelines for seismic event monitoring at the IDC and NDCs, this processing step usually follows seismic phase detection in continuous data and precedes event association and location. The apparent velocity is used to classify the type of the detected phase, while the measured back-azimuth is assumed to point towards the event epicentre. Phase type and back-azimuth are usually determined under the plane wave assumption using Frequency-Wavenumber (FK) analysis or other wave front fitting algorithms such as Progressive Multi-Channel Correlation (PMCC). However, local inhomogeneities below the seismic array as well as regional sub-surface structures can lead to deviations from the plane wave character and to differences between the measured back-azimuth and the actual source direction. This can also affect the slowness estimates and, thus, the accuracy of phase type classification. Previous attempts to take these issues into account were based for example on empirical array-dependent slowness vector corrections. Here, we suggest a neural network architecture to learn from past observations and to determine the seismic phase type and back-azimuth directly from the arrival time differences between all combinations of stations of a given array (the co-array), without assuming a certain wavefield geometry. In particular, input data are phase differences measured for multiple frequencies from the cross-spectrum of each co-array element. The neural network is a combined classification (phase type) and regression (back-azimuth) network and is trained using P and S arrivals of over 30,000 seismic events from the reviewed regional bulletins in Scandinavia of the past three decades and seismic noise examples. Hence, phase types are classified without first measuring the apparent velocity and without using pre-set velocity thresholds, and an unbiased back-azimuth is determined pointing directly towards the source. Training data are selected based on coherency thresholds to avoid training with too noisy arrivals included in the bulletins where for example the analysist placed a pick based on additional information. Furthermore, we test augmenting training data with time differences corresponding to plane waves to add source directions which are underrepresented in the bulletins. Models are trained and evaluated for regional seismic phase observations at the ARCES, NORES and SPITS arrays. Very good performance for seismic phase type classification (97% accuracy) and low source back-azimuth misfits were obtained. A systematic and careful test of the performance compared to FK analysis in NORSAR’s automatic processing (FKX) was conducted to evaluate potential improvements for event association and location. Taking the reviewed bulletins as reference, our first results suggest that the machine learning phase classifier performs equally well as FKX processing when it comes to phase classification and better for source back-azimuth estimation.
Global estimates for future growth indicate that city inhabitation will increase by 13% due to a gradual shift in residence from rural to urban areas. The continuous increase in urban population has caused many cities to upgrade their infrastructures and embrace the vision of a “smart-city”. Data collection through sensors represents the base layer of every smart-city solution. Large datasets are processed, and relevant information is transferred to the police, local authorities, and the general public to facilitate decisions and to optimize the performance of cities in areas such as transport, health care, safety, natural resources and energy. The objective of the GEObyIT project is to provide a real-time risk reduction system in an urban environment by applying machine learning methodologies to automatically identify and categorise different types of geodata, i.e., seismic events and geological structures. The project focusses on the city of Oslo, Norway, addressing the common need of two departments of the municipality, i.e., the Emergency Department and the Water and Sewage Department. In the present work, we focus on passive seismic records acquired with the objective to quickly locate urban events as well as to continuous monitor changes in the near surface. For this purpose, a seismic network of Raspberry Shake 3D sensors connected to GSM modems, to facilitate real-time data transfer, was deployed in target areas within the city of Oslo in 2021. We present preliminary results of three approaches applied to the continuous data: (1) automatic detection of metro trains, (2) automatic identification of outlier events such as construction and mining blasts, and (3) noise interferometry to monitor the near sub-surface in an area with quick clay. We use a supervised method based on convolutional neural networks trained with visually identified seismic signals on three sensors distributed along a busy metro track (1). Application to continuous data allowed us the reliably detect trains as well as their direction, while not triggering other events. Further development of this approach will be useful to either sort out known repeating seismic signals or to monitor traffic in an urban environment. In approach (2) we aim to detect rare or unusual seismic events using an outlier detection method. A convolutional autoencoder was trained to create dense features from continuous signals for each sensor. These features are used in a one-class support vector machine to detect anomalies. We were able to identify a series of construction and mine blasts, a meteor signal as well as two earthquakes. Finally, we apply seismic noise interferometry to close-by sensor pairs to measure temporal variations in the shallow ground (3). We observe clear seismic velocity variations during periods of strong frost in winter 2021/2022. This opens up for the potential to detect also non-seasonal changes in the ground, for example related to instabilities in quick clay deposits located within the city of Oslo.
SUMMARY Seismic signals generated by iceberg calving can be used to monitor ice loss at tidewater glaciers with high temporal resolution and independent of visibility. We combine the empirical matched field (EMF) method and machine learning using convolutional neural networks (CNNs) for calving event detection at the Spitsbergen (SPITS) seismic array and the single broad-band station KBS on the Arctic Archipelago of Svalbard. EMF detection with seismic arrays seeks to identify all signals generated by events in a confined target region similar to single P and/or S phase templates by assessing the beam power obtained using empirical phase delays between the array stations. The false detection rate depends on threshold settings and therefore needs appropriate tuning or, alternatively, post-processing. We combine the EMF detector at the SPITS array, as well as an STA/LTA (short term average/long term average) detector at the KBS station, with a post-detection classification step using CNNs. The CNN classifier uses waveforms of the three-component record at KBS as input. We apply the methodology to detect and classify calving events at tidewater glaciers close to the KBS station in the Kongsfjord region in Northwestern Svalbard. In a previous study, a simpler method was implemented to find these calving events in KBS data, and we use it as the baseline in our attempt to improve the detection and classification performance. The CNN classifier is trained using classes of confirmed calving signals from four different glaciers in the Kongsfjord region, seismic noise examples and regional tectonic seismic events. Subsequently, we process continuous data of six months in 2016. We test different CNN architectures and data augmentations to deal with the limited training data set available. Targeting Kronebreen, one of the most active glaciers in the Kongsfjord region, we show that the best performing models significantly improve the baseline classifier. This result is achieved for both the STA/LTA detection at KBS followed by CNN classification, as well as EMF detection at SPITS combined with a CNN classifier at KBS, despite of SPITS being located at 100 km distance from the target glacier in contrast to KBS at 15 km distance. Our results will further increase confidence in estimates of ice loss at Kronebreen derived from seismic observations which in turn can help to better understand the impact of climate change in Svalbard.
We have created a knowledge graph based on major data sources used in ecotoxicological risk assessment. We have applied this knowledge graph to an important task in risk assessment, namely chemical effect prediction. We have evaluated nine knowledge graph embedding models from a selection of geometric, decomposition, and convolutional models on this prediction task. We show that using knowledge graph embeddings can increase the accuracy of effect prediction with neural networks. Furthermore, we have implemented a fine-tuning architecture which adapts the knowledge graph embeddings to the effect prediction task and leads to a better performance. Finally, we evaluate certain characteristics of the knowledge graph embedding models to shed light on the individual model performance.
Extrapolation of adverse biological (toxic) effects of chemicals is an important contribution to expand available hazard data in (eco)toxicology without the use of animals in laboratory experiments. In this work, we extrapolate effects based on a knowledge graph (KG) consisting of the most relevant effect data as domain-specific background knowledge. An effect prediction model, with and without background knowledge, was used to predict mean adverse biological effect concentration of chemicals as a prototypical type of stressors. The background knowledge improves the model prediction performance by up to 40\% in terms of $R^2$ (\ie coefficient of determination). We use the KG and KG embeddings to provide quantitative and qualitative insights into the predictions. These insights are expected to improve the confidence in effect prediction. Larger scale implementation of such extrapolation models should be expected to support hazard and risk assessment, by simplifying and reducing testing needs.
Fast detection and characterization of seismic sources is crucial for decision-making and warning systems that monitor natural and induced seismicity. However, besides the laying out of ever denser monitoring networks of seismic instruments, the incorporation of new sensor technologies such as Distributed Acoustic Sensing (DAS) further challenges our processing capabilities to deliver short turnaround answers from seismic monitoring. In response, this work describes a methodology for the learning of the seismological parameters: location and moment tensor from compressed seismic records. In this method, data dimensionality is reduced by applying a general encoding protocol derived from the principles of compressive sensing. The data in compressed form is then fed directly to a convolutional neural network that outputs fast predictions of the seismic source parameters. Thus, the proposed methodology can not only expedite data transmission from the field to the processing center, but also remove the decompression overhead that would be required for the application of traditional processing methods. An autoencoder is also explored as an equivalent alternative to perform the same job. We observe that the CS-based compression requires only a fraction of the computing power, time, data and expertise required to design and train an autoencoder to perform the same task. Implementation of the CS-method with a continuous flow of data together with generalization of the principles to other applications such as classification are also discussed.
The Toxicological and Risk Assessment Knowledge Graph (TERA) [1] integrates several disparate datasets relevant to ecological risk assessment and effect prediction. TERA is being used in conjunction with knowledge graph embedding models to improve the extrapolation of chemical effect data in the Norwegian Institute for Water Research (Norsk institutt for vannforskning, NIVA) [1].1 The largest publicly available repository of effect data is the ECOTOXicology knowledge base (ECOTOX) developed by the US Environmental Protection Agency [2]. The dataset consists of 940k experiments using 12k compounds and 13k species. ECOTOX contains a taxonomy (of species), however, this only considers the species represented in the ECOTOX effect data. Hence, to enable extrapolation of effects across a larger taxonomic domain, an alignment to the NCBI taxonomy have to be established. However, there does not exist a complete and public mapping set between the 47,785 ECOTOX taxa and the 2,140,344 NCBI taxa. In this paper we present the ECOTOX-NCBI alignment results of three ontology matching algorithms.
The usefulness and usability of knowledge bases (KBs) is often limited by quality issues. One common issue is the presence of erroneous assertions, often caused by lexical or semantic confusion. We study the problem of correcting such assertions, and present a general correction framework which combines lexical matching, semantic embedding, soft constraint mining and semantic consistency checking. The framework is evaluated using DBpedia and an enterprise medical KB.
Exploring the effects a chemical compound has on a species takes a considerable experimental effort. Appropriate methods for estimating and suggesting new effects can dramatically reduce the work needed to be done by a laboratory. In this paper we explore the suitability of using a knowledge graph embedding approach for ecotoxicological effect prediction. A knowledge graph has been constructed from publicly available data sets, including a species taxonomy and chemical classification and similarity. The publicly available effect data is integrated to the knowledge graph using ontology alignment techniques. Our experimental results show that the knowledge graph based approach improves the selected baselines.
Ecological risk assessment requires large amounts of chemical effect data from laboratory experiments. Due to experimental effort and animal welfare concerns it is desired to extrapolate data from existing sources. To cover the required chemical effect data several data sources need to be integrated to enable their interoperability. In this paper we introduce the Toxicological Effect and Risk Assessment (TERA) knowledge graph, which aims at providing such integrated view, and the data preparation and steps followed to construct this knowledge graph. We also present the applications of TERA for chemical effect prediction and the potential applications within the Semantic Web community.
In this paper, we present a preliminary study to compute embeddings for OWL 2 ontologies by projecting the ontology axioms into a graph and performing (random) walks over the ontology graph to create a corpus of sentences. This corpus is then given to a neural language model to create concept embeddings. The conducted preliminary evaluation shows promising results.
Exploring the effects a chemical compound has on a species takes a considerable experimental effort. Appropriate methods for estimating and suggesting new effects can dramatically reduce the work needed to be done by a laboratory. In this PhD research we aim at exploring the suitability of using a knowledge graph embedding approach for ecotoxicological effect prediction. A knowledge graph is being constructed from publicly available data sets, including a species taxonomy and chemical classification and similarity. We use ontology alignment techniques to integrate the effect data into the knowledge graph. Our preliminary experimental results show that the knowledge graph based approach improves the selected baselines.