
Climate-induced disasters are and will continue to be on the rise, and thus search-and-rescue (SAR) operations, where the task is to localize and assist one or several people who are missing, become increasingly relevant. In many cases the rough location may be known and a UAV can be deployed to explore a given, confined area to precisely localize the missing people. Due to time and battery constraints it is often critical that localization is performed as efficiently as possible. In this work we approach this type of problem by abstracting it as an aerial view goal localization task in a framework that emulates a SAR-like setup without requiring access to actual UAVs. In this framework, an agent operates on top of an aerial image (proxy for a search area) and is tasked with localizing a goal that is described in terms of visual cues. To further mimic the situation on an actual UAV, the agent is not able to observe the search area in its entirety, not even at low resolution, and thus it has to operate solely based on partial glimpses when navigating towards the goal. To tackle this task, we propose AiRLoc, a reinforcement learning (RL)-based model that decouples exploration (searching for distant goals) and exploitation (localizing nearby goals). Extensive evaluations show that AiRLoc outperforms heuristic search methods as well as alternative learnable approaches, and that it generalizes across datasets, e.g. to disaster-hit areas without seeing a single disaster scenario during training. We also conduct a proof-of-concept study which indicates that the learnable methods outperform humans on average. Code and models have been made publicly available at https://github.com/aleksispi/airloc.
The one-pixel attack is an image attack method for creating adversarial instances with minimal perturbations, i.e., pixel modification. The attack method makes the adversarial instances difficult to detect as it only manipulates a single pixel in the image. In this paper, we study four different defense approaches against adversarial attacks, and more specifically the one-pixel attack, over three different models. The defense methods used are: data augmentation, spatial smoothing, and Gaussian data augmentation used during both training and testing. The empirical experiments involve the following three models: all convolutional network (CNN), network in network (NiN), and the convolutional neural network VGG16. Experiments were executed and the results show that Gaussian data augmentation performs quite poorly when applied during the prediction phase. When used during the training phase, we see a reduction in the number of instances that could be perturbed by the NiN model. However, the CNN model shows an overall significantly worse performance compared to no defense technique. Spatial smoothing shows an ability to reduce the effectiveness of the one-pixel attack, and it is on average able to defend against half of the adversarial examples. Data augmentation also shows promising results, reducing the number of successfully perturbed images for both the CNN and NiN models. However, data augmentation leads to slightly worse overall model performance for the NiN and VGG16 models. Interestingly, it significantly improves the performance for the CNN model. We conclude that the most suitable defense is dependent on the model used. For the CNN model, our results indicate that a combination of data augmentation and spatial smoothing is a suitable defense setup. For the NiN and VGG16 models, a combination of Gaussian data augmentation together with spatial smoothing is more promising. Finally, the experiments indicate that applying Gaussian noise during the prediction phase is not a workable defense against the one-pixel attack.
During the last decade, we have witnessed a rapid development of extended reality (XR) technologies such as augmented reality (AR) and virtual reality (VR). Further, there have been tremendous advancements in artificial intelligence (AI) and machine learning (ML). These two trends will have a significant impact on future digital societies. The vision of an immersive, ubiquitous, and intelligent virtual space opens up new opportunities for creating an enhanced digital world in which the users are at the center of the development process, so-called intelligent realities (IRs). The “Human-Centered Intelligent Realities” (HINTS) profile project will develop concepts, principles, methods, algorithms, and tools for human-centered IRs, thus leading the way for future immersive, user-aware, and intelligent interactive digital environments. The HINTS project is centered around an ecosystem combining XR and communication paradigms to form novel intelligent digital systems. HINTS will provide users with new ways to understand, collaborate with, and control digital systems. These novel ways will be based on visual and data-driven platforms which enable tangible, immersive cognitive interactions within real and virtual realities. Thus, exploiting digital systems in a more efficient, effective, engaging, and resource-aware condition. Moreover, the systems will be equipped with cognitive features based on AI and ML, which allow users to engage with digital realities and data in novel forms. This paper describes the HINTS profile project and its initial results.
Manufacturers of high-end professional products are committed to delivering outstanding customer-quality experiences. They maintain databases of customer complaints and repair service jobs data to monitor product quality. Analyzing the text data from service jobs can help identify common problems, recurring issues, and patterns that impact customer satisfaction, and aid manufacturers in taking corrective actions to improve product design, manufacturing processes, and customer support services. However, distinguishing legitimate quality issues from a brief, domain-specific text in service jobs remains a challenge. This study aims to automate the classification of technical service repair job data into legitimate quality issues or non-issues to assist individuals in the quality field department in a large company. To achieve this goal, we developed a comprehensive pipeline based on natural language processing and machine learning techniques including raw text preprocessing, dealing with imbalance class distribution, feature extraction, and classification. In this study, We evaluate several feature extraction and machine learning classification methods and perform the Friedman test followed by Nemenyi post-hoc analysis to find the best-performing model. Our results show that the passive-aggressive classifier achieved the highest average accuracy of 94% and 89% average macro F1-score when trained on TF-IDF vectors.
A recent report by the Swedish Authority for Privacy Protection (IMY) evaluates the potential of jointly training and exchangingmachine learningmodels between two healthcare providers. In relation to the privacy problems identified therein, this article explores the trade-off between utility and privacy when using privacyenhancing technologies (PETs) in combination with federated learning. Results are reported from numerical experiments with standard text-book machine learning models under both differential privacy (DP) and FullyHomomorphic Encryption (FHE). The results indicate that FHE is a promising approach for privacy-preserving federated learning, with the CKKS scheme being more favorable in terms of computational performance due to its support of SIMD operations and compact representation of encrypted vectors. The results for DP are more inconclusive. The article briefly discusses the current regulatory context and aspects that lawmakers may consider to enable an AI leap in Swedish healthcare while maintaining data protection.
This paper is motivated by Floridi’s recent claim that Large Language Models like ChatGPT can be seen as ‘intelligence-free’ agents. Where I do not agree with Floridi that such systems are intelligence-free, my paper does question whether they can be called agents, and if so, what kind. I argue for the adoption of a more restricted understanding of agent in AI-research, one that comes closer in its meaning to how the term is used in the philosophies of mind, action, and agency. I propose such a more narrowing understanding of agent, suggesting that an agent can be seen as entity or system that things can be ‘up to’, that can act autonomously in a way that is best understood on the basis of Husserl’s notion of indeterminate determinability.
During the last decade we have witnessed how artificial intelligence (AI) have changed businesses all over the world. The customer life cycle framework is widely used in businesses and AI plays a role in each stage. However, implementing and generating value from AI in the customer life cycle is not always simple. When evaluating the AI against business impact and value it is critical to consider both the model performance and the policy outcome. Proper analysis of AI-derived policies must not be overlooked in order to ensure ethical and trustworthy AI. This paper presents a comprehensive analysis of the literature on AI in customer life cycles (CLV) from an industry perspective. The study included 31 of 224 analyzed peer-reviewed articles from Scopus search result. The results show a significant research gap regarding outcome evaluations of AI implementations in practice. This paper proposes that policy evaluation is an important tool in the AI pipeline and empathizes the significance of validating both policy outputs and outcomes to ensure reliable and trustworthy AI.
The Urdarbrunnen project is a Saab-led exploratory initiative that aims to develop an operator-assisted AI-enabled mission system for basic autonomous functions. In its first iteration, presented in this project paper, the system is designed to be capable of performing the search task of a combat search and rescue mission in a complex and dynamic environment, while providing basic human machine interaction support for remote operators. The system enables a team of agents to cooperatively plan and execute a search mission while also interfacing with the WARA-PS core system that allows human operators and other agents to monitor activities and interact with each other. The aim of the project is to develop the system iteratively, with each iteration incorporating feedback from simulations and real-world experiments. In future work, the capability of the system will be extended to incorporate additional tasks for other scenarios, making it a promising starting point for the integration of autonomous capabilities in a future air force.
By identifying and characterising the narratives told in news media we can better understand political and societal processes. The problem is challenging from the perspective of natural language processing because it requires a combination of quantitative and qualitative methods. This paper reports on work in progress, which aims to build a human-in-the-loop pipeline for analysing how the variation of narrative themes across different domains, based on topic modelling and word embeddings. As an illustration, we study the language associated with the threat narrative in British news media.
Intensifying climate change will lead to more extreme weather events, including heavy rainfall and drought. Accurate streamflow prediction models which are adaptable and robust to new circumstances in a changing climate will be an important source of information for decisions on climate adaptation efforts, especially regarding mitigation of the risks of and damages associated with flooding. In this work we propose a machine learning-based approach for predicting water flow intensities in inland watercourses based on the physical characteristics of the catchment areas, obtained from geospatial data (including elevation and soil maps, as well as satellite imagery), in addition to temporal information about past rainfall quantities and temperature variations. We target the one-day-ahead regime, where a fully convolutional neural network model receives spatio-temporal inputs and predicts the water flow intensity in every coordinate of the spatial input for the subsequent day. To the best of our knowledge, we are the first to tackle the task of dense water flow intensity prediction; earlier works have considered the prediction of flow intensities at a sparse set of locations at a time. An extensive set of model evaluations and ablations are performed, which empirically justify our various design choices. Code and preprocessed data have been made publicly available at https://github.com/aleksispi/fcn-water-flow.
Neural networks are very successful tools in for example advanced classification. From a statistical point of view, fitting a neural network may be seen as a kind of regression, where we seek a function from the input space to a space of classification probabilities that follows the “general” shape of the data, but avoids overfitting by avoiding memorization of individual data points. In statistics, this can be done by controlling the geometric complexity of the regression function. We propose to do something similar when fitting neural networks by controlling the slope of the network.After defining the slope and discussing some of its theoretical properties, we go on to show empirically in examples, using ReLU networks, that the distribution of the slope of a well-trained neural network classifier is generally independent of the width of the layers in a fully connected network, and that the mean of the distribution only has a weak dependence on the model architecture in general. We discuss possible applications of the slope concept, such as using it as a part of the loss function or stopping criterion during network training, or ranking data sets in terms of their complexity.
Ecosystem models can be used for understanding general phenomena of evolution, ecology, and ethology. They can also be used for analyzing and predicting the ecological consequences of human activities on specific ecosystems, e.g., the effects of agriculture, forestry, construction, hunting, and fishing. We argue that powerful ecosystem models need to include reasonable models of the physical environment and of animal behavior. We also argue that several well-known ecosystem models are unsatisfactory in this regard. Then we present the open-source ecosystem simulator Ecotwin, which is built on top of the game engine Unity. To model a specific ecosystem in Ecotwin, we first generate a 3D Unity model of the physical environment, based on topographic or bathymetric data. Then we insert digital 3D models of the organisms of interest into the environment model. Each organism is equipped with a genome and capable of sexual or asexual reproduction. An organism dies if it runs out of some vital resource or reaches its maximum age. The animal models are equipped with behavioral models that include sensors, actions, reward signals, and mechanisms of learning and decision-making. Finally, we illustrate how Ecotwin works by building and running one terrestrial and one marine ecosystem model.
In recent years, machine learning (ML) algorithms have been used to minimize maintenance costs and identify problems early in the automotive sector. The determination of an asset’s residual useful life of a component at a specific time is known as “remaining useful life” (RUL). The extensive evolution of data makes it challenging to analyze and interpret high-level and valuable features from the data. The issue arises in all disciplines, and the automotive industry is no exception, given the large number of sensors to consider. Existing RUL research has not given much thought to the influence of high dimensionality data on component maintenance and deterioration. The fundamental purpose of feature selection (FS) is to select a subset of features from the data without compromising model performance. This work proposes a hybrid approach to the FS problem that combines Ant Colony Optimization (ACO) and Particle Swarm Optimization (PSO). When tested on public datasets, our results demonstrate a rise in regression accuracy and a reduction in the number of selected features.
Temporal patterns are encoded within the time-series data, and neural networks, with their unique feature extraction ability, process those patterns to provide a better predictive response. Ensembles of neural networks have proven to be very effective Human Activity Recognition (HAR) tasks with time-series data, e.g., wearable sensors. The combination of predictions coming from the individual models in the ensemble helps boost the overall classification metric through efficient temporal pattern recognition. Currently, the most common strategy for combining the predictions coming from the individual models is simple averaging. However, since each ensemble model learns different temporal patterns of the time-series classification problem, a simple averaging strategy is sub-optimal. This sub-optimality is addressed in this paper through a neural network-based adaptive learning framework. The method’s core is training a neural gate that ingests the same input time-series data fed to the other temporal models. The goal of the training process is to adaptively learn scaler values against each temporal model by looking at the input data. These scaler values weigh each temporal model while combining the ensemble. The framework obtains superior predictive performance as compared to the standard ensembling techniques. The framework is evaluated on a benchmark HAR dataset called PAMAP2 [3] with two popular state-of-the-art ensemble architectures namely DTE [1] and LSTM-ensemble [2]. In both cases, the classification performance of the framework in HAR tasks surpasses the state-of-the-art models.
With the growing interest of the research community in making deep learning (DL) robust and reliable, detecting out-of-distribution (OOD) data has become critical. Detecting OOD inputs during test/prediction allows the model to account for discriminative features unknown to the model. This capability increases the model's reliability since this model provides a class prediction solely at incoming data similar to the training one. OOD detection is well established in computer vision problems. However, it remains relatively under-explored in other domains such as time series (i.e., Human Activity Recognition (HAR)). Since uncertainty has been a critical driver for OOD in vision-based models, the same component has proven effective in time-series applications. We plan to address the OOD detection problem in HAR with time-series data in this work. To test the capability of the proposed method, we define different types of OOD for HAR that arise from realistic scenarios. We apply an ensemble-based temporal learning framework that incorporates uncertainty and detects OOD for the defined HAR workloads. In particular, we extract OODs from popular benchmark HAR datasets and use the framework to separate those OODs from the indistribution (ID) data. Across all the datasets, the ensemble framework outperformed the traditional deep-learning method (our baseline) on the OOD detection task.
Code Search is a practical tool that helps developers navigate growing source code repositories by connecting natural language queries with code snippets. Platforms such as StackOverflow resolve coding questions and answers; however, they cannot perform a semantic search through the code. Moreover, poorly documented code adds more complexity to search for code snippets in repositories. To tackle this challenge, this paper presents Siambert, a BERT-based model that gets the question in natural language and returns relevant code snippets. The Siambert architecture consists of two stages, where the first stage, inspired by Siamese Neural Network, returns the top K relevant code snippets to the input questions, and the second stage ranks the given snippets by the first stage. The experiments show that Siambert outperforms non-BERT-based models having improvements that range from 12% to 39% on the Recall@1 metric and improves the inference time performance, making it 15x faster than standard BERT models.
People living with type 1 diabetes often use several apps and devices that help them collect and analyse data for a better monitoring and management of their disease. When such health related data is analysed in the cloud, one must always carefully consider privacy protection and adhere to laws regulating the use of personal data. In this paper we present our experience at the pilot Vinter competition 2021–22 organised by Vinnova. The competition focused on digital services that handle sensitive diabetes related data. The architecture that we proposed for the competition is discussed in the context of a hypothetical cloud-based service that calculates diabetes self-care metrics under strong privacy preservation. It is based on Fully Homomorphic Encryption (FHE) - a technology that makes computation on encrypted data possible. Our solution promotes safe key management and data life-cycle control. Our benchmarking experiment demonstrates execution times that scale well for the implementation of personalised health services. We argue that this technology has great potentials for AI-based health applications and opens up new markets for third-party providers of such services, and will ultimately promote patient health and a trustworthy digital society.
Extracting speaker-dependent paralinguistic information out of a person’s voice, provides an opportunity for adaptive behaviour related to speaker information in speech processing applications. For instance, in audio-based conversational applications, adapting responses to the attributes of the correspondent is an integral part in making the conversations effective. Two speaker attributes that humans can estimate quite well, based solely on hearing a person speak, is the gender and age of that person. However, in the field of speech processing, age and gender classification are relatively unexplored tasks, especially in a multilingual setting. In most cases, hand-crafted features, such as MFCCs, have been used with some success. However, recently large transformer networks, utilizing self-supervised pre-training, have shown promise in creating general speech embeddings for various speech processing tasks. We present a baseline for gender and age detection, in both monolingual and multilingual settings, for multiple state-of-the-art speech processing models, fine-tuned for age classification. We created four different datasets with data extracted from the Common Voice project to compare monolingual and multilingual performances. For gender classification, we could reach a macro average F1 score of ~96% in both a monolingual and multilingual setting. For age classification, using classes with a size of 10 years, we obtained a macro average mean absolute class error (MACE) of 0.68 and 0.86 on monolingual and multilingual datasets, respectively. For the English TIMIT dataset, we improve upon the previous state of the art for both age regression and gender classification. Our fine-tuned WavLM model reaches a mean absolute error (MAE) of 4.11 years for males and 4.44 for females in age estimation and our fine-tuned UniSpeech-SAT model reaches an accuracy of 99.8% for gender classification. All the models were deemed fast enough on a GPU to be used in real-time settings, and accurate enough, using only a small amount of speech, to be applicable in multilingual speech processing applications.
“Magic” is referred to here and there in the robotics literature, from “magical moments” afforded by a mobile bubble machine, to “spells” intended to entertain and motivate children–but what exactly could this concept mean for designers? Here, we present (1) some theoretical discussion on how magic could inform interaction designs based on reviewing the literature, followed by (2) a practical description of using such ideas to develop a simplified prototype, which received an award in an international robot magic competition. Although this topic can be considered unusual and some negative connotations exist (e.g., unrealistic thinking can be referred to as magical), our results seem to suggest that magic, in the experiential, supernatural, and illusory senses of the term, could be useful to consider in various robot design contexts, also for artifacts like home assistants and autonomous vehicles–thus, inviting further discussion and exploration.
Sign gesture recognition is the field that models sign gestures in order to facilitate communication with hearing and speech impaired people. Sign gestures are recorded with devices like a video camera or a depth camera. Palm gestures are also recorded with the Leap motion sensor. In this paper, we address palm sign gesture recognition using the Leap motion sensor. We extract geometric features from Leap motion recordings. Next, we encode the Genetic Algorithm (GA) for feature selection. Genetically selected features are fed to different classifiers for gesture recognition. Here we have used Support Vector Machine (SVM), Random Forest (RF), and Naive Bayes (NB) classifiers to have their comparative results. The gesture recognition accuracy of 74.00% is recorded with RF classifier on the Leap motion sign gesture dataset.