Ethical tools are frequently proposed as a means to promote the design and implementation of responsible Artificial Intelligence (AI). Yet many organizations designing and deploying AI make only limited use of ethical tools. This study explores the application of ethical tools for responsible AI in the media sector through four case studies conducted at three Dutch media organizations. Each case study involves the application of an ethical tool to improve the responsible design, development, or deployment of an AI application. The findings reveal that successfully implementing ethical tools is highly contextual, and requires more than their mere availability. Tools must be selected, adapted, or even partly developed to align with specific challenges. Additionally, successful adoption of ethical tools necessitates organizational awareness of AI ethics, knowledge of mitigation strategies, and an organizational governance that supports responsible AI. These insights thus highlight the importance of contextualization and organizational readiness for establishing a responsible AI practice.
While there is much focus on interventions to foster ethical reflection in the design process of AI, there is less focus on fostering ethical reflection for (end)users. Yet, with the rise of genAI, AI technologies are no longer confined to expert users; non-experts are widely using these technologies. In this case study in a governmental organization in the Netherlands, we investigated a bottom-up approach to foster ethical reflection on the use of genAI tools. An approach of guided experimentation, including an intervention with a serious game, allowed civil servants to experiment, to understand the technology and its associated risks. The case study demonstrates that this approach enhances the awareness of possibilities and limitations, and the ethical considerations, of genAI usage. By analyzing usage statistics, we estimated the organization's energy consumption.
It is crucial that ASR systems can handle the wide range of variations in speech of speakers from different demographic groups, with different speaking styles, and of speakers with (dis)abilities. A potential quality-of-service harm arises when ASR systems do not perform equally well for everyone. ASR systems may exhibit bias against certain types of speech, such as non-native accents, different age groups and gender. In this study, we evaluate two widely-used neural network-based architectures: Wav2vec2 and Whisper on potential biases for Dutch speakers. We used the Dutch speech corpus JASMIN as a test set containing read and conversational speech in a human-machine interaction setting. The results reveal a significant bias against non-natives, children and elderly and some regional dialects. The ASR systems generally perform slightly better for women than for men.
Artificial Intelligence (AI) is increasingly used in the media industry, for instance, for the automatic creation, personalization, and distribution of media content. This development raises concerns in society and the media sector itself about the responsible use of AI. This study examines how different stakeholders in media organizations perceive ethical issues in their work concerning AI development and application, and how they interpret and put them into practice. We conducted an empirical study consisting of 14 semi-structured qualitative interviews with different stakeholders in public and private media organizations, and mapped the results of the interviews on stakeholder journeys to specify how AI applications are initiated, designed, developed, and deployed in the different media organizations. This results in insights into the current situation and challenges regarding responsible AI practices in media organizations.
With the proliferation of misinformation on the web, automatic methods for detecting misinformation are becoming an increasingly important subject of study. If automatic misinformation detection is applied in a real-world setting, it is necessary to validate the methods being used. Large language models (LLMs) have produced the best results among text-based methods. However, fine-tuning such a model requires a significant amount of training data, which has led to the automatic creation of large-scale misinformation detection datasets. In this paper, we explore the biases present in one such dataset for misinformation detection in English, NELA-GT-2019. We find that models are at least partly learning the stylistic and other features of different news sources rather than the features of unreliable news. Furthermore, we use SHAP to interpret the outputs of a fine-tuned LLM and validate the explanation method using our inherently interpretable baseline. We critically analyze the suitability of SHAP for text applications by comparing the outputs of SHAP to the most important features from our logistic regression models.
In the marketing area, new trends are emerging, as customers are not only interested in the quality of the products or delivered services, but also in a stimulating shopping experience. Creating and influencing customers' experiences has become a valuable differentiation strategy for retailers. Therefore, understanding and assessing the customers' emotional response in relation to products/services represents an important asset. The purpose of this paper consists of investigating whether the customer's facial expressions shown during product appreciation are positive or negative and also which types of emotions are related to product appreciation. We collected a database of emotional facial expressions, by presenting a set of forty product related pictures to a number of test subjects. Next, we analysed the obtained facial expressions, by extracting both geometric and appearance features. Furthermore, we modeled them both in an unsupervised and supervised manner. Clustering techniques proved to be efficient at differentiating between positive and negative facial expressions in 78% of the cases. Next, we performed more refined analysis of the different types of emotions, by employing different classification methods and we achieved 84% accuracy for seven emotional classes and 95% for the positive vs. negative.
Since the rehabilitation process after a hip surgery is increasingly performed at home, professional caregivers experience a lack of insight in the rehabilitation of their patients. In this demo we present a portable sensor system for remote monitoring, consisting of both wearable and ambient sensor
This paper describes the Care4Balance (C4B) system for better facilitating communication and task coordination between formal and informal caregivers, and older adults as care receivers. Field-tests with older adults (n=3) and user studies (n=9) were conducted to evaluate the system and the perceived usefulness of the system. A review of related work and the study findings show that (1) the perceived benefit for the older target group was very low. The main motivation for using the system was triggered by the perceived benefit for their closest informal caregivers; (2) Informal caregivers do not regularly seek help for themselves, and (3) Introducing a C4B-like system is more than solving hardware and usability issues. The study suggests that more flexibility in the organizational structure of formal care (in The Netherlands and beyond) is needed.
Due to their advantages over conventional n-gram language models, recurrent neural network language models (RNNLMS) recently have attracted a fair amount of research attention in the speech recognition community. In this paper, we explore one advantage of RNNLMS, namely, the ease with which they allow the integration of additional knowledge sources. We concentrate on features that provide complementary information w.r.t. the lexical identities of the words. We refer to such information as meta-information. We single out three cases and investigate their merits by means of N-best list re-scoring experiments on a challenging corpus of spoken Dutch (referred to as CGN) as well as on the English Wall Street Journal (WSJ) corpus. First, we look at Parts of Speech (POS) tags and lemmas, two sources of word-level linguistic information that are known to make a contribution to the performance of conventional language models. We confirm that RNNLMS can benefit from these sources as well. Second, we investigate socio-situational settings (SSSs) and topics, two sources of discourse-level information that are also known to benefit language models. SSSs are present in the CGN data, and can be seen as a proxy for the language register. For the purposes of our investigation, we assume that information on the SSS can be captured at the moment at which speech is recorded. Topics, i.e., treatments of different subjects, are present in the WSJ data. In order to predict POS, lemmas, sss and topic, a second RNNLM is coupled to the main RNNLM. We refer to this architecture as a recurrent neural network tandem language model (RNNTLM). Our experimental findings show that if high-quality meta-information labels are available, both word-level and discourse-level information improve performance of language models. Third, we investigate sentence length and word length (i.e., token size), two sources of intrinsic information that are readily available for exploitation because they are known at the time of re-scoring. Intrinsic information has been largely overlooked by language modeling research. The results of both experiments on CGN data and WSJ data show that integrating sentence length and word length can achieve improvement. RNNLMS allow these features to be incorporated with ease, and obtain improved performance. (C) 2015 Elsevier B.V. All rights reserved.
Humanoid robot navigation in domestic environments remains a challenging task. In this paper, we present an approach for navigating such environments for the humanoid robot Nao. We assume that a map of the environment is given and focus on the localization task. The approach is based on the use only of odometry and a single camera. The camera is used to correct for the drift of odometry estimates. Additionally, scene-classification is used to obtain information about the robot's position when it gets close to the destination. The approach is tested in an office environment to demonstrate that it can be reliably used for navigation in a domestic environment.
To test whether synthetic emotions expressed by a virtual human elicit positive or negative emotions in a human conversation partner and affect satisfaction towards the conversation, an experiment was conducted where the emotions of a virtual human were manipulated during both the listening and speaking phase of the dialogue. Twenty-four participants were recruited and were asked to have a real conversation with the virtual human on six different topics. For each topic the virtual human’s emotions in the listening and speaking phase were different, including positive, neutral and negative emotions. The results support our hypotheses that (1) negative compared to positive synthetic emotions expressed by a virtual human can elicit a more negative emotional state in a human conversation partner, (2) synthetic emotions expressed in the speaking phase have more impact on a human conversation partner than emotions expressed in the listening phase, (3) humans with less speaking confidence also experience a conversation with a virtual human as less positive, and (4) random positive or negative emotions of a virtual human have a negative effect on the satisfaction with the conversation. These findings have practical implications for the treatment of social anxiety as they allow therapists to control the anxiety evoking stimuli, i.e., the expressed emotion of a virtual human in a virtual reality exposure environment of a simulated conversation. In addition, these findings may be useful to other virtual applications that include conversations with a virtual human.
Automatic understanding of customers' shopping behavior and acting according to their needs is relevant in the marketing domain and is attracting a lot of attention lately. In this work, we propose a multi-level framework for the automatic assessment of customers' shopping behavior. The low level input to the framework is obtained from different types of cameras, which are synchronized, facilitating efficient processing of information. A fish-eye camera is used for tracking people, while a high-definition one serves for the action recognition task. The experiments are performed on both laboratory and real-life recordings in a supermarket. From the video recordings, we extract features related to the spatio-temporal behavior of trajectories, the dynamics and the time spent in each region of interest (ROI) in the shop and regarding the customer-products interaction patterns. Next we analyze the shopping sequences using a Hidden Markov Model (HMM). We conclude that it is possible to accurately classify trajectories (93%), discriminate between different shopping related actions (91.6%), and recognize shopping behavioral types by means of our proposed reasoning model in 95% of the cases. (C) 2012 Elsevier B.V. All rights reserved.
Taking important life decisions is a complex task leading to long-lasting consequences. It requires balancing one's own needs and those of other stakeholders. Current digital decision support focuses little on the human decision-making capabilities. Systems are designed as analytic tools to find optimal outcomes assuming stable and known preferences. However, insights from psychology and behavioral decision research show that people construct preferences during an adaptive decision-making process and are less rational than assumed by current tools. It has been suggested that a stronger focus on personal values could lead to improved decision making, but reflection on values is difficult for people. This paper presents a first exploration of how to aid people in reflecting on their values. It serves as a starting point to develop digital value-focused decision support tools. We describe the design of a probe for value reflection and several studies with experts and end-users that led to a first set of considerations for such tools.
Conventional n-gram language models are known for their limited ability to capture long-distance dependencies and their brittleness with respect to within-domain variations. In this paper, we propose a k-component recurrent neural network language model using curriculum learning (CL-KRNNLM) to address within-domain variations. Based on a Dutch-language corpus, we investigate three methods of curriculum learning that exploit dedicated component models for specific sub-domains. Under an oracle situation in which context information is known during testing, we experimentally test three hypotheses. The first is that domain-dedicated models perform better than general models on their specific domains. The second is that curriculum learning can be used to train recurrent neural network language models (RNNLMs) from general patterns to specific patterns. The third is that curriculum learning, used as an implicit weighting method to adjust the relative contributions of general and specific patterns, outperforms conventional linear interpolation. Under the condition that context information is unknown during testing, the CL-KRNNLM also achieves improvement over conventional RNNLM by 13% relative in terms of word prediction accuracy. Finally, the CL-KRNNLM is tested in an additional experiment involving N-best rescoring on a standard data set. Here, the context domains are created by clustering the training data using Latent Dirichlet Allocation and k-means clustering.
Having a free-speech conversation with avatars in a virtual environment can be desirable in virtual reality applications, such as virtual therapy and serious games. However, recognizing and processing free speech seems too ambitious to realize with the current technology. As an alternative, pre-scripted conversations with keyword detection can handle a number of goal-oriented situations, as well as some scenarios in which the conversation content is of secondary importance. This is, for example, the case in virtual exposure therapy for the treatment of people with social phobia, where conversation is for exposure and anxiety arousal only. A drawback of pre-scripted dialog is the limited scope of the user's answers. The system cannot handle a user's response that does not match the pre-defined content, other than by providing a default reply. A new method, which uses priming material to restrict the possibility of the user's response, is proposed in this paper to solve this problem. Two studies were conducted to investigate whether people can be guided to mention specific keywords with video and/or picture primings. Study 1 was a two-by-two experiment in which participants (n = 20) were asked to answer a number of open questions. Prior to the session, participants watched priming videos or unrelated videos. During the session, they could see priming pictures or unrelated pictures on a whiteboard behind the person who asked the questions. The results showed that participants tended to mention more keywords both with priming videos and pictures. Study 2 shared the same experimental setting but was carried out in virtual reality instead of in the real world. Participants (n = 20) were asked to answer questions of an avatar when they were exposed to priming material, before and/or during the conversation session. The same results were found: the surrounding media content had a guidance effect. Furthermore, when priming pictures appeared in the environment, people sometimes forgot to mention the content they typically would mention.
Automatic understanding and recognition of human shopping behavior has many potential applications, attracting an increasing interest in the marketing domain. The reliability and performance of the automatic recognition system is highly influenced by the adopted theoretical model of behavior. In this work, we address the analogy between human shopping behavior and a natural language. The adopted methodology associates low-level information extracted from video data with semantic information using the proposed behavior language model. Our contribution on the action recognition level consists of proposing a new feature set which fuses Histograms of Optical Flow (HOF) with directional features. On the behavior level we propose combining smoothed bi-grams with the maximum dependency in a chain of conditional probabilities. The experiments are performed on both laboratory and real-life datasets. The introduced behavior language model achieves an accuracy of 87% on the laboratory data and 76% on the real-life dataset, an improvement of 11% and 8% respectively over the baseline model, by incorporating semantic knowledge and capturing correlations between the basic actions. (C) 2012 Elsevier B.V. All rights reserved.
In automatic speech recognition, conventional language models recognize the current word using only information from preceding words. Recently, Recurrent Neural Network Language Models ( RNNLM s) have drawn increased research attention because of their ability to outperform conventional n-gram language models. The superiority of RNNLM s is based in their ability to capture long-distance word dependencies. RNNLM s are, in practice, applied in an N-best rescoring framework, which of-fers new possibilities for information integration. In particular, it becomes interesting to extend the ability of RNNLM s to capture long distance information by also allowing them to exploit information from succeeding words during the rescoring pro-cess. This paper proposes three approaches for exploiting succeeding word information in RNNLM s. The first is a forward-backward model that combines RNNLM s exploiting preceding and succeeding words. The second is an extension of a Maximum Entropy RNNLM ( RNNME ) that incorporates succeeding word information. The third is an approach that combines language models using two-pass alternating rescoring. Experi-mental results demonstrate the ability of succeeding word information to improve RNNLM performance, both in terms of perplexity and Word Error Rate ( WER ). The best performance is achieved by a combined model that exploits the three words succeeding the current word.
Virtual reality applications with virtual humans, such as virtual reality exposure therapy, health coaches and negotiation simulators, are developed for different contexts and usually for users from different countries. The emphasis on a virtual human’s emotional expression depends on the application; some virtual reality applications need an emotional expression of the virtual human during the speaking phase, some during the listening phase and some during both speaking and listening phases. Although studies have investigated how humans perceive a virtual human’s emotion during each phase separately, few studies carried out a parallel comparison between the two phases. This study aims to fill this gap, and on top of that, includes an investigation of the cultural interpretation of the virtual human’s emotion, especially with respect to the emotion’s valence. The experiment was conducted with both Chinese and non-Chinese participants. These participants were asked to rate the valence of seven different emotional expressions (ranging from negative to neutral to positive during speaking and listening) of a Chinese virtual lady. The results showed that there was a high correlation in valence rating between both groups of participants, which indicated that the valence of the emotional expressions was as easily recognized by people from a different cultural background as the virtual human. In addition, participants tended to perceive the virtual human’s expressed valence as more intense in the speaking phase than in the listening phase. The additional vocal emotional expression in the speaking phase is put forward as a likely cause for this phenomenon.
In this paper, we investigate automatic classification of the socio-situational settings of transcripts of a spoken discourse. Knowledge of the socio-situational setting can be used to search for content recorded in a particular setting or to select context-dependent models for example in speech recognition. The subjective experiment we report on in this paper shows that people correctly classify 68% the socio-situational settings. Based on the cues that participants mentioned in the experiment, we developed two types of automatic socio-situational setting classification methods; a static socio-situational setting classification method using support vector machines (s3c-svm), and a dynamic socio-situational classification method applying dynamic Bayesian networks (s3c-dbn). Using these two methods, we developed classifiers applying various features and combinations of features. The s3c-svm method with sentence length, function word ratio, single occurrence word ratio, part of speech (pos) and words as features results in a classification accuracy of almost 90%. Using a bigram s3c-dbn with pos tag and word features results in a dynamic classifier which can obtain nearly 89% classification accuracy. The dynamic classifiers not only can achieve similar results as the static classifiers, but also can track the socio-situational setting while processing a transcript or conversation. On discourses with a static social situational setting, the dynamic classifiers only need the initial 25% of data to achieve a classification accuracy close to the accuracy achieved when all data of a transcript is used.
Surveillance systems in shopping malls or supermarkets are usually used for detecting abnormal behavior. We used the distributed video cameras system to design digital shopping assistants which assess the behavior of customers while shopping, detect when they need assistance, and offer their support in case there is a selling opportunity. In this paper we propose a system for analyzing human behavior patterns related to products interaction, such as browse through a set of products, examine, pick products, try on, interact with the shopping cart, and look for support by waiving one hand. We used the Kinect sensor to detect the silhouettes of people and extracted discriminative features for basic action detection. Next we analyzed different classification methods, statistical and also spatio-temporal ones, which capture relations between frames, features, and basic actions. By employing feature level fusion of appearance and movement information we obtained an accuracy of 80% for the mentioned six basic actions.
Siska Fitrianie合作论文数Output Fission at CHIM - Computer Human Interaction Modelling2