We present DIJIT, a novel binocular robotic head expressly designed for mobile agents that behave as active observers. DIJIT's unique breadth of functionality enables active vision research and the study of human-like eye and head-neck motions, their interrelationships, and how each contributes to visual ability. DIJIT is also being used to explore the differences between how human vision employs eye/head movements to solve visual tasks and current computer vision methods. DIJIT's design features nine mechanical degrees of freedom, while the cameras and lenses provide an additional four optical degrees of freedom. The ranges and speeds of the mechanical design are comparable to human performance. DIJIT attains 85% of the peak human saccade speed. Our design includes the ranges of motion required for convergent stereo, namely, vergence, version, and cyclotorsion. Here, we present DIJIT and some aspects of its performance. We also present a novel method for saccadic camera movements, using a direct relationship between camera orientation and motor values. The resulting saccadic camera movements are close to human movements in terms of their accuracy, with 1.17(degrees) and 1.14(degrees) mean error for the left and right cameras, respectively.
ABSTRACT The remote operation of underwater vehicles at depth is complicated by the presence of invisible and unpredictable environmental disturbances such as cross‐currents. Communicating the presence of these disturbances to an operator on the surface is made more difficult by the nature of the disturbances and the lack of visible features to highlight in the visual display presented to the operator. Here we explore the use of a novel interactive soft haptic touchpad that utilizes vibration and particle jamming to provide information about the presence and direction of cross‐currents to the operator of an ROV (remotely operated vehicle). An in‐water experiment using a thruster‐based ROV and artificially generated cross‐current was performed with nonexpert ROV operators to evaluate the effectiveness of multimodal haptic feedback to communicate complex environmental information during high‐risk operations. Advanced haptic displays can signal both the presence of external factors as well as their direction, information that can enhance operational performance as well as reduce operator cognitive load. Using haptic feedback resulted in a statistically significant reduction in cognitive load of 24.3% and an increase in positioning accuracy of 28.3% for novice operators. Deviation from an ideal path was also reduced by 29.5% for experienced operators when using haptic feedback compared to without. While this experiment took place in controlled conditions with a fixed direction cross‐current and haptic interface, this approach could be extended to communicate real‐time environmental information in real‐world unstructured environments.
Gesture-based communication is a standard underwater communication strategy that is taught to divers as part of their regular diver training and it would seem a natural mechanism to leverage for diver to robot communication underwater. Enabling an unmanned underwater vehicle (UUV) to understand such sequences would involve having the robot learn the large set of gestures that divers use and the way they are combined. As perfect transcription of gestures is unlikely, the communication process also requires an error-correcting framework to ensure that communication is clear and correct. Here we describe an interactive process that provides this infrastructure. A weakly supervised transfer learning approach is used to recognize standard SCUBA gestures in individual video frames and within a Sim2Real process to train a LSTM to recognize gesture sequences. This process is placed within a per-gesture and per-sequence interaction process to assist and confirm the recognition of individual gestures and to confirm entire gesture sequences. Individual aspects of this process and complete end-to-end operation are demonstrated using an unmanned underwater vehicle.
Accurate electric load forecasting is of critical importance for modern power grids. It can help optimize energy management, reduce operational costs, and enhance grid stability. Existing load forecasting tools typically perform well when modelling long-term trends with substantive data upon which to build a model, but can perform poorly for short-term load forecasting or when data is sparse or incomplete. Diffusion models have recently emerged as powerful generative tools that excel in modelling complex distributions, making them a promising approach for electric load forecasting. In this paper, we build upon recent diffusion models for time series forecasting and explore the potential of combining diffusion models with prior models to improve performance. Additionally, we propose a metric-based meta-learning approach for fast data adaptation. Experimental results with this metric-based meta-learning approach on real-world load forecasting datasets outperform state-of-the-art baselines, showcasing the potential of diffusion-based refinement in practical forecasting applications.
Short-term load forecasting (STLF) for residential households has become of critical importance for the secure operation of power grids as well as home energy management systems. While machine learning is effective for residential STLF, data and resource limitations hinder individual household predictions operated on local devices. In contrast, utility companies have access to broader sets of data as well as to better computational resources, and thus have the potential to deploy complex forecasting models such as Graph neural network-based models to explore the spatial-temporal relationships between households for achieving impressive STLF performance. In this work, we propose an efficient and privacy-conservative knowledge distillation-based STLF framework. This framework can improve the STLF forecasting accuracy of lightweight individual household forecasting models via leveraging the benefits of knowledge distillation and graph neural networks (GNN). Specifically, we distill the knowledge learned from a GNN model pre-trained on utility data sets into individual models without the need to access data sets of other households. Extensive experiments on real-world residential electric load datasets demonstrate the effectiveness of the proposed method.
Speech emotion recognition (SER) is the task of automatically recognizing emotions expressed in spoken language. Current approaches focus on analyzing isolated speech segments to identify a speaker’s emotional state. Meanwhile, recent text-based emotion recognition methods have effectively shifted towards emotion recognition in conversation (ERC) that considers conversational context. Motivated by this shift, here we propose SERC-GCN, a method for speech emotion recognition in conversation (SERC) that predicts a speaker’s emotional state by incorporating conversational context, speaker interactions, and temporal dependencies between utterances. SERC-GCN is a two-stage method. First, emotional features of utterance-level speech signals are extracted. Then, these features are used to form conversation graphs that are used to train a graph convolutional network to perform SERC. We empirically evaluate the effectiveness of SERC-GCN and show that it outperforms the current state-of-the-art methods on the IEMOCAP benchmark dataset.
Traffic signal control (TSC) has seen substantial advancements through the application of reinforcement learning (RL) algorithms, which have shown remarkable potential in enhancing traffic flow efficiency. These RL-based approaches often surpass traditional rule-based methods, particularly in dynamic traffic environments. However, current RL solutions for TSC predominantly rely on model-free methods, necessitating extensive environmental interactions during training. This requirement can be prohibitively expensive or unfeasible in real-world implementations. Furthermore, existing methods have frequently neglected the issue of fairness in multi-intersection control, resulting in unbalanced congestion across different intersections. To address these challenges, we present FM2Light, a fairness-aware model-based multi-agent RL framework for TSC. Our approach leverages an ensemble of global world models for generating synthetic samples to enhance sample efficiency, thereby mitigating the data-intensive nature of the training process. Additionally, FM2Light incorporates a refined reward structure to promote fairness and improve coordination across multiple intersections. Extensive evaluations conducted in diverse real-world scenarios demonstrate that FM2Light achieves performance comparable to or exceeding that of model-free RL (MFRL) methods, while significantly reducing sample requirements and ensuring more equitable control among multiple agents.
This Element reviews the current state of what is known about the visual and vestibular contributions to our perception of self-motion and orientation with an emphasis on the central role that gravity plays in these perceptions. The Element then reviews the effects of impoverished challenging environments that do not provide full information that would normally contribute to these perceptions (such as driving a car or piloting an aircraft) and inconsistent challenging environments where expected information is absent, such as the microgravity experienced on the International Space Station.
Large language models (LLMs), including ChatGPT, Bard, and Llama, have achieved remarkable successes over the last two years in a range of different applications. In spite of these successes, there exist concerns that limit the wide application of LLMs. A key problem is the problem of hallucination. Hallucination refers to the fact that in addition to correct responses, LLMs can also generate seemingly correct but factually incorrect responses. This report aims to present a comprehensive review of the current literature on both hallucination detection and hallucination mitigation. We hope that this report can serve as a good reference for both engineers and researchers who are interested in LLMs and applying them to real world tasks.
Now in its third edition, this textbook is a comprehensive introduction to the multidisciplinary field of mobile robotics, which lies at the intersection of artificial intelligence, computational vision, and traditional robotics. Written for advanced undergraduates and graduate students in computer science and engineering, the book covers algorithms for a range of strategies for locomotion, sensing, and reasoning. The new edition includes recent advances in robotics and intelligent machines, including coverage of human-robot interaction, robot ethics, and the application of advanced AI techniques to end-to-end robot control and specific computational tasks. This book also provides support for a number of algorithms using ROS 2, and includes a review of critical mathematical material and an extensive list of sample problems. Researchers as well as students in the field of mobile robotics will appreciate this comprehensive treatment of state-of-the-art methods and key technologies.
Traffic congestion is a pervasive challenge in urban areas, contributing significantly to greenhouse gas emissions. Efficient traffic signal control stands as a pivotal factor in enhancing modern transportation systems. The complexity of this decision-making process is heightened by the dynamic nature of traffic patterns. Reinforcement Learning (RL) approaches have showcased promise in addressing traffic signal control, exhibiting notable performance gains over traditional methods. However, the practical utility of many RL-based solutions is constrained by their substantial data requirements, limiting applicability to real-world scenarios. This paper introduces an innovative model-based meta-reinforcement learning framework, ModelLight, designed for traffic signal control. In ModelLight, world models capturing the dynamics of signalized intersection are acquired and employed to generate imaginary trajectories within an optimization-based meta-learning approach, thereby enhancing overall sample efficiency. Experimental evaluations on real-world scenarios demonstrate that ModelLight surpasses RL-based traffic signal control baselines while demanding only a fraction of the interactions with the environment. Our datasets and code can be found at https://github.com/XingshuaiHuang/ModelLight.
We present AirChair, a semi-autonomous human transportation system composed of multiple wheelchairs operating as a convoy. The first wheelchair follows an on-foot human guide, the second wheelchair follows the first, and so on. Each wheelchair independently tracks its target with the help of an RGBD camera, and performs motion planning to follow along while steering clear of obstacles. The guide manages the convoy through a mobile control interface, allowing them to intervene as needed to ensure passenger safety. The effectiveness of combining automation technologies with an engaged operator is demonstrated by experiments in uncontrolled, real-world environments, which also suggest directions for further development.
With the global aim of reducing carbon emissions, energy saving for communication systems has gained tremendous attention. Efficient energy-saving solutions are not only required to accommodate the fast growth in communication demand but solutions are also challenged by the complex nature of the load dynamics. Recent reinforcement learning (RL)-based methods have shown promising performance for network optimization problems, such as base station energy saving. However, a major limitation of these methods is the requirement of online exploration of potential solutions using a high-fidelity simulator or the need to perform exploration in a real-world environment. We circumvent this issue by proposing an offline reinforcement learning energy saving (ORES) framework that allows us to learn an efficient control policy using previously collected data. We first deploy a behavior energy-saving policy on base stations and generate a set of interaction experiences. Then, using a robust deep offline reinforcement learning algorithm, we learn an energy-saving control policy based on the collected experiences. Results from experiments conducted on a diverse collection of communication scenarios with different behavior policies showcase the effectiveness of the proposed energy-saving algorithms.
Understanding and predicting the emotional trajectory in multi-party multi-turn conversations is of great significance. Such information can be used, for example, to generate empathetic response in human-machine interaction or to inform models of pre-emptive toxicity detection. In this work, we introduce the novel problem of Predicting Emotions in Conversations (PEC) for the next turn (n+1), given combinations of textual and/or emotion input up to turn n. We systematically approach the problem by modeling three dimensions inherently connected to evoked emotions in dialogues, including (i) sequence modeling, (ii) self-dependency modeling, and (iii) recency modeling. These modeling dimensions are then incorporated into two deep neural network architectures, a sequence model and a graph convolutional network model. The former is designed to capture the sequence of utterances in a dialogue, while the latter captures the sequence of utterances and the network formation of multi-party dialogues. We perform a comprehensive empirical evaluation of the various proposed models for addressing the PEC problem. The results indicate (i) the importance of the self-dependency and recency model dimensions for the prediction task, (ii) the quality of simpler sequence models in short dialogues, (iii) the importance of the graph neural models in improving the predictions in long dialogues.
Online conversations are particularly susceptible to derailment, which can manifest itself in the form of toxic communication patterns including disrespectful comments and abuse. Forecasting conversation derailment predicts signs of derailment in advance enabling proactive moderation of conversations. State-of-the-art approaches to conversation derailment forecasting sequentially encode conversations and use graph neural networks to model dialogue user dynamics. However, existing graph models are not able to capture complex conversational characteristics such as context propagation and emotional shifts. The use of common sense knowledge enables a model to capture such characteristics, thus improving performance. Following this approach, here we derive commonsense statements from a knowledge base of dialogue contextual information to enrich a graph neural network classification architecture. We fuse the multi-source information on utterance into capsules, which are used by a transformer-based forecaster to predict conversation derailment. Our model captures conversation dynamics and context propagation, outperforming the state-of-the-art models on the CGA and CMV benchmark datasets
With the increasing use of data-intensive mobile applications and the number of mobile users, the demand for wireless data services has been increasing exponentially in recent years. In order to address this demand, a large number of new cellular base stations are being deployed around the world, leading to a significant increase in energy consumption and greenhouse gas emission. Consequently, energy consumption has emerged as a key concern in the fifth-generation (5G) network era and beyond. Reinforcement learning (RL), which aims to learn a control policy via interacting with the environment, has been shown to be effective in addressing network optimization problems. However, for reinforcement learning, especially deep reinforcement learning, a large number of interactions with the environment are required. This often limits its applicability in the real world. In this work, to better deal with dynamic traffic scenarios and improve real-world applicability, we propose a transfer deep reinforcement learning framework for energy optimization in cellular communication networks. Specifically, we first pre-train a set of RL-based energy-saving policies on source base stations and then transfer the most suitable policy to the given target base station in an unsupervised learning manner. Experimental results demonstrate that base station energy consumption can be reduced significantly using this approach.
The amount of cellular communication network traffic has increased dramatically in recent years, and this increase has led to a demand for enhanced network performance. Communication load balancing aims to balance the load across available network resources and thus improve the quality of service for network users. Most existing load balancing algorithms are manually designed and tuned rule-based methods where near-optimality is almost impossible to achieve. Furthermore, rule-based methods are difficult to adapt to quickly changing traffic patterns in real-world environments. Reinforcement learning (RL) algorithms, especially deep reinforcement learning algorithms, have achieved impressive successes in many application domains and offer the potential of good adaptabiity to dynamic changes in network load patterns. This survey presents a systematic overview of RL-based communication load-balancing methods and discusses related challenges and opportunities. We first provide an introduction to the load balancing problem and to RL from fundamental concepts to advanced models. Then, we review RL approaches that address emerging communication load balancing issues important to next generation networks, including 5G and beyond. Finally, we highlight important challenges, open issues, and future research directions for applying RL for communication load balancing.
The perceptual upright results from the multisensory integration of the directions indicated by vision and gravity as well as a prior assumption that upright is towards the head. The direction of gravity is signalled by multiple cues, the predominant of which are the otoliths of the vestibular system and somatosensory information from contact with the support surface. Here, we used neutral buoyancy to remove somatosensory information while retaining vestibular cues, thus "splitting the gravity vector" leaving only the vestibular component. In this way, neutral buoyancy can be used as a microgravity analogue. We assessed spatial orientation using the oriented character recognition test (OChaRT, which yields the perceptual upright, PU) under both neutrally buoyant and terrestrial conditions. The effect of visual cues to upright (the visual effect) was reduced under neutral buoyancy compared to on land but the influence of gravity was unaffected. We found no significant change in the relative weighting of vision, gravity, or body cues, in contrast to results found both in long-duration microgravity and during head-down bed rest. These results indicate a relatively minor role for somatosensation in determining the perceptual upright in the presence of vestibular cues. Short-duration neutral buoyancy is a weak analogue for microgravity exposure in terms of its perceptual consequences compared to long-duration head-down bed rest.
The need for social robot systems has become even more critical as a result of the ongoing pandemic. Labour shortages in the services sector and public health concerns around infection transmission combine to favour the deployment of autonomous systems in a number of traditional roles including server robots in restaurants, companion robots in long-term care homes and security robots in public spaces, to identify but a few examples. To be successful, social robots must communicate with a wide range of individuals under a wide range of different scenarios. Understanding and reacting to the sentiment being expressed by an individual is key in human-human interaction, especially in critical situations that require de-escalation. This paper takes as a starting point that user sentiment is also critical for the successful deployment of social robot systems. Although much can be learned from experiments performed in simulation, real-world experiments in the development of sentiment-aware social robots requires an infrastructure upon which to explore questions related to the role of sentiment in social robotics. This includes the development of an appropriate robot morphology and user/robot interface. This paper reports early results in the development of sentiment and display technologies as part of the development of a sentiment-informed social robot named Sentrybot, an autonomous robot intended for deployment in the security domain.
Jarek Gryz合作论文数Department of Computer Science and Engineering;York University3