In this paper, we study the problem of gesture recognition as a method for divers to communicate with an underwater robot. Gesture is a common method of communication between divers, and yet autonomous underwater vehicles have very limited capacity to understand gesture given lighting and visibility constraints (e.g., from water turbidity and diver depth). Traditional deep learning methods are limited in this domain because of a lack of sufficient training data. We show that it is not enough to learn a gesture in a laboratory setting, because the appearance changes dramatically underwater. We show how hyperdimensional computing can solve this problem by permitting hypervectors to serve as abstract representations of gestures, yielding rapid adaptation to new environments and new gestures. We experimentally verify this approach using a novel dataset of 6 diving relevant gestures. We show that we can accurately adapt to a gesture learned in a laboratory setting to work with a gesture observed underwater. Our approach compares favorably to a ResNet-18, which performs well in laboratory conditions (91.9% accuracy), but performs poorly underwater (53.9% accuracy). Our proposed approach is capable of rapid adaptation, resulting in an accuracy of 83.8% on underwater gestures with just one additional example from each class added to the support set. Finally, we also show the ability to adapt to new gestures not present in our original training set. We use hypervectors to learn new gestures from the Sign Language MNIST dataset, providing a high level of accuracy with a limited amount of training data.
Vision language models (VLMs) are designed to extract relevant visuospatial information from images. Some research suggests that VLMs can exhibit humanlike scene understanding, while other investigations reveal difficulties in their ability to process relational information. To achieve widespread applicability, VLMs must perform reliably, yielding comparable competence across a wide variety of related tasks. We sought to test how reliable these architectures are at engaging in trivial spatial cognition, e.g., recognizing whether one object is left of another in an uncluttered scene. We developed a benchmark dataset – TableTest – whose images depict 3D scenes of objects arranged on a table, and used it to evaluate state-of-the-art VLMs. Results show that performance could be degraded by minor variations of prompts that use logically equivalent descriptions. These analyses suggest limitations in how VLMs may reason about spatial relations in real-world applications. They also reveal novel opportunities for bolstering image caption corpora for more efficient training and testing.
Our autonomous robot, Bight, can be a reliable teammate that is capable of assisting in performing routine maintenance tasks on a Naval vessel. In this paper, we consider the task of maintaining the electrical panel. A vital first step is putting the robot into the correct position to view all of the parts of the electrical panel. The robot can get close, but the arm of the robot will need to move to where it can see everything. Here, we propose to solve this using a sigma delta spiking network that is trained using deep Q learning. Our approach is able to successfully solve this problem at varying distances. While we show how this works on this specific problem, we believe this approach to be general enough to be applied to any similar problem.
Action-effect predictions are believed to facilitate movement based on its association with sensory objectives and suppress the neurophysiological response to self- versus externally-generated stimuli (i.e., sensory attenuation). However, research is needed to explore theorised differences in the use of action-effect prediction based on whether movement is uncued (i.e., volitional) or in response to external cues (i.e., stimulus-driven). While much of the sensory attenuation literature has examined effects involving the auditory N1, evidence is also conflicted regarding this component’s sensitivity to action-effect prediction. In this study (N = 64), we explored the influence of action-effect contingency on event-related potentials associated with visually-cued and uncued movement, as well as resultant stimuli. Our findings replicate recent evidence demonstrating reduced N1 amplitude for tones produced by stimulus-driven movement. Despite influencing motor preparation, action-effect contingency was not found to affect N1 amplitudes. Instead, we explore electrophysiological markers suggesting that attentional mechanisms may suppress the neurophysiological response to sound produced by stimulus-driven movement. Our findings demonstrate lateralised parieto-occipital activity that coincides with the auditory N1, corresponds to a reduction in its amplitude, and is topographically consistent with documented effects of attentional suppression. These results provide new insights into sensorimotor coordination and potential mechanisms underlying sensory attenuation.
Learning to recognize new objects in real time in unconstrained environments presents significant challenges for robotic platforms. We present a meta-learning solution to this problem as well as a registered image and events dataset to facilitate work in this domain. Our solution uses interactive motion to isolate the object, and motion-based saliency (from events) to select relevant keypoints from a high-resolution RGB image. Salient keypoints are then passed to a meta-learner to classify the object type. We show that using our interactive isolation and keypoint selection approach, we outperform existing techniques by 6-20%.
Research into juror perceptions regarding the impact of intoxication on eyewitness memory and credibility is scarce for substances other than alcohol. However, jurors are frequently told to draw on their personal beliefs and experience with intoxicating substances to infer their impact on the case. It is therefore important to investigate laypeople’s perceptions regarding witness and victim intoxication across a range of substances, and whether these perceptions are associated with substance familiarity. Participants (n = 470) completed a survey assessing familiarity and use of different substances, as well as perceptions regarding effects on the memory and credibility of intoxicated victims and witnesses. While participants most frequently reported believing that alcohol, hallucinogens, and polysubstance use of alcohol and cannabis have large negative effects on memory, they more frequently reported that they do not know the extent to which cannabis and cocaine affect memory. In addition, attitudes were found to vary based on substance familiarity. Differences with respect to the perceived impact on memory and credibility of various substances have relevance to court proceedings, particularly in terms of voir dire procedures and whether an expert witness may be required to educate the court on the impacts of different forms of intoxication.
Victims, witnesses, and suspects of crime are frequently intoxicated by Alcohol or Other Drugs (AOD) during the event. How intoxication is perceived by investigating officers, and the manner in which this is handled during interview procedures, can affect the quality of information obtained and therefore investigative outcomes. Various factors are likely to contribute to how intoxication is handled during the investigation of a crime, including standard procedures, familiarity with the effects of different substances, and cultural attitudes. While findings with respect to the effect of different substances on memory are still emerging, it is important to investigate whether police beliefs are consistent with available evidence. In this study, Australian and Indonesian police officers were surveyed about their perceptions of memory accuracy and credibility of victims and witnesses intoxicated with various substances (e.g., alcohol, cannabis, amphetamines, and opioids). A higher proportion of Australian police identified larger negative memory effects associated with alcohol intoxication. At the same time, Indonesian police were found to be more likely to report that intoxication with alcohol would make a victim or witness less credible. With regard to timing, across multiple substances, larger proportions of Australian police reported believing that information obtained from witnesses that were still intoxicated would be more accurate than if interviewed after they became sober. It is concluded that, in order to rectify misconceptions about the impact of AOD intoxication on memory and improve investigative practices, both Australian and Indonesian police would benefit from additional training on the effects of intoxication.
Background: We present a novel account of delusion propensity that integrates the roles of working memory (WM), decision criteria, and information gathering biases. This framework emphasises the role of aberrant correlation detection, which leads to the spurious perception of relationships between one's experiences. The frequency of such outcomes is moderated by the scaling of one's decision criteria which, for reasons discussed, must also account for WM capacity. The proposed dysregulated correlation detection account posits that propensity for delusional ideation is influenced by disturbances in this mechanism. Methods: Hypotheses were tested using a novel task that required participants (N = 92) to identify correlation between binary manipulations of simple shapes, presented as sequential pairs. Decision criteria and correlation detection were assessed under a Signal Detection Theory framework, while WM capacity was assessed through the Automated Operation Span Task and delusion propensity was measured using the Peters Delusion Inventory. Structural equation modeling was conducted to evaluate the proposed model. Results: Consistent with the central hypothesis, an interaction between decision criteria and WM was found to contribute significantly to delusion propensity through its effect on correlation detection accuracy. Greater delusion propensity was observed among participants with more liberal decision criteria, which was also in accordance with hypotheses. At the same time, the total effect of WM on delusion propensity was not found to be significant. Conclusions: These findings provide preliminary support for the proposed dysregulated correlation detection account of propensity for delusional ideation.
Fatigue is a common and highly debilitating symptom of multiple sclerosis (MS). This meta-analytic systematic review with detailed narrative synthesis examined randomised-controlled (RCTs) and controlled trials of behavioural and exercise interventions targeting fatigue in adults with MS to assess which treatments offer the most promise in reducing fatigue severity/impact. Medline, EMBASE and PsycInfo electronic databases, amongst others, were searched through to August 2018. Thirty-four trials (12 exercise, 16 behavioural and 6 combined; n = 2,434 participants) met inclusion criteria. Data from 31 studies (n = 1,991 participants) contributed to the meta-analysis. Risk of bias (using the Cochrane tool) and study quality (GRADE) were assessed. The pooled (SMD) end-of-treatment effects on self-reported fatigue were: exercise interventions (n = 13) -.84 (95% CI -1.20 to -.47); behavioural interventions (n = 16) -.37 (95% CI -.53 to -.22); combined interventions (n = 5) -.16 (95% CI: -.36 to .04). Heterogeneity was high overall. Study quality was very low for exercise interventions and moderate for behavioural and combined interventions. Considering health care professional time, subgroup results suggest web-based cognitive behavioural therapy for fatigue, balance and/or multicomponent exercise interventions may be the cost-efficient therapies. These need testing in large RCTs with long-term follow-up to help define an implementable fatigue management pathway in MS.
Self-generated stimuli have been found to elicit a reduced sensory response compared with externally-generated stimuli. However, much of the literature has not adequately controlled for differences in the temporal predictability and temporal control of stimuli. In two experiments, we compared the N1 (and P2) components of the auditory-evoked potential to self- and externally-generated tones that differed with respect to these two factors. In Experiment 1 (n = 42), we found that increasing temporal predictability reduced N1 amplitude in a manner that may often account for the observed reduction in sensory response to self-generated sounds. We also observed that reducing temporal control over the tones resulted in a reduction in N1 amplitude. The contrasting effects of temporal predictability and temporal control on N1 amplitude meant that sensory attenuation prevailed when controlling for each. Experiment 2 (n = 38) explored the potential effect of selective attention on the results of Experiment 1 by modifying task requirements such that similar levels of attention were allocated to the visual stimuli across conditions. The results of Experiment 2 replicated those of Experiment 1, and suggested that the observed effects of temporal control and sensory attenuation were not driven by differences in attention. Given that self- and externally-generated sensations commonly differ with respect to both temporal predictability and temporal control, findings of the present study may necessitate a re-evaluation of the experimental paradigms used to study sensory attenuation.
Background: Fatigue is a common, debilitating symptom of multiple sclerosis (MS) without a current standardised treatment. Objective: The aim of this systematic review with network meta-analyses was to estimate the relative effectiveness of both fatigue-targeted and non-targeted exercise, behavioural and combined (behavioural and exercise) interventions. Methods: Nine electronic databases up to August 2018 were searched, and 113 trials (n = 6909) were included: 34 were fatigue-targeted and 79 non-fatigue-targeted trials. Intervention characteristics were extracted using the Template for Intervention Description and Replication guidelines. Certainty of evidence was assessed using GRADE. Results: Pairwise meta-analyses showed that exercise interventions demonstrated moderate to large effects across subtypes regardless of treatment target, with the largest effect for balance exercise (SMD = 0.84). Cognitive behavioural therapies (CBTs) showed moderate to large effects (SMD = 0.60), with fatigue-targeted treatments showing larger effects than those targeting distress. Network meta-analysis showed that balance exercise performed significantly better compared to other exercise and behavioural intervention subtypes, except CBT. CBT was estimated to be superior to energy conservation and other behavioural interventions. Combined exercise also had a moderate to large effect. Conclusion: Treatment recommendations for balance and combined exercise are tentative as the certainty of the evidence was moderate. The certainty of the evidence for CBT was high.
Joint attention has been identified as a critical component of successful human machine teams. Teaching robots to develop awareness of human cues is an important first step towards attaining and maintaining joint attention. We present a joint attention estimator that creates many possible candidates for joint attention and chooses the most likely object based on a human teammate's hand cues. Our system works within natural human interaction time (< 3 seconds) and above 80% accuracy. Our joint attention estimator provides a meaningful step towards ensuring robots enable human social skills for successful human machine teaming.
Antiretroviral therapy (ART) has significantly improved immune health and survival rates in HIV, but these outcomes rely on near perfect adherence. While many psychosocial factors are related to sub-optimal adherence, effectiveness of associated interventions are modest or inconsistent. The Psychological Flexibility (PF) model underlying Acceptance and Commitment Therapy (ACT) identifies a core set of broadly applicable transdiagnostic processes that may be useful to explain and improve non-adherence. However, PF has not previously been examined in relation to ART adherence. Therefore, this cross-sectional study (n = 275) explored relationships between PF and intentional/unintentional ART non-adherence in people with HIV. Adults with HIV prescribed ART were recruited online. Participants completed online questionnaires assessing self-reported PF, adherence and emotional and general functioning. Logistic regressions examined whether PF processes were associated with intentional/unintentional non-adherence. Fifty-eight percent of participants were classified as nonadherent according to the Medication Adherence Rating Scale, of which 41.0% reported intentional and 94.0% unintentional non-adherence. Correlations between PF and adherence were small. PF did not significantly explain intentional/unintentional non-adherence after controlling for demographic and disease factors. Further clarification of the utility of PF in understanding ART non-adherence is warranted using prospective or experimental designs in conjunction with more objective adherence measures.
ACT-R/S: A Computational and Neurologically Inspired Model of Spatial Reasoning Anthony M. Harrison (anh23@pitt.edu) Christian D. Schunn (schunn@pitt.edu) Department of Psychology, University of Pittsburgh 3939 O’Hara St, Pittsburgh, PA 15260 USA The field of cognitive modeling has seen a recent push in two major areas: embodied cognition, and neurological re- alism. No longer is it sufficient to show that a model of cognition can produce a specific behavior without actually interacting with its environment in someway, be it a real environment or simulated. Nor can psychologists ignore the fact that for every system, representation, rule, and compu- tation proposed there must be some underlying neurological reality behind it. With both these constraints in mind, we set out to develop an extension to ACT-R (Anderson & Le- biere, 1998) allowing it to enter into a three-dimensional world in a neurologically plausible manner. ACT-R/S (spatial) relies specifically upon three process- ing modules, only two of which are new to the architecture. Each of these modules has been shown to be both behavior- ally and neurologically separate. The representations and computations of each of the systems are similarly distinct. Three Visiospatial Systems Visual System The primary function of a visual system is to identify a set of visual features as an object. The visual sys- tem needs to be able to take fine-grained detail and through special processing, recognize an object. A feature of this system is that it is able to perform its task based off of basic two-dimensional retinotopic information. An object’s depth or spatial extent is not typically necessary for its accurate identification. This functionality is currently available in Mike Byrne’s ACT-R/PM (perceptual & motor extension). Neurologically, the visual system is seated in the primary visual areas as well as the ventral visual processing stream which limits processing to fine detail, color perception, local form perception, visual scanning and visual feature analysis (see Previc, 1997 for review). Manipulative System When it comes to grasping and ma- nipulating objects, we need to be able to represent them in a manner that will enable us to effectively prepare the motor system for the task ahead. The manipulative system is con- cerned entirely with a metric, geon-based (Biederman, 1987), three-dimensional representation of objects. These representations are then typically fed to the motor system permitting the development of complex motor programs. The manipulative system can represent almost any three- dimensional object, but its primary purpose is to support actual manual manipulation. The manipulative system relies upon the dorso-lateral visual stream as well as the parietal cortex. The involvement of the parietal cortex is not surprising given that these tasks often involve actual manipulation. However, when subjects are asked to imagine object rotations, the parietal cortex is still often activated (see Previc, 1997 for review). Configural System The configural system is concerned with representing objects in space to facilitate navigation. It represents the world around us as spatial blobs that need to be navigated around, above, or below. Its representations are nowhere near as precise as those found in the manipulative system. It encodes the locations of objects in terms of ego- centric vectors that are continuously updated through path- integration. The utilization of multiple landmarks allows the system to uniquely position itself in space and return to lo- cations at later points in time. The discovery of “place-cells” in the rat hippocampus has been viewed as the definitive location of cognitive-maps in the brain (O’Keefe & Nadel, 1978). Recent research has shown that the parahippocampal regions are more important in primate navigation but they still represent some form of a map of the environment. Our own meta-analysis brings the “egocentric” assumption of “place-cells” into question, hence our usage of egocentric vectors in the configural rep- resentations. Summary With the proposal of two additional processing systems that specialize specifically in three-dimensional processing, it is hoped that we will be able to expand the range of phenome- non that computational cognitive models can represent. We present this not only for the ACT-R architecture, but also so that other architectures might get a foothold in three- dimensional embodiment. References Anderson, J. R., & Lebiere, C. (1998). Atomic components of thought. Mahwah, NJ: Erlbaum. Biederman, I. (1987). Recognition-by-components: A the- ory of human image understanding. Psychological Re- view, 94, 115-117. O’Keefe, J., & Nadel, L. (1978). The hippocampus as a cognitive map. Oxford: Clarendon. Previc, F. H. (1997). The neuropsychology of 3-D space. Psychological bulletin, 124, 123-164.
Social perceivers often view a human agent's norm-violating behavior as diagnostic of that person's mental states, while behaviors that conform to norms are viewed as less informative. We developed a series of stimulus videos depicting a DRC-HUBO robot engaging in norm-violating and norm-conforming behaviors. We explored the hypothesis that robots' norm-violating actions may invite social perceivers to increase their mental state attributions in a similar manner as they do in humans. Surprisingly, we found that norm-conforming behaviors appear to be at least as conducive as norm-violating behaviors, and perhaps even moreso, to mental state attribution to robotic agents.
Automated Detection of Strategies in Free Text Responses Anthony Harrison (anh23@pitt.edu) Lelyn Saner (les53@pitt.edu) Celestine Cookson (clcst70@pitt.edu) Darcie Kunder (dakst67@pitt.edu) Christian D. Schunn (schunn@pitt.edu) Learning Research and Development Center, University of Pittsburgh 3939 O’Hara St., Pittsburgh, PA 15260 USA When solving problems, people often use a wide array of different strategies. Effective teaching often requires isolating what strategies students are using (or not using) in order to more effectively structure the instructional intervention. Nowhere is this truer than in the realm of intelligent adaptive tutors. The classification of strategy use in complex domains presents an interesting challenge to intelligent tutors. This is made even greater if the strategies are to be extracted from free text responses given by the students. To this end, we have been using Latent Semantic Analysis (LSA) as an automatic strategy classification tool. LSA is a computational tool that extracts the co-occurrence of words in a corpus. Through high-dimensional matrix decomposition, LSA is able to produce a “semantic-space” allowing all experienced words, phrases, and sentences to be represented as vectors within that space. The more similar the vectors are to each other, the more similar their meanings. As LSA has matured, some have suggested that it may be a psychologically plausible theory of semantic learning. We remain noncommittal in this regard, choosing instead to rely upon LSA in its original capacity as a fast and efficient text-processing tool. novices to the same vignettes, transform them into vectors in semantic space, and then each of these vectors is compared against those in the databases. Since they are vectors, the cosine between the two serves as a simple similarity score. As similarity increases, the cosine value will approach one. This process yields a ranking of similarities to the descriptive database, where the classified strategy is merely the most similar. Additionally, since we have a sample of strategy exemplars, we can also look at the distribution of similarity scores across strategies. This yields a simple measure of confidence: the greater the number of high similarity matches within a given strategy gets, the more confident we can be that it is representative of that strategy. At this early stage in the development of the system, we were pleased to see that LSA was classifying strategies about as well as our human coders, with almost equivalent inter-rater reliabilities. This is a significant accomplishment given how limited our semantic space is currently (only 100,000+ words, in comparison to the millions of most other LSA corpora), and the limited scale of our descriptive database (10 strategies, approx. 16 exemplars each). Strategy Classification Future Directions Our current endeavor is to use LSA to intelligently classify strategy use in day-to-day military operations. The hope is that by accurately classifying young officers’ strategy uses, we can develop tutoring systems to broaden their range of strategies as well as train them to more appropriately apply the strategies. The strategy classification system relies upon a series of key steps. First the LSA semantic space was generated based on a set of military handbooks, training documents, and pedagogical examples. Free text responses to military scenarios were collected from officers in training as well as experienced military officers. These were then human coded into different strategy categories. The responses were then fed into LSA to generate their vector representations in the semantic space. These two sources yielded two databases of semantically coded (vectors in semantic space) strategies. The novice database (officers in training) is used as a descriptive reference, while the expert database (experienced officers) provides the normative references. The final steps are to take the free text responses of other Aside from increasing the scale of both the semantic space and the reference databases, we hope to begin working on the tutoring system proper. This will mean developing a training system that adapts to the strategy use of the individual to provide sufficient scaffolding to enable them to explore alternative strategies, as well as to learn how to appropriately apply them. Then, as the student progresses through the tutor, the normative database (provided by experienced military officers) will come into greater play. References Landauer, T. K. & Dumais, S. T. (1997). A solution to Plato’s problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge. Psychological Review, 104(2). 211-240. Landauer, T. K., Foltz, P. W., & Laham, D. (1998). Introduction to latent semantic analysis. Discourse Processes, 25, 259-284.
The Sensing, Computing, Interactive Platform for Research Robotics (SCIPRR), is a humanoid, head-shaped sensor package designed to balance both sensing and interaction needs within HRI. Designed to maximize reconfigurability and modularity, SCIPRR can accommodate most sensors and computers with little effort. We report the process of adapting SCIPRR for audio and point-cloud sensors through the use of 3D modeling and printing
We developed and evaluated a novel humanoid head, SCIPRR (Sensing, Computing, Interacting Platform for Robotics Research). SCIPRR is a head shell that was iteratively created with additive manufacturing. SCIPRR contains internal scaffolding that allows sensors, small form computers, and a back-projection system to display an animated face on a front-facing screen. SCIPRR was developed using User Centered Design principles and evaluated using three different methods. First, we created multiple, small-scale prototypes through additive manufacturing and performed polling and refinement of the overall head shape. Second, we performed usability evaluations of expert HRI mechanics as they swapped sensors and computers within the the SCIPRR head. Finally, we ran and analyzed an experiment to evaluate how much novices would like a robot with our head design to perform different social and traditional robot tasks. We made both major and minor changes after each evaluation and iteration. Overall, expert users liked the SCIPRR head and novices wanted a robot with the SCIPRR head to perform more tasks (including social tasks) than a more traditional robot.