Robots in shared workspaces must interpret human actions from partial, ambiguous observations, where overconfident early predictions can lead to unsafe or disruptive interaction. This challenge is amplified in egocentric views, where viewpoint changes and occlusions increase perceptual noise and ambiguity. As a result, downstream human-robot interaction modules require not only an action hypothesis but also a trustworthy estimate of confidence under partial observation. Recent vision-language model-based approaches have been proposed for short-term action recognition due to their open-vocabulary and context-aware reasoning, but their uncertainty reliability in the temporal-prefix regime is largely uncharacterized. We present the first systematic evaluation of uncertainty in vision-language model-based short-term action recognition for human-robot interaction. We introduce a temporal-prefix evaluation protocol and metrics for calibration and selective prediction. We also characterize miscalibration patterns and failure modes under partial observations. Our study provides the missing reliability evidence needed to use vision-language model predictions in confidence-gated human-robot interaction modules.
Intent inferencing in teleoperation has been instrumental in aligning operator goals and coordinating actions with robotic partners. However, current intent inference methods often ignore subtle motion that can be strong indicators for a sudden change in intent. Specifically, we aim to tackle 1) if we can detect sudden jumps in operator trajectories, 2) how to appropriately use these sudden jump motions to infer an operators goal state, and 3) how to incorporate these discontinuous and continuous dynamics to infer operator motion. Our framework, called Psychic, models these small indicative motions through a jump-drift-diffusion stochastic differential equation to cover discontinuous and continuous dynamics. Kramers-Moyal (KM) coefficients allow us to detect jumps with a trajectory which we pair with a statistical outlier detection algorithm to nominate goal transitions. Through identifying jumps, we can perform early detection of existing goals and discover undefined goals in unstructured scenarios. Our framework then applies a Sparse Identification of Nonlinear Dynamics (SINDy) model using KM coefficients with the goal transitions as a control input to infer an operators motion behavior in unstructured scenarios. We demonstrate Psychic can produce probabilistic reachability sets and compare our strategy to a negative log-likelihood model fit. We perform a retrospective study on 600 operator trajectories in a hands-free teleoperation task to evaluate the efficacy of our opensource package, Psychic, in both offline and online learning.
ABSTRACT:Acute myeloid leukemia (AML) is a multiclonal disease, existing as a milieu of clones with unique but related genotypes as initiating clones acquire subsequent mutations. However, bulk sequencing cannot fully capture AML clonal architecture or the clonal evolution that occurs as patients undergo therapy. To interrogate clonal evolution, we performed simultaneous single-cell molecular profiling and immunophenotyping on 43 samples from 32 patients with NPM1 (nucleophosmin 1)-mutated AML at different time points in disease progression. Here, we show that diagnosis and relapse AML samples display similar clonal architecture patterns, but signaling mutations drive increased clonal complexity, specifically at relapse, which correlates with overall survival. We uncovered unique genotype-immunophenotype relationships regardless of disease state, suggesting leukemic lineage trajectories can be hard-wired by the mutations present. Analysis of longitudinal samples from patients on front-line AML therapy identified dynamic clonal and immunophenotypic changes consistent with the genotype-immunophenotype relationships we identified.
Advances in single cell multi-omics technologies have allowed for investigation into genotype-immunophenotype relationships and dynamic clonal changes in cancer patient samples. These technologies provide rich insights into genetic profiling, drivers of disease progression, and subclonal dynamics. Increased adoption and throughput of these technologies, necessitate accessible computational tools for processing and analysis. Toward developing easy to use computational tools, we introduce the Single Cell DNA (scDNA) package that allows for rapid analysis of single-cell molecular profiling. Our platform aims to provide a diagnostic summary tool for sample quality control, representative subclonal architecture within a sample, demultiplexing functionality for multi-sample processing, copy number variation, and mutation trajectory analysis to identify order of subclonal mutation acquisition. We showcase a series of vignettes on hematopoiesis datasets for each aim that reflects recently deployed uses of scDNA. Additionally, we show scDNA provides a modular, user-friendly framework that readily feeds into other standard software pipelines.
Humans directly completing tasks in dangerous or hazardous conditions is not always possible where these tasks are increasingly be performed remotely by teleoperated robots. However, teleoperation is difficult since the operator feels a disconnect with the robot caused by missing feedback from several senses, including touch, and the lack of depth in the video feedback presented to the operator. To overcome this problem, the proposed system actively infers the operator's intent and provides assistance based on the predicted intent. Furthermore, a novel method of calculating confidence in the inferred intent modifies the human-in-the-loop control. The operator's gaze is employed to intuitively indicate the target before the manipulation with the robot begins. A potential field method is used to provide a guiding force towards the intended target, and a safety boundary reduces risk of damage. Modifying these assistances based on the confidence level in the operator's intent makes the control more natural, and gives the robot an intuitive understanding of its human master. Initial validation results show the ability of the system to improve accuracy, execution time, and reduce operator error.
Cancer evolution is a multifaceted process leading to dysregulation of cellular expansion and differentiation through somatic mutations and epigenetic dysfunction. Clonal expansion and evolution is driven by cell-intrinsic and -extrinsic selective pressures, which can be captured with increasing resolution by single-cell and bulk DNA sequencing. Despite the extensive genomic alterations revealed in profiling studies, there remain limited experimental systems to model and perturb evolutionary processes. Here, we integrate multi-recombinase tools for reversible, sequential mutagenesis from premalignancy to leukemia. We demonstrate that inducible Flt3 mutations differentially cooperate with Dnmt3a, Idh2, and Npm1 mutant alleles, and that changing the order of mutations influences cellular and transcriptional landscapes. We next use a generalizable, reversible approach to demonstrate that mutation reversion results in rapid leukemic regression with distinct differentiation patterns depending upon co-occurring mutations. These studies provide a path to experimentally model sequential mutagenesis, investigate mechanisms of transformation and probe oncogenic dependency in disease evolution.
Acute myeloid leukemia (AML) with mutations in the tumor suppressor gene, TP53 ( TP53 mut AML), is fatal with a median survival of only 6 months. RNA sequencing on purified AML patient samples show TP53 mut AML has higher expression of mevalonate pathway genes. We retrospectively identified a survival benefit in TP53 mut AML patients who received chemotherapy concurrently with a statin, which inhibits the mevalonate pathway. Mechanistically, TP53 mut AML resistance to standard AML chemotherapy, cytarabine (AraC), correlates with increased mevalonate pathway activity and a mitochondria stress response with increased mitochondria mass and oxidative phosphorylation. Pretreatment with a statin reverses these effects and chemosensitizes TP53 mut AML cell lines and primary samples in vitro and in vivo . Mitochondria-dependent chemoresistance requires the geranylgeranyl pyrophosphate (GGPP) branch of the mevalonate pathway and novel GGPP-dependent synthesis of glutathione to manage AraC-induced reactive oxygen species (ROS). Overall, we show that the mevalonate pathway is a novel therapeutic target in TP53 mut AML. Significance Chemotherapy-persisting TP53 mut AML cells induce a mitochondria stress response that requires mevalonate byproduct, GGPP, through its novel role in glutathione synthesis and regulation of mitochondria metabolism. We provide insight into prior failures of the statin family of mevalonate pathway inhibitors in AML. We identify clinical settings and strategies to successfully target the mevalonate pathway, particularly to address the unmet need of TP53 mut AML. ### Competing Interest Statement CL is a consultant for Bristol Myers Squibb (BMS), AbbVie and Astellas; CL is on advisory boards for Servier, Daiichi, AbbVie, Rigel and Genentech; CL receives research funding from JAZZ and BMS. DAF is a consultant for AbbVie.
The unique challenges in shared control for telemanipulation-beyond teleoperation-include physical discrepancy between human and robot hands and the fine manipulation constraints needed for task success. We present an intuitive shared-control strategy that generates robotic grasp poses better suited for human perception of success and feeling of control while ensuring a stable grasp for task success. The robot's motion follows an arbitration between following the user's motion constraints and accomplishing the inferred task. The arbitration adapts based on the physical discrepancy between the human and robot hands. We have conducted a user study with a telemanipulation scenario to analyze the effects of task predictability, following, and user preference. The results demonstrated that intent-based approaches provide advantages over direct motion mapping strategies.
In-hand manipulation is challenging for a multi-finger robotic hand due to its high degrees of freedom and complex interaction with the object. To enable in-hand manipulation, existing deep reinforcement learning-based approaches mainly focus on training a single robot-structure-specific policy through the centralized learning mechanism, lacking adaptability to changes like robot malfunction. To solve this limitation, this work treats each finger as an individual agent and trains multiple agents to control their assigned fingers to complete the in-hand manipulation task cooperatively. We propose the Multi-Agent Global-Observation Critic and Local-Observation Actor (MAGCLA) method, where the critic can observe all agents' actions globally, and the actor only locally observes its neighbors' actions. Besides, conventional individual experience replay may cause unstable cooperation due to the asynchronous performance increment of each agent, which is critical for in-hand manipulation tasks. To solve this issue, we propose the Synchronized Hindsight Experience Replay (SHER) method to synchronize and efficiently reuse the replayed experience across all agents. The methods are evaluated in two in-hand manipulation tasks on the Shadow dexterous hand. The results show that SHER helps MAGCLA achieve comparable learning efficiency to a single policy, and the MAGCLA approach is more generalizable in different tasks. The trained policies have higher adaptability in the robot malfunction test compared to the baseline multi-agent and single-agent approaches.
One of the fundamental questions in shared control is how to allocate control power to the human and robot effectively. Conventional arbitration policies often define a uniform singular scalar for all 6 DOFs to blend human input and robot assistance. However, this singular scalar over-dominates some dimensions of the inputs and provides insufficient assistance in other dimensions. Thus, current shared control can support simple telemanipulation tasks such as pushing, pressing, and simple positional control but is limited in tasks with more DOFs like rotational motion. A dimension-specific arbitration policy is developed to customize the control arbitration along each DOF to fill the gap. It looks at whether the robotic assistance is too timid or aggressive along each DOF and determines the arbitration magnitude according to disagreement levels of control allocation and the user's willingness to accept assistance. The user's willingness is estimated from a feedback psychology model. The method has higher similarity and ratio of agreement between the human and robot (lower over-dominance) over existing methods and, simultaneously, improves the task performance. This arbitration strategy is expected to increase the adoption of teleoperation for object manipulation.
Filter-based shared control aims to accept and augment an operator's ability to control a robot. Current solutions accept actions based on their direction aligning with the robot's optimal policy. These strategies reject a human's small corrective actions if they conflict with the robot's direction and accept too aggressive actions as long as they are consistent with the robot's direction. Such strategies may cause task failures and the operator's feeling of loss of control. To close the gap, we propose WE-Filter, which has flexible, adaptive criteria allowing the operator's small corrective actions and tempering too aggressive ones. Inspired by classical work-energy impact problems between two dynamic, interactive bodies, both inputs' properties (direction and magnitude) are inherently considered, creating intuitive, adaptive bounds to accept sensible actions. The model identifies behaviors before and after impact. The rationale is that each timestep of shared control acts as an impact between the operator's and the robot's policies, where post-impact behaviors depend on their previous behaviors. As time continues, a series of impacts occur. The aim is to minimize impacts that occur to reach an agreement faster and reduce strong reactionary behaviors. Our model determines flexible acceptance criteria to bound a mismatch of magnitude and finds a replacement action for conflicting policies. The WE-Filter achieves better task performance, the ratio of accepted actions, and action similarity than the existing methods.
Abstract — Learning-based grasping can afford real-time motion planning of multi-fingered robotics hands thanks to its high computational efficiency. However, it needs to explore large search spaces during its learning process. The search space causes low learning efficiency, which has been the main barrier to its practical adoption. In addition, the generalizability of the trained policy is limited unless they are identical or similar to the trained objects. In this work, we develop a novel Physics-Guided Deep Reinforcement Learning with a Hierarchical Reward Mechanism to improve the learning efficiency and generalizability for learning-based autonomous grasping. Unlike conventional observation-based grasp learning, physics-informed metrics are utilized to convey correlations between features associated with hand structures and objects to improve learning efficiency and outcomes. Further, a hierarchical reward mechanism is developed to enable the robot to learn the grasping task in a prioritized way. It is validated in grasping tasks with a MICO robot arm in both simulation and physical experiments. The results show that our method outperformed the standard Deep Reinforcement learning method in task performance by 48% and learning efficiency by 40%.
Learning-based grasping can afford real-time grasp motion planning of multi-fingered robotics hands thanks to its high computational efficiency. However, learning-based methods are required to explore large search spaces during the learning process. The search space causes low learning efficiency, which has been the main barrier to its practical adoption. In addition, the trained policy lacks a generalizable outcome unless objects are identical to the trained objects. In this work, we develop a novel Physics-Guided Deep Reinforcement Learning with a Hierarchical Reward Mechanism to improve learning efficiency and generalizability for learning-based autonomous grasping. Unlike conventional observation-based grasp learning, physics-informed metrics are utilized to convey correlations between features associated with hand structures and objects to improve learning efficiency and outcomes. Further, the hierarchical reward mechanism enables the robot to learn prioritized components of the grasping tasks. Our method is validated in robotic grasping tasks with a 3-finger MICO robot arm. The results show that our method outperformed the standard Deep Reinforcement Learning methods in various robotic grasping tasks.
In human-robot cooperation, the robot cooperates with humans to accomplish the task together. Existing approaches assume the human has a specific goal during the cooperation, and the robot infers and acts toward it. However, in real-world environments, a human usually only has a general goal (e.g., general direction or area in motion planning) at the beginning of the cooperation, which needs to be clarified to a specific goal (i.e., an exact position) during cooperation. The specification process is interactive and dynamic, which depends on the environment and the partner’s behavior. The robot that does not consider the goal specification process may cause frustration to the human partner, elongate the time to come to an agreement, and compromise team performance. This work presents the Evolutionary Value Learning approach to model the dynamics of the goal specification process with State-based Multivariate Bayesian Inference and goal specificity-related features. This model enables the robot to enhance the process of the human’s goal specification actively and find a cooperative policy in a Deep Reinforcement Learning manner. Our method outperforms existing methods with faster goal specification processes and better team performance in a dynamic ball balancing task with real human subjects.
Enabling robots to provide effective assistance yet still accommodating the operator’s commands for telemanipulation of an object is very challenging because robot’s assistance is not always intuitive for human operators and human behaviors and preferences are sometimes ambiguous for the robot to interpret. Due to the difference in hand structures, some motion assistance from the robot may surprise the operator with counter-intuitive movements, which could introduce more burden to the human to correct the actions and/or reduce the operator’s sense of system control. To address these problems, we developed a novel preference-aware assistance knowledge learning approach. An assistance preference model learns what assistance is preferred by a human, and a stage-wise model updating method ensures the learning stability while dealing with the ambiguity of human preference data. Such a preference-aware assistance knowledge enables a teleoperated robot hand to provide more active yet preferred assistance toward manipulation success. We also developed knowledge transfer methods to transfer the preference knowledge across different robot hand structures to avoid extensive robot-specific training. Experiments to telemanipulate a 3-finger hand and 2-finger hand, respectively, to use, move, and hand over a cup have been conducted. Results demonstrated that the methods enabled the robots to effectively learn the preference knowledge and allowed knowledge transfer between robots with less training effort.
Although data-driven motion mapping methods are promising to allow intuitive robot control and teleoperation that generate human-like robot movement, they normally require tedious pair-wise training for each specific human and robot pair. This paper proposes a transferability-based mapping scheme to allow new robot and human input systems to leverage the mapping of existing trained pairs to form a mapping transfer chain, which will reduce the number of new pair-specific mappings that need to be generated. The first part of the mapping schematic is the development of a Synergy Mapping via Dual-Autoencoder (SyDa) method. This method uses the latent features from two autoencoders to extract the common synergy of the two agents. Secondly, a transferability metric is created that approximates how well the mapping between a pair of agents will perform compared to another pair before creating the motion mapping models. Thus, it can guide the formation of an optimal mapping chain for the new human-robot pair. Experiments with human subjects and a Pepper robot demonstrated 1) The SyDa method improves the accuracy and generalizability of the pair mappings, 2) the SyDa method allows for bidirectional mapping that does not prioritize the direction of mapping motion, and 3) the transferability metric measures how compatible two agents are for accurate teleoperation. The combination of the SyDa method and transferability metric creates generalizable and accurate mapping need to create the transfer mapping chain.
Autonomous grasping is challenging due to the high computational cost caused by multi-fingered robotic hands and their interactions with objects. Various analytical methods have been developed yet their high computational cost limits the adoption in real-world applications. Learning-based grasping can afford real-time motion planning thanks to its high computational efficiency. However, it needs to explore large search spaces during its learning process. The search space causes low learning efficiency, which has been the main barrier to its practical adoption. In this work, we develop a novel Physics-Guided Deep Reinforcement Learning with a Hierarchical Reward Mechanism, which combines the benefits of both analytical methods and learning-based methods for autonomous grasping. Different from conventional observation-based grasp learning, physics-informed metrics are utilized to convey correlations between features associated with hand structures and objects to improve learning efficiency and learning outcomes. Further, a hierarchical reward mechanism is developed to enable the robot to learn the grasping task in a prioritized way. It is validated in a grasping task with a MICO robot arm in simulation and physical experiments. The results show that our method outperformed the baseline in task performance by 48% and learning efficiency by 40%.
Applying Deep Reinforcement Learning (DRL) to Human-Robot Cooperation (HRC) in dynamic control problems is promising yet challenging as the robot needs to learn the dynamics of the controlled system and dynamics of the human partner. In existing research, the robot powered by DRL adopts coupled observation of the environment and the human partner to learn both dynamics simultaneously. However, such a learning strategy is limited in terms of learning efficiency and team performance. This work proposes a novel task decomposition method with a hierarchical reward mechanism that enables the robot to learn the hierarchical dynamic control task separately from learning the human partner's behavior. The method is validated with a hierarchical control task in a simulated environment with human subject experiments. Our method also provides insight into the design of the learning strategy for HRC. The results show that the robot should learn the task first to achieve higher team performance and learn the human first to achieve higher learning efficiency.
Seok Won Lee合作论文数University of Nebraska Lincoln3