Volume 21, Number 1, January 2007 Masaki Hamamoto, Yoshiji Ohta, Keita Hara, Toshiaki Hisada: Application of fluid-structure interaction analysis to flapping flight of insects with deformable wings. 1–21 Guillermo Heredia, Anı́bal Ollero: Stability of autonomous vehicle path tracking with pure delays in the control loop. 23–50 Dong-Hoon Yang, Suk-Kyo Hong: A roadmap construction algorithm for mobile robot path planning using skeleton maps. 51–63 Norihiro Kamamichi, Masaki Yamakita, Takahiro Kozuki, Kinji Asaka, Zhi Wei Luo: Doping effects on robotic systems with ionic polymermetal composite actuators. 65–85 Keehoon Kim, Wan Kyun Chung, Sang Yep Nam: Accurate force reflection method for a multi-d.o.f. haptic interface using instantaneous restriction space without a force sensor in an unstructured environment. 87–104 Amit Agarwal, Meng-Hiot Lim, Meng Joo Er, Nguyen Trung Nghia: Rectilinear workspace partitioning for parallel coverage using multiple unmanned aerial vehicles. 105–120 Yasuyoshi Yokokohji, Satoshi Chaen, Tsuneo Yoshikawa: Evaluation of traversability of wheeled mobile robots on uneven terrains by a fractal terrain model. 121–142 Sehoon Park, Yun-Jung Lee: Discontinuous zigzag gait planning of a quadruped walking robot with a waist-joint. 143–164 Fakhreddine Ababsa, Malik Mallem: Hybrid three-dimensional camera pose estimation using particle filter sensor fusion. 165–181 Heng Wang, K. H. Low, Michael Yu Wang: Virtual circle mapping for master-slave hand systems. 183–208 Johan Bos, Tetsushi Oka: Meaningful conversation with mobile robots. 209–232
In many situations, a set of hard constraints encodes the feasible configurations of some system or product over which multiple users have distinct preferences. However, making suitable decisions requires that the preferences of a specific user for different configurations be articulated or elicited, something generally acknowledged to be onerous. We address two problems associated with preference elicitation: computing a best feasible solution when the user's utilities are imprecisely specified; and developing useful elicitation procedures that reduce utility uncertainty, with minimal user interaction, to a point where (approximately) optimal decisions can be made. Our main contributions are threefold. First, we propose the use of minimax regret as a suitable decision criterion for decision making in the presence of such utility function uncertainty. Second, we devise several different procedures, all relying on mixed integer linear programs, that can be used to compute minimax regret and regret-optimizing solutions effectively. In particular, our methods exploit generalized additive structure in a user's utility function to ensure tractable computation. Third, we propose various elicitation methods that can be used to refine utility uncertainty in such a way as to quickly (i.e., with as few questions as possible) reduce minimax regret. Empirical study suggests that several of these methods are quite successful in minimizing the number of user queries, while remaining computationally practical so as to admit real-time user interaction.
Autonomic (self-managing) computing systems face the critical problem of resource allocation to different computing elements. Adopting a recent model, we view the problem of provisioning resources as involving utility elicitation and optimization to allocate resources given imprecise utility information. In this paper, we propose a new algorithm for regret-based optimization that performs significantly faster than that proposed in earlier work. We also explore new regret-based elicitation heuristics that are able to find near-optimal allocations while requiring a very small amount of utility information from the distributed computing elements. Since regret-computation is intensive, we compare these to the more tractable Nelder-Mead optimization technique w.r.t. amount of utility information required.
We propose new methods of preference elicitation for constraint-based optimization problems based on the use of minimax regret. Specifically, we assume a constraint-based optimization problem (e.g., product configuration) in which the objective function (e.g., consumer preferences) are unknown or imprecisely specified. Assuming a graphical utility model, we describe several elicitation strategies that require the user to answer only binary (bound) queries on the utility model parameters. While a theoretically motivated algorithm can provably reduce regret quickly (in terms of number of queries), we demonstrate that, in practice, heuristic strategies perform much better, and are able to find optimal (or near-optimal) configurations with far fewer queries.
In many situations, a set of hard constraints encodes the feasible configurations of some system or product over which users have preferences. We consider the problem of computing a best feasible solution when the user's utilities are partially known. Assuming bounds on utilities, efficient mixed integer linear programs are devised to compute the solution with minimax regret while exploiting generalized additive structure in a user's utility function.
One of the central challenges in reinforcement learning is to balance the exploration/exploitation tradeoff while scaling up to large problems. Although model-based reinforcement learning has been less prominent than value-based methods in addressing these challenges, recent progress has generated renewed interest in pursuing modelbased approaches: Theoretical work on the exploration/exploitation tradeoff has yielded provably sound model-based algorithms such as E 3 and Rmax, while work on factored MDP representations has yielded model-based algorithms that can scale up to large problems. Recently the benefits of both achievements have been combined in the Factored E3 algorithm of Kearns and Koller. In this paper, we address a significant shortcoming of Factored E3: namely that it requires an oracle planner that cannot be feasibly implemented. We propose an alternative approach that uses a practical approximate planner, approximate linear programming, that maintains desirable properties. Further, we develop an exploration strategy that is targeted toward improving the performance of the linear programming algorithm, rather than an oracle planner. This leads to a simple exploration strategy that visits states relevant to tightening the LP solution, and achieves sample efficiency logarithmic in the size of the problem description. Our experimental results show that the targeted approach performs better than using approximate planning for implementing either Factored E3 or Factored Rmax.
The study of value estimation in Markov reward processes has been dominated by research on temporal difference methods since the introduction of TD(0) in 1988. Temporal difference methods are often contrasted with a maximum likelihood approach where the transition matrix and reward vector are estimated explicitly and converted into a value estimate by solving a matrix equation. It is often asserted that maximum likelihood estimation yields more accurate values, but the temporal difference...
The wedging action fixture to which the invention relates has a wedging action means, a pair of pressing members, wedging action releasing means and other members for limiting the relative movement and transmitting forces between these parts. This fixture can be secured to any desired portion of an elongated supporting member by a wedging action performed by the wedging action means. The external force to be borne is applied to the wedging action means so as to obtain a wedging force which is transmitted to a pair of pressing members. The pressing members acts on both surface of the elongated supporting member so as to cramp the latter therebetween or, alternatively, on both opposing walls of a channeled supporting member to urge the walls away from each other, thereby to bear the external force at any desired position on the elongated supporting member. The fixture can easily be unfastened simply by operating the wedging action releasing means, so that it can be easily moved to any desired position on the elongated supporting member or detached from the latter.
The exploration/exploitation trade-off is a difficult problem for a reinforcement learning agent. A non-stationary environment coupled with current connectionist implementations of reinforcement learning algorithms is a recipe for disaster. Towards a solution for such situations we introduce a novel technique, called past-success directed exploration, and an implementation of reinforcement learning algorithms based on the fuzzy ARTMAP architecture. We compare through experimentation features of a traditional approach with our own
David A. Cohen合作论文数Royal Holloway College, University of London1
Silvana Badaloni合作论文数Department of Information Engineering, University of Padova - Italy1