Arrow Research search

Author name cluster

Emanuel Todorov

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

35 papers
2 author rows

Possible papers

35

ICRA Conference 2019 Conference Paper

Learning Deep Visuomotor Policies for Dexterous Hand Manipulation

  • Divye Jain
  • Andrew Li
  • Shivam Singhal
  • Aravind Rajeswaran
  • Vikash Kumar
  • Emanuel Todorov

Multi-fingered dexterous hands are versatile and capable of acquiring a diverse set of skills such as grasping, in-hand manipulation, and tool use. To fully utilize their versatility in real-world scenarios, we require algorithms and policies that can control them using on-board sensing capabilities, without relying on external tracking or motion capture systems. Cameras and tactile sensors are the most widely used on-board sensors that do not require instrumentation of the world. In this work, we demonstrate an imitation learning based approach to train deep visuomotor policies for a variety of manipulation tasks with a simulated five fingered dexterous hand. These policies directly control the hand using high dimensional visual observations of the world and propreoceptive observations from the robot, and can be trained efficiently with a few hundred expert demonstration trajectories. We also find that using touch sensing information enables faster learning and better asymptotic performance for tasks with high degree of occlusions. Video demonstration of our results are available at: https://sites.google.com/view/hand-vil/

ICRA Conference 2018 Conference Paper

Goal Directed Dynamics

  • Emanuel Todorov

We develop a general control framework where a low-level optimizer is built into the robot dynamics. This optimizer together with the robot constitute a goal directed dynamical system, controlled on a higher level. The high level command is a cost function. It can encode desired accelerations, end-effector poses, center of pressure, and other intuitive features that have been studied before. Unlike the currently popular quadratic programming framework, which comes with performance guarantees at the expense of modeling flexibility, the optimization problem we solve at each time step is non-convex and non-smooth. Nevertheless, by exploiting the unique properties of the soft-constraint physics model we have recently developed, we are able to design an efficient solver for goal directed dynamics. It is only two times slower than the forward dynamics solver, and is much faster than real time. The simulation results reveal that complex movements can be generated via greedy optimization of simple costs. This new computational infrastructure can facilitate teleoperation, feature-based control, deep learning of control policies, and trajectory optimization. It will become a standard feature in future releases of the MuJoCo simulator.

NeurIPS Conference 2017 Conference Paper

Towards Generalization and Simplicity in Continuous Control

  • Aravind Rajeswaran
  • Kendall Lowrey
  • Emanuel Todorov
  • Sham Kakade

The remarkable successes of deep learning in speech recognition and computer vision have motivated efforts to adapt similar techniques to other problem domains, including reinforcement learning (RL). Consequently, RL methods have produced rich motor behaviors on simulated robot tasks, with their success largely attributed to the use of multi-layer neural networks. This work is among the first to carefully study what might be responsible for these recent advancements. Our main result calls this emerging narrative into question by showing that much simpler architectures -- based on linear and RBF parameterizations -- achieve comparable performance to state of the art results. We not only study different policy representations with regard to performance measures at hand, but also towards robustness to external perturbations. We again find that the learned neural network policies --- under the standard training scenarios --- are no more robust than linear (or RBF) policies; in fact, all three are remarkably brittle. Finally, we then directly modify the training scenarios in order to favor more robust policies, and we again do not find a compelling case to favor multi-layer architectures. Overall, this study suggests that multi-layer architectures should not be the default choice, unless a side-by-side comparison to simpler architectures shows otherwise. More generally, we hope that these results lead to more interest in carefully studying the architectural choices, and associated trade-offs, for training generalizable and robust policies.

ICRA Conference 2016 Conference Paper

Design of a highly biomimetic anthropomorphic robotic hand towards artificial limb regeneration

  • Zhe Xu 0002
  • Emanuel Todorov

A wide range of research areas, from telemanipulation in robotics to limb regeneration in tissue engineering, could benefit from an anthropomorphic robotic hand that mimics the salient features of the human hand. The challenges of designing such a robotic hand are mainly resulted from our limited understanding of the human hand from engineering point of view and our ability to replicate the important biomechanical features with conventional mechanical design. We believe that the biomechanics of human hand is an essential component of the hand dexterity and can be replicated with highly biomimetic design. To this end, we reinterpret the important biomechanical advantages of the human hand from roboticist's perspective and design a biomimetic robotic hand that closely mimics its human counterpart with artificial joint capsules, crocheted ligaments and tendons, laser-cut extensor hood, and elastic pulley mechanisms. We experimentally identify the workspaces of the fingertips and successfully demonstrate that our proofof- concept design can be teleoperated to grasp and manipulate daily objects with a variety of natural hand postures based on hand taxonomy.

ICRA Conference 2016 Conference Paper

Optimal control with learned local models: Application to dexterous manipulation

  • Vikash Kumar
  • Emanuel Todorov
  • Sergey Levine

We describe a method for learning dexterous manipulation skills with a pneumatically-actuated tendon-driven 24-DoF hand. The method combines iteratively refitted time-varying linear models with trajectory optimization, and can be seen as an instance of model-based reinforcement learning or as adaptive optimal control. Its appeal lies in the ability to handle challenging problems with surprisingly little data. We show that we can achieve sample-efficient learning of tasks that involve intermittent contact dynamics and under-actuation. Furthermore, we can control the hand directly at the level of the pneumatic valves, without the use of a prior model that describes the relationship between valve commands and joint torques. We compare results from learning in simulation and on the physical system. Even though the learned policies are local, they are able to control the system in the face of substantial variability in initial state.

IROS Conference 2015 Conference Paper

Ensemble-CIO: Full-body dynamic motion planning that transfers to physical humanoids

  • Igor Mordatch
  • Kendall Lowrey
  • Emanuel Todorov

While a lot of progress has recently been made in dynamic motion planning for humanoid robots, much of this work has remained limited to simulation. Here we show that executing the resulting trajectories on a Darwin-OP robot, even with local feedback derived from the optimizer, does not result in stable movements. We then develop a new trajectory optimization method, adapting our earlier CIO algorithm to plan through ensembles of perturbed models. This makes the plan robust to model uncertainty, and leads to successful execution on the robot. We obtain a high rate of task completion without trajectory divergence (falling) in dynamic forward walking, sideways walking, and turning, and a similarly high success rate in getting up from the floor (the robot broke before we could quantify the latter). Even though the planning is still done offline, the present work represents a significant step towards automating the tedious scripting of complex movements.

NeurIPS Conference 2015 Conference Paper

Interactive Control of Diverse Complex Characters with Neural Networks

  • Igor Mordatch
  • Kendall Lowrey
  • Galen Andrew
  • Zoran Popovic
  • Emanuel Todorov

We present a method for training recurrent neural networks to act as near-optimal feedback controllers. It is able to generate stable and realistic behaviors for a range of dynamical systems and tasks -- swimming, flying, biped and quadruped walking with different body morphologies. It does not require motion capture or task-specific features or state machines. The controller is a neural network, having a large number of feed-forward units that learn elaborate state-action mappings, and a small number of recurrent units that implement memory states beyond the physical system state. The action generated by the network is defined as velocity. Thus the network is not learning a control policy, but rather the dynamics under an implicit policy. Essential features of the method include interleaving supervised learning with trajectory optimization, injecting noise during training, training for unexpected changes in the task specification, and using the trajectory optimizer to obtain optimal feedback gains in addition to optimal actions.

ICRA Conference 2015 Conference Paper

Simulation tools for model-based robotics: Comparison of Bullet, Havok, MuJoCo, ODE and PhysX

  • Tom Erez
  • Yuval Tassa
  • Emanuel Todorov

There is growing need for software tools that can accurately simulate the complex dynamics of modern robots. While a number of candidates exist, the field is fragmented. It is difficult to select the best tool for a given project, or to predict how much effort will be needed and what the ultimate simulation performance will be. Here we introduce new quantitative measures of simulation performance, focusing on the numerical challenges that are typical for robotics as opposed to multi-body dynamics and gaming. We then present extensive simulation results, obtained within a new software framework for instantiating the same model in multiple engines and running side-by-side comparisons. Overall we find that each engine performs best on the type of system it was designed and optimized for: MuJoCo wins the robotics-related tests, while the gaming engines win the gaming-related tests without a clear leader among them. The simulations are illustrated in the accompanying movie.

IROS Conference 2015 Conference Paper

Whole-body model-predictive control applied to the HRP-2 humanoid

  • Jonas Koenemann
  • Andrea Del Prete
  • Yuval Tassa
  • Emanuel Todorov
  • Olivier Stasse
  • Maren Bennewitz
  • Nicolas Mansard

Controlling the robot with a permanently-updated optimal trajectory, also known as model predictive control, is the Holy Grail of whole-body motion generation. Before obtaining it, several challenges should be faced: computation cost, non-linear local minima, algorithm stability, etc. In this paper, we address the problem of applying the updated optimal control in real-time on the physical robot. In particular, we focus on the problems raised by the delays due to computation and by the differences between the real robot and the simulated model. Based on the optimal-control solver MuJoCo, we implemented a complete model-predictive controller and we applied it in real-time on the physical HRP-2 robot. It is the first time that such a whole-body model predictive controller is applied in real-time on a complex dynamic robot. Aside from the technical contributions cited above, the main contribution of this paper is to report the experimental results of this première implementation.

ICRA Conference 2014 Conference Paper

Control-limited differential dynamic programming

  • Yuval Tassa
  • Nicolas Mansard
  • Emanuel Todorov

Trajectory optimizers are a powerful class of methods for generating goal-directed robot motion. Differential Dynamic Programming (DDP) is an indirect method which optimizes only over the unconstrained control-space and is therefore fast enough to allow real-time control of a full humanoid robot on modern computers. Although indirect methods automatically take into account state constraints, control limits pose a difficulty. This is particularly problematic when an expensive robot is strong enough to break itself. In this paper, we demonstrate that simple heuristics used to enforce limits (clamping and penalizing) are not efficient in general. We then propose a generalization of DDP which accommodates box inequality constraints on the controls, without significantly sacrificing convergence quality or computational effort. We apply our algorithm to three simulated problems, including the 36-DoF HRP-2 robot. A movie of our results can be found here goo. gl/eeiMnn.

ICRA Conference 2014 Conference Paper

Convex and analytically-invertible dynamics with contacts and constraints: Theory and implementation in MuJoCo

  • Emanuel Todorov

We describe a full-featured simulation pipeline implemented in the MuJoCo physics engine. It includes multi-joint dynamics in generalized coordinates, holonomic constraints, dry joint friction, joint and tendon limits, frictionless and frictional contacts that can have sliding, torsional and rolling friction. The forward dynamics of a 27-dof humanoid with 10 contacts are evaluated in 0. 1 msec. Since the simulation is stable at 10 msec timesteps, it can run 100 times faster than real-time on a single core of a desktop processor. Furthermore the entire simulation pipeline can be inverted analytically, an order-of-magnitude faster than the corresponding forward dynamics. We soften all constraints, in a way that avoids instabilities and unrealistic penetrations associated with earlier spring-damper methods and yet is sufficient to allow inversion. Constraints are imposed via impulses, using an extended version of the velocity-stepping approach. For holomonic constraints the extension involves a soft version of the Gauss principle. For all other constraints we extend our earlier work on complementarity-free contact dynamics — which were already known to be invertible via an iterative solver — and develop a new formulation allowing analytical inversion.

ICRA Conference 2014 Conference Paper

Design, optimization, calibration, and a case study of a 3D-printed, low-cost fingertip sensor for robotic manipulation

  • Zhe Xu 0002
  • Svetoslav Kolev
  • Emanuel Todorov

We describe a low-cost 3-axis fingertip force sensor for robotic manipulation. Our design makes the most of 3D printing technology, and takes important factors such as maintainability and modification into consideration. The resulting sensor features a detachable fingertip made of 3D-printed materials, and a cantilever mechanism that allows the detection of contact forces via three off-the-shelf, low-cost force sensors. To improve our design concept, optimization on the configuration of the fingertip sensor is performed under statistical analysis of the hysteresis performance. The optimized fingertip sensor is experimentally investigated and calibrated. At the end, through a case-study, we demonstrate that our proposed design can measure the direction of contact forces in the radial plane of the fingertip sensor.

IROS Conference 2014 Conference Paper

Physically-consistent sensor fusion in contact-rich behaviors

  • Kendall Lowrey
  • Svetoslav Kolev
  • Yuval Tassa
  • Tom Erez
  • Emanuel Todorov

We describe an accurate approach to state estimation which fuses any available sensor data with physical consistency priors. This is done by combining the advantages of recursive estimation and fixed-lag smoothing: at each step we re-estimate the trajectory over a time window into the past, but also use a recursive prior obtained from the previous time step via internal simulation. We also incorporate a physics engine into the estimator, which makes it possible to adjust the state estimates so that the inferred contact interactions are consistent with the observed accelerations. The estimator can utilize contact sensors to improve accuracy, but even in the absence of such sensors it reasons correctly about contact forces. Estimation speed and accuracy are demonstrated on a 28-DOF humanoid robot (Darwin) in a walking task. Timing tests and leave-one-out cross-validation show that the proposed approach can be used in real-time and is substantially more accurate than the EKF, without any over-fitting. A video of our results is attached.

ICRA Conference 2014 Conference Paper

Real-time behaviour synthesis for dynamic hand-manipulation

  • Vikash Kumar
  • Yuval Tassa
  • Tom Erez
  • Emanuel Todorov

Dexterous hand manipulation is one of the most complex types of biological movement, and has proven very difficult to replicate in robots. The usual approaches to robotic control — following pre-defined trajectories or planning online with reduced models — are both inapplicable. Dexterous manipulation is so sensitive to small variations in contact force and object location that it seems to require online planning without any simplifications. Here we demonstrate for the first time online planning (or model-predictive control) with a full physics model of a humanoid hand, with 28 degrees of freedom and 48 pneumatic actuators. We augment the actuation space with motor synergies which speed up optimization without removing dexterity. Most of our results are in simulation, showing non-prehensile object manipulation as well as typing. In both cases the input to the system is a high level task description, while all details of the hand movement emerge online from fully automated numerical optimization.

UAI Conference 2014 Conference Paper

Universal Convexification via Risk-Aversion

  • Krishnamurthy Dvijotham
  • Maryam Fazel
  • Emanuel Todorov

We develop a framework for convexifying a general class of optimization problems. We analyze the suboptimality of the solution to the convexified problem relative to the original nonconvex problem, and prove additive approximation guarantees under some assumptions. In simple settings, the convexification procedure can be applied directly and standard optimization methods can be used. In the general case we rely on stochastic gradient algorithms, whose convergence rate can be bounded using the convexity of the underlying optimization problem. We then extend the framework to a general class of discretetime dynamical systems where our convexification approach falls under the paradigm of risk-sensitive Markov Decision Processes. We derive the first model-based and modelfree policy gradient optimization algorithms with guaranteed convergence to the optimal solution. We also present numerical results in different machine learning applications.

ICRA Conference 2013 Conference Paper

Fast, strong and compliant pneumatic actuation for dexterous tendon-driven hands

  • Vikash Kumar
  • Zhe Xu 0002
  • Emanuel Todorov

We describe a pneumatic actuation system for dexterous robotic hands. It was motivated by our desire to improve the ShadowHand system, yet it is quite universal and indeed we are already using it with a second robotic hand we have developed. Our actuation system allows us to move the ShadowHand skeleton faster than a human hand (70 msec limit-to-limit movement, 30 msec overall reflex latency), generate sufficient forces (40 N at each finger tendon, 125N at each wrist tendon), and achieve high compliance on the mechanism level (6 grams of external force at the fingertip displaces the finger when the system is powered.) This combination of speed, force and compliance is a prerequisite for dexterous manipulation, yet it has never before been achieved with a tendon-driven system, let alone a system with 24 degrees of freedom and 40 tendons.

IROS Conference 2012 Conference Paper

MuJoCo: A physics engine for model-based control

  • Emanuel Todorov
  • Tom Erez
  • Yuval Tassa

We describe a new physics engine tailored to model-based control. Multi-joint dynamics are represented in generalized coordinates and computed via recursive algorithms. Contact responses are computed via efficient new algorithms we have developed, based on the modern velocity-stepping approach which avoids the difficulties with spring-dampers. Models are specified using either a high-level C++ API or an intuitive XML file format. A built-in compiler transforms the user model into an optimized data structure used for runtime computation. The engine can compute both forward and inverse dynamics. The latter are well-defined even in the presence of contacts and equality constraints. The model can include tendon wrapping as well as actuator activation states (e. g. pneumatic cylinders or muscles). To facilitate optimal control applications and in particular sampling and finite differencing, the dynamics can be evaluated for different states and controls in parallel. Around 400, 000 dynamics evaluations per second are possible on a 12-core machine, for a 3D homanoid with 18 dofs and 6 active contacts. We have already used the engine in a number of control applications. It will soon be made publicly available.

ICRA Conference 2012 Conference Paper

Reduced dimensionality control for the ACT hand

  • Mark Malhotra
  • Eric Rombokas
  • Evangelos A. Theodorou
  • Emanuel Todorov
  • Yoky Matsuoka

Redundant tendon-driven systems such as the human hand or the ACT robotic hand are high-dimensional and nonlinear systems that make traditional control strategies ineffective. The synergy hypothesis from neuroscience suggests that employing dimensionality reduction techniques can simplify the system without a major loss in function. We define a dimensionality reduction framework consisting of separate observation and activation synergies, a first-order model, and an optimal controller. The framework is implemented for two example tasks: adaptive control of thumb posture and hybrid position/force control to enable dynamic handwriting.

IROS Conference 2012 Conference Paper

Synthesis and stabilization of complex behaviors through online trajectory optimization

  • Yuval Tassa
  • Tom Erez
  • Emanuel Todorov

We present an online trajectory optimization method and software platform applicable to complex humanoid robots performing challenging tasks such as getting up from an arbitrary pose on the ground and recovering from large disturbances using dexterous acrobatic maneuvers. The resulting behaviors, illustrated in the attached video, are computed only 7 × slower than real time, on a standard PC. The video also shows results on the acrobot problem, planar swimming and one-legged hopping. These simpler problems can already be solved in real time, without pre-computing anything.

ICRA Conference 2012 Conference Paper

Tendon-driven control of biomechanical and robotic systems: A path integral reinforcement learning approach

  • Eric Rombokas
  • Evangelos A. Theodorou
  • Mark Malhotra
  • Emanuel Todorov
  • Yoky Matsuoka

We apply path integral reinforcement learning to a biomechanically accurate dynamics model of the index finger and then to the Anatomically Correct Testbed (ACT) robotic hand. We illustrate the applicability of Policy Improvement with Path Integrals (PI 2 ) to parameterized and non-parameterized control policies. This method is based on sampling variations in control, executing them in the real world, and minimizing a cost function on the resulting performance. Iteratively improving the control policy based on real-world performance requires no direct modeling of tendon network nonlinearities and contact transitions, allowing improved task performance.

IROS Conference 2012 Conference Paper

Trajectory optimization for domains with contacts using inverse dynamics

  • Tom Erez
  • Emanuel Todorov

This paper presents an algorithm for direct trajectory optimization in domains with contact. Since contacts and other unilateral constraints may introduce non-smooth dynamics, many standard algorithms of optimal control and reinforcement learning cannot be directly applied to such domains. We use a smooth contact model that can compute inverse dynamics through the contact, thereby avoiding hybrid representation of the non-smooth contact state. This allows us to formulate an unconstrained, continuous trajectory optimization problem, which can be solved using standard optimization tools. We demonstrate our approach by optimizing a running gait for a 31-dimensional simulated humanoid. The resulting gait is demonstrated in a movie attached as supplementary material. The optimization result exhibits a synchronous motion of the arm and the opposite leg, eliminating undesired angular momentum; this is a key feature of bipedal running, and its emergence attests to the power of the optimization process.

ICRA Conference 2011 Conference Paper

A convex, smooth and invertible contact model for trajectory optimization

  • Emanuel Todorov

Trajectory optimization is done most efficiently when an inverse dynamics model is available. Here we develop the first model of contact dynamics defined in both the forward and inverse directions. The contact impulse is the solution to a convex optimization problem: minimize kinetic energy in contact space subject to non-penetration and friction-cone constraints. We use a custom interior-point method to make the optimization problem unconstrained; this is key to defining the forward and inverse dynamics in a consistent way. The resulting model has a parameter which sets the amount of contact smoothing, facilitating continuation methods for optimization. We implemented the proposed contact solver in our new physics engine (MuJoCo). A full Newton step of trajectory optimization for a 3D walking gait takes only 160 msec, on a 12-core PC.

ICRA Conference 2011 Conference Paper

Design and analysis of an artificial finger joint for anthropomorphic robotic hands

  • Zhe Xu 0002
  • Emanuel Todorov
  • Brian Dellon
  • Yoky Matsuoka

In order to further understand what physiological characteristics make a human hand irreplaceable for many dexterous tasks, it is necessary to develop artificial joints that are anatomically correct while sharing similar dynamic features. In this paper, we address the problem of designing a two degree of freedom metacarpophalangeal (MCP) joint of an index finger. The artificial MCP joint is composed of a ball joint, crocheted ligaments, and a silicon rubber sleeve which as a whole provides the functions required of a human finger joint. We quantitatively validate the efficacy of the artificial joint by comparing its dynamic characteristics with that of two human subjects' index fingers by analyzing their impulse response with linear regression. Design parameters of the artificial joint are varied to highlight their effect on the joint's dynamics. A modified, second-order model is fit which accounts for non-linear stiffness and damping, and a higher order model is considered. Good fits are observed both in the human (R 2 = 0. 97) and the artificial joint of the index finger (R 2 = 0. 95). Parameter estimates of stiffness and damping for the artificial joint are found to be similar to those in the literature, indicating our new joint is a good approximation for an index finger's MCP joint.

ICRA Conference 2011 Conference Paper

First-exit model predictive control of fast discontinuous dynamics: Application to ball bouncing

  • Paul Kulchenko
  • Emanuel Todorov

We extend model-predictive control so as to make it applicable to robotic tasks such as legged locomotion, hand manipulation and ball bouncing. The online optimal control problem is defined in a first-exit rather than the usual finite-horizon setting. The exit manifold corresponds to changes in contact state. In this way the need for online optimization through dynamic discontinuities is avoided. Instead the effects of discontinuities are incorporated in a final cost which is tuned offline. The new method is demonstrated on the task of 3D ball bouncing. Even though our robot is mechanically limited, it bounces one ball robustly and recovers from a wide range of disturbances, and can also bounce two balls with the same paddle. This is possible due to intelligent responses computed online, without relying on pre-existing plans.

ICRA Conference 2011 Conference Paper

Modular bio-mimetic robots that can interact with the world the way we do

  • Alex Simpkins
  • Michael S. Kelley
  • Emanuel Todorov

The study of human sensorimotor control and learning through robotic devices requires systems that possess bio-mimetic characteristics which allow them to interact with the world in a similar fashion dynamically (this includes backdrivability, high bandwidth, implementability of complex control algorithms, force feedback, low friction, low inertia, robustness, autonomy, appropriate strength to weight ratio in the case of locomotion, safety, among others). Traditionally, engineered systems have evolved based upon industrial requirements, which are quite different. From the context of studying and mimicking the way humans perform control of their bodies, we state that there are specific groups of characteristics that are essential for use in researching sensorimotor control and learning. Designing robotic systems from this perspective necessitates solving new constrained design challenges that integrate with, adapt, and improve upon known approaches. A new design for a bio-mimetic backdrivable modular robot finger which addresses these challenges is presented. Results are presented demonstrating its effectiveness. The novelty of this robot is not only in the integrative design approach, which meets all the constraints presented without compromise, but also as the first bio-mimetic robot to integrate modularity, and it is highly compact. Additionally, the system has the capacity and bandwidth to run real-time control algorithms that can eventually model human behavior and perform complex manipulation tasks.

ICRA Conference 2010 Conference Paper

Implicit nonlinear complementarity: A new approach to contact dynamics

  • Emanuel Todorov

Contact dynamics are commonly formulated as a linear complementarity problem. While this approach is superior to earlier spring-damper models, it can be inaccurate due to pyramid approximations to the friction cone, and inefficient due to lack of convexity coupled with a large number of auxiliary variables. Here we propose a new approach: implicit complementarity. Instead of treating contact velocities and forces as independent variables subject to explicit complementarity constraints, we express them as functions of a minimal set of unconstrained variables, and design these functions so that the complementarity constraints are automatically satisfied. We then solve the equations of motion via a non-smooth Gauss-Newton method augmented with an original linesearch procedure which exploits the problem structure. This enables us to represent the friction cone exactly and to reduce the number of unknowns by about a factor of 3. Numerical tests suggest that, in usage scenarios typical for robotics, the solver takes only about 5 iterations even without warm starts. More extensive tests and side-by-side comparisons remain to be done, but nevertheless the potential of the new approach is clear.

NeurIPS Conference 2010 Conference Paper

Policy gradients in linearly-solvable MDPs

  • Emanuel Todorov

We present policy gradient results within the framework of linearly-solvable MDPs. For the first time, compatible function approximators and natural policy gradients are obtained by estimating the cost-to-go function, rather than the (much larger) state-action advantage function as is necessary in traditional MDPs. We also develop the first compatible function approximators and natural policy gradients for continuous-time stochastic systems.

NeurIPS Conference 2009 Conference Paper

Compositionality of optimal control laws

  • Emanuel Todorov

We present a theory of compositionality in stochastic optimal control, showing how task-optimal controllers can be constructed from certain primitives. The primitives are themselves feedback controllers pursuing their own agendas. They are mixed in proportion to how much progress they are making towards their agendas and how compatible their agendas are with the present task. The resulting composite control law is provably optimal when the problem belongs to a certain class. This class is rather general and yet has a number of unique properties - one of which is that the Bellman equation can be made linear even for non-linear or discrete dynamics. This gives rise to the compositionality developed here. In the special case of linear dynamics and Gaussian noise our framework yields analytical solutions (i. e. non-linear mixtures of linear-quadratic regulators) without requiring the final cost to be quadratic. More generally, a natural set of control primitives can be constructed by applying SVD to Greens function of the Bellman equation. We illustrate the theory in the context of human arm movements. The ideas of optimality and compositionality are both very prominent in the field of motor control, yet they are hard to reconcile. Our work makes this possible.

NeurIPS Conference 2006 Conference Paper

Linearly-solvable Markov decision problems

  • Emanuel Todorov

We introduce a class of MPDs which greatly simplify Reinforcement Learning. They have discrete state spaces and continuous control spaces. The controls have the effect of rescaling the transition probabilities of an underlying Markov chain. A control cost penalizing KL divergence between controlled and uncontrolled transition probabilities makes the minimization problem convex, and allows analytical computation of the optimal controls given the optimal value function. An exponential transformation of the optimal value function makes the minimized Bellman equation linear. Apart from their theoretical signi cance, the new MDPs enable ef cient approximations to traditional MDPs. Shortest path problems are approximated to arbitrary precision with largest eigenvalue problems, yielding an O (n) algorithm. Accurate approximations to generic MDPs are obtained via continuous embedding reminiscent of LP relaxation in integer programming. Offpolicy learning of the optimal value function is possible without need for stateaction values; the new algorithm (Z-learning) outperforms Q-learning. This work was supported by NSF grant ECS0524761.

NeurIPS Conference 2002 Conference Paper

A Minimal Intervention Principle for Coordinated Movement

  • Emanuel Todorov
  • Michael Jordan

Behavioral goals are achieved reliably and repeatedly with movements rarely reproducible in their detail. Here we offer an explanation: we show that not only are variability and goal achievement compatible, but indeed that allowing variability in redundant dimensions is the optimal control strategy in the face of uncertainty. The optimal feedback control laws for typical motor tasks obey a “minimal intervention” principle: deviations from the average trajectory are only corrected when they interfere with the task goals. The resulting behavior exhibits task-constrained variabil- ity, as well as synergetic coupling among actuators—which is another unexplained empirical phenomenon.

NeurIPS Conference 1996 Conference Paper

A Model of Recurrent Interactions in Primary Visual Cortex

  • Emanuel Todorov
  • Athanassios Siapas
  • David Somers

A general feature of the cerebral cortex is its massive intercon(cid: 173) nectivity - it has been estimated anatomically [19] that cortical neurons receive upwards of 5, 000 synapses, the majority of which originate from other nearby cortical neurons. Numerous experi(cid: 173) ments in primary visual cortex (VI) have revealed strongly nonlin(cid: 173) ear interactions between stimulus elements which activate classical and non-classical receptive field regions. Recurrent cortical con(cid: 173) nections likely contribute substantially to these effects. However, most theories of visual processing have either assumed a feedfor(cid: 173) ward processing scheme [7], or have used recurrent interactions to account for isolated effects only [1, 16, 18]. Since nonlinear sys(cid: 173) tems cannot in general be taken apart and analyzed in pieces, it is not clear what one learns by building a recurrent model that only accounts for one, or very few phenomena. Here we develop a relatively simple model of recurrent interactions in VI, that re(cid: 173) flects major anatomical and physiological features of intracortical connectivity, and simultaneously accounts for a wide range of phe(cid: 173) nomena observed physiologically. All phenomena we address are strongly nonlinear, and cannot be explained by linear feedforward models.

NeurIPS Conference 1994 Conference Paper

Catastrophic Interference in Human Motor Learning

  • Tom Brashers-Krug
  • Reza Shadmehr
  • Emanuel Todorov

Biological sensorimotor systems are not static maps that transform input (sensory information) into output (motor behavior). Evi(cid: 173) dence from many lines of research suggests that their representa(cid: 173) tions are plastic, experience-dependent entities. While this plastic(cid: 173) ity is essential for flexible behavior, it presents the nervous system with difficult organizational challenges. If the sensorimotor system adapts itself to perform well under one set of circumstances, will it then perform poorly when placed in an environment with different demands (negative transfer)? Will a later experience-dependent change undo the benefits of previous learning (catastrophic inter(cid: 173) ference)? We explore the first question in a separate paper in this volume (Shadmehr et al. 1995). Here we present psychophysical and computational results that explore the question of catastrophic interference in the context of a dynamic motor learning task. Un(cid: 173) der some conditions, subjects show evidence of catastrophic inter(cid: 173) ference. Under other conditions, however, subjects appear to be immune to its effects. These results suggest that motor learning can undergo a process of consolidation. Modular neural networks are well suited for the demands of learning multiple input/output mappings. By incorporating the notion of fast- and slow-changing connections into a modular architecture, we were able to account for the psychophysical results. 20 Tom Brashers-Krug, Reza Shadmelzr, Emanuel Todorov

NeurIPS Conference 1994 Conference Paper

Factorial Learning by Clustering Features

  • Joshua Tenenbaum
  • Emanuel Todorov

We introduce a novel algorithm for factorial learning, motivated by segmentation problems in computational vision, in which the underlying factors correspond to clusters of highly correlated input features. The algorithm derives from a new kind of competitive clustering model, in which the cluster generators compete to ex(cid: 173) plain each feature of the data set and cooperate to explain each input example, rather than competing for examples and cooper(cid: 173) ating on features, as in traditional clustering algorithms. A natu(cid: 173) ral extension of the algorithm recovers hierarchical models of data generated from multiple unknown categories, each with a differ(cid: 173) ent, multiple causal structure. Several simulations demonstrate the power of this approach.

v2026.09.13