Arrow Research search
Back to ICRA

ICRA 2018

SE3-Pose-Nets: Structured Deep Dynamics Models for Visuomotor Control

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

In this work, we present an approach to deep visuomotor control using structured deep dynamics models. Our model, a variant of SE3-Nets, learns a low-dimensional pose embedding for visuomotor control via an encoder-decoder structure. Unlike prior work, our model is structured: given an input scene, our network explicitly learns to segment salient parts and predict their pose embedding and motion, modeled as a change in the pose due to the applied actions. We train our model using a pair of point clouds separated by an action and show that given supervision only through point-wise data associations between the frames our network is able to learn a meaningful segmentation of the scene along with consistent poses. We further show that our model can be used for closed-loop control directly in the learned low-dimensional pose space, where the actions are computed by minimizing pose error using gradient-based methods, similar to traditional model-based control. We present results on controlling a Baxter robot from raw depth data in simulation and RGBD data in the real world and compare against two baseline deep networks. We also test the robustness and generalization performance of our controller under changes in camera pose, lighting, occlusion, and motion. Our method is robust, runs in real-time, achieves good prediction of scene dynamics, and outperforms baselines on multiple control runs. Video results can be found at: https://rse-lab.cs.washington.edu/se3-structured-deep-ctrl/.

Authors

Keywords

  • Three-dimensional displays
  • Predictive models
  • Transforms
  • Computational modeling
  • Data models
  • Aerospace electronics
  • Training
  • Visuomotor Control
  • Simulated Data
  • Point Cloud
  • Low-dimensional Space
  • Camera Pose
  • Dynamic Scenes
  • Pose Changes
  • RGB-D Data
  • Raw Depth
  • Prediction Model
  • Control Performance
  • Efficient Control
  • Latent Space
  • Depth Images
  • Transition Model
  • External System
  • Future Work In This Area
  • Reactive Control
  • Input Point
  • Observation Space
  • Visual Servoing
  • Arm Motion
  • Kinematic Chain
  • Input Point Cloud
  • Consistency Loss
  • Target Pose
  • Real Robot
  • Part Of The Scene

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
270302207447087215
v2026.09.13