Arrow Research search

Author name cluster

Joseph Ortiz

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
2 author rows

Possible papers

6

TMLR Journal 2025 Journal Article

Diffusion Model Predictive Control

  • Guangyao Zhou
  • Sivaramakrishnan Swaminathan
  • Rajkumar Vasudeva Raju
  • J Swaroop Guntupalli
  • Wolfgang Lehrach
  • Joseph Ortiz
  • Antoine Dedieu
  • Miguel Lazaro-Gredilla

We propose Diffusion Model Predictive Control (D-MPC), a novel MPC approach that learns a multi-step action proposal and a multi-step dynamics model, both using diffusion models, and combines them for use in online MPC. On the popular D4RL benchmark, we show performance that is significantly better than existing model-based offline planning methods using MPC (e.g. MBOP) and competitive with state-of-the-art (SOTA) model-based and model-free reinforcement learning methods. We additionally illustrate D-MPC’s ability to optimize novel reward functions at run time and adapt to novel dynamics, and highlight its advantages compared to existing diffusion-based planning baselines.

ICML Conference 2025 Conference Paper

Improving Transformer World Models for Data-Efficient RL

  • Antoine Dedieu
  • Joseph Ortiz
  • Xinghua Lou
  • Carter Wendelken
  • J. Swaroop Guntupalli
  • Wolfgang Lehrach
  • Miguel Lázaro-Gredilla
  • Kevin Murphy 0002

We present an approach to model-based RL that achieves a new state of the art performance on the challenging Craftax-classic benchmark, an open-world 2D survival game that requires agents to exhibit a wide range of general abilities—such as strong generalization, deep exploration, and long-term reasoning. With a series of careful design choices aimed at improving sample efficiency, our MBRL algorithm achieves a reward of 69. 66% after only 1M environment steps, significantly outperforming DreamerV3, which achieves $53. 2%$, and, for the first time, exceeds human performance of 65. 0%. Our method starts by constructing a SOTA model-free baseline, using a novel policy architecture that combines CNNs and RNNs. We then add three improvements to the standard MBRL setup: (a) "Dyna with warmup", which trains the policy on real and imaginary data, (b) "nearest neighbor tokenizer" on image patches, which improves the scheme to create the transformer world model (TWM) inputs, and (c) "block teacher forcing", which allows the TWM to reason jointly about the future tokens of the next timestep.

ICML Conference 2024 Conference Paper

A Touch, Vision, and Language Dataset for Multimodal Alignment

  • Letian Fu
  • Gaurav Datta
  • Huang Huang
  • William Chung-Ho Panitch
  • Jaimyn Drake
  • Joseph Ortiz
  • Mustafa Mukadam
  • Mike Lambeta

Touch is an important sensing modality for humans, but it has not yet been incorporated into a multimodal generative language model. This is partially due to the difficulty of obtaining natural language labels for tactile data and the complexity of aligning tactile readings with both visual observations and language descriptions. As a step towards bridging that gap, this work introduces a new dataset of 44K in-the-wild visiontouch pairs, with English language labels annotated by humans (10%) and textual pseudo-labels from GPT-4V (90%). We use this dataset to train a vision-language-aligned tactile encoder for open-vocabulary classification and a touch-visionlanguage (TVL) model for text generation using the trained encoder. Results suggest that by incorporating touch, the TVL model improves (+29% classification accuracy) tactile-vision-language alignment over existing models trained on any pair of those modalities. Although only a small fraction of the dataset is human labeled, the TVL model demonstrates improved visual-tactile understanding over GPT-4V (+12%) and open-source vision-language models (+32%) on a new touch-vision understanding benchmark. Code, checkpoints and data are available on https: //tactile-vlm. github. io.

NeurIPS Conference 2024 Conference Paper

DMC-VB: A Benchmark for Representation Learning for Control with Visual Distractors

  • Joseph Ortiz
  • Antoine Dedieu
  • Wolfgang Lehrach
  • J. S. Guntupalli
  • Carter Wendelken
  • Ahmad Humayun
  • Guangyao Zhou
  • Sivaramakrishnan Swaminathan

Learning from previously collected data via behavioral cloning or offline reinforcement learning (RL) is a powerful recipe for scaling generalist agents by avoiding the need for expensive online learning. Despite strong generalization in some respects, agents are often remarkably brittle to minor visual variations in control-irrelevant factors such as the background or camera viewpoint. In this paper, we present theDeepMind Control Visual Benchmark (DMC-VB), a dataset collected in the DeepMind Control Suite to evaluate the robustness of offline RL agents for solving continuous control tasks from visual input in the presence of visual distractors. In contrast to prior works, our dataset (a) combines locomotion and navigation tasks of varying difficulties, (b) includes static and dynamic visual variations, (c) considers data generated by policies with different skill levels, (d) systematically returns pairs of state and pixel observation, (e) is an order of magnitude larger, and (f) includes tasks with hidden goals. Accompanying our dataset, we propose three benchmarks to evaluate representation learning methods for pretraining, and carry out experiments on several recently proposed methods. First, we find that pretrained representations do not help policy learning on DMC-VB, and we highlight a large representation gap between policies learned on pixel observations and on states. Second, we demonstrate when expert data is limited, policy learning can benefit from representations pretrained on (a) suboptimal data, and (b) tasks with stochastic hidden goals. Our dataset and benchmark code to train and evaluate agents are available at https: //github. com/google-deepmind/dmc vision benchmark.

ICRA Conference 2022 Conference Paper

Incremental Abstraction in Distributed Probabilistic SLAM Graphs

  • Joseph Ortiz
  • Talfan Evans
  • Edgar Sucar
  • Andrew J. Davison

Scene graphs represent the key components of a scene in a compact and semantically rich way, but are difficult to build during incremental SLAM operation because of the challenges of robustly identifying abstract scene elements and optimising continually changing, complex graphs. We present a distributed, graph-based SLAM framework for incrementally building scene graphs based on two novel components. First, we propose an incremental abstraction framework in which a neural network proposes abstract scene elements that are incorporated into the factor graph of a feature-based monocular SLAM system. Scene elements are confirmed or rejected through optimisation and incrementally replace the points yielding a more dense, semantic and compact representation. Second, enabled by our novel routing procedure, we use Gaussian Belief Propagation (GBP) for distributed inference on a graph processor. The time per iteration of GBP is structure-agnostic and we demonstrate the speed advantages over direct methods for inference of heterogeneous factor graphs. We run our system on real indoor datasets using planar abstractions and recover the major planes with significant compression.

NeurIPS Conference 2022 Conference Paper

Theseus: A Library for Differentiable Nonlinear Optimization

  • Luis Pineda
  • Taosha Fan
  • Maurizio Monge
  • Shobha Venkataraman
  • Paloma Sodhi
  • Ricky T. Q. Chen
  • Joseph Ortiz
  • Daniel DeTone

We present Theseus, an efficient application-agnostic open source library for differentiable nonlinear least squares (DNLS) optimization built on PyTorch, providing a common framework for end-to-end structured learning in robotics and vision. Existing DNLS implementations are application specific and do not always incorporate many ingredients important for efficiency. Theseus is application-agnostic, as we illustrate with several example applications that are built using the same underlying differentiable components, such as second-order optimizers, standard costs functions, and Lie groups. For efficiency, Theseus incorporates support for sparse solvers, automatic vectorization, batching, GPU acceleration, and gradient computation with implicit differentiation and direct loss minimization. We do extensive performance evaluation in a set of applications, demonstrating significant efficiency gains and better scalability when these features are incorporated. Project page: https: //sites. google. com/view/theseus-ai/

v2026.09.13