Arrow Research search

Author name cluster

Govind Thattai

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
2 author rows

Possible papers

3

NeurIPS Conference 2023 Conference Paper

Alexa Arena: A User-Centric Interactive Platform for Embodied AI

  • Qiaozi Gao
  • Govind Thattai
  • Suhaila Shakiah
  • Xiaofeng Gao
  • Shreyas Pansare
  • Vasu Sharma
  • Gaurav Sukhatme
  • Hangjie Shi

We introduce Alexa Arena, a user-centric simulation platform to facilitate research in building assistive conversational embodied agents. Alexa Arena features multi-room layouts and an abundance of interactable objects. With user-friendly graphics and control mechanisms, the platform supports the development of gamified robotic tasks readily accessible to general human users, allowing high-efficiency data collection and EAI system evaluation. Along with the platform, we introduce a dialog-enabled task completion benchmark with online human evaluations.

IROS Conference 2022 Conference Paper

Learning to Act with Affordance-Aware Multimodal Neural SLAM

  • Zhiwei Jia
  • Kaixiang Lin
  • Yizhou Zhao
  • Qiaozi Gao
  • Govind Thattai
  • Gaurav S. Sukhatme

Recent years have witnessed an emerging paradigm shift toward embodied artificial intelligence, in which an agent must learn to solve challenging tasks by interacting with its environment. There are several challenges in solving embodied multimodal tasks, including long-horizon planning, vision-and-language grounding, and efficient exploration. We focus on a critical bottleneck, namely the performance of planning and navigation. To tackle this challenge, we propose a Neural SLAM approach that, for the first time, utilizes several modalities for exploration, predicts an affordance-aware semantic map, and plans over it at the same time. This signif-icantly improves exploration efficiency, leads to robust long-horizon planning, and enables effective vision-and-language grounding. With the proposed Affordance-aware Multimodal Neural SLAM (AMSLAM) approach, we obtain more than 40% improvement over prior published work on the ALFRED benchmark and set a new state-of-the-art generalization per-formance at a success rate of 23. 48% on the test unseen scenes.

TMLR Journal 2022 Journal Article

Learning Two-Step Hybrid Policy for Graph-Based Interpretable Reinforcement Learning

  • Tongzhou Mu
  • Kaixiang Lin
  • Feiyang Niu
  • Govind Thattai

We present a two-step hybrid reinforcement learning (RL) policy that is designed to generate interpretable and robust hierarchical policies on the RL problem with graph-based input. Unlike prior deep reinforcement learning policies parameterized by an end-to-end black-box graph neural network, our approach disentangles the decision-making process into two steps. The first step is a simplified classification problem that maps the graph input to an action group where all actions share a similar semantic meaning. The second step implements a sophisticated rule-miner that conducts explicit one-hop reasoning over the graph and identifies decisive edges in the graph input without the necessity of heavy domain knowledge. This two-step hybrid policy presents human-friendly interpretations and achieves better performance in terms of generalization and robustness. Extensive experimental studies on four levels of complex text-based games have demonstrated the superiority of the proposed method compared to the state-of-the-art.

v2026.09.13