Arrow Research search
Back to ICRA

ICRA 2023

ADAPT: Action-aware Driving Caption Transformer

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

End-to-end autonomous driving has great potential in the transportation industry. However, the lack of transparency and interpretability of the automatic decision-making process hinders its industrial adoption in practice. There have been some early attempts to use attention maps or cost volume for better model explainability which is difficult for ordinary passengers to understand. To bridge the gap, we propose an end-to-end transformer-based architecture, ADAPT (Action-aware Driving cAPtion Transformer), which provides user-friendly natural language narrations and reasoning for each decision making step of autonomous vehicular control and action. ADAPT jointly trains both the driving caption task and the vehicular control prediction task, through a shared video representation. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate state-of-the-art performance of the ADAPT framework on both automatic metrics and human evaluation. To illustrate the feasibility of the proposed framework in real-world applications, we build a novel deployable system that takes raw car videos as input and outputs the action narrations and reasoning in real time. The code, models and data are available at https://github.com/jxbbb/ADAPT.

Authors

Keywords

  • Measurement
  • Training
  • Adaptation models
  • Transportation industry
  • Decision making
  • Streaming media
  • Transformers
  • Transformer
  • Narrative
  • Natural Language
  • Autonomous Vehicles
  • Autonomic Control
  • Attention Map
  • Deployment Of Systems
  • Joint Training
  • Raw Video
  • Cost Volume
  • Root Mean Square Error
  • Autonomic System
  • Control Signal
  • Sampling Frame
  • Simulation Environment
  • Video Frames
  • Traffic Light
  • Multi-task Learning
  • Tokenized
  • Syntactic Structure
  • Video Captioning
  • Text Generation
  • Multi-task Framework
  • Video Features
  • Video Encoding
  • Bird’s Eye
  • Prediction Head
  • Natural Sentences
  • Inductive Bias

Context

Venue
IEEE International Conference on Robotics and Automation
Archive span
1984-2025
Indexed papers
30179
Paper id
768790220034480329
v2026.09.13