Arrow Research search
Back to IROS

IROS 2025

GraspMAS: Zero-Shot Language-driven Grasp Detection with Multi-Agent System

Conference Paper Accepted Paper Artificial Intelligence ยท Robotics

Abstract

Language-driven grasp detection has the potential to revolutionize human-robot interaction by allowing robots to understand and execute grasping tasks based on natural language commands. However, existing approaches face two key challenges. First, they often struggle to interpret complex text instructions or operate ineffectively in densely cluttered environments. Second, most methods require a training or fine-tuning step to adapt to new domains, limiting their generation in real-world applications. In this paper, we introduce GraspMAS, a new multi-agent system framework for language-driven grasp detection. GraspMAS is designed to reason through ambiguities and improve decision-making in real-world scenarios. Our framework consists of three specialized agents: Planner, responsible for strategizing complex queries; Coder, which generates and executes source code; and Observer, which evaluates the outcomes and provides feedback. Intensive experiments on two large-scale datasets demonstrate that our GraspMAS significantly outperforms existing baselines. Additionally, robot experiments conducted in both simulation and real-world settings further validate the effectiveness of our approach. Our project page is available at https://zquang2202.github.io/GraspMAS.

Authors

Keywords

  • Training
  • Limiting
  • Foundation models
  • Source coding
  • Natural languages
  • Human-robot interaction
  • Grasping
  • Observers
  • Intelligent robots
  • Multi-agent systems
  • Natural Language
  • Real-world Setting
  • Real-world Scenarios
  • Text Complexity
  • Complex Queries
  • Robot Experiments
  • Convolutional Neural Network
  • Simulation Experiments
  • Simulation Environment
  • Video Analysis
  • Language Model
  • Variety Of Scenarios
  • Python Code
  • Input Text
  • Real Robot
  • Composite Approach
  • Foundation Model
  • Code Execution
  • Semantic Understanding
  • Text Query
  • Visual Question Answering

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
1035129531054960977
v2026.09.13