Arrow Research search
Back to IROS

IROS 2025

PhyGrasp: Generalizing Robotic Grasping with Physics-informed Large Multimodal Models

Conference Paper Accepted Paper Artificial Intelligence · Robotics

Abstract

Robotic grasping, crucial for robot interaction with objects, still struggles with counter-intuitive or long-tailed scenarios like uncommon materials and shapes. Humans, however, intuitively adjust grasps with their physics-informed interpretations of the object, using visual and linguistic cues. This work introduces PhyGrasp, a large multimodal model and dataset that enhance robotic manipulation by combining natural language and 3D point clouds using a bridge module to integrate these inputs. The language modality exhibits robust reasoning capabilities concerning the impacts of diverse physical properties on grasping, while the 3D modality comprehends object shapes and parts. With these two capabilities, PhyGrasp is able to accurately assess the physical properties of object parts and determine optimal grasping poses. Additionally, the model’s language comprehension enables human instruction interpretation, generating grasping poses that align with human preferences. To train PhyGrasp, we construct a dataset PhyPartNet with 195K object instances with varying physical properties and human preferences, alongside their corresponding language descriptions. Extensive experiments conducted in the simulation and on the real robots demonstrate that PhyGrasp achieves state-of-the-art performance, particularly in long-tailed cases, e. g. , about 10% improvement in success rate over GraspNet. More demos and information are available on https://sites.google.com/view/phygrasp.

Authors

Keywords

  • Point cloud compression
  • Bridges
  • Solid modeling
  • Visualization
  • Three-dimensional displays
  • Heavily-tailed distribution
  • Shape
  • Grasping
  • Linguistics
  • Robots
  • Multimodal Model
  • Robotic Grasping
  • Natural Language
  • Point Cloud
  • Object Parts
  • Robot Manipulator
  • 3D Point Cloud
  • Description Language
  • Human Preferences
  • Properties Of Parts
  • Object Instances
  • Improve Success Rates
  • Language Mode
  • Center Of Mass
  • Local Features
  • Visual Features
  • Analytical Solutions
  • Global Features
  • Multilayer Perceptron
  • Kullback-Leibler
  • Physical Reasons
  • Object Pose
  • Normal Scenario
  • Language Model
  • Object Point Cloud
  • Linguistic Features
  • Visual Input
  • Bridge Network
  • Manipulation Tasks
  • Object Properties

Context

Venue
IEEE/RSJ International Conference on Intelligent Robots and Systems
Archive span
1988-2025
Indexed papers
26578
Paper id
615643671900784760
v2026.09.13