Arrow Research search

Author name cluster

Xiaotong Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

ICLR Conference 2025 Conference Paper

EditRoom: LLM-parameterized Graph Diffusion for Composable 3D Room Layout Editing

  • Kaizhi Zheng
  • Xiaotong Chen
  • Xuehai He
  • Jing Gu
  • Linjie Li
  • Zhengyuan Yang
  • Kevin Lin
  • Jianfeng Wang

Given the steep learning curve of professional 3D software and the time- consuming process of managing large 3D assets, language-guided 3D scene editing has significant potential in fields such as virtual reality, augmented reality, and gaming. However, recent approaches to language-guided 3D scene editing either require manual interventions or focus only on appearance modifications without supporting comprehensive scene layout changes. In response, we propose EditRoom, a unified framework capable of executing a variety of layout edits through natural language commands, without requiring manual intervention. Specifically, EditRoom leverages Large Language Models (LLMs) for command planning and generates target scenes using a diffusion-based method, enabling six types of edits: rotate, translate, scale, replace, add, and remove. To address the lack of data for language-guided 3D scene editing, we have developed an automatic pipeline to augment existing 3D scene synthesis datasets and introduced EditRoom-DB, a large-scale dataset with 83k editing pairs, for training and evaluation. Our experiments demonstrate that our approach consistently outperforms other baselines across all metrics, indicating higher accuracy and coherence in language-guided scene layout editing.

ICRA Conference 2022 Conference Paper

Composable Causality in Semantic Robot Programming

  • Emily Sheetz
  • Xiaotong Chen
  • Zhen Zeng
  • Kaizhi Zheng
  • Qiuyu Shi
  • Odest Chadwicke Jenkins

Assembly tasks are challenging for robot manipulation because the robot must reason over the composed effects of actions and execute multi-objective behaviors. Robots typically use predefined priorities provided by users to determine how to compose controller behaviors, but we want the robot to autonomously select these compositions based on their composed effects within the task. We present Composable Causality in Semantic Robot Programming to allow robots to reason over the composed effects of controllers when executing multi-objective actions and autonomously compose controllers without predefined priorities. Our proposed causal control basis combines controller behaviors with causal information about how the behaviors can be used to execute high-level symbolic actions. The robot uses the causal control basis to predict the transition probability of achieving the composed effects of a multi-objective action. The composed causality estimates are used to select which action to execute within the context of a furniture assembly task. We evaluate the robot's transition probability estimates in different furniture assembly trials in simulation on the Baxter robot. The robot's ability to assemble furniture using different multi-objective connection actions demonstrates the usefulness of the composed causality estimates from our causal control basis.

IROS Conference 2022 Conference Paper

ProgressLabeller: Visual Data Stream Annotation for Training Object-Centric 3D Perception

  • Xiaotong Chen
  • Huijie Zhang
  • Zeren Yu
  • Stanley R. Lewis
  • Odest Chadwicke Jenkins

Visual perception tasks often require vast amounts of labelled data, including 3D poses and image space segmen-tation masks. The process of creating such training data sets can prove difficult or time-intensive to scale up to efficacy for general use. Consider the task of pose estimation for rigid objects. Deep neural network based approaches have shown good performance when trained on large, public datasets. However, adapting these networks for other novel objects, or fine-tuning existing models for different environments, requires significant time investment to generate newly labelled instances. Towards this end, we propose ProgressLabeller as a method for more efficiently generating large amounts of 6D pose training data from color images sequences for custom scenes in a scalable manner. ProgressLabeller is intended to also support transparent or translucent objects, for which the previous methods based on depth dense reconstruction will fail. We demonstrate the effectiveness of ProgressLabeller by rapidly create a dataset of over 1M samples with which we fine-tune a state-of-the-art pose estimation network in order to markedly improve the downstream robotic grasp success rates. Progresslabeller is open-source at https://github.com/huijieZH/ProgressLabeller

NeurIPS Conference 2022 Conference Paper

VLMbench: A Compositional Benchmark for Vision-and-Language Manipulation

  • Kaizhi Zheng
  • Xiaotong Chen
  • Odest Chadwicke Jenkins
  • Xin Wang

Benefiting from language flexibility and compositionality, humans naturally intend to use language to command an embodied agent for complex tasks such as navigation and object manipulation. In this work, we aim to fill the blank of the last mile of embodied agents---object manipulation by following human guidance, e. g. , “move the red mug next to the box while keeping it upright. ” To this end, we introduce an Automatic Manipulation Solver (AMSolver) system and build a Vision-and-Language Manipulation benchmark (VLMbench) based on it, containing various language instructions on categorized robotic manipulation tasks. Specifically, modular rule-based task templates are created to automatically generate robot demonstrations with language instructions, consisting of diverse object shapes and appearances, action types, and motion constraints. We also develop a keypoint-based model 6D-CLIPort to deal with multi-view observations and language input and output a sequence of 6 degrees of freedom (DoF) actions. We hope the new simulator and benchmark will facilitate future research on language-guided robotic manipulation.

IROS Conference 2019 Conference Paper

GRIP: Generative Robust Inference and Perception for Semantic Robot Manipulation in Adversarial Environments

  • Xiaotong Chen
  • Rui Chen
  • Zhiqiang Sui
  • Zhefan Ye
  • Yanqi Liu
  • R. Iris Bahar
  • Odest Chadwicke Jenkins

Recent advancements have led to a proliferation of machine learning systems used to assist humans in a wide range of tasks. However, we are still far from accurate, reliable, and resource-efficient operations of these systems. For robot perception, convolutional neural networks (CNNs) for object detection and pose estimation are recently coming into widespread use. However, neural networks are known to suffer from overfitting during the training process and are less robust under unforeseen conditions (which makes them especially vulnerable to adversarial scenarios). In this work, we propose Generative Robust Inference and Perception (GRIP) as a two-stage object detection and pose estimation system that aims to combine the relative strengths of discriminative CNNs and generative inference methods to achieve robust estimation. Our results show that a second stage of sample-based generative inference is able to recover from false object detections by CNNs, and produce robust estimations in adversarial conditions. We demonstrate the efficacy of GRIP robustness through comparison with state-of-the-art learning-based pose estimators and pick-and-place manipulation in dark and cluttered environments.

ICRA Conference 2017 Conference Paper

A two-level approach for solving the inverse kinematics of an extensible soft arm considering viscoelastic behavior

  • Hao Jiang 0015
  • Zhanchi Wang
  • Xinghua Liu
  • Xiaotong Chen
  • Yusong Jin
  • Xuanke You
  • Xiaoping Chen

Soft compliant materials and novel actuation mechanisms ensure flexible motions and high adaptability for soft robots, but also increase the difficulty and complexity of constructing control systems. In this work, we provide an efficient control algorithm for a multi-segment extensible soft arm in 2D plane. The algorithm separate the inverse kinematics into two levels. The first level employs gradient descent to select optimized arm's pose (from task space to configuration space) according to designed cost functions. With consideration of viscoelasticity, the second level utilizes neural networks to figure out the pressures from each segment's pose (from configuration space to actuation space). In experiments with a physical prototype, the control accuracy and effectiveness are validated, where the control algorithm is further improved by an optional feedback strategy.

IROS Conference 2017 Conference Paper

Model-free control for soft manipulators based on reinforcement learning

  • Xuanke You
  • Yixiao Zhang
  • Xiaotong Chen
  • Xinghua Liu
  • Zhanchi Wang
  • Hao Jiang 0015
  • Xiaoping Chen

Most control methods of soft manipulators are developed based on physical models derived from mathematical analysis or learning methods. However, due to internal nonlinearity and external uncertain disturbances, it is difficult to build an accurate model, further, these methods lack robustness and portability among different prototypes. In this work, we propose a model-free control method based on reinforcement learning and implement it on a multi-segment soft manipulator in 2D plane, which focuses on the learning of control strategy rather than the physical model. The control strategy is validated to be effective and robust in prototype experiments, where we design a simulation method to speed up the training process.

IROS Conference 2017 Conference Paper

Model-less feedback control for soft manipulators

  • Yusong Jin
  • Yufei Wang
  • Xiaotong Chen
  • Zhanchi Wang
  • Xinghua Liu
  • Hao Jiang 0015
  • Xiaoping Chen

Soft manipulators have been a rising focus of soft robotics research. Taking advantage of soft materials and flexible, continuous movements, they have promising applicable prospect. However, their highly internal nonlinearity and unpredictable deformation caused by environmental effects make it difficult to build an exact model for control. In this work, we propose a generalized controller for soft manipulators using an estimated Jacobian-based model derived from structural analysis. The model can be simplified from reasonable assumptions of manipulator structure, and updated to balance conformity to reality and stability. In prototype experiments on an 3D multi-segment soft manipulator, the control method exhibits accuracy as well as adaptability to self gravity and external loads.

v2026.09.13