Arrow Research search

Author name cluster

Bin Wang 0034

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
1 author row

Possible papers

6

ICLR Conference 2024 Conference Paper

Towards Faithful XAI Evaluation via Generalization-Limited Backdoor Watermark

  • Mengxi Ya
  • Yiming Li 0004
  • Tao Dai 0001
  • Bin Wang 0034
  • Yong Jiang 0001
  • Shu-Tao Xia

Saliency-based representation visualization (SRV) ($e.g.$, Grad-CAM) is one of the most classical and widely adopted explainable artificial intelligence (XAI) methods for its simplicity and efficiency. It can be used to interpret deep neural networks by locating saliency areas contributing the most to their predictions. However, it is difficult to automatically measure and evaluate the performance of SRV methods due to the lack of ground-truth salience areas of samples. In this paper, we revisit the backdoor-based SRV evaluation, which is currently the only feasible method to alleviate the previous problem. We first reveal its \emph{implementation limitations} and \emph{unreliable nature} due to the trigger generalization of existing backdoor watermarks. Given these findings, we propose a generalization-limited backdoor watermark (GLBW), based on which we design a more faithful XAI evaluation. Specifically, we formulate the training of watermarked DNNs as a min-max problem, where we find the `worst' potential trigger (with the highest attack effectiveness and differences from the ground-truth trigger) via inner maximization and minimize its effects and the loss over benign and poisoned samples via outer minimization in each iteration. In particular, we design an adaptive optimization method to find desired potential triggers in each inner maximization. Extensive experiments on benchmark datasets are conducted, verifying the effectiveness of our generalization-limited watermark. Our codes are available at \url{https://github.com/yamengxi/GLBW}.

ICLR Conference 2024 Conference Paper

Tree-Planner: Efficient Close-loop Task Planning with Large Language Models

  • Mengkang Hu
  • Yao Mu 0001
  • Xinmiao Yu
  • Mingyu Ding
  • Shiguang Wu 0004
  • Wenqi Shao
  • Qiguang Chen
  • Bin Wang 0034

This paper studies close-loop task planning, which refers to the process of generating a sequence of skills (a plan) to accomplish a specific goal while adapting the plan based on real-time observations. Recently, prompting Large Language Models (LLMs) to generate actions iteratively has become a prevalent paradigm due to its superior performance and user-friendliness. However, this paradigm is plagued by two inefficiencies: high token consumption and redundant error correction, both of which hinder its scalability for large-scale testing and applications. To address these issues, we propose Tree-Planner, which reframes task planning with LLMs into three distinct phases: plan sampling, action tree construction, and grounded deciding. Tree-Planner starts by using an LLM to sample a set of potential plans before execution, followed by the aggregation of them to form an action tree. Finally, the LLM performs a top-down decision-making process on the tree, taking into account real-time environmental information. Experiments show that Tree-Planner achieves state-of-the-art performance while maintaining high efficiency. By decomposing LLM queries into a single plan-sampling call and multiple grounded-deciding calls, a considerable part of the prompt are less likely to be repeatedly consumed. As a result, token consumption is reduced by 92.2\% compared to the previously best-performing model. Additionally, by enabling backtracking on the action tree as needed, the correction process becomes more flexible, leading to a 40.5\% decrease in error corrections.

ICML Conference 2023 Conference Paper

ChiPFormer: Transferable Chip Placement via Offline Decision Transformer

  • Yao Lai
  • Jinxin Liu
  • Zhentao Tang
  • Bin Wang 0034
  • Jianye Hao
  • Ping Luo 0002

Placement is a critical step in modern chip design, aiming to determine the positions of circuit modules on the chip canvas. Recent works have shown that reinforcement learning (RL) can improve human performance in chip placement. However, such an RL-based approach suffers from long training time and low transfer ability in unseen chip circuits. To resolve these challenges, we cast the chip placement as an offline RL formulation and present ChiPFormer that enables learning a transferable placement policy from fixed offline data. ChiPFormer has several advantages that prior arts do not have. First, ChiPFormer can exploit offline placement designs to learn transferable policies more efficiently in a multi-task setting. Second, ChiPFormer can promote effective finetuning for unseen chip circuits, reducing the placement runtime from hours to minutes. Third, extensive experiments on 32 chip circuits demonstrate that ChiPFormer achieves significantly better placement quality while reducing the runtime by 10x compared to recent state-of-the-art approaches in both public benchmarks and realistic industrial tasks. The deliverables are released at https: //sites. google. com/view/chipformer/home.

ICLR Conference 2023 Conference Paper

Learnable Behavior Control: Breaking Atari Human World Records via Sample-Efficient Behavior Selection

  • Jiajun Fan
  • Yuzheng Zhuang
  • Yuecheng Liu
  • Jianye Hao
  • Bin Wang 0034
  • Jiangcheng Zhu
  • Hao Wang 0049
  • Shu-Tao Xia

The exploration problem is one of the main challenges in deep reinforcement learning (RL). Recent promising works tried to handle the problem with population-based methods, which collect samples with diverse behaviors derived from a population of different exploratory policies. Adaptive policy selection has been adopted for behavior control. However, the behavior selection space is largely limited by the predefined policy population, which further limits behavior diversity. In this paper, we propose a general framework called Learnable Behavioral Control (LBC) to address the limitation, which a) enables a significantly enlarged behavior selection space via formulating a hybrid behavior mapping from all policies; b) constructs a unified learnable process for behavior selection. We introduce LBC into distributed off-policy actor-critic methods and achieve behavior control via optimizing the selection of the behavior mappings with bandit-based meta-controllers. Our agents have achieved 10077.52% mean human normalized score and surpassed 24 human world records within 1B training frames in the Arcade Learning Environment, which demonstrates our significant state-of-the-art (SOTA) performance without degrading the sample efficiency.

ICML Conference 2023 Conference Paper

MetaDiffuser: Diffusion Model as Conditional Planner for Offline Meta-RL

  • Fei Ni 0001
  • Jianye Hao
  • Yao Mu 0001
  • Yifu Yuan
  • Yan Zheng 0002
  • Bin Wang 0034
  • Zhixuan Liang

Recently, diffusion model shines as a promising backbone for the sequence modeling paradigm in offline reinforcement learning(RL). However, these works mostly lack the generalization ability across tasks with reward or dynamics change. To tackle this challenge, in this paper we propose a task-oriented conditioned diffusion planner for offline meta-RL(MetaDiffuser), which considers the generalization problem as conditional trajectory generation task with contextual representation. The key is to learn a context conditioned diffusion model which can generate task-oriented trajectories for planning across diverse tasks. To enhance the dynamics consistency of the generated trajectories while encouraging trajectories to achieve high returns, we further design a dual-guided module in the sampling process of the diffusion model. The proposed framework enjoys the robustness to the quality of collected warm-start data from the testing task and the flexibility to incorporate with different task representation method. The experiment results on MuJoCo benchmarks show that MetaDiffuser outperforms other strong offline meta-RL baselines, demonstrating the outstanding conditional generation ability of diffusion architecture.

ICRA Conference 2021 Conference Paper

Relational Navigation Learning in Continuous Action Space among Crowds

  • Xueyou Zhang
  • Wei Xi 0002
  • Xian Guo
  • Yongchun Fang
  • Bin Wang 0034
  • Wulong Liu
  • Jianye Hao

In this paper, a novel navigation learning method in continuous action space among crowds based on relational graph is proposed which can be directly deployed on differential-drive mobile robots without any change. More specifically, in order to increase generalization ability in crowd sizes, Graph Convolutional Network (GCN) is at first adopted to extract the relationships between robot and pedestrians. Then the relation features are further utilized as the inputs of the pedestrian state prediction network, the actor network, and the critic network. To efficiently and safely learn the navigation policy, all networks are pretrained through imitating ORCA which is a state-of-the-art algorithm in crowd navigation, and then a model-based reinforcement learning (RL) method which combines the model prediction and the clipped advantage-weighted regression is proposed to finetune the networks. Finally, simulation experiments are performed and it’s verified that the proposed learning method performs significantly better than ORCA and the other state-of-the-art RL methods.

v2026.09.13