Arrow Research search

Author name cluster

Wenbo Ye

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
2 author rows

Possible papers

3

AAAI Conference 2025 Conference Paper

Consistency of Compositional Generalization Across Multiple Levels

  • Chuanhao Li
  • Zhen Li
  • Chenchen Jing
  • Xiaomeng Fan
  • Wenbo Ye
  • Yuwei Wu
  • Yunde Jia

Compositional generalization is the capability of a model to understand novel compositions composed of seen concepts. There are multiple levels of novel compositions including phrase-phrase level, phrase-word level, and word-word level. Existing methods achieve promising compositional generalization, but the consistency of compositional generalization across multiple levels of novel compositions remains unexplored. The consistency refers to that a model should generalize to a phrase-phrase level novel composition, and phrase-word/word-word level novel compositions that can be derived from it simultaneously. In this paper, we propose a meta-learning based framework, for achieving consistent compositional generalization across multiple levels. The basic idea is to progressively learn compositions from simple to complex for consistency. Specifically, we divide the original training set into multiple validation sets based on compositional complexity, and introduce multiple meta-weight-nets to generate sample weights for samples in different validation sets. To fit the validation sets in order of increasing compositional complexity, we optimize the parameters of each meta-weight-net independently and sequentially in a multilevel optimization manner. We build a GQA-CCG dataset to quantitatively evaluate the consistency. Experimental results on visual question answering and temporal video grounding, demonstrate the effectiveness of the proposed framework.

IJCAI Conference 2025 Conference Paper

Multi-Sourced Compositional Generalization in Visual Question Answering

  • Chuanhao Li
  • Wenbo Ye
  • Zhen Li
  • Yuwei Wu
  • Yunde Jia

Compositional generalization is the ability of generalizing novel compositions from seen primitives, and has received much attention in vision-and-language (V&L) recently. Due to the multi-modal nature of V&L tasks, the primitives composing compositions source from different modalities, resulting in multi-sourced novel compositions. However, the generalization ability over multi-sourced novel compositions, i. e. , multi-sourced compositional generalization (MSCG) remains unexplored. In this paper, we explore MSCG in the context of visual question answering (VQA), and propose a retrieval-augmented training framework to enhance the MSCG ability of VQA models by learning unified representations for primitives from different modalities. Specifically, semantically equivalent primitives are retrieved for each primitive in the training samples, and the retrieved features are aggregated with the original primitive to refine the model. This process helps the model learn consistent representations for the same semantic primitives across different modalities. To evaluate the MSCG ability of VQA models, we construct a new GQA-MSCG dataset based on the GQA dataset, in which samples include three types of novel compositions composed of primitives from different modalities. The GQA-MSCG dataset is available at https: //github. com/NeverMoreLCH/MSCG.

IROS Conference 2006 Conference Paper

Structure-Constrained Obstacles Recognition for Power Transmission Line Inspection Robot

  • Si-Yao Fu
  • Weiming Li
  • Yun-Chu Zhang
  • Zi-ze Liang
  • Zeng-Guang Hou
  • Min Tan 0001
  • Wenbo Ye
  • Lian Bo

Inspection robot must plan its behavior to detect the obstacles from the complex background according to their types when it is crawling along the power transmission line in order to negotiate reliably. However, in most instances, detecting the obstacles from the complex background is a hard task. For this purpose, a novel and fast visual obstacle recognition algorithm is designed based on the structure of the 220 KV power transmission line. Basic principle and architecture of the algorithm are given. By this approach, three typical obstacles on the power transmission line such as insulator strings, counterweights and suspension clamps can be recognized with high accuracy. Experiments in the real power transmission line show its effectiveness. This method can contribute to the process of the mobile robot negotiating obstacles

v2026.09.13