Arrow Research search

Author name cluster

Bin Sheng

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
1 author row

Possible papers

10

EAAI Journal 2026 Journal Article

Group morphological adaptation via adversarial imitation learning

  • Liming Xin
  • Zhen Wang
  • Jinlin Peng
  • Bin Sheng

Reinforcement learning has achieved significant success in training policies for specific agents. However, the vast diversity of potential robotic designs makes the replication of the training process for each individual design highly impractical. To address this challenge, this paper presents a novel group adversarial imitation learning framework that trains policies capable of seamlessly adapting to diverse robotic morphologies. The proposed approach leverages experience sharing and utilizes a unified actor–critic architecture to develop a cohesive policy for a group of agents. Additionally, a group feature alignment module is integrated to stabilize performance across disparate morphologies by aligning state–action representations. Empirical evaluations demonstrate that our group adversarial imitation learning approach outperforms baseline methods, achieving approximately a 20% improvement in mean reward and a tenfold increase in minimum reward. These results highlight the framework’s robustness and adaptability, positioning it as an ideal candidate for developing real-world robots with diverse morphological variations, such as modular robots for warehouse automation, adaptive multi-legged robots for search-and-rescue, and heterogeneous robotic teams in manufacturing.

EAAI Journal 2026 Journal Article

Slimmable neural architecture design based on cross architecture and token distillation

  • Guhao Qiu
  • Zhihua Chen
  • Lei Dai
  • Ping Li
  • Bin Sheng

The manually designed neural networks have the drawbacks of requiring a large amount of training data and high computational costs. In this paper, we propose the masked autoencoder based lightweight network search algorithm which leverages the efficient channel search algorithm and specific distillation strategy to obtain the optimal architecture. During SuperNet training process, we design the cross-token distillation and cross-architecture strategy. Token distillation strategy is used to enforce the similar representation obtained from different masks in one image. Architecture distillation strategy is used to fully utilize the representation from the sampled subnetwork and use the feature maps from one image but the same token. In the subnetwork searching process, we further pretrain the selected network considering the component dependency. Comprehensive experiments verify that our proposed method is efficient and flexible than baseline self-supervised learning algorithm and structured pruning algorithms. For example, our method obtains 4. 4% improvement in TOP-1 metrics compared with the classic Masked Autoencoder algorithm designed for lightweight Transformer architecture with less than 10M parameter.

JBHI Journal 2025 Journal Article

MHFNet: A Multimodal Hybrid-Embedding Fusion Network for Automatic Sleep Staging

  • Ruhan Liu
  • Jiajia Li
  • Yang Wen
  • Xian Huang
  • Bin Sheng
  • David Dagan Feng
  • Ping Zhang

Scoring sleep stages is essential for evaluating the status of sleep continuity and comprehending its structure. Despite previous attempts, automating sleep scoring remains challenging. First, most existing works did not fuse local and global temporal information. Second, the correlation for special waves in different signals is rarely used in sleep staging modeling. Third, the logic of scoring rules based on adjacent epochs is not considered in developing sleep staging models. This paper introduces a multimodal hybrid-embedding fusion network (MHFNet), which aims to tackle these challenges in automating sleep stage scoring. MHFNet comprises multi-stream Xception blocks to extract wave characteristics, a hybrid time-embedding module to combine local and global temporal information, a dual-path gate transformer to fuse and enhance attention features, and a refined output header to reconstruct sleep scoring. We perform experiments using three publicly available datasets (SleepEDF-ST, SleepEDF-SC, and SHHS). Experimental results indicate the superiority of MHFNet over baseline approaches in cross-validation. Moreover, at the individual level, MHFNet yielded an average $R^{2}$ score improvement of 9 $\%$ in the testing dataset compared to state-of-the-art models, paving the way for its applications in real-world sleep medicine.

AIIM Journal 2024 Journal Article

SSM-Net: Semi-supervised multi-task network for joint lesion segmentation and classification from pancreatic EUS images

  • Jiajia Li
  • Pingping Zhang
  • Xia Yang
  • Lei Zhu
  • Teng Wang
  • Ping Zhang
  • Ruhan Liu
  • Bin Sheng

Pancreatic cancer does not show specific symptoms, which makes the diagnosis of early stages difficult with established image-based screening methods and therefore has the worst prognosis among all cancers. Although endoscopic ultrasonography (EUS) has a key role in diagnostic algorithms for pancreatic diseases, B-mode imaging of the pancreas can be affected by confounders such as chronic pancreatitis, which can make both pancreatic lesion segmentation and classification laborious and highly specialized. To address these challenges, this work proposes a semi-supervised multi-task network (SSM-Net) to leverage unlabeled and labeled EUS images for joint pancreatic lesion classification and segmentation. Specifically, we first devise a saliency-aware representation learning module (SRLM) on a large number of unlabeled images to train a feature extraction encoder network for labeled images by computing a contrastive loss with a semantic saliency map, which is obtained by our spectral residual module (SRM). Moreover, for labeled EUS images, we devise channel attention blocks (CABs) to refine the features extracted from the pre-trained encoder on unlabeled images for segmenting lesions, and then devise a merged global attention module (MGAM) and a feature similarity loss (FSL) for obtaining a lesion classification result. We collect a large-scale EUS-based pancreas image dataset (LS-EUSPI) consisting of 9, 555 pathologically proven labeled EUS images (499 patients from four categories) and 15, 500 unlabeled EUS images. Experimental results on the LS-EUSPI dataset and a public thyroid gland lesion dataset show that our SSM-Net clearly outperforms state-of-the-art methods.

AAAI Conference 2024 Conference Paper

Text2City: One-Stage Text-Driven Urban Layout Regeneration

  • Yiming Qin
  • Nanxuan Zhao
  • Bin Sheng
  • Rynson W.H. Lau

Regenerating urban layout is an essential process for urban regeneration. In this paper, we propose a new task called text-driven urban layout regeneration, which provides an intuitive input modal - text - for users to specify the regeneration, instead of designing complex rules. Given the target region to be regenerated, we propose a one-stage text-driven urban layout regeneration model, Text2City, to jointly and progressively regenerate the urban layout (i.e., road and building layouts) based on textual layout descriptions and surrounding context (i.e., urban layouts and functions of the surrounding regions). Text2City first extracts road and building attributes from the textual layout description to guide the regeneration. It includes a novel one-stage joint regenerator network based on the conditioned denoising diffusion probabilistic models (DDPMs) and prior knowledge exchange. To harmonize the regenerated layouts through joint optimization, we propose the interactive & enhanced guidance module for self-enhancement and prior knowledge exchange between road and building layouts during the regeneration. We also design a series of constraints from attribute-, geometry- and pixel-levels to ensure rational urban layout generation. To train our model, we build a large-scale dataset containing urban layouts and layout descriptions, covering 147K regions. Qualitative and quantitative evaluations show that our proposed method outperforms the baseline methods in regenerating desirable urban layouts that meet the textual descriptions.

AAAI Conference 2023 Short Paper

AsT: An Asymmetric-Sensitive Transformer for Osteonecrosis of the Femoral Head Detection (Student Abstract)

  • Haoyang Chen
  • Shuai Liu
  • Feng Lu
  • Wei Li
  • Bin Sheng
  • Mi Li
  • Hai Jin
  • Albert Y. Zomaya

Early diagnosis of osteonecrosis of the femoral head (ONFH) can inhibit the progression and improve femoral head preservation. The radiograph difference between early ONFH and healthy ones is not apparent to the naked eye. It is also hard to produce a large dataset to train the classification model. In this paper, we propose Asymmetric-Sensitive Transformer (AsT) to capture the uneven development of the bilateral femoral head to enable robust ONFH detection. Our ONFH detection is realized using the self-attention mechanism to femoral head regions while conferring sensitivity to the uneven development by the attention-shared transformer. The real-world experiment studies show that AsT achieves the best performance of AUC 0.9313 in the early diagnosis of ONFH and can find out misdiagnosis cases firmly.

TCS Journal 2023 Journal Article

Fixed parameterized algorithms for generalized feedback vertex set problems

  • Bin Sheng
  • Gregory Gutin

A graph is an r-pseudoforest if every connected component of it has a feedback edge set of size at most r. A graph is a d-quasi-forest if every connected component of it has a feedback vertex set of size at most d. The r-Pseudoforest Deletion problem (d-Quasi-Forest Deletion problem) asks to delete a minimum number of vertices to get an r-pseudoforest (a d-quasi-forest, respectively). The well-studied feedback vertex set problem is the special case of r-Pseudoforest Deletion (d-Quasi-Forest Deletion, respectively) in which r = 0 ( d = 0, respectively). We provide an improved FPT algorithm and a smaller kernel for r -Pseudoforest Deletion when parameterized by the solution size (when r is fixed). For d-Quasi-Forest Deletion, we show that it is FPT as well when parameterized by d and the solution size.

EAAI Journal 2023 Journal Article

Global-and-local aware network for low-light image enhancement

  • Xufeng He
  • Zhihua Chen
  • Lei Dai
  • Lei Liang
  • Jianfa Wu
  • Bin Sheng

Photos taken under nighttime or backlit conditions often suffer from complex and unpredictable degradation, such as low visibility, messy noise, and distorted color. Previous methods mainly focused on global brightness and contrast while ignoring structural and textural details, or they handled the fusion of features without adequately considering their intrinsic association, resulting in incomplete feature representations. To address this issue, we propose a global-and-local aware network (GLAN) by projecting the features into the frequency domain and incorporating them in a knowledge-sharing manner. This method effectively integrates the global modeling capability of the transformer and the local sensitivity of the convolutional neural network to represent structure and texture. First, the global branch, which is comprised of transformer blocks, performs feature extraction under the global receptive field, while the local branch constructs multi-scale features to provide local fine-grained details. Then, we design a novel adaptive multi-scale feature block (AMSFB) that deploys channel split operation to decrease the calculation amount. To better learn the channel and spatial correlations of intermediate features, we introduce a multi-scale channel attention module (MSCAM) and a pixel attention module (PAM) into the AMSFB. Finally, a frequency-aware interaction module (FAIM) is developed for bidirectional information supplementation, which builds feature descriptors simultaneously covering low-frequency and high-frequency information based on the discrete cosine transform (DCT). Through extensive quantitative and qualitative experiments, our method can achieve competitive results compared with over ten state-of-the-art image enhancement methods on eight benchmark datasets.

AAAI Conference 2022 Conference Paper

Input-Specific Robustness Certification for Randomized Smoothing

  • Ruoxin Chen
  • Jie Li
  • Junchi Yan
  • Ping Li
  • Bin Sheng

Although randomized smoothing has demonstrated high certified robustness and superior scalability to other certified defenses, the high computational overhead of the robustness certification bottlenecks the practical applicability, as it depends heavily on the large sample approximation for estimating the confidence interval. In existing works, the sample size for the confidence interval is universally set and agnostic to the input for prediction. This Input-Agnostic Sampling (IAS) scheme may yield a poor Average Certified Radius (ACR)-runtime trade-off which calls for improvement. In this paper, we propose Input-Specific Sampling (ISS) acceleration to achieve the cost-effectiveness for robustness certification, in an adaptive way of reducing the sampling size based on the input characteristic. Furthermore, our method universally controls the certified radius decline from the ISS sample size reduction. The empirical results on CIFAR-10 and ImageNet show that ISS can speed up the certification by more than three times at a limited cost of 0. 05 certified radius. Meanwhile, ISS surpasses IAS on the average certified radius across the extensive hyperparameter settings. Specifically, ISS achieves ACR=0. 958 on ImageNet in 250 minutes, compared to ACR=0. 917 by IAS under the same condition. We release our code in https: //github. com/roy-ch/Input-Specific-Certification.

AAAI Conference 2018 Conference Paper

Action Recognition With Coarse-to-Fine Deep Feature Integration and Asynchronous Fusion

  • Weiyao Lin
  • Chongyang Zhang
  • Ke Lu
  • Bin Sheng
  • Jianxin Wu
  • Bingbing Ni
  • Xin Liu
  • Hongkai Xiong

Action recognition is an important yet challenging task in computer vision. In this paper, we propose a novel deepbased framework for action recognition, which improves the recognition accuracy by: 1) deriving more precise features for representing actions, and 2) reducing the asynchrony between different information streams. We first introduce a coarse-to-fine network which extracts shared deep features at different action class granularities and progressively integrates them to obtain a more accurate feature representation for input actions. We further introduce an asynchronous fusion network. It fuses information from different streams by asynchronously integrating stream-wise features at different time points, hence better leveraging the complementary information in different streams. Experimental results on action recognition benchmarks demonstrate that our approach achieves the state-of-the-art performance.

v2026.09.13