Arrow Research search

Author name cluster

Peng Gao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

30 papers
2 author rows

Possible papers

30

JBHI Journal 2026 Journal Article

Phage Host Prediction Using Deep Neural Network With Multi-Source Protein Language Models and Squeeze-and-Excitation Attention Mechanism

  • Peng Gao
  • Long Xu
  • Yuan Bai
  • Qiuzhen Lin
  • Junkai Ji
  • Lijia Ma

Phage therapy (PT) has become a promising alternative for treating infections with the increase of antimicrobial resistance. PT utilizes phages to bind to specific receptors on bacterial surfaces via receptor-binding proteins (RBPs), enabling precise destruction of targeted hosts. In PT, a key issue is the phage host prediction (PHP), which tries to match therapeutic phages to pathogenic hosts. However, traditional PHP methods are often hindered by the time-consuming and expensive wet-lab experiments, while recent computational methods neglect the evolutionary diversity and local feature patterns of RBPs. In this article, we propose a novel deep neural network (called PHPRBP) for PHP based on phage RBPs. In PHPRBP, we first utilize pre-trained protein language models (i. e. , ESM2 and ProtT5) to learn the multi-source embedding representations from these RBPs, revealing diverse and complementary features. Then, we employ an adaptive synthetic technique to augment minority class samples, addressing the data scarcity issue. Subsequently, we design a deep neural network architecture, which uses a convolutional neural network to capture local sequence features, and applies a squeeze-and-excitation attention mechanism to enhance the contribution of important features. Finally, a fully connected network is used for host prediction. Experimental results show that PHPRBP outperforms the state-of-the-arts in host prediction at both genus and species levels.

AAAI Conference 2026 Conference Paper

Remember Me: Bridging the Long-Range Gap in LVLMs with Three-Step Inference-Only Decay Resilience Strategies

  • Peng Gao
  • Yujian Lee
  • Xiaofeng Zhang
  • Zailong Chen
  • Hui Zhang

Large Vision-Language Models (LVLMs) have achieved impressive performance across a wide range of multimodal tasks. However, they still face critical challenges in modeling long-range dependencies under the usage of Rotary Positional Encoding (ROPE). Although it can facilitate precise modeling of token positions, it induces progressive attention decay as token distance increases, especially with progressive attention decay over distant token pairs, which severely impairs the model's ability to remember global context. To alleviate this issue, we propose inference-only Three-step Decay Resilience Strategies (T-DRS), comprising (1) Semantic-Driven DRS (SD-DRS), amplifying semantically meaningful but distant signals via content-aware residuals, (2) Distance-aware Control DRS (DC-DRS), which can purify attention by smoothly modulating weights based on positional distances, suppressing noise while preserving locality, and (3) re-Reinforce Distant DRS (reRD-DRS), consolidating the remaining informative remote dependencies to maintain global coherence. Together, the T-DRS recover suppressed long-range token pairs without harming local inductive biases. Extensive experiments on Vision Question Answering (VQA) benchmarks demonstrate that T-DRS can consistently improve performance in an inference-only manner.

AAAI Conference 2026 Conference Paper

TIDE: Temporal-Aware Sparse Autoencoders for Interpretable Diffusion Transformers in Image Generation

  • Victor Shea-Jay Huang
  • Le Zhuo
  • Yi Xin
  • Zhaokai Wang
  • Fu-Yun Wang
  • Yuchi Wang
  • Renrui Zhang
  • Peng Gao

Diffusion Transformers (DiTs) are a powerful yet underexplored class of generative models compared to U-Net-based diffusion architectures. We propose TIDE—Temporal-aware sparse autoencoders for Interpretable Diffusion transformErs—a framework designed to extract sparse, interpretable activation features across timesteps in DiTs. TIDE effectively captures temporally-varying representations and reveals that DiTs naturally learn hierarchical semantics (e.g., 3D structure, object class, and fine-grained concepts) during large-scale pretraining. Experiments show that TIDE enhances interpretability and controllability while maintaining reasonable generation quality, enabling applications such as safe image editing and style transfer.

AAAI Conference 2025 Conference Paper

A Multi-Focus-Driven Multi-Branch Network for Robust Multimodal Sentiment Analysis

  • Chuanqi Tao
  • Jiaming Li
  • Tianzi Zang
  • Peng Gao

Multimodal sentiment analysis aims to integrate diverse modalities for precise emotional interpretation. However, external factors such as sensor malfunctions or network issues may disrupt certain modalities. This may lead to missing data, which poses challenges in real-world deployment. Most existing approaches focus on designing feature reconstruction strategies, overlooking the collaborative integration of reconstruction and fusion strategies. Moreover, they fail to capture the relationships between features in the global dimension and those in the local dimension. These limitations hinder the full capture of the complex nature of multimodal data, especially in scenarios involving missing modalities. To address the above issues, this paper proposes a robust model named MFMB-Net with multiple branches for feature multi-focus fusion and reconstruction. We design a two-stream fusion branch where macro-fusion focuses on the fusion of features in the global dimension and micro-fusion targets local dimension features. This dual-stream fusion branch distributes multi-focus across both pathways, simultaneously capturing global coarse-grained and local fine-grained features. Additionally, the reconstruction branch interacts collaboratively with the fusion branch to reconstruct and enhance the missing data. It integrates the reconstructed feature information with the fused information thus refining the representation fidelity of the missing information. Experiments performed on two benchmarks show that our approach obtains results superior to state-of-the-art models.

NeurIPS Conference 2025 Conference Paper

CAML: Collaborative Auxiliary Modality Learning for Multi-Agent Systems

  • Rui Liu
  • Yu Shen
  • Peng Gao
  • Pratap Tokekar
  • Ming C. Lin

Multi-modal learning has emerged as a key technique for improving performance across domains such as autonomous driving, robotics, and reasoning. However, in certain scenarios, particularly in resource-constrained environments, some modalities available during training may be absent during inference. While existing frameworks effectively utilize multiple data sources during training and enable inference with reduced modalities, they are primarily designed for single-agent settings. This poses a critical limitation in dynamic environments such as connected autonomous vehicles (CAV), where incomplete data coverage can lead to decision-making blind spots. Conversely, some works explore multi-agent collaboration but without addressing missing modality at test time. To overcome these limitations, we propose Collaborative Auxiliary Modality Learning (CAML), a novel multi-modal multi-agent framework that enables agents to collaborate and share multi-modal data during training, while allowing inference with reduced modalities during testing. Experimental results in collaborative decision-making for CAV in accident-prone scenarios demonstrate that CAML achieves up to a 58. 1% improvement in accident detection. Additionally, we validate CAML on real-world aerial-ground robot data for collaborative semantic segmentation, achieving up to a 10. 6% improvement in mIoU.

NeurIPS Conference 2025 Conference Paper

EnerVerse: Envisioning Embodied Future Space for Robotics Manipulation

  • Siyuan Huang
  • Liliang Chen
  • Pengfei Zhou
  • Shengcong Chen
  • Yue Liao
  • Zhengkai Jiang
  • Yue Hu
  • Peng Gao

We introduce EnerVerse, a generative robotics foundation model that constructs and interprets embodied spaces. EnerVerse employs a chunk-wise autoregressive video diffusion framework to predict future embodied spaces from instructions, enhanced by a sparse context memory for long-term reasoning. To model the 3D robotics world, we adopt a multi-view video representation, providing rich perspectives to address challenges like motion ambiguity and 3D grounding. Additionally, EnerVerse-D, a data engine pipeline combining generative modeling with 4D Gaussian Splatting, forms a self-reinforcing data loop to reduce the sim-to-real gap. Leveraging these innovations, EnerVerse translates 4D world representations into physical actions via a policy head (EnerVerse-A), achieving state-of-the-art performance in both simulation and real-world tasks. For efficiency, EnerVerse-A reuses features from the first denoising step and predicts action chunks, achieving about 280 ms per 8-step action chunk on a single RTX 4090. Further video demos, dataset samples could be found in our project page.

EAAI Journal 2025 Journal Article

Learning multi-level graph attentional representation for thermal infrared object tracking

  • Peng Gao
  • Shi-Min Li
  • Fei Wang
  • Hamido Fujita
  • Hanan Aljuaid
  • Ru-Yue Yuan

Thermal infrared (TIR) object tracking is a fundamental task in computer vision that is not affected by changes in lighting conditions. It performs better than visible light trackers in extreme environments such as nighttime, heavy rain, haze, and sandstorms. However, TIR object tracking also faces challenges such as occlusion, thermal crossover, motion blur, and similarity interference. Unlike visual tracking, TIR images lack color information and texture features. These factors make it challenging to learn detailed and shape features of the targets, making it hard to distinguish between targets and interference effectively. In this study, we propose a graph-based deep learning model, SiamMLGR, within the Siamese framework for stable TIR object tracking to address these issues. Specifically, to extract more fine-grained features of TIR targets, we propose a multiple graph attention module (MGAM) to replace the global matching information transmission method in the Siamese framework. This module constructs a graph structure to establish local and global connections between the target and the search area. Furthermore, to retain more of the features learned by the MGAM, we propose a spatial graph convolutional module (SGCM), which uses an explicit graph adjacency matrix to propagate information between the attention graphs. Additionally, we incorporate large-scale datasets from the visual tracking field into the model training process. By mixing these with TIR datasets, we address the sample imbalance issue present in pure TIR datasets. Extensive experimental results indicate that the proposed method achieves state-of-the-art performance.

AAAI Conference 2025 Conference Paper

LiDAR-LLM: Exploring the Potential of Large Language Models for 3D LiDAR Understanding

  • Senqiao Yang
  • Jiaming Liu
  • Renrui Zhang
  • Mingjie Pan
  • Ziyu Guo
  • Xiaoqi Li
  • Zehui Chen
  • Peng Gao

Recently, Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs) have shown promise in instruction following and image understanding. While these models are powerful, they have not yet been developed to comprehend the more challenging 3D geometric and physical scenes, especially when it comes to the sparse outdoor LiDAR data. In this paper, we introduce LiDAR-LLM, which takes raw LiDAR data as input and harnesses the remarkable reasoning capabilities of LLMs to gain a comprehensive understanding of outdoor 3D scenes. The central insight of our LiDAR-LLM is the reformulation of 3D outdoor scene cognition as a language modeling problem, encompassing tasks such as 3D captioning, 3D grounding, 3D question answering, etc. Specifically, due to the scarcity of 3D LiDAR-text pairing data, we introduce a three-stage training strategy and generate relevant datasets, progressively aligning the 3D modality with the language embedding of LLM. Furthermore, we design a Position-Aware Transformer (PAT) to connect the 3D encoder with the LLM, which effectively bridges the modality gap and enhances the LLM's spatial orientation comprehension of visual features. Our experiments demonstrate that LiDAR-LLM effectively comprehends a wide range of instructions related to 3D scenes, achieving a 40.9 BLEU-1 score on the 3D captioning dataset, a Grounded Captioning accuracy of 63.1%, and a BEV mIoU of 14.3%.

IROS Conference 2025 Conference Paper

Non-Overlap-Aware Egocentric Pose Estimation for Collaborative Perception in Connected Autonomy

  • Hong Huang
  • Dongkuan Xu
  • Hao Zhang
  • Peng Gao

Egocentric pose estimation is a fundamental capability for multi-robot collaborative perception in connected autonomy, such as connected autonomous vehicles. During multi-robot operations, a robot needs to know the relative pose between itself and its teammates with respect to its own coordinates. However, different robots usually observe completely different views that contains similar objects, which leads to wrong pose estimation. In addition, it is unrealistic to allow robots to share their raw observations to detect overlap due to the limited communication bandwidth constraint. In this paper, we introduce a novel method for Non-Overlap-Aware Egocentric Pose Estimation (NOPE), which performs egocentric pose estimation in a multi-robot team while identifying the non-overlap views and satifying the communication bandwidth constraint. NOPE is built upon an unified hierarchical learning framework that integrates two levels of robot learning: (1) high-level deep graph matching for correspondence identification, which allows to identify if two views are overlapping or not, (2) low-level position-aware cross-attention graph learning for egocentric pose estimation. To evaluate NOPE, we conduct extensive experiments in both high-fidelity simulation and real-world scenarios. Experimental results have demonstrated that NOPE enables the novel capability for non-overlapping-aware egocentric pose estimation and achieves state-of-art performance compared with the existing methods.

NeurIPS Conference 2024 Conference Paper

Lumina-Next : Making Lumina-T2X Stronger and Faster with Next-DiT

  • Le Zhuo
  • Ruoyi Du
  • Han Xiao
  • Yangguang Li
  • Dongyang Liu
  • Rongjie Huang
  • Wenze Liu
  • Xiangyang Zhu

Lumina-T2X is a nascent family of Flow-based Large Diffusion Transformers (Flag-DiT) that establishes a unified framework for transforming noise into various modalities, such as images and videos, conditioned on text instructions. Despite its promising capabilities, Lumina-T2X still encounters challenges including training instability, slow inference, and extrapolation artifacts. In this paper, we present Lumina-Next, an improved version of Lumina-T2X, showcasing stronger generation performance with increased training and inference efficiency. We begin with a comprehensive analysis of the Flag-DiT architecture and identify several suboptimal components, which we address by introducing the Next-DiT architecture with 3D RoPE and sandwich normalizations. To enable better resolution extrapolation, we thoroughly compare different context extrapolation methods applied to text-to-image generation with 3D RoPE, and propose Frequency- and Time-Aware Scaled RoPE tailored for diffusion transformers. Additionally, we introduce a sigmoid time discretization schedule for diffusion sampling, which achieves high-quality generation in 5-10 steps combined with higher-order ODE solvers. Thanks to these improvements, Lumina-Next not only improves the basic text-to-image generation but also demonstrates superior resolution extrapolation capabilities as well as multilingual generation using decoder-based LLMs as the text encoder, all in a zero-shot manner. To further validate Lumina-Next as a versatile generative framework, we instantiate it on diverse tasks including visual recognition, multi-views, audio, music, and point cloud generation, showcasing strong performance across these domains. By releasing all codes and model weights at https: //github. com/Alpha-VLLM/Lumina-T2X, we aim to advance the development of next-generation generative AI capable of universal modeling.

NeurIPS Conference 2024 Conference Paper

Phased Consistency Models

  • Fu-Yun Wang
  • Zhaoyang Huang
  • Alexander W. Bergman
  • Dazhong Shen
  • Peng Gao
  • Michael Lingelbach
  • Keqiang Sun
  • Weikang Bian

Consistency Models (CMs) have made significant progress in accelerating the generation of diffusion models. However, their application to high-resolution, text-conditioned image generation in the latent space remains unsatisfactory. In this paper, we identify three key flaws in the current design of Latent Consistency Models~(LCMs). We investigate the reasons behind these limitations and propose Phased Consistency Models (PCMs), which generalize the design space and address the identified limitations. Our evaluations demonstrate that PCMs outperform LCMs across 1--16 step generation settings. While PCMs are specifically designed for multi-step refinement, they achieve comparable 1-step generation results to previously state-of-the-art specifically designed 1-step methods. Furthermore, we show the methodology of PCMs is versatile and applicable to video generation, enabling us to train the state-of-the-art few-step text-to-video generator. Our code is available at https: //github. com/G-U-N/Phased-Consistency-Model.

AAAI Conference 2024 Conference Paper

Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

  • Shilin Yan
  • Renrui Zhang
  • Ziyu Guo
  • Wenchao Chen
  • Wei Zhang
  • Hongyang Li
  • Yu Qiao
  • Hao Dong

Recently, video object segmentation (VOS) referred by multi-modal signals, e.g., language and audio, has evoked increasing attention in both industry and academia. It is challenging for exploring the semantic alignment within modalities and the visual correspondence across frames. However, existing methods adopt separate network architectures for different modalities, and neglect the inter-frame temporal interaction with references. In this paper, we propose MUTR, a Multi-modal Unified Temporal transformer for Referring video object segmentation. With a unified framework for the first time, MUTR adopts a DETR-style transformer and is capable of segmenting video objects designated by either text or audio reference. Specifically, we introduce two strategies to fully explore the temporal relations between videos and multi-modal signals. Firstly, for low-level temporal aggregation before the transformer, we enable the multi-modal references to capture multi-scale visual cues from consecutive video frames. This effectively endows the text or audio signals with temporal knowledge and boosts the semantic alignment between modalities. Secondly, for high-level temporal interaction after the transformer, we conduct inter-frame feature communication for different object embeddings, contributing to better object-wise correspondence for tracking along the video. On Ref-YouTube-VOS and AVSBench datasets with respective text and audio references, MUTR achieves +4.2% and +8.7% J&F improvements to state-of-the-art methods, demonstrating our significance for unified multi-modal VOS. Code is released at https://github.com/OpenGVLab/MUTR.

EAAI Journal 2023 Journal Article

Real-time segmentation network for accurate weld detection in large weldments

  • Zijian Wu
  • Peng Gao
  • Jing Han
  • Lianfa Bai
  • Jun Lu
  • Zhuang Zhao

Aiming at the defects of inaccurate weld extraction and high matching error rate in automatic welding system of large weldments currently. We propose a multi task detection model based on CNN architecture, which integrates the semantic segmentation technology required for weldment merging as well as the edge detection technology needed for weld matching. In particular, for the purpose of predicting smoother edges and welds, we carefully construct a new segment head, which adopts the sub-pixel convolution technology for up-sampling. Furthermore, a joint optimization loss function is explored to alleviate the imbalance of category distribution in large-scale weldment datasets. To verify the effectiveness of the model, abundant groups of data are collected for training and testing. The experimental results indicate that the proposed method has achieved the optimal trade-off between detection accuracy (83. 35% mIoU, 95. 15% F-score of welds and edges) as well as speed (74FPS) on a 2080Ti GPU compared with other state-of-the-arts, which greatly improves the robustness of the automatic welding system for large weldments.

AAAI Conference 2023 Conference Paper

Resilient Binary Neural Network

  • Sheng Xu
  • Yanjing Li
  • Teli Ma
  • Mingbao Lin
  • Hao Dong
  • Baochang Zhang
  • Peng Gao
  • Jinhu Lu

Binary neural networks (BNNs) have received ever-increasing popularity for their great capability of reducing storage burden as well as quickening inference time. However, there is a severe performance drop compared with {real-valued} networks, due to its intrinsic frequent weight oscillation during training. In this paper, we introduce a Resilient Binary Neural Network (ReBNN) to mitigate the frequent oscillation for better BNNs' training. We identify that the weight oscillation mainly stems from the non-parametric scaling factor. To address this issue, we propose to parameterize the scaling factor and introduce a weighted reconstruction loss to build an adaptive training objective. For the first time, we show that the weight oscillation is controlled by the balanced parameter attached to the reconstruction loss, which provides a theoretical foundation to parameterize it in back propagation. Based on this, we learn our ReBNN by calculating the balanced parameter based on its maximum magnitude, which can effectively mitigate the weight oscillation with a resilient training process. Extensive experiments are conducted upon various network models, such as ResNet and Faster-RCNN for computer vision, as well as BERT for natural language processing. The results demonstrate the overwhelming performance of our ReBNN over prior arts. For example, our ReBNN achieves 66.9% Top-1 accuracy with ResNet-18 backbone on the ImageNet dataset, surpassing existing state-of-the-arts by a significant margin. Our code is open-sourced at https://github.com/SteveTsui/ReBNN.

YNICL Journal 2022 Journal Article

Altered frequency-specific/universal amplitude characteristics of spontaneous brain oscillations in patients with bipolar disorder

  • Zhi-Fang Zhang
  • Qi-Jing Bo
  • Feng Li
  • Lei Zhao
  • Peng Gao
  • Yun Wang
  • Rui Liu
  • Xiong-Ying Chen

The human brain is a dynamic system with intrinsic oscillations in spontaneous neural activity. Whether the dynamic characteristics of these spontaneous oscillations are differentially altered across different frequency bands in patients with bipolar disorder (BD) remains unclear. This study recruited 65 patients with BD and 85 healthy controls (HCs). The entire frequency range of resting-state fMRI data was decomposed into four frequency intervals. Two-way repeated-measures ANCOVA was employed to detect frequency-specific/universal alterations in the dynamic oscillation amplitude in BD. The patients were then divided into two subgroups according to their mood states to explore whether these alterations were independent of their mood states. Finally, other window sizes, step sizes, and window types were tested to replicate all analyses. Frequency-specific abnormality of the dynamic oscillation amplitude was detected within the posterior medial parietal cortex (centered at the precuneus extending to the posterior cingulate cortex). This specific profile indicates decreased amplitudes in the lower frequency bands (slow-5/4) and no amplitude changes in the higher frequency bands (slow-3/2) compared with HCs. Frequency-universal abnormalities of the dynamic oscillation amplitude were also detectable, indicating increased amplitudes in the thalamus and left cerebellum anterior lobe but decreased amplitudes in the medial superior frontal gyrus. These alterations were independent of the patients' mood states and replicable across multiple analytic and parametric settings. In short, frequency-specific/universal amplitude characteristics of spontaneous oscillations were observed in patients with BD. These abnormal characteristics have important implications for specific functional changes in BD from multiple frequency and dynamic perspectives.

JBHI Journal 2022 Journal Article

Automatic Lung Nodule Segmentation and Intra-Nodular Heterogeneity Image Generation

  • Jiangdian Song
  • Shih-Cheng Huang
  • Brendan Kelly
  • Guanqun Liao
  • Jingyun Shi
  • Ning Wu
  • Weimin Li
  • Zaiyi Liu

Automatic segmentation of lung nodules on computed tomography (CT) images is challenging owing to the variability of morphology, location, and intensity. In addition, few segmentation methods can capture intra-nodular heterogeneity to assist lung nodule diagnosis. In this study, we propose an end-to-end architecture to perform fully automated segmentation of multiple types of lung nodules and generate intra-nodular heterogeneity images for clinical use. To this end, a hybrid loss is considered by introducing a Faster R-CNN model based on generalized intersection over union loss in generative adversarial network. The Lung Image Database Consortium image collection dataset, comprising 2, 635 lung nodules, was combined with 3, 200 lung nodules from five hospitals for this study. Compared with manual segmentation by radiologists, the proposed model obtained an average dice coefficient (DC) of 82. 05% on the test dataset. Compared with U-net, NoduleNet, nnU-net, and other three models, the proposed method achieved comparable performance on lung nodule segmentation and generated more vivid and valid intra-nodular heterogeneity images, which are beneficial in radiological diagnosis. In an external test of 91 patients from another hospital, the proposed model achieved an average DC of 81. 61%. The proposed method effectively addresses the challenges of inevitable human interaction and additional pre-processing procedures in the existing solutions for lung nodule segmentation. In addition, the results show that the intra-nodular heterogeneity images generated by the proposed model are suitable to facilitate lung nodule diagnosis in radiology.

NeurIPS Conference 2022 Conference Paper

MCMAE: Masked Convolution Meets Masked Autoencoders

  • Peng Gao
  • Teli Ma
  • Hongsheng Li
  • Ziyi Lin
  • Jifeng Dai
  • Yu Qiao

Vision Transformers (ViT) become widely-adopted architectures for various vision tasks. Masked auto-encoding for feature pretraining and multi-scale hybrid convolution-transformer architectures can further unleash the potentials of ViT, leading to state-of-the-art performances on image classification, detection and semantic segmentation. In this paper, our MCMAE framework demonstrates that multi-scale hybrid convolution-transformer can learn more discriminative representations via the mask auto-encoding scheme. However, directly using the original masking strategy leads to the heavy computational cost and pretraining-finetuning discrepancy. To tackle the issue, we adopt the masked convolution to prevent information leakage in the convolution blocks. A simple block-wise masking strategy is proposed to ensure computational efficiency. We also propose to more directly supervise the multi-scale features of the encoder to boost multi-scale features. Based on our pretrained MCMAE models, MCMAE-Base improves ImageNet-1K finetuning accuracy by 1. 4% compared with MAE-Base. On object detection, MCMAE-Base finetuned for only 25 epochs surpasses MAE-Base fined-tuned for 100 epochs by 2. 9% box AP and 2. 2% mask AP respectively. Code and pretrained models are available at \url{https: //github. com/Alpha-VL/ConvMAE}.

NeurIPS Conference 2022 Conference Paper

Point-M2AE: Multi-scale Masked Autoencoders for Hierarchical Point Cloud Pre-training

  • Renrui Zhang
  • Ziyu Guo
  • Peng Gao
  • Rongyao Fang
  • Bin Zhao
  • Dong Wang
  • Yu Qiao
  • Hongsheng Li

Masked Autoencoders (MAE) have shown great potentials in self-supervised pre-training for language and 2D image transformers. However, it still remains an open question on how to exploit masked autoencoding for learning 3D representations of irregular point clouds. In this paper, we propose Point-M2AE, a strong Multi-scale MAE pre-training framework for hierarchical self-supervised learning of 3D point clouds. Unlike the standard transformer in MAE, we modify the encoder and decoder into pyramid architectures to progressively model spatial geometries and capture both fine-grained and high-level semantics of 3D shapes. For the encoder that downsamples point tokens by stages, we design a multi-scale masking strategy to generate consistent visible regions across scales, and adopt a local spatial self-attention mechanism during fine-tuning to focus on neighboring patterns. By multi-scale token propagation, the lightweight decoder gradually upsamples point tokens with complementary skip connections from the encoder, which further promotes the reconstruction from a global-to-local perspective. Extensive experiments demonstrate the state-of-the-art performance of Point-M2AE for 3D representation learning. With a frozen encoder after pre-training, Point-M2AE achieves 92. 9% accuracy for linear SVM on ModelNet40, even surpassing some fully trained methods. By fine-tuning on downstream tasks, Point-M2AE achieves 86. 43% accuracy on ScanObjectNN, +3. 36% to the second-best, and largely benefits the few-shot classification, part segmentation and 3D object detection with the hierarchical pre-training scheme. Code is available at https: //github. com/ZrrSkywalker/Point-M2AE.

NeurIPS Conference 2022 Conference Paper

Q-ViT: Accurate and Fully Quantized Low-bit Vision Transformer

  • Yanjing Li
  • Sheng Xu
  • Baochang Zhang
  • Xianbin Cao
  • Peng Gao
  • Guodong Guo

The large pre-trained vision transformers (ViTs) have demonstrated remarkable performance on various visual tasks, but suffer from expensive computational and memory cost problems when deployed on resource-constrained devices. Among the powerful compression approaches, quantization extremely reduces the computation and memory consumption by low-bit parameters and bit-wise operations. However, low-bit ViTs remain largely unexplored and usually suffer from a significant performance drop compared with the real-valued counterparts. In this work, through extensive empirical analysis, we first identify the bottleneck for severe performance drop comes from the information distortion of the low-bit quantized self-attention map. We then develop an information rectification module (IRM) and a distribution guided distillation (DGD) scheme for fully quantized vision transformers (Q-ViT) to effectively eliminate such distortion, leading to a fully quantized ViTs. We evaluate our methods on popular DeiT and Swin backbones. Extensive experimental results show that our method achieves a much better performance than the prior arts. For example, our Q-ViT can theoretically accelerates the ViT-S by 6. 14x and achieves about 80. 9% Top-1 accuracy, even surpassing the full-precision counterpart by 1. 0% on ImageNet dataset. Our codes and models are attached on https: //github. com/YanjingLi0202/Q-ViT

NeurIPS Conference 2021 Conference Paper

Container: Context Aggregation Networks

  • Peng Gao
  • Jiasen Lu
  • Hongsheng Li
  • Roozbeh Mottaghi
  • Aniruddha Kembhavi

Convolutional neural networks (CNNs) are ubiquitous in computer vision, with a myriad of effective and efficient variations. Recently, Transformers -- originally introduced in natural language processing -- have been increasingly adopted in computer vision. While early adopters continued to employ CNN backbones, the latest networks are end-to-end CNN-free Transformer solutions. A recent surprising finding now shows that a simple MLP based solution without any traditional convolutional or Transformer components can produce effective visual representations. While CNNs, Transformers and MLP-Mixers may be considered as completely disparate architectures, we provide a unified view showing that they are in fact special cases of a more general method to aggregate spatial context in a neural network stack. We present the \model (CONText AggregatIon NEtwoRk), a general-purpose building block for multi-head context aggregation that can exploit long-range interactions \emph{a la} Transformers while still exploiting the inductive bias of the local convolution operation leading to faster convergence speeds, often seen in CNNs. Our \model architecture achieves 82. 7 \% Top-1 accuracy on ImageNet using 22M parameters, +2. 8 improvement compared with DeiT-Small, and can converge to 79. 9 \% Top-1 accuracy in just 200 epochs. In contrast to Transformer-based methods that do not scale well to downstream tasks that rely on larger input image resolutions, our efficient network, named \modellight, can be employed in object detection and instance segmentation networks such as DETR, RetinaNet and Mask-RCNN to obtain an impressive detection mAP of 38. 9, 43. 8, 45. 1 and mask mAP of 41. 3, providing large improvements of 6. 6, 7. 3, 6. 9 and 6. 6 pts respectively, compared to a ResNet-50 backbone with a comparable compute and parameter size. Our method also achieves promising results on self-supervised learning compared to DeiT on the DINO framework. Code is released at https: //github. com/allenai/container.

NeurIPS Conference 2021 Conference Paper

Dual-stream Network for Visual Recognition

  • Mingyuan Mao
  • Peng Gao
  • Renrui Zhang
  • Honghui Zheng
  • Teli Ma
  • Yan Peng
  • Errui Ding
  • Baochang Zhang

Transformers with remarkable global representation capacities achieve competitive results for visual tasks, but fail to consider high-level local pattern information in input images. In this paper, we present a generic Dual-stream Network (DS-Net) to fully explore the representation capacity of local and global pattern features for image classification. Our DS-Net can simultaneously calculate fine-grained and integrated features and efficiently fuse them. Specifically, we propose an Intra-scale Propagation module to process two different resolutions in each block and an Inter-Scale Alignment module to perform information interaction across features at dual scales. Besides, we also design a Dual-stream FPN (DS-FPN) to further enhance contextual information for downstream dense predictions. Without bells and whistles, the proposed DS-Net outperforms DeiT-Small by 2. 4\% in terms of top-1 accuracy on ImageNet-1k and achieves state-of-the-art performance over other Vision Transformers and ResNets. For object detection and instance segmentation, DS-Net-Small respectively outperforms ResNet-50 by 6. 4\% and 5. 5 \% in terms of mAP on MSCOCO 2017, and surpasses the previous state-of-the-art scheme, which significantly demonstrates its potential to be a general backbone in vision tasks. The code will be released soon.

AAAI Conference 2021 Conference Paper

Dynamic Graph Representation Learning for Video Dialog via Multi-Modal Shuffled Transformers

  • Shijie Geng
  • Peng Gao
  • Moitreya Chatterjee
  • Chiori Hori
  • Jonathan Le Roux
  • Yongfeng Zhang
  • Hongsheng Li
  • Anoop Cherian

Given an input video, its associated audio, and a brief caption, the audio-visual scene aware dialog (AVSD) task requires an agent to indulge in a question-answer dialog with a human about the audio-visual content. This task thus poses a challenging multi-modal representation learning and reasoning scenario, advancements into which could influence several human-machine interaction applications. To solve this task, we introduce a semantics-controlled multi-modal shuffled Transformer reasoning framework, consisting of a sequence of Transformer modules, each taking a modality as input and producing representations conditioned on the input question. Our proposed Transformer variant uses a shuffling scheme on their multi-head outputs, demonstrating better regularization. To encode fine-grained visual information, we present a novel dynamic scene graph representation learning pipeline that consists of an intra-frame reasoning layer producing spatio-semantic graph representations for every frame, and an inter-frame aggregation module capturing temporal cues. Our entire pipeline is trained end-to-end. We present experiments on the benchmark AVSD dataset, both on answer generation and selection tasks. Our results demonstrate state-of-the-art performances on all evaluation metrics.

IJCAI Conference 2021 Conference Paper

Pairwise Half-graph Discrimination: A Simple Graph-level Self-supervised Strategy for Pre-training Graph Neural Networks

  • Pengyong Li
  • Jun Wang
  • Ziliang Li
  • Yixuan Qiao
  • Xianggen Liu
  • Fei Ma
  • Peng Gao
  • Sen Song

Self-supervised learning has gradually emerged as a powerful technique for graph representation learning. However, transferable, generalizable, and robust representation learning on graph data still remains a challenge for pre-training graph neural networks. In this paper, we propose a simple and effective self-supervised pre-training strategy, named Pairwise Half-graph Discrimination (PHD), that explicitly pre-trains a graph neural network at graph-level. PHD is designed as a simple binary classification task to discriminate whether two half-graphs come from the same source. Experiments demonstrate that the PHD is an effective pre-training strategy that offers comparable or superior performance on 13 graph classification tasks compared with state-of-the-art strategies, and achieves notable improvements when combined with node-level strategies. Moreover, the visualization of learned representation revealed that PHD strategy indeed empowers the model to learn graph-level knowledge like the molecular scaffold. These results have established PHD as a powerful and effective self-supervised learning strategy in graph-level representation learning.

JBHI Journal 2020 Journal Article

Aceso: PICO-Guided Evidence Summarization on Medical Literature

  • Xiang Zhang
  • Ping Geng
  • Tengteng Zhang
  • Qian Lu
  • Peng Gao
  • Jing Mei

Evidence-Based Medicine (EBM) aims to apply the best available evidence gained from scientific methods to clinical decision making. A generally accepted criterion to formulate evidence is to use the PICO framework, where PICO stands for Problem/Population, Intervention, Comparison, and Outcome. Automatic extraction of PICO-related sentences from medical literature is crucial to the success of many EBM applications. In this work, we present our Aceso 1 system, which automatically generates PICO-based evidence summaries from medical literature. In Aceso, we adopt an active learning paradigm, which helps to minimize the cost of manual labeling and to optimize the quality of summarization with limited labeled data. An UMLS2Vec model is proposed to learn a vector representation of medical concepts in UMLS, 2 and we fuse the embedding of medical knowledge with textual features in summarization. The evaluation shows that our approach is better on identifying PICO sentences against state-of-the-art studies and outperforms baseline methods on producing high-quality evidence summaries.

AAAI Conference 2020 Conference Paper

Long-Term Loop Closure Detection through Visual-Spatial Information Preserving Multi-Order Graph Matching

  • Peng Gao
  • Hao Zhang

Loop closure detection is a fundamental problem for simultaneous localization and mapping (SLAM) in robotics. Most of the previous methods only consider one type of information, based on either visual appearances or spatial relationships of landmarks. In this paper, we introduce a novel visual-spatial information preserving multi-order graph matching approach for long-term loop closure detection. Our approach constructs a graph representation of a place from an input image to integrate visual-spatial information, including visual appearances of the landmarks and the background environment, as well as the second and third-order spatial relationships between two and three landmarks, respectively. Furthermore, we introduce a new formulation that formulates loop closure detection as a multi-order graph matching problem to compute a similarity score directly from the graph representations of the query and template images, instead of performing conventional vectorbased image matching. We evaluate the proposed multi-order graph matching approach based on two public long-term loop closure detection benchmark datasets, including the St. Lucia and CMU-VL datasets. Experimental results have shown that our approach is effective for long-term loop closure detection and it outperforms the previous state-of-the-art methods.

AAAI Conference 2020 Conference Paper

Region Focus Network for Joint Optic Disc and Cup Segmentation

  • Ge Li
  • Changsheng Li
  • Chan Zeng
  • Peng Gao
  • Guotong Xie

Glaucoma is one of the three leading causes of blindness in the world and is predicted to affect around 80 million people by 2020. The optic cup (OC) to optic disc (OD) ratio (CDR) in fundus images plays a pivotal role in the screening and diagnosis of glaucoma. Existing methods usually crop the optic disc region first, and subsequently perform segmentation in this region. However, these approaches come up with high complexities due to the separate operations. To remedy this issue, we propose a Region Focus Network (RF-Net) that innovatively integrates detection and multi-class segmentation into a unified architecture for end-to-end joint optic disc and cup segmentation with global optimization. The key idea of our method is designing a novel multi-class mask branch which generates a high-quality segmentation in the detected region for both disc and cup. To bridge the connection between the backbone and multi-class mask branch, a Fusion Feature Pooling (FFP) structure is presented to extract features from each level of the pyramid network and fuse them into a final feature representation for segmentation. Extensive experimental results on the REFUGE-2018 challenge dataset and the Drishti-GS dataset show that the proposed method achieves the best performance, compared with competitive approaches reported in the literature and the official leaderboard. Our code will be released soon.

AAAI Conference 2019 Conference Paper

Video Object Detection with Locally-Weighted Deformable Neighbors

  • Zhengkai Jiang
  • Peng Gao
  • Chaoxu Guo
  • Qian Zhang
  • Shiming Xiang
  • Chunhong Pan

Deep convolutional neural networks have achieved great success on various image recognition tasks. However, it is nontrivial to transfer the existing networks to video due to the fact that most of them are developed for static image. Frame-byframe processing is suboptimal because temporal information that is vital for video understanding is totally abandoned. Furthermore, frame-by-frame processing is slow and inefficient, which can hinder the practical usage. In this paper, we propose LWDN (Locally-Weighted Deformable Neighbors) for video object detection without utilizing time-consuming optical flow extraction networks. LWDN can latently align the high-level features between keyframes and keyframes or nonkeyframes. Inspired by (Zhu et al. 2017a) and (Hetang et al. 2017) who propose to aggregate features between keyframes and keyframes, we adopt brain-inspired memory mechanism to propagate and update the memory feature from keyframes to keyframes. We call this process Memory-Guided Propagation. With such a memory mechanism, the discriminative ability of features in keyframes and non-keyframes are both enhanced, which helps to improve the detection accuracy. Extensive experiments on VID dataset demonstrate that our method achieves superior performance in a speed and accuracy trade-off, i. e. , 76. 3% on the challenging VID dataset while maintaining 20fps in speed on Titan X GPU.

IJCAI Conference 2018 Conference Paper

Dynamic Bayesian Logistic Matrix Factorization for Recommendation with Implicit Feedback

  • Yong Liu
  • Lifan Zhao
  • Guimei Liu
  • Xinyan Lu
  • Peng Gao
  • Xiao-li Li
  • Zhihui Jin

Matrix factorization has been widely adopted for recommendation by learning latent embeddings of users and items from observed user-item interaction data. However, previous methods usually assume the learned embeddings are static or homogeneously evolving with the same diffusion rate. This is not valid in most scenarios, where users’ preferences and item attributes heterogeneously drift over time. To remedy this issue, we have proposed a novel dynamic matrix factorization model, named Dynamic Bayesian Logistic Matrix Factorization (DBLMF), which aims to learn heterogeneous user and item embeddings that are drifting with inconsistent diffusion rates. More specifically, DBLMF extends logistic matrix factorization to model the probability a user would like to interact with an item at a given timestamp, and a diffusion process to connect latent embeddings over time. In addition, an efficient Bayesian inference algorithm has also been proposed to make DBLMF scalable on large datasets. The effectiveness of the proposed method has been demonstrated by extensive experiments on real datasets, compared with the state-of-the-art methods.

AAAI Conference 2017 Conference Paper

SCOPE: Scalable Composite Optimization for Learning on Spark

  • Shen-Yi Zhao
  • Ru Xiang
  • Ying-Hao Shi
  • Peng Gao
  • Wu-Jun Li

Many machine learning models, such as logistic regression (LR) and support vector machine (SVM), can be formulated as composite optimization problems. Recently, many distributed stochastic optimization (DSO) methods have been proposed to solve the large-scale composite optimization problems, which have shown better performance than traditional batch methods. However, most of these DSO methods might not be scalable enough. In this paper, we propose a novel DSO method, called scalable composite optimization for learning (SCOPE), and implement it on the fault-tolerant distributed platform Spark. SCOPE is both computation-efficient and communication-efficient. Theoretical analysis shows that SCOPE is convergent with linear convergence rate when the objective function is strongly convex. Furthermore, empirical results on real datasets show that SCOPE can outperform other state-of-the-art distributed learning methods on Spark, including both batch learning methods and DSO methods.

ICRA Conference 2014 Conference Paper

Motion planning with Satisfiability Modulo Theories

  • William N. N. Hung
  • Xiaoyu Song
  • Jindong Tan
  • Xiaojuan Li
  • Jie Zhang 0074
  • Rui Wang 0024
  • Peng Gao

Motion planning is an important problem with many applications in robotics. In this paper, we focus on motion planning with rectangular obstacles parallel to the X, Y or Z axis. We formulate motion planning using Satisfiability Modulo Theories (SMT) and use SMT solvers to find a feasible path from the source to the goal. Our formulation decompose the robotic path into N path segments where the two ends of each path segment can be constrained using difference logic. Our SMT approach will find a solution if and only if a feasible path exists for the given constraints. We present extensive experimental results to demonstrate the scalability of our approach.

v2026.09.13