Arrow Research search

Author name cluster

Bo Xu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

58 papers
2 author rows

Possible papers

58

AAAI Conference 2026 Conference Paper

KNNDA: A New Perspective of Alignment Recovery for Partially View-Aligned Clustering

  • Liang Zhao
  • Tianqi Yue
  • Shubin Ma
  • Ziyue Wang
  • Zhiyuan Liu
  • Bo Xu

In multi-view clustering (MVC), complementary and consistent information from multiple views is integrated to improve clustering performance. However, inter-view sample correspondences may be partially missing in practice, making it difficult to learn cross-view consistency, which leads to the partially view-aligned problem (PVP). Most existing partially view-aligned clustering (PVC) methods first learn cross-view consistent representations based on known alignments, and then recover missing correspondences by measuring cross-view similarity between samples. However, such an indirect alignment recovery process depends on high-quality consistent representations and lacks effective utilization of known alignments, often resulting in sub-optimal outcomes. To address this, we propose a novel direct alignment recovery perspective, instantiated as K-Nearest Neighbors Direct Alignment (KNNDA). Specifically, we first construct an alignment domain by mapping the aligned neighbors of each unaligned sample into the aligned view. Then, we compute alignment confidence based on the similarity between known aligned pairs of neighbors. In particular, we use a dynamic threshold to filter out unreliable alignments. Finally, new alignments are generated within the high-confidence alignment domain. Contrastive loss is used to learn consistent representations for clustering. Comprehensive experiments on several real-world datasets show the effectiveness and superiority of our module in partially view-aligned clustering.

AAAI Conference 2026 Conference Paper

MetaGPT: A Large Vision-Language Model for Meme Metaphor Understanding

  • Bo Xu
  • Chenyuan Wang
  • Xinyu Chen
  • Hongfei Lin
  • Feng Xia

Meme is an expressive medium that often conveys rich emotions and intentions. Recent studies have confirmed the critical role of metaphors in meme understanding. However, existing metaphor research heavily relies on manual annotations, and mainstream vision-language models (VLMs) still struggle with the recognition and comprehension of metaphors. To address these challenges, we introduce MetaGPT, the first vision-language model specifically designed for meme metaphor understanding. MetaGPT is capable of identifying and extracting metaphors in memes, and generating accurate meme interpretations. Furthermore, we construct a dedicated dataset for meme understanding, MUnd, which comprises approximately 32,000 high-quality question-answer (QA) pairs across three core tasks: metaphor detection, metaphor domain extraction, and meme interpretation. Based on MUnd, we further propose an evaluation benchmark for meme understanding and conduct a comprehensive assessment of existing VLMs. Experimental results reveal that current models still face challenges in metaphor comprehension, while MetaGPT consistently outperforms them across all tasks, highlighting its potential in advancing meme understanding.

AAAI Conference 2026 Conference Paper

MrCoM: A Meta-Regularized World-Model Generalizing Across Multi-Scenarios

  • Xuantang Xiong
  • Ni Mu
  • Runpeng Xie
  • Senhao Yang
  • Yaqing Wang
  • Lexiang Wang
  • Yao Luan
  • Siyuan Li

Model-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL methods focus primarily on building world models for single tasks and rarely address generalization across different scenarios. Building on the insight that dynamics within the same simulation engine share inherent properties, we attempt to construct a unified world model capable of generalizing across different scenarios, named Meta-Regularized Contextual World-Model (MrCoM). This method first decomposes the latent state space into various components based on the dynamic characteristics, thereby enhancing the accuracy of world-model prediction. Further, MrCoM adopts meta-state regularization to extract unified representation of scenario-relevant information, and meta-value regularization to align world-model optimization with policy learning across diverse scenario objectives. We theoretically analyze the generalization error upper bound of MrCoM in multi-scenario settings. We systematically evaluate our algorithm's generalization ability across diverse scenarios, demonstrating significantly better performance than previous state-of-the-art methods.

AAAI Conference 2026 Conference Paper

Speech-Aware Long Context Pruning and Integration for Contextualized Automatic Speech Recognition

  • Yiming Rong
  • Yixin Zhang
  • Ziyi Wang
  • Deyang Jiang
  • Yunlong Zhao
  • Haoran Wu
  • Shiyu Zhou
  • Bo Xu

Automatic speech recognition (ASR) systems have achieved remarkable performance in common conditions but often struggle to leverage long-context information in contextualized scenarios that require domain-specific knowledge, such as conference presentations. This challenge arises primarily due to constrained model context windows and the sparsity of relevant information within extensive contextual noise. To solve this, we propose the SAP^2 method, a novel framework that dynamically prunes and integrates relevant contextual keywords in two stages. Specifically, each stage leverages our proposed Speech-Driven Attention-based Pooling mechanism, enabling efficient compression of context embeddings while preserving speech-salient information. Experimental results demonstrate state-of-the-art performance of SAP^2 on the SlideSpeech and LibriSpeech datasets, achieving word error rates (WER) of 7.71% and 1.12%, respectively. On SlideSpeech, our method notably reduces biased keyword error rates (B-WER) by 41.1% compared to non-contextual baselines. SAP^2 also exhibits robust scalability, consistently maintaining performance under extensive contextual input conditions on both datasets.

TMLR Journal 2026 Journal Article

SpikingBrain: Spiking Brain-inspired Large Models

  • Yuqi Pan
  • Yupeng Feng
  • JingHao Zhuang
  • siyu ding
  • Han Xu
  • Zehao Liu
  • Bohan Sun
  • Yuhong Chou

Mainstream Transformer-based large language models (LLMs) face significant efficiency bottlenecks: training computation scales quadratically with sequence length, and inference memory grows linearly. These constraints limit their ability to process long sequences effectively. In addition, building large models on non-NVIDIA computing platforms poses major challenges in achieving stable and efficient training and deployment. To address these issues, we introduce SpikingBrain, a new family of brain-inspired models designed for efficient long-context training and inference. SpikingBrain leverages the MetaX GPU cluster and focuses on three core aspects: (1) Model Architecture: linear and hybrid-linear attention architectures with adaptive spiking neurons; (2) Algorithmic Optimizations: an efficient, conversion-based training pipeline compatible with existing LLMs, along with a dedicated spike coding framework; (3) System Engineering: customized training frameworks, operator libraries, and parallelism strategies tailored to the MetaX hardware. Using these techniques, we develop two models: SpikingBrain-7B, a linear LLM, and SpikingBrain-76B, a hybrid-linear MoE LLM. These models demonstrate the feasibility of large-scale LLM development on non-NVIDIA platforms, and our training framework supports weeks of stable training on hundreds of MetaX GPUs with Model FLOPs Utilization (MFU) at expected levels. SpikingBrain achieves performance comparable to open-source Transformer baselines while using exceptionally low data resources (continual pre-training of approximately 150B tokens). Our models also significantly improve long-context efficiency and deliver inference with (partially) constant memory and event-driven spiking behavior. For example, SpikingBrain-7B achieves more than 100× speedup in Time to First Token (TTFT) for 4M-token sequences. Furthermore, the proposed spiking scheme achieves 69.15% sparsity, enabling low-power operation. Overall, this work demonstrates the potential of brain-inspired mechanisms to drive the next generation of efficient and scalable large model design.

AAAI Conference 2026 Conference Paper

TinyChemVL: Advancing Chemical Vision-Language Models via Efficient Visual Token Reduction and Complex Reaction Tasks

  • Xuanle Zhao
  • Shuxin Zeng
  • Xinyuan Cai
  • Xiang Cheng
  • Duzhen Zhang
  • Xiuyi Chen
  • Bo Xu

While Vision Language Models (VLMs) have demonstrated remarkable capabilities in general visual understanding, their application in the chemical domain has been limited, with previous works predominantly focusing on text and thus overlooking critical visual information, such as molecular structures. Current approaches that directly adopt standard VLMs for chemical tasks suffer from two primary issues: (i) computational inefficiency of processing entire chemical images with non-informative backgrounds. (ii) a narrow scope on molecular-level tasks that restricts progress in chemical reasoning. In this work, we propose TinyChemVL, an efficient and powerful chemical VLM that leverages visual token reduction and reaction-level tasks to improve model efficiency and reasoning capacity. Also, we propose ChemRxn-V, a reaction-level benchmark for assessing vision-based reaction recognition and prediction tasks. Directly predicting reaction products from molecular images poses a non-trivial challenge, as it requires models to integrate both recognition and reasoning capacities. Our results demonstrate that, with only 4B parameters, TinyChemVL achieves superior performance on both molecular and reaction tasks, while also demonstrating faster inference and training speeds compared to existing models. Notably, TinyChemVL outperforms ChemVLM while utilizing only 1/16th of the visual tokens. This work builds efficient yet powerful VLMs for chemical domains by co-designing model architecture and task complexity.

NeurIPS Conference 2025 Conference Paper

4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular Videos

  • Mengqi Guo
  • Bo Xu
  • Yanyan Li
  • Gim Hee Lee

Novel view synthesis from monocular videos of dynamic scenes with unknown camera poses remains a fundamental challenge in computer vision and graphics. While recent advances in 3D representations such as Neural Radiance Fields (NeRF) and 3D Gaussian Splatting (3DGS) have shown promising results for static scenes, they struggle with dynamic content and typically rely on pre-computed camera poses. We present 4D3R, a pose-free dynamic neural rendering framework that decouples static and dynamic components through a two-stage approach. Our method first leverages 3D foundational models for initial pose and geometry estimation, followed by motion-aware refinement. 4D3R introduces two key technical innovations: (1) a motion-aware bundle adjustment (MA-BA) module that combines transformer-based learned priors with SAM2 for robust dynamic object segmentation, enabling more accurate camera pose refinement; and (2) an efficient Motion-Aware Gaussian Splatting (MA-GS) representation that uses control points with a deformation field MLP and linear blend skinning to model dynamic motion, significantly reducing computational cost while maintaining high-quality reconstruction. Extensive experiments on real-world dynamic datasets demonstrate that our approach achieves up to 1. 8dB PSNR improvement over state-of-the-art methods, particularly in challenging scenarios with large dynamic objects, while reducing computational requirements by 5× compared to previous dynamic scene representations.

IJCAI Conference 2025 Conference Paper

Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search

  • Shubin Ma
  • Liang Zhao
  • Mingdong Lu
  • Yifan guo
  • Bo Xu

Multi-modal representation is faithful and highly effective in describing real-world data samples' characteristics by describing their complementary information. However, the collected data often exhibits incomplete and misaligned characteristics due to factors such as inconsistent sensor frequencies and device malfunctions. Existing research has not effectively addressed the issue of filling missing data in scenarios where multiview data are both imbalanced and misaligned. Instead, it relies on class-level alignment of the available data. Thus, it results in some data samples not being well-matched, thereby affecting the quality of data fusion. In this paper, we propose the Consistency-Aware Padding for Incomplete Multi-Modal Alignment Clustering Based on Self-Repellent Greedy Anchor Search(CAPIMAC) to tackle the problem of filling imbalanced and misaligned data in multi-modal datasets. Specifically, we propose a self-repellent greedy anchor search module(SRGASM), which employs a self-repellent random walk combined with a greedy algorithm to identify anchor points for re-representing incomplete and misaligned multi-modal data. Subsequently, based on noise-contrastive learning, we design a consistency-aware padding module (CAPM) to effectively interpolate and align imbalanced and misaligned data, thereby improving the quality of multi-modal data fusion. Experimental results demonstrate the superiority of our method over benchmark datasets. The code will be publicly released at https: //github. com/bestow09090/-CAPIMAC. git.

NeurIPS Conference 2025 Conference Paper

DAIL: Beyond Task Ambiguity for Language-Conditioned Reinforcement Learning

  • Runpeng Xie
  • Quanwei Wang
  • Hao Hu
  • Zherui Zhou
  • Ni Mu
  • Xiyun Li
  • Yiqin Yang
  • Shuang Xu

Comprehending natural language and following human instructions are critical capabilities for intelligent agents. However, the flexibility of linguistic instructions induces substantial ambiguity across language-conditioned tasks, severely degrading algorithmic performance. To address these limitations, we present a novel method named DAIL (Distributional Aligned Learning), featuring two key components: distributional policy and semantic alignment. Specifically, we provide theoretical results that the value distribution estimation mechanism enhances task differentiability. Meanwhile, the semantic alignment module captures the correspondence between trajectories and linguistic instructions. Extensive experimental results on both structured and visual observation benchmarks demonstrate that DAIL effectively resolves instruction ambiguities, achieving superior performance to baseline methods. Our implementation is available at https: //github. com/RunpengXie/Distributional-Aligned-Learning.

EAAI Journal 2025 Journal Article

Digital twin-driven reinforcement learning-based operational management for customized manufacturing

  • Hao Tang
  • Minghao Cheng
  • Uzair Aslam Bhatti
  • Bo Xu
  • Nan Zhou
  • Rong Guo
  • Bing Wei

Due to the increasing complexity of customer demands for different batches and types of products, manufacturing operations management has been facing the challenge of uncertain product arrival times and resource processing times in customized manufacturing (CM). This paper proposes a dynamic scheduling method to solve the uncertainty in CM via the integration of the digital twin and fuzzy reinforcement learning methods. In this study, a digital twin-driven framework is first designed to describe the operation management system (OMS) hierarchies. Then a semi-Markov decision process (MDP) model with fuzzy definition is built by abstracting the stochastic scheduling process. To solve the semi-MDP model, an asynchronous multi-edge co-training method is presented to train a fuzzy deep neural network through closed-loop control of virtual commissioning, illustrating how the digital twin-driven OMS adapts to dynamic production requirements. Finally, the proposed method is verified by the performance of comparative experiment. Experimental results show that for randomly arriving products, the proposed method guarantees timely training and scheduling decisions and has the highest total system profit compared to other competing methods (Hybrid Multi-Agent System Negotiation and Ant Colony Optimization (HMA), Onto_MDP, and Deep Q Networks (DQN)). Also, the proposed method shows better scheduling performance in terms of average decision time, average training time and number of finished products when resources are abnormal.

IJCAI Conference 2025 Conference Paper

Dual Robust Unbiased Multi-View Clustering for Incomplete and Unpaired Information

  • Liang Zhao
  • Ziyue Wang
  • Chuanye He
  • Qingchen Zhang
  • Bo Xu

Recently, multi-view data has gradually attracted attention. However, real-world applications often face Partial View-aligned Problem (PVP) and Partially Sample-missing Problem (PSP) due to data loss or corruption. Existing methods addressing PVP typically focus only on learning from the information of aligned data, while ignoring unaligned data where samples exist but lack alignment relationships. This introduces PSP, which does not inherently exist in the data, leading to biased learning of the data's information. For PSP, due to varying degrees of missing data, incomplete spatial structures can cause clustering centers-shifted problem, resulting in the model learning incorrect correspondences and biased spatial structures. To tackle them, we propose a novel method called Dual Robust Unbiased Multi-View Clustering for Incomplete and Unpaired Information (DRUMVC). To our knowledge, this is the first noise-robust and unbiased multi-view clustering method capable of simultaneously addressing both PVP and PSP. Specifically, DRUMVC leverages aligned and complete samples as a bridge to construct high-quality correspondences for samples lacking cross-view relationship information due to PVP or PSP. Additionally, we employ a dual noise-robust contrastive learning loss to mitigate the impact of noise potentially introduced during the pair construction. Experiments on several challenging datasets demonstrate the superiority of our proposed method.

IJCAI Conference 2025 Conference Paper

EchoGPT: An Interactive Cardiac Function Assessment Model for Echocardiogram Videos

  • Bo Xu
  • Quanhao Zhu
  • Qingchen Zhang
  • Mengmeng Wang
  • Liang Zhao
  • Hongfei Lin
  • Jing Ren
  • Feng Xia

With the development of wearable cardiac ultrasound devices, it is no longer sufficient to solely rely on doctors for diagnosing long-term echocardiogram videos. Automated diagnosis of echocardiogram videos has now become a research hotspot. Existing studies only analyze echocardiogram video through discriminative models, which have limited question-answering capabilities. Therefore, this study innovatively proposes a large language model with cardiac ultrasound diagnostic capabilities—EchoGPT. EchoGPT integrates the robust communication and comprehension capabilities of large language models (LLMs) with the diagnostic prowess of traditional medical models, empowering patients to obtain accurate medical indicator data and comprehend their health conditions through interactive questioning with the model. The model is capable of local deployment on personal computers, effectively safe guarding user privacy. EchoGPT operates through three main components: left ventricle segmentation, left ventricular ejection fraction LVEF prediction, and finetuning of video-text LLMs. Experimental results demonstrate EchoGPT’s superior accuracy in predicting LVEF compared to other models, and positive feedback from professional physicians through questionnaire surveys, validating its potential in practical applications. The demo is available at https: //github. com/zhuqh19/EchoGPT.

AAAI Conference 2025 Conference Paper

Efficient 3D Recognition with Event-driven Spike Sparse Convolution

  • Xuerui Qiu
  • Man Yao
  • Jieyuan Zhang
  • Yuhong Chou
  • Ning Qiao
  • Shibo Zhou
  • Bo Xu
  • Guoqi Li

Spiking Neural Networks (SNNs) provide an energy-efficient way to extract 3D spatio-temporal features. Point clouds are sparse 3D spatial data, which suggests that SNNs should be well-suited for processing them. However, when applying SNNs to point clouds, they often exhibit limited performance and fewer application scenarios. We attribute this to inappropriate preprocessing and feature extraction methods. To address this issue, we first introduce the Spike Voxel Coding (SVC) scheme, which encodes the 3D point clouds into a sparse spike train space, reducing the storage requirements and saving time on point cloud preprocessing. Then, we propose a Spike Sparse Convolution (SSC) model for efficiently extracting 3D sparse point cloud features. Combining SVC and SSC, we design an efficient 3D SNN backbone (E-3DSNN), which is friendly with neuromorphic hardware. For instance, SSC can be implemented on neuromorphic chips with only minor modifications to the addressing function of vanilla spike convolution. Experiments on ModelNet40, KITTI, and Semantic KITTI datasets demonstrate that E-3DSNN achieves state-of-the-art (SOTA) results with remarkable efficiency. Notably, our E-3DSNN (1.87M) obtained 91.7% top-1 accuracy on ModelNet40, surpassing the current best SNN baselines (14.3M) by 3.0%. To our best knowledge, it is the first direct training 3D SNN backbone that can simultaneously handle various 3D computer vision tasks (e.g., classification, detection, and segmentation) with an event-driven nature.

AAAI Conference 2025 Conference Paper

GuideNER: Annotation Guidelines Are Better than Examples for In-Context Named Entity Recognition

  • Shizhou Huang
  • Bo Xu
  • Yang Yu
  • Changqun Li
  • Xin Alex Lin

Large language models (LLMs) demonstrate impressive performance on downstream tasks through in-context learning(ICL). However, there is a significant gap between their performance in Named Entity Recognition (NER) and in fine-tuning methods. We believe this discrepancy is due to inconsistencies in labeling definitions in NER. In addition, recent research indicates that LLMs do not learn the specific input-label mappings from the demonstrations. Therefore, we argue that using examples to implicitly capture the mapping between inputs and labels in in-context learning is not suitable for NER. Instead, it requires explicitly informing the model of the range of entities contained in the labels, such as annotation guidelines. In this paper, we propose GuideNER, which uses LLMs to summarize concise annotation guidelines as contextual information in ICL. We have conducted experiments on widely used NER datasets, and the experimental results indicate that our method can consistently and significantly outperform state-of-the-art methods, while using shorter prompts. Especially on the GENIA dataset, our model outperforms the previous state-of-the-art model by 12.63 F1 scores.

AAAI Conference 2025 Conference Paper

Incomplete and Unpaired Multi-View Graph Clustering with Cross-View Feature Fusion

  • Liang Zhao
  • Ziyue Wang
  • Xiao Wang
  • Zhikui Chen
  • Bo Xu

Due to its effectiveness and efficiency, graph-based multi-view clustering has recently attracted much attention. However, the multi-view data are often incomplete and unpaired in real-world applications as a consequence of data loss or corruption. Although efforts have been made through a series of methods to address the problems of incomplete or unpaired multi-view data, the following issues still persist: 1) Most existing methods only focus on the incomplete multi-view data or unpaired multi-view data, and exhibit weaknesses when addressing both incomplete and unpaired multi-view data simultaneously. 2) Some methods neglect the graph information of the data from different views during the learning process. To tackle these issues, we propose the Multi-view Graph Clustering framework with Cross-view Feature Fusion (MGCCFF), a novel approach for clustering incomplete and unpaired multi-view data. Specifically, MGCCFF learns soft clustering label information from complete data and utilizes this to capture category-level cross-view correspondences. It then learns latent representation enriched with cross-view information based on the established mappings. To obtain a multi-view graph structure under conditions of incomplete and unpaired data, MGCCFF innovatively integrates the concept of self-expression with the autoencoder architecture and exploits the latent relationships between labels and the graph structure, thereby enabling the generation of sparse and accurate graphical structure under multi-view conditions for the final clustering task. The experiments on incomplete and unpaired multi-view datasets demonstrate that MGCCFF outperforms state-of-the-art methods.

AAMAS Conference 2025 Conference Paper

Integrating Large Language Models with Reinforcement Learning for Generalization in Strategic Card Games

  • Wannian Xia
  • Meng Fang
  • Zihao Guo
  • Yali Du
  • Bo Xu

Strategic card games, such as Hearthstone, offer a rich environment for exploring decision-making in reinforcement learning (RL). Yet, achieving generalization across diverse and evolving game scenarios remains a significant challenge. In this paper, we introduce a novel framework that integrates large language models (LLMs) with RL agents to improve generalization in strategic card games. Our approach leverages a fine-tuned T5 model to encode and interpret card strategies expressed in natural language, facilitating efficient policy learning across a large and continually expanding set of cards. Employing a self-play RL framework augmented with an auxiliary transition loss in the latent space, our agent captures and generalizes the complex, dynamic nature of card interactions. Experimental results show that our method not only enhances learning efficiency but also significantly improves the agent’s ability to generalize, maintaining robust performance when encountering new cards.

AAAI Conference 2025 Conference Paper

Leveraging Attention to Effectively Compress Prompts for Long-Context LLMs

  • Yunlong Zhao
  • Haoran Wu
  • Bo Xu

Prompt compression is increasingly studied for its potential to reduce computational costs and alleviate the burden on language models when processing lengthy prompts. Prior research has assessed token retention and removal by calculating information entropy. However, prompt compression encounters two significant challenges: (1) Information entropy, while widely used, may not be the optimal compression metric; and (2) The semantic significance of tokens is context-dependent, which renders independent token retention decisions inadequate. We posit that the solution to these challenges lies in the intrinsic mechanism of language models. Large language models (LLMs) exhibit robust contextual processing capabilities, with recent studies on their internal dynamics revealing that the attention mechanism plays a crucial role in modeling how LLMs leverage long contexts. Building on this insight, we introduce AttnComp, a novel approach that exploits the attention mechanism within language models to guide prompt compression. Our method employs causal cross-attention from the query to the context to evaluate the significance of each token, and we develop a graph-based algorithm to efficiently cluster tokens into semantic units, thus mitigating the issue of independent dependencies. We conduct experiments on datasets for retrieval-augmented generation and multiple long tasks involving single or multi-document QA. Our proposed method, AttnComp, outperforms previous baselines and validates the contributions of our components through analytical experiments. Compared to other methods that use a causal LM for prompt compression, our approach results in shorter latency and improved performance.

TCS Journal 2025 Journal Article

Model training oriented mobile crowdsourcing tasks: Worker selection and incentive mechanism design

  • Bo Xu
  • Zexing Wang
  • Zhi-Ping Fan

This paper investigates the problem of worker selection and incentive mechanism for model training oriented mobile crowdsourcing tasks. Firstly, based on the reverse auction theory, a ex-ante selection scheme is designed to select high-quality workers from a pool of candidates. Secondly, a ex-post payment incentive mechanism is designed based on Stackelberg game theory, which can evaluate the results of task execution to motivate participating workers to complete model training oriented crowdsourcing tasks with high quality. Finally, the effectiveness of the mechanism is demonstrated through numerical analysis. The results indicate that the ex-ante selection scheme and ex-post payment incentive mechanism can effectively select high-quality crowdsourcing workers to complete model training oriented crowdsourcing tasks and ensure the quality of training task completion.

JBHI Journal 2025 Journal Article

Multi-Scale Dynamic Sparse Token Multi-Instance Learning for Pathology Image Classification

  • Dajiang Lei
  • Yuqi Zhang
  • Haodong Wang
  • Xiaomin Xiong
  • Bo Xu
  • Guoyin Wang

In many challenging breast cancer pathology images, the proportion of truly informative tumor regions is extremely limited. The disparity between the essential information required for clinical diagnosis (Tumor area less than 10 $\%$ ) and the vast amount of data within Whole Slide Images (WSIs) makes it exceedingly difficult for pathologists to identify subtle lesions. To address the labor-intensive task imposed by this information gap, this paper proposes a dynamic sparse token based multi-instance learning framework. This framework incorporates a dynamic sparse layer into the transformer architecture, gradually adapting to selectively filter key instances beneficial for the task. Furthermore, to tackle complex scenarios in pathology image tasks, we introduce a weakly supervised cross-scale contrastive learning framework. This framework leverages pathology image features at different scales to perform contrastive learning at the bag-level representation to overcome existing challenges in multi-scale feature fusion in pathology image tasks. To validate the effectiveness and transferability of the model, we conducted various single-scale and multi-scale experiments across four cancer datasets and conducted interpretable analyses. Compared to other state-of-the-art methods, our classification model demonstrates superior performance across six evaluation metrics.

IJCAI Conference 2025 Conference Paper

S-EPOA: Overcoming the Indistinguishability of Segments with Skill-Driven Preference-Based Reinforcement Learning

  • Ni Mu
  • Yao Luan
  • Yiqin Yang
  • Bo Xu
  • Qing-Shan Jia

Preference-based reinforcement learning (PbRL) stands out by utilizing human preferences as a direct reward signal, eliminating the need for intricate reward engineering. However, despite its potential, traditional PbRL methods are often constrained by the indistinguishability of segments, which impedes the learning process. In this paper, we introduce Skill-Enhanced Preference Optimization Algorithm (S-EPOA), which addresses the segment indistinguishability issue by integrating skill mechanisms into the preference learning framework. Specifically, we first conduct the unsupervised pretraining to learn useful skills. Then, we propose a novel query selection mechanism to balance the information gain and distinguishability over the learned skill space. Experimental results on a range of tasks, including robotic manipulation and locomotion, demonstrate that S-EPOA significantly outperforms conventional PbRL methods in terms of both robustness and learning efficiency. The results highlight the effectiveness of skill-driven learning in overcoming the challenges posed by segment indistinguishability.

NeurIPS Conference 2025 Conference Paper

Self-Verifying Reflection Helps Transformers with CoT Reasoning

  • Zhongwei Yu
  • Wannian Xia
  • Xue Yan
  • Bo Xu
  • Haifeng Zhang
  • Yali Du
  • Jun Wang

Advanced large language models (LLMs) frequently reflect in reasoning chain-of-thoughts (CoTs), where they self-verify the correctness of current solutions and explore alternatives. However, given recent findings that LLMs detect limited errors in CoTs, how reflection contributes to empirical improvements remains unclear. To analyze this issue, in this paper, we present a minimalistic reasoning framework to support basic self-verifying reflection for small transformers without natural language, which ensures analytic clarity and reduces the cost of comprehensive experiments. Theoretically, we prove that self-verifying reflection guarantees improvements if verification errors are properly bounded. Experimentally, we show that tiny transformers, with only a few million parameters, benefit from self-verification in both training and reflective execution, reaching remarkable LLM-level performance in integer multiplication and Sudoku. Similar to LLM results, we find that reinforcement learning (RL) improves in-distribution performance and incentivizes frequent reflection for tiny transformers, yet RL mainly optimizes shallow statistical patterns without faithfully reducing verification errors. In conclusion, integrating generative transformers with discriminative verification inherently facilitates CoT reasoning, regardless of scaling and natural language.

AAAI Conference 2025 Conference Paper

Spike2Former: Efficient Spiking Transformer for High-performance Image Segmentation

  • Zhenxin Lei
  • Man Yao
  • Jiakui Hu
  • Xinhao Luo
  • Yanye Lu
  • Bo Xu
  • Guoqi Li

Spiking Neural Networks (SNNs) have a low-power advantage but perform poorly in image segmentation tasks. The reason is that directly converting neural networks with complex architectural designs for segmentation tasks into spiking versions leads to performance degradation and non-convergence. To address this challenge, we first identify the modules in the architecture design that lead to the severe reduction in spike firing, make targeted improvements, and propose Spike2Former architecture. Second, we propose normalized integer spiking neurons to solve the training stability problem of SNNs with complex architectures. We set a new state-of-the-art for SNNs in various semantic segmentation datasets, with a significant improvement of +12.7% mIoU and 5.0x efficiency on ADE20K, +14.3% mIoU and 5.2x efficiency on VOC2012, and +9.1% mIoU and 6.6x efficiency on CityScapes.

NeurIPS Conference 2025 Conference Paper

STAIR: Addressing Stage Misalignment through Temporal-Aligned Preference Reinforcement Learning

  • Yao Luan
  • Ni Mu
  • Yiqin Yang
  • Bo Xu
  • Qing-Shan Jia

Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning rewards directly from human preferences, enabling better alignment with human intentions. However, its effectiveness in multi-stage tasks, where agents sequentially perform sub-tasks (e. g. , navigation, grasping), is limited by stage misalignment: Comparing segments from mismatched stages, such as movement versus manipulation, results in uninformative feedback, thus hindering policy learning. In this paper, we validate the stage misalignment issue through theoretical analysis and empirical experiments. To address this issue, we propose ST age- A l I gned R eward learning (STAIR), which first learns a stage approximation based on temporal distance, then prioritizes comparisons within the same stage. Temporal distance is learned via contrastive learning, which groups temporally close states into coherent stages, without predefined task knowledge, and adapts dynamically to policy changes. Extensive experiments demonstrate STAIR's superiority in multi-stage tasks and competitive performance in single-stage tasks. Furthermore, human studies show that stages approximated by STAIR are consistent with human cognition, confirming its effectiveness in mitigating stage misalignment.

NeurIPS Conference 2025 Conference Paper

STAR: Efficient Preference-based Reinforcement Learning via Dual Regularization

  • Fengshuo Bai
  • Rui Zhao
  • Hongming Zhang
  • Sijia Cui
  • Shao Zhang
  • Bo Xu
  • Lei Han
  • Ying Wen

Preference-based reinforcement learning (PbRL) bypasses complex reward engineering by learning from human feedback. However, due to the high cost of obtaining feedback, PbRL typically relies on a limited set of preference-labeled samples. This data scarcity introduces two key inefficiencies: (1) the reward model overfits to the limited feedback, leading to poor generalization to unseen samples, and (2) the agent exploits the learned reward model, exacerbating overestimation of action values in temporal difference (TD) learning. To address these issues, we propose STAR, an efficient PbRL method that integrates preference margin regularization and policy regularization. Preference margin regularization mitigates overfitting by introducing a bounded margin in reward optimization, preventing excessive bias toward specific feedback. Policy regularization bootstraps a conservative estimate $\widehat{Q}$ from well-supported state-action pairs in the replay memory, reducing overestimation during policy learning. Experimental results show that STAR improves feedback efficiency, achieving 34. 8\% higher performance in online settings and 29. 7\% in offline settings compared to state-of-the-art methods. Ablation studies confirm that STAR facilitates more robust reward and value function learning. The videos of this project are released at https: //sites. google. com/view/pbrl-star.

AAAI Conference 2025 Conference Paper

Text-Guided Fine-grained Counterfactual Inference for Short Video Fake News Detection

  • Linlin Zong
  • Wenmin Lin
  • Jiahui Zhou
  • Xinyue Liu
  • Xianchao Zhang
  • Bo Xu
  • Shimin Wu

Detecting fake news in short videos is crucial for combating misinformation. Existing methods utilize topic modeling and co-attention mechanism, overlooking the modality heterogeneity and resulting in suboptimal performance. To address this issue, we introduce Text-Guided Fine-grained Counterfactual Inference for Short Video Fake News detection (TGFC-SVFN). TGFC-SVFN leverages modality bias removal and teacher-model-enhanced inter-modal knowledge distillation to integrate the heterogeneous modalities in short videos. Specifically, we use causality-based reasoning prompts guided text as teacher model, which then transfers knowledge to the video and audio student models. Subsequently, a multi-head attention mechanism is employed to fuse information from different modalities. In each module, we utilize fine-grained counterfactual inference based on a diffusion model to eliminate modality bias. Experimental results on publicly available fake short video news datasets demonstrate that our method outperforms state-of-the-art techniques.

IJCAI Conference 2025 Conference Paper

Unveiling Maternity and Infant Care Conversations: A Chinese Dialogue Dataset for Enhanced Parenting Support

  • Bo Xu
  • Liangzhi Li
  • Junlong Wang
  • Xuening Qiao
  • Erchen Yu
  • Yiming Qian
  • Linlin Zong
  • Hongfei Lin

The rapid development of large language models has greatly advanced human-computer dialogue research. However, applying these models to specialized fields like maternity and infant care often leads to subpar performance due to a lack of domain-specific datasets. To address this problem, we have created MicDialogue, a Chinese dialogue dataset for maternity and infant care. MicDialogue involves a wide range of specialized topics, including gynecological health, pediatric care, pregnancy preparation, emotional counseling and other related topics. This dataset is curated from two types of Chinese social media: short videos and blog posts. Short videos capture real-time interactions and pragmatic dialogue patterns, while blog posts offer comprehensive coverage of various topics within the domain. We have also included detailed annotations for topics, diseases, symptoms, and causes, enabling in-depth research. Additionally, we developed a knowledge-driven benchmark model using LLM-based prompt learning and multiple knowledge graphs to address diverse dialogue topics. Experiments validate MicDialogue's usability, providing benchmarks for future research and essential data for fine-tuning language models in maternity and infant care.

AAMAS Conference 2025 Conference Paper

β-DQN: Improving Deep Q-Learning By Evolving the Behavior

  • Hongming Zhang
  • Fengshuo Bai
  • Chenjun Xiao
  • Chao Gao
  • Bo Xu
  • Martin Müller

While many sophisticated exploration methods have been proposed, their lack of generality and high computational cost often lead researchers to favor simpler methods like 𝜖-greedy. Motivated by this, we introduce 𝛽-DQN, a simple and efficient exploration method that augments the standard DQN with a behavior function 𝛽. This function estimates the probability that each action has been taken at each state. By leveraging 𝛽, we generate a population of diverse policies that balance exploration between state-action coverage and overestimation bias correction. An adaptive meta-controller is designed to select an effective policy for each episode, enabling flexible and explainable exploration. 𝛽-DQN is straightforward to implement and adds minimal computational overhead to the standard DQN. Experiments on both simple and challenging exploration domains show that 𝛽-DQN outperforms existing baseline methods across a wide range of tasks, providing an effective solution for improving exploration in deep reinforcement learning.

TMLR Journal 2024 Journal Article

A Distance-based Anomaly Detection Framework for Deep Reinforcement Learning

  • Hongming Zhang
  • Ke Sun
  • Bo Xu
  • Linglong Kong
  • Martin Müller

In deep reinforcement learning (RL) systems, abnormal states pose significant risks by potentially triggering unpredictable behaviors and unsafe actions, thus impeding the deployment of RL systems in real-world scenarios. It is crucial for reliable decision-making systems to have the capability to cast an alert whenever they encounter unfamiliar observations that they are not equipped to handle. In this paper, we propose a novel Mahalanobis distance-based (MD) anomaly detection framework, called \textit{MDX}, for deep RL algorithms. MDX simultaneously addresses random, adversarial, and out-of-distribution (OOD) state outliers in both offline and online settings. It utilizes Mahalanobis distance within class-conditional distributions for each action and operates within a statistical hypothesis testing framework under the Gaussian assumption. We further extend it to robust and distribution-free versions by incorporating Robust MD and conformal inference techniques. Through extensive experiments on classical control environments, Atari games and autonomous driving scenarios, we demonstrate the effectiveness of our MD-based detection framework. MDX offers a simple, unified, and practical tool for enhancing the safety and reliability of RL systems in real-world applications.

EAAI Journal 2024 Journal Article

A novel reconstruction method for displacement missing data of arch dam via hierarchical clustering and deep learning

  • Hu Zhang
  • Bo Xu
  • Zeyuan Chen

The absence of displacement monitoring data presents a challenge to real-time dam safety monitoring. This article introduces a novel method for reconstructing missing displacements, comprehensively taking into account the spatiotemporal correlation and causal effect mechanism of displacement. Firstly, adaptive weighted derivative dynamic time warping (AWDDTW) and adaptive weighted dynamic time warping (AWDTW), in conjunction with clustering algorithms, are proposed to extract spatiotemporal correlation among dam displacements. Secondly, by stacking residual blocks within the residual network and densifying the shortcut connections between residual blocks, the deep stacked residual network (DS-ResNet) is established to effectively capture the intricate mappings between displacements and various factors. Finally, using two arch dams, as examples, we simulated scenarios of continuous long-term and multiple short-term missing data to validate the new method proposed in this study. The results indicate that the proposed clustering algorithm can accurately compute the similarity between displacement sequences of different lengths and monitoring frequencies, thereby providing more precise displacement point partitioning results. Compared to models such as random forests, relevance vector machines, multilayer perceptron, ResNet and Transformer, the DS-ResNet demonstrates superior capability in identifying and extracting effective information from displacement sequences under both continuous long-term and multiple short-term missing data scenarios. Compared to methods that solely consider spatiotemporal correlation or causal effect mechanism of displacement, the proposed method can more comprehensively reflect the operating patterns of arch dams and more accurately reconstruct missing displacement values. The research presented herein offers new technical means and solution approaches for dam safety monitoring and operational management.

NeurIPS Conference 2024 Conference Paper

Exploiting the Replay Memory Before Exploring the Environment: Enhancing Reinforcement Learning Through Empirical MDP Iteration

  • Hongming Zhang
  • Chenjun Xiao
  • Chao Gao
  • Han Wang
  • Bo Xu
  • Martin Müller

Reinforcement learning (RL) algorithms are typically based on optimizing a Markov Decision Process (MDP) using the optimal Bellman equation. Recent studies have revealed that focusing the optimization of Bellman equations solely on in-sample actions tends to result in more stable optimization, especially in the presence of function approximation. Upon on these findings, in this paper, we propose an Empirical MDP Iteration (EMIT) framework. EMIT constructs a sequence of empirical MDPs using data from the growing replay memory. For each of these empirical MDPs, it learns an estimated Q-function denoted as $\widehat{Q}$. The key strength is that by restricting the Bellman update to in-sample bootstrapping, each empirical MDP converges to a unique optimal $\widehat{Q}$ function. Furthermore, gradually expanding from the empirical MDPs to the original MDP induces a monotonic policy improvement. Instead of creating entirely new algorithms, we demonstrate that EMIT can be seamlessly integrated with existing online RL algorithms, effectively acting as a regularizer for contemporary Q-learning methods. We show this by implementing EMIT for two representative RL algorithms, DQN and TD3. Experimental results on Atari and MuJoCo benchmarks show that EMIT significantly reduces estimation errors and substantially improves the performance of both algorithms.

IS Journal 2024 Journal Article

Improved Small Object Detection Algorithm Based on YOLOv5

  • Bo Xu
  • Bin Gao
  • Yunhu Li

YOLOv5 is a popular object detection algorithm that is widely used in various industrial fields, especially in the field of autonomous driving. However, this algorithm has problems, such as false positives and false negatives when detecting small targets. The article proposes an improved method for small object detection using YOLOv5s. First, a multilevel feature fusion detection head is proposed to extract larger feature maps from the backbone of the model, improving the ability to extract features of small objects. Second, a decoupled attention mechanism is introduced at each detection head, which separates the detection of object box position, object box confidence, and class probability to reduce confusion between different feature information. Finally, the focal minimum points distance intersection over union loss function is adopted to mitigate the effects of class imbalance and poor-quality object pixels.

NeurIPS Conference 2024 Conference Paper

MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map

  • Yuhong Chou
  • Man Yao
  • Kexin Wang
  • Yuqi Pan
  • Ruijie Zhu
  • Yiran Zhong
  • Yu Qiao
  • Jibin Wu

Various linear complexity models, such as Linear Transformer (LinFormer), State Space Model (SSM), and Linear RNN (LinRNN), have been proposed to replace the conventional softmax attention in Transformer structures. However, the optimal design of these linear models is still an open question. In this work, we attempt to answer this question by finding the best linear approximation to softmax attention from a theoretical perspective. We start by unifying existing linear complexity models as the linear attention form and then identify three conditions for the optimal linear attention design: (1) Dynamic memory ability; (2) Static approximation ability; (3) Least parameter approximation. We find that none of the current linear models meet all three conditions, resulting in suboptimal performance. Instead, we propose Meta Linear Attention (MetaLA) as a solution that satisfies these conditions. Our experiments on Multi-Query Associative Recall (MQAR) task, language modeling, image classification, and Long-Range Arena (LRA) benchmark demonstrate that MetaLA is more effective than the existing linear models.

AAAI Conference 2024 Conference Paper

Privileged Prior Information Distillation for Image Matting

  • Cheng Lyu
  • Jiake Xie
  • Bo Xu
  • Cheng Lu
  • Han Huang
  • Xin Huang
  • Ming Wu
  • Chuang Zhang

Performance of trimap-free image matting methods is limited when trying to decouple the deterministic and undetermined regions, especially in the scenes where foregrounds are semantically ambiguous, chromaless, or high transmittance. In this paper, we propose a novel framework named Privileged Prior Information Distillation for Image Matting (PPID-IM) that can effectively transfer privileged prior environment-aware information to improve the performance of trimap-free students in solving hard foregrounds. The prior information of trimap regulates only the teacher model during the training stage, while not being fed into the student network during actual inference. To achieve effective privileged cross-modality (i.e. trimap and RGB) information distillation, we introduce a Cross-Level Semantic Distillation (CLSD) module that reinforces the students with more knowledgeable semantic representations and environment-aware information. We also propose an Attention-Guided Local Distillation module that efficiently transfers privileged local attributes from the trimap-based teacher to trimap-free students for the guidance of local-region optimization. Extensive experiments demonstrate the effectiveness and superiority of our PPID on image matting. The code will be released soon.

ECAI Conference 2024 Conference Paper

Semantic Similarity Driven Multi-Modal Model for Rumor Detection

  • Chenyang Li
  • Bo Xu
  • Meng Wang 0039
  • Kun He 0001

The wide spread of rumors with images and texts on social media has attracted broad attention in the academy and industry. Existing models focus on utilizing powerful feature extractors to obtain multi-modal features and introducing various external knowledge. However, the intrinsic semantic similarity of different modalities is either simply ignored in most models or far from adequate in others. The insufficiency of semantic similarity information suppresses the potential of rumor detection models severely. To address this issue, we propose a novel model termed the Semantic Similarity driven Multi-modal model (SemSim) for rumor detection, which deeply captures the semantic similarity through more comprehensive fusion between different modalities and designs a new classification method consequently. Specifically, the proposed SemSim first integrates the raw image and raw text into a virtual image, which fuses information at a new view, i. e. , via the diffusion process inside stable diffusion models. Then SemSim captures the semantic similarity score between virtual image and raw image as the intrinsic information to drive SemSim. Besides, co-attention mechanism is employed to further perceive consistency and enhance interaction between the raw text-image pair. The fused representations via co-attention are utilized to evaluate the multi-modal feature score. In the end, SemSim balances the above two scores for final classification. Experiments on two typical real-world datasets show that SemSim can effectively detect rumors and outperform state-of-the-art methods.

NeurIPS Conference 2024 Conference Paper

Towards Comprehensive Detection of Chinese Harmful Memes

  • Junyu Lu
  • Bo Xu
  • Xiaokun Zhang
  • Hongbo Wang
  • Haohao Zhu
  • Dongyu Zhang
  • Liang Yang
  • Hongfei Lin

Harmful memes have proliferated on the Chinese Internet, while research on detecting Chinese harmful memes significantly lags behind due to the absence of reliable datasets and effective detectors. To this end, we present the comprehensive detection of Chinese harmful memes. We introduce ToxiCN MM, the first Chinese harmful meme dataset, which consists of 12, 000 samples with fine-grained annotations for meme types. Additionally, we propose a baseline detector, Multimodal Knowledge Enhancement (MKE), designed to incorporate contextual information from meme content, thereby enhancing the model's understanding of Chinese memes. In the evaluation phase, we conduct extensive quantitative experiments and qualitative analyses on multiple baselines, including LLMs and our MKE. Experimental results indicate that detecting Chinese harmful memes is challenging for existing models, while demonstrating the effectiveness of MKE.

AAAI Conference 2024 Conference Paper

Video-Context Aligned Transformer for Video Question Answering

  • Linlin Zong
  • Jiahui Wan
  • Xianchao Zhang
  • Xinyue Liu
  • Wenxin Liang
  • Bo Xu

Video question answering involves understanding video content to generate accurate answers to questions. Recent studies have successfully modeled video features and achieved diverse multimodal interaction, yielding impressive outcomes. However, they have overlooked the fact that the video contains richer instances and events beyond the scope of the stated question. Extremely imbalanced alignment of information from both sides leads to significant instability in reasoning. To address this concern, we propose the Video-Context Aligned Transformer (V-CAT), which leverages the context to achieve semantic and content alignment between video and question. Specifically, the video and text are encoded into a shared semantic space initially. We apply contrastive learning to global video token and context token to enhance the semantic alignment. Then, the pooled context feature is utilized to obtain corresponding visual content. Finally, the answer is decoded by integrating the refined video and question features. We evaluate the effectiveness of V-CAT on MSVD-QA and MSRVTT-QA dataset, both achieving state-of-the-art performance. Extended experiments further analyze and demonstrate the effectiveness of each proposed module.

AAAI Conference 2023 Conference Paper

Complex Dynamic Neurons Improved Spiking Transformer Network for Efficient Automatic Speech Recognition

  • Qingyu Wang
  • Tielin Zhang
  • Minglun Han
  • Yi Wang
  • Duzhen Zhang
  • Bo Xu

The spiking neural network (SNN) using leaky-integrated-and-fire (LIF) neurons has been commonly used in automatic speech recognition (ASR) tasks. However, the LIF neuron is still relatively simple compared to that in the biological brain. Further research on more types of neurons with different scales of neuronal dynamics is necessary. Here we introduce four types of neuronal dynamics to post-process the sequential patterns generated from the spiking transformer to get the complex dynamic neuron improved spiking transformer neural network (DyTr-SNN). We found that the DyTr-SNN could handle the non-toy automatic speech recognition task well, representing a lower phoneme error rate, lower computational cost, and higher robustness. These results indicate that the further cooperation of SNNs and neural dynamics at the neuron and network scales might have much in store for the future, especially on the ASR tasks.

AAMAS Conference 2023 Conference Paper

M3: Modularization for Multi-task and Multi-agent Offline Pre-training

  • Linghui Meng
  • Jingqing Ruan
  • Xuantang Xiong
  • Xiyun Li
  • Xi Zhang
  • Dengpeng Xing
  • Bo Xu

Learning a multi-task policy is crucial in multi-agent reinforcement learning (MARL). Recent work has focused on learning in the context of online multi-task reinforcement learning, where a policy is jointly trained from scratch, aiming to generalize well to few-shot or even zero-shot tasks. However, existing online methods require tremendous interactions and are therefore unsuitable for environments where interactions are expensive. In this work, we novelly introduce the modularization for multi-task and multi-agent offline pre-training (M3) to learn high-level transferable policy representations. We claim that the discrete policy representation is critical for multi-task offline learning and accordingly leverage contexts as a task prompt to enhance the adaptability of pre-trained models to various tasks. To disentangle multiple agents of variation under heterogeneous and non-stationary properties even though they receive the same task, we employ an agent-invariant VQ-VAE to identify each of the multiple agents. We encapsulate the pretrained model as part of an online MARL algorithm and fine-tune it * These authors contribute equally to this work. † Corresponding authors. Proc. of the 22nd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2023), A. Ricci, W. Yeoh, N. Agmon, B. An (eds.), May 29 – June 2, 2023, London, United Kingdom. © 2023 International Foundation for Autonomous Agents and Multiagent Systems (www. ifaamas. org). All rights reserved. to improve generalization. We also theoretically analyze the generalization error of our method. We test the proposed method on the challenging StarCraft Multi-Agent Challenge (SMAC) tasks, and empirical results show that it can achieve supreme performance in few-shot or even zero-shot settings across multiple tasks over state-of-the-art MARL methods.

NeurIPS Conference 2023 Conference Paper

ODE-based Recurrent Model-free Reinforcement Learning for POMDPs

  • Xuanle Zhao
  • Duzhen Zhang
  • Han Liyuan
  • Tielin Zhang
  • Bo Xu

Neural ordinary differential equations (ODEs) are widely recognized as the standard for modeling physical mechanisms, which help to perform approximate inference in unknown physical or biological environments. In partially observable (PO) environments, how to infer unseen information from raw observations puzzled the agents. By using a recurrent policy with a compact context, context-based reinforcement learning provides a flexible way to extract unobservable information from historical transitions. To help the agent extract more dynamics-related information, we present a novel ODE-based recurrent model combines with model-free reinforcement learning (RL) framework to solve partially observable Markov decision processes (POMDPs). We experimentally demonstrate the efficacy of our methods across various PO continuous control and meta-RL tasks. Furthermore, our experiments illustrate that our method is robust against irregular observations, owing to the ability of ODEs to model irregularly-sampled time series.

AAAI Conference 2023 Conference Paper

PiCor: Multi-Task Deep Reinforcement Learning with Policy Correction

  • Fengshuo Bai
  • Hongming Zhang
  • Tianyang Tao
  • Zhiheng Wu
  • Yanna Wang
  • Bo Xu

Multi-task deep reinforcement learning (DRL) ambitiously aims to train a general agent that masters multiple tasks simultaneously. However, varying learning speeds of different tasks compounding with negative gradients interference makes policy learning inefficient. In this work, we propose PiCor, an efficient multi-task DRL framework that splits learning into policy optimization and policy correction phases. The policy optimization phase improves the policy by any DRL algothrim on the sampled single task without considering other tasks. The policy correction phase first constructs an adaptive adjusted performance constraint set. Then the intermediate policy learned by the first phase is constrained to the set, which controls the negative interference and balances the learning speeds across tasks. Empirically, we demonstrate that PiCor outperforms previous methods and significantly improves sample efficiency on simulated robotic manipulation and continuous control tasks. We additionally show that adaptive weight adjusting can further improve data efficiency and performance.

NeurIPS Conference 2023 Conference Paper

Spike-driven Transformer

  • Man Yao
  • Jiakui Hu
  • Zhaokun Zhou
  • Li Yuan
  • Yonghong Tian
  • Bo Xu
  • Guoqi Li

Spiking Neural Networks (SNNs) provide an energy-efficient deep learning option due to their unique spike-based event-driven (i. e. , spike-driven) paradigm. In this paper, we incorporate the spike-driven paradigm into Transformer by the proposed Spike-driven Transformer with four unique properties: (1) Event-driven, no calculation is triggered when the input of Transformer is zero; (2) Binary spike communication, all matrix multiplications associated with the spike matrix can be transformed into sparse additions; (3) Self-attention with linear complexity at both token and channel dimensions; (4) The operations between spike-form Query, Key, and Value are mask and addition. Together, there are only sparse addition operations in the Spike-driven Transformer. To this end, we design a novel Spike-Driven Self-Attention (SDSA), which exploits only mask and addition operations without any multiplication, and thus having up to $87. 2\times$ lower computation energy than vanilla self-attention. Especially in SDSA, the matrix multiplication between Query, Key, and Value is designed as the mask operation. In addition, we rearrange all residual connections in the vanilla Transformer before the activation functions to ensure that all neurons transmit binary spike signals. It is shown that the Spike-driven Transformer can achieve 77. 1\% top-1 accuracy on ImageNet-1K, which is the state-of-the-art result in the SNN field.

AAMAS Conference 2022 Conference Paper

GCS: Graph-Based Coordination Strategy for Multi-Agent Reinforcement Learning

  • Jingqing Ruan
  • Yali Du
  • Xuantang Xiong
  • Dengpeng Xing
  • Xiyun Li
  • Linghui Meng
  • Haifeng Zhang
  • Jun Wang

Many real-world scenarios involve a team of agents that have to coordinate their policies to achieve a shared goal. Previous studies mainly focus on decentralized control to maximize a common reward and barely consider the coordination among control policies, which is critical in dynamic and complicated environments. In this work, we propose factorizing the joint team policy into a graph generator and graph-based coordinated policy to enable coordinated behaviours among agents. The graph generator adopts an encoder-decoder framework that outputs directed acyclic graphs (DAGs) to capture the underlying dynamic decision structure. We also apply the DAGness-constrained and DAG depth-constrained optimization in the graph generator to balance efficiency and performance. The graph-based coordinated policy exploits the generated decision structure. The graph generator and coordinated policy are trained simultaneously to maximize the discounted return. Empirical evaluations on Collaborative Gaussian Squeeze, Cooperative Navigation, and Google Research Football demonstrate the superiority of the proposed method. The code is available at https: //github. com/Amanda-1997/GCS_aamas337.

AAAI Conference 2022 Conference Paper

Multi-Sacle Dynamic Coding Improved Spiking Actor Network for Reinforcement Learning

  • Duzhen Zhang
  • Tielin Zhang
  • Shuncheng Jia
  • Bo Xu

With the help of deep neural networks (DNNs), deep reinforcement learning (DRL) has achieved great success on many complex tasks, from games to robotic control. Compared to DNNs with partial brain-inspired structures and functions, spiking neural networks (SNNs) consider more biological features, including spiking neurons with complex dynamics and learning paradigms with biologically plausible plasticity principles. Inspired by the efficient computation of cell assembly in the biological brain, whereby memorybased coding is much more complex than readout, we propose a multiscale dynamic coding improved spiking actor network (MDC-SAN) for reinforcement learning to achieve effective decision-making. The population coding at the network scale is integrated with the dynamic neurons coding (containing 2nd-order neuronal dynamics) at the neuron scale towards a powerful spatial-temporal state representation. Extensive experimental results show that our MDC-SAN performs better than its counterpart deep actor network (based on DNNs) on four continuous control tasks from OpenAI gym. We think this is a significant attempt to improve SNNs from the perspective of efficient coding towards effective decisionmaking, just like that in biological networks.

AAAI Conference 2021 Conference Paper

Consecutive Decoding for Speech-to-text Translation

  • Qianqian Dong
  • Mingxuan Wang
  • Hao Zhou
  • Shuang Xu
  • Bo Xu
  • Lei Li

Speech-to-text translation (ST), which directly translates the source language speech to the target language text, has attracted intensive attention recently. However, the combination of speech recognition and machine translation in a single model poses a heavy burden on the direct cross-modal crosslingual mapping. To reduce the learning difficulty, we propose COnSecutive Transcription and Translation (COSTT), an integral approach for speech-to-text translation. The key idea is to generate source transcript and target translation text with a single decoder. It benefits the model training so that additional large parallel text corpus can be fully exploited to enhance the speech translation training. Our method is verified on three mainstream datasets, including Augmented LibriSpeech English-French dataset, TED English-German dataset, and TED English-Chinese dataset. Experiments show that our proposed COSTT outperforms the previous state-ofthe-art methods. The code is available at https: //github. com/ dqqcasia/st.

AAAI Conference 2021 Conference Paper

Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text Translation

  • Qianqian Dong
  • Rong Ye
  • Mingxuan Wang
  • Hao Zhou
  • Shuang Xu
  • Bo Xu
  • Lei Li

An end-to-end speech-to-text translation (ST) takes audio in a source language and outputs the text in a target language. Existing methods are limited by the amount of parallel corpus. Can we build a system to fully utilize signals in a parallel ST corpus? We are inspired by human understanding system which is composed of auditory perception and cognitive processing. In this paper, we propose Listen-Understand- Translate, (LUT), a unified framework with triple supervision signals to decouple the end-to-end speech-to-text translation task. LUT is able to guide the acoustic encoder to extract as much information from the auditory input. In addition, LUT utilizes a pre-trained BERT model to enforce the upper encoder to produce as much semantic information as possible, without extra data. We perform experiments on a diverse set of speech translation benchmarks, including Librispeech English-French, IWSLT English-German and TED English-Chinese. Our results demonstrate LUT achieves the state-of-the-art performance, outperforming previous methods. The code is available at https: //github. com/dqqcasia/st.

AAAI Conference 2020 Conference Paper

DMRM: A Dual-Channel Multi-Hop Reasoning Model for Visual Dialog

  • Feilong Chen
  • Fandong Meng
  • Jiaming Xu
  • Peng Li
  • Bo Xu
  • Jie Zhou

Visual Dialog is a vision-language task that requires an AI agent to engage in a conversation with humans grounded in an image. It remains a challenging task since it requires the agent to fully understand a given question before making an appropriate response not only from the textual dialog history, but also from the visually-grounded information. While previous models typically leverage single-hop reasoning or single-channel reasoning to deal with this complex multimodal reasoning task, which is intuitively insufficient. In this paper, we thus propose a novel and more powerful Dual-channel Multi-hop Reasoning Model for Visual Dialog, named DMRM. DMRM synchronously captures information from the dialog history and the image to enrich the semantic representation of the question by exploiting dual-channel reasoning. Specifically, DMRM maintains a dual channel to obtain the question- and history-aware image features and the question- and image-aware dialog history features by a mulithop reasoning process in each channel. Additionally, we also design an effective multimodal attention to further enhance the decoder to generate more accurate responses. Experimental results on the VisDial v0. 9 and v1. 0 datasets demonstrate that the proposed model is effective and outperforms compared models by a significant margin.

IJCAI Conference 2020 Conference Paper

LISNN: Improving Spiking Neural Networks with Lateral Interactions for Robust Object Recognition

  • Xiang Cheng
  • Yunzhe Hao
  • Jiaming Xu
  • Bo Xu

Spiking Neural Network (SNN) is considered more biologically plausible and energy-efficient on emerging neuromorphic hardware. Recently backpropagation algorithm has been utilized for training SNN, which allows SNN to go deeper and achieve higher performance. However, most existing SNN models for object recognition are mainly convolutional structures or fully-connected structures, which only have inter-layer connections, but no intra-layer connections. Inspired by Lateral Interactions in neuroscience, we propose a high-performance and noise-robust Spiking Neural Network (dubbed LISNN). Based on the convolutional SNN, we model the lateral interactions between spatially adjacent neurons and integrate it into the spiking neuron membrane potential formula, then build a multi-layer SNN on a popular deep learning framework, i. \, e. , PyTorch. We utilize the pseudo-derivative method to solve the non-differentiable problem when applying backpropagation to train LISNN and test LISNN on multiple standard datasets. Experimental results demonstrate that the proposed model can achieve competitive or better performance compared to current state-of-the-art spiking neural networks on MNIST, Fashion-MNIST, and N-MNIST datasets. Besides, thanks to lateral interactions, our model processes stronger noise-robustness than other SNN. Our work brings a biologically plausible mechanism into SNN, hoping that it can help us understand the visual information processing in the brain.

NeurIPS Conference 2020 Conference Paper

Sequence to Multi-Sequence Learning via Conditional Chain Mapping for Mixture Signals

  • Jing Shi
  • Xuankai Chang
  • Pengcheng Guo
  • Shinji Watanabe
  • Yusuke Fujita
  • Jiaming Xu
  • Bo Xu
  • Lei Xie

Neural sequence-to-sequence models are well established for applications which can be cast as mapping a single input sequence into a single output sequence. In this work, we focus on one-to-many sequence transduction problems, such as extracting multiple sequential sources from a mixture sequence. We extend the standard sequence-to-sequence model to a conditional multi-sequence model, which explicitly models the relevance between multiple output sequences with the probabilistic chain rule. Based on this extension, our model can conditionally infer output sequences one-by-one by making use of both input and previously-estimated contextual output sequences. This model additionally has a simple and efficient stop criterion for the end of the transduction, making it able to infer the variable number of output sequences. We take speech data as a primary test field to evaluate our methods since the observed speech data is often composed of multiple sources due to the nature of the superposition principle of sound waves. Experiments on several different tasks including speech separation and multi-speaker speech recognition show that our conditional multi-sequence models lead to consistent improvements over the conventional non-conditional models.

AAAI Conference 2019 Conference Paper

Adapting Translation Models for Transcript Disfluency Detection

  • Qianqian Dong
  • Feng Wang
  • Zhen Yang
  • Wei Chen
  • Shuang Xu
  • Bo Xu

Transcript disfluency detection (TDD) is an important component of the real-time speech translation system, which arouses more and more interests in recent years. This paper presents our study on adapting neural machine translation (NMT) models for TDD. We propose a general training framework for adapting NMT models to TDD task rapidly. In this framework, the main structure of the model is implemented similar to the NMT model. Additionally, several extended modules and training techniques which are independent of the NMT model are proposed to improve the performance, such as the constrained decoding, denoising autoencoder initialization and a TDD-specific training object. With the proposed training framework, we achieve significant improvement. However, it is too slow in decoding to be practical. To build a feasible and production-ready solution for TDD, we propose a fast non-autoregressive TDD model following the non-autoregressive NMT model emerged recently. Even we do not assume the specific architecture of the NMT model, we build our TDD model on the basis of Transformer, which is the state-of-the-art NMT model. We conduct extensive experiments on the publicly available set, Switchboard, and in-house Chinese set. Experimental results show that the proposed model significantly outperforms previous state-ofthe-art models.

AIIM Journal 2019 Journal Article

Detection of protein complexes from multiple protein interaction networks using graph embedding

  • Xiaoxia Liu
  • Zhihao Yang
  • Shengtian Sang
  • Hongfei Lin
  • Jian Wang
  • Bo Xu

Cellular processes are typically carried out by protein complexes rather than individual proteins. Identifying protein complexes is one of the keys to understanding principles of cellular organization and function. Also, protein complexes are a group of interacting genes underlying similar diseases, which points out the therapeutic importance of protein complexes. With the development of life science and computing science, an increasing amount of protein–protein interaction (PPI) data becomes available, which makes it possible to predict protein complexes from PPI networks. However, most PPI data produced by high-throughput experiments often has many false positive interactions and false negative edge loss, which makes it difficult to predict complexes accurately. In this paper, we present a new method, named as MEMO (Multiple network Embedding for coMplex detectiOn), to detect protein complexes. MEMO integrates multiple PPI datasets from different species into a single PPI network by using functional orthology information across multiple species and then uses a graph embedding technology to embed protein nodes of the network into continuous vector spaces, so as to quantify the relationships between nodes and better guild the protein complex detection process. Finally, it utilizes a seed-and-extend strategy to identify protein complexes from multiple PPI networks based on the similarities of their corresponding protein representations. As part of our approach, we also define a new quality measure which combines the cluster cohesiveness and cluster density to measure the likelihood of a detected protein complex being a real protein complex. Extensive experimental results demonstrate the proposed method outperforms state-of-the-art complex detection techniques.

IJCAI Conference 2018 Conference Paper

Brain-inspired Balanced Tuning for Spiking Neural Networks

  • Tielin Zhang
  • Yi Zeng
  • Dongcheng Zhao
  • Bo Xu

Due to the nature of Spiking Neural Networks (SNNs), it is challenging to be trained by biologically plausible learning principles. The multi-layered SNNs are with non-differential neurons, temporary-centric synapses, which make them nearly impossible to be directly tuned by back propagation. Here we propose an alternative biological inspired balanced tuning approach to train SNNs. The approach contains three main inspirations from the brain: Firstly, the biological network will usually be trained towards the state where the temporal update of variables are equilibrium (e. g. membrane potential); Secondly, specific proportions of excitatory and inhibitory neurons usually contribute to stable representations; Thirdly, the short-term plasticity (STP) is a general principle to keep the input and output of synapses balanced towards a better learning convergence. With these inspirations, we train SNNs with three steps: Firstly, the SNN model is trained with three brain-inspired principles; then weakly supervised learning is used to tune the membrane potential in the final layer for network classification; finally the learned information is consolidated from membrane potential into the weights of synapses by Spike-Timing Dependent Plasticity (STDP). The proposed approach is verified on the MNIST hand-written digit recognition dataset and the performance (the accuracy of 98. 64%) indicates that the ideas of balancing state could indeed improve the learning ability of SNNs, which shows the power of proposed brain-inspired approach on the tuning of biological plausible SNNs.

IJCAI Conference 2018 Conference Paper

Listen, Think and Listen Again: Capturing Top-down Auditory Attention for Speaker-independent Speech Separation

  • Jing Shi
  • Jiaming Xu
  • Guangcan Liu
  • Bo Xu

Recent deep learning methods have made significant progress in multi-talker mixed speech separation. However, most existing models adopt a driftless strategy to separate all the speech channels rather than selectively attend the target one. As a result, those frameworks may be failed to offer a satisfactory solution in complex auditory scene where the number of input sounds is usually uncertain and even dynamic. In this paper, we present a novel neural network based structure motivated by the top-down attention behavior of human when facing complicated acoustical scene. Different from previous works, our method constructs an inference-attention structure to predict interested candidates and extract each speech channel of them. Our work gets rid of the limitation that the number of channels must be given or the high computation complexity for label permutation problem. We evaluated our model on the WSJ0 mixed-speech tasks. In all the experiments, our model gets highly competitive to reach and even outperform the baselines.

AAAI Conference 2018 Conference Paper

Modeling Attention and Memory for Auditory Selection in a Cocktail Party Environment

  • Jiaming Xu
  • Jing Shi
  • Guangcan Liu
  • Xiuyi Chen
  • Bo Xu

Developing a computational auditory model to solve the cocktail party problem has long bedeviled scientists, especially for a single microphone recording. Although recent deep learning based frameworks have made significant progress in multi-talker mixed speech separation, most existing deep learning based methods, focusing on separating all the speech channels rather than selectively attending the target speech and ignoring other sounds, may fail to offer a satisfactory solution in a complex auditory scene where the number of input sounds is usually uncertain and even dynamic. In this work, we employ ideas from auditory selective attention of behavioral and cognitive neurosciences and from recent advances of memory-augmented neural networks. Specifically, a unified Auditory Selection framework with Attention and Memory (dubbed ASAM) is proposed. Our ASAM first accumulates the prior knowledge (that is the acoustic feature to one specific speaker) into a life-long memory during the training phase, meanwhile a speech perceptor is trained to extract the temporal acoustic feature and update the memory online when a salient speech is given. Then, the learned memory is utilized to interact with the mixture input to attend and filter the target frequency out from the mixture stream. Finally, the network is trained to minimize the reconstruction error of the attended speech. We evaluate the proposed approach on WSJ0 and THCHS-30 datasets and the experimental results demonstrate that our approach successfully conducts two auditory selection tasks: the top-down task-specific attention (e. g. to follow a conversation with friend) and the bottom-up stimulus-driven attention (e. g. be attracted by a salient speech). Compared with deep clustering based methods, our method conducts competitive advantages especially in a real noise environment (e. g. street junction). Our code is available at https: //github. com/jacoxu/ASAM.

IJCAI Conference 2016 Conference Paper

Learning Defining Features for Categories

  • Bo Xu
  • Chenhao Xie
  • Yi Zhang
  • Yanghua Xiao
  • Haixun Wang
  • Wei Wang

Categories play a fundamental role in human cognition. Defining features (short for DFs) are the key elements to define a category, which enables machines to categorize objects. Categories enriched with their DFs significantly improve the machine's ability of categorization and benefit many applications built upon categorization. However, defining features can rarely be found for categories in current knowledge bases. Traditional efforts such as manual construction by domain experts are not practical to find defining features for millions of categories. In this paper, we make the first attempt to automatically find defining features for millions of categories in the real world. We formalize the defining feature learning problem and propose a bootstrapping solution to learn defining features from the features of entities belonging to a category. Experimental results show the effectiveness and efficiency of our method. Finally, we find defining features for overall 60, 247 categories with acceptable accuracy.

IJCAI Conference 2015 Conference Paper

Convolutional Neural Networks for Text Hashing

  • Jiaming Xu
  • Peng Wang
  • Guanhua Tian
  • Bo Xu
  • Jun Zhao
  • Fangyuan Wang
  • Hongwei Hao

Hashing, as a popular approximate nearest neighbor search, has been widely used for large-scale similarity search. Recently, a spectrum of machine learning methods are utilized to learn similarity-preserving binary codes. However, most of them directly encode the explicit features, keywords, which fail to preserve the accurate semantic similarities in binary code beyond keyword matching, especially on short texts. Here we propose a novel text hashing framework with convolutional neural networks. In particular, we first embed the keyword features into compact binary code with a locality preserving constraint. Meanwhile word features and position features are together fed into a convolutional network to learn the implicit features which are further incorporated with the explicit features to fit the pre-trained binary code. Such base method can be successfully accomplished without any external tags/labels, and other three model variations are designed to integrate tags/labels. Experimental results show the superiority of our proposed approach over several state-of-the-art hashing methods when tested on one short text dataset as well as one normal text dataset.

IJCAI Conference 2013 Conference Paper

Joint and Coupled Bilingual Topic Model Based Sentence Representations for Language Model Adaptation

  • Shixiang Lu
  • Xiaoyin Fu
  • Wei Wei
  • Xingyuan Peng
  • Bo Xu

This paper is concerned with data selection for adapting language model (LM) in statistical machine translation (SMT), and aims to find the LM training sentences that are topic similar to the translation task. Although the traditional approaches have gained significant performance, they ignore the topic information and the distribution information of words when selecting similar training sentences. In this paper, we present two bilingual topic model (BLTM) (joint and coupled BLTM) based sentence representations for cross-lingual data selection. We map the data selection task into cross-lingual semantic representations that are language independent, then rank and select sentences in the target language LM training corpus for a sentence in the translation task by the semanticsbased likelihood. The semantic representations are learned from the parallel corpus, with the assumption that the bilingual pair shares the same or similar distribution over semantic topics. Largescale experimental results demonstrate that our approaches significantly outperform the state-of-theart approaches on both LM perplexity and translation performance, respectively.

TAAS Journal 2009 Journal Article

Machine learning in disruption-tolerant MANETs

  • Bo Xu
  • Ouri Wolfson
  • Channah Naiman

In this article we study the data dissemination problem in which data items are flooded to all the moving objects in a mobile ad hoc network by peer-to-peer transfer. We show that if memory and bandwidth are bounded at moving objects, then the problem of determining whether a set of data items can be disseminated to all the moving objects is NP-complete. For a heuristic solution we postulate that a moving object should save and transmit the data items that are most likely to be new (i.e., previously unknown) to future encountered moving objects. We propose a method to be used by each moving object to prioritize data items based on their probabilities of being new to future receivers. The method employs a machine learning system for estimation of the novelty probability and the machine learning system is progressively trained by received data items. Through simulations based on real mobility traces, we show the superiority of the method against some natural alternatives.

IJCAI Conference 2007 Conference Paper

  • Xiangyu Duan
  • Jun Zhao
  • Bo Xu

Currently most word sense disambiguation (WSD) systems are relatively individual word sense experts. Scarcely do these systems take word sense transitions between senses of linearly consecutive words or syntactically dependent words into consideration. Word sense transitions are very important. They embody the fluency of semantic expression and avoid sparse data problem effectively. In this paper, HowNet knowledge base is used to decompose every word sense into several sememes. Then one transition between two words' senses becomes multiple transitions between sememes. Sememe transitions are much easier to be captured than word sense transitions due to much less sememes. When sememes are labeled, WSD is done. In this paper, multi-layered conditional random fields (MLCRF) is proposed to model sememe transitions. The experiments show that MLCRF performs better than a base-line system and a maximum entropy model. Syntactic and hypernym features can enhance the performance significantly.

v2026.09.13