Arrow Research search

Author name cluster

Bin Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

66 papers
2 author rows

Possible papers

66

EAAI Journal 2026 Journal Article

Advanced fetal cerebellar vermis segmentation and gestational age prediction in ultrasound imaging for prenatal neural development assessment

  • Qifeng Wang
  • Dan Zhao
  • Hao Ma
  • Bin Liu

In prenatal diagnostics, accurate segmentation of the Cerebellar Vermis (CV) in fetal brain ultrasound images is essential for assessing fetal neural development. Traditional manual segmentation methods are prone to omissions and misdiagnoses, and even minor measurement errors can significantly affect fetal health and diagnostic accuracy. To address these challenges, we propose the Fetal Brain Feature Enhanced UNet (FB-FEUNet), an advanced segmentation network enhancing CV segmentation precision through specialized modules: the Week Embedding Module (WEM) for incorporating gestational timing, the Fusion Feature Attention Module (FFAM) for multisource feature integration, the Week Conditional Attention Module (WCAM) for gestational age-aware adjustments, and the Fusion Constraint Module (FCM) to enhance segmentation accuracy. The model was trained on the Fetal Brain Cerebellar Vermis Dataset (FB-CV), a curated collection of fetal brain ultrasound images designed to support robust evaluation. Experimental results demonstrate that FB-FEUNet achieves a Dice Coefficient of 0. 8670 and an Intersection over Union of 0. 7686, outperforming state-of-the-art methods in both accuracy and stability, while also providing faster inference times. These findings confirm the effectiveness of FB-FEUNet in addressing segmentation challenges and highlight its clinical potential to improve diagnostic accuracy, reduce manual errors, and improve fetal neural development assessments.

EAAI Journal 2026 Journal Article

Cross-domain attention guided multi-source domain adaptation method for machinery fault diagnosis

  • Jie Wang
  • Jianning Gou
  • Haidong Shao
  • Yiming Xiao
  • Ying Peng
  • Bin Liu

Compared with single-source approaches, multi-source domain adaptation (MSDA) for fault diagnosis integrates complementary information from various domains. This avoids the subjectivity and arbitrariness associated with selecting a single source. However, existing MSDA methods for fault diagnosis typically enforce global distribution alignment between the feature of source and target domains. Such alignment often leads to the loss of discriminative fault features in the target domain, resulting in negative transfer. To address aforementioned issues, a cross-domain attention guided MSDA model (CDA-MSDA) is proposed in this paper. In this framework, a cross-domain attention module is constructed to dynamically fuse source and target domain features. This module effectively enhances the transfer of task-relevant features in the source domain and preserves discriminative features in the target domain. Then, a fault knowledge distillation module is developed to guide the feature extractor and classifier in achieving cross-domain fault category alignment. Finally, a multi-model dynamic collaborative decision module is designed. By aggregating prediction results from multiple classifiers, it addresses prediction conflicts arising from the varying reliability of different source domains. Extensive experiments on three benchmark datasets across 16 transfer tasks validate the effectiveness of the proposed method. Specifically, CDA-MSDA achieves an average diagnostic accuracy of 94. 99 %, outperforming state-of-the-art baselines by 2–10 %, demonstrating superior robustness and stability in complex fault diagnosis scenarios.

AAAI Conference 2026 Conference Paper

Discriminative Graph Embedding Framework via Label-Free Marginal Fisher Analysis

  • Qianqian Wang
  • Mengping Jiang
  • Wei Feng
  • Haixi Zhang
  • Bin Liu

Marginal Fisher Analysis (MFA) is a classical dimensionality reduction (DR) method that leverages dual graphs to capture intra-class compactness and inter-class separability. However, MFA’s reliance on high-quality labels limits its practical application. For another, existing unsupervised DR methods neglect data’s local manifold relationship, resulting in poor discriminativeness. To address these limitations, we propose a novel DR method named Discriminative Graph Embedding Framework (DGEF) via Label-Free Marginal Fisher Analysis. Our approach uses the adjacency matrix and cluster indicator matrix derived from centerless K-Means to construct intrinsic graph and penalty graph, which preserve the local manifold structure of the data. Additionally, we have derived the convertible relationship between centerless K-Means and Manifold learning and unified them within a graph embedding framework. By adopting the intrinsic graph and penalty graph, our DGEF avoids centroid initialization and ensures robustness and discriminativeness. This method achieves dimensionality reduction adaptively without relying on labeled data. Extensive experiments on benchmark datasets show that our approach outperforms conventional methods in clustering performance.

EAAI Journal 2026 Journal Article

Dual-stage interpretable domain generalization fault diagnosis: integrating prior knowledge and gradient-weighted class activation mapping

  • Ying Peng
  • Haidong Shao
  • Yiming Xiao
  • Jie Wang
  • Bin Liu

Recent advancements in domain generalization methods for fault diagnosis have achieved excellent performance. However, its inherent black-box characteristics seriously hinder its practical deployment in critical industrial scenarios. In addition, current cross-domain interpretability research often focuses on a single stage, resulting in an incomplete and unreliable understanding of model behavior. To overcome the above bottlenecks, this article proposes a dual-stage interpretable domain generalization fault diagnosis framework. In the first stage, a prior knowledge-guided feature extractor is constructed to extract steady-state and transient features from low- and high-frequency directions, thereby improving the model's ante-hoc interpretability. In the second stage, gradient-weighted class activation mapping is employed to visualize the class activation maps, revealing the attention regions during signal processing and enabling post-hoc interpretability analysis. The proposed method is validated using two distinct gearbox datasets, demonstrating superior performance in diagnostic accuracy and model interpretability compared to conventional domain generalization fault diagnosis approaches. In addition, the prior knowledge-guided feature extractor proves effective when integrated into other domain generalization models, and gradient-weighted class activation mapping proves to be a valuable tool for post-hoc interpretability assessment in the field of domain generalization fault diagnosis.

TCS Journal 2026 Journal Article

Dynamic algorithms for maximizing a DR-submodular function subtracted by a linear function over the integer lattice

  • Yuanyuan Qiang
  • Bin Liu

Submodular maximization plays a fundamental role in combinatorial optimization. Recently, driven by real-world applications, there has been growing interest in studying submodularity over the integer lattice. In this paper, we propose dynamic algorithms for two distinct problems. First, for maximizing a monotone DR-submodular function minus a linear function over a bounded integer lattice c = ( c 1, c 2, ⋯, c n ) ∈ Z + n, we develop a ( 1 2, 1 ) -bicriteria approximation algorithm, with an amortized update query complexity of O ( r ^ | | c | | 1 log | | c | | 1 log | | c | | ∞ ϵ 2 ), where r ^ represents the average number of elements inserted or deleted per update, and the amortized update time corresponds to the amortized number of oracle queries per update. Second, we extend this problem to include an additional cardinality constraint k, and propose a different dynamic algorithm that achieves a ( 3 − 5 2, 1 ) -bicriteria approximation with an amortized update query complexity of O ( r ^ k log | | c | | 1 log | | c | | ∞ ϵ 2 ). Unlike classical dynamic models, our framework allows for the simultaneous insertions or deletions of multiple identical elements. This paper presents a unified approach that extends beyond conventional set-based methods, offering new perspectives on the efficiency and behavior of dynamic algorithms in a multiset setting.

AAAI Conference 2026 Conference Paper

EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding

  • Kai Zou
  • Hongbo Liu
  • Dian Zheng
  • Jianxiong Gao
  • Zhiwei Zhao
  • Bin Liu

In this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with both accurate layout and high fidelity to the text description.(e.g., spatial relationship), and grounding the image robustly at the same time. We believe that image grounding possesses strong text and layout understanding abilities, which can compensate for the corresponding limitations in layout-to-image generation. At the same time, images generated from layouts exhibit high diversity in content, thereby enhancing the robustness of image grounding. Jointly training both tasks within a unified model can promote performance improvements for each. However, we identify that this joint training paradigm encounters several optimization challenges and results in restricted performance. To address these issues, we propose progressive training strategies. First, the Parallel Multi-Task Pre-training (PMTP) stage equips the model with basic abilities for both tasks, leveraging shared tokens to accelerate training. Next, the Dual Joint Optimization (DJO) stage exploits task duality to sequentially integrate the two tasks, enabling unified optimization. Finally, the Cycle RL stage eliminates reliance on visual supervision by using consistency constraints as rewards, significantly enhancing the model’s unified capabilities via the GRPO strategy. Extensive experiments demonstrate state-of-the-art results on both layout-to-image generation and image grounding benchmarks, and reveal clear synergistic gains from optimizing the two tasks together.

EAAI Journal 2026 Journal Article

Exploiting anchor-free and graph reasoning framework for dense tea bud detection and picking point identification

  • Zhiye Shen
  • Yinghu Cai
  • Kaile Yuan
  • Bin Liu
  • Wenbin Zhen
  • Ruijun Ma
  • Long Qi

Accurate detection of tea buds and identification of picking points are essential for automated tea harvesting. However, these tasks remain difficult in natural plantation environments due to dense clustering of buds, irregular spatial patterns, and frequent occlusions. To address these challenges, this study presents a two-stage perception framework that combines anchor-free dense tea bud detection with a graph reasoning approach to accurately identify picking points. In the first stage, an anchor-free dense tea bud detection strategy is adopted to avoid unstable anchor assignment in crowded scenes. It incorporates bounding box refinement with an Intersection over Union (IoU) class score to align detection confidence with geometric precision. In the second stage, the refined detections are exploited as structural cues for the occluded picking point identification module. A graph aware layer with relative position loss is employed to model spatial dependencies among picking points and auxiliary landmarks, enabling the inference of occluded targets based on learned structural cues. Experiments on a custom dataset of 5001 images demonstrate that the proposed framework achieves competitive performance compared with representative methods under the same evaluation protocol. Specifically, it achieves a mean Average Precision (mAP) of 60. 9% for dense detection and 96. 2% for identification of occluded picking points. The proposed framework has been deployed on embedded computing devices, achieving an inference speed of 28. 50 frames per second (FPS) for detection and 8. 88 ms per bud for picking point identification. These results demonstrate its feasibility for real-world tea harvesting applications.

AAAI Conference 2026 Conference Paper

Federated Incomplete Multi-View Clustering with Tensorized Low-Rank Constraint

  • Wei Feng
  • Danting Liu
  • Qianqian Wang
  • Mengping Jiang
  • Bin Liu

Federated Multi-View Clustering has gained increasing attention for its ability to discover complementary clustering structures of distributed multi-view data while preserving data privacy. However, real-world clients often only have access to partial views, and the view incompleteness poses great challenges to federated multi-view feature fusion to exploit consistent and complementary information. Moreover, efficiency is highly expected in federated scenarios due to the limited resources of each client. To alleviate these issues, we propose Federated Incomplete Multi-View Clustering with Tensorized Low-Rank Constraint (FIMVC-TLRC), which incorporates anchors to improve efficiency and is able to address prevalent view incompleteness issue in federated scenarios. FIMVC-TLRC aligns the local anchor graphs and employs a tensorized low-rank constraint based on the tensor Schatten p-norm to enforce the consistency of the data representations learned by each client. Besides, a federated optimization framework is developed to jointly optimize the construction and alignment of anchor graphs, thus enabling collaborative and privacy-preserving training. Experimental results on multiple datasets demonstrate its effectiveness.

AAAI Conference 2025 Conference Paper

Batch Selection for Multi-Label Classification Guided by Uncertainty and Dynamic Label Correlations

  • Ao Zhou
  • Bin Liu
  • Jin Wang
  • Grigorios Tsoumakas

The accuracy of deep neural networks is significantly influenced by the effectiveness of mini-batch construction during training. In single-label scenarios, such as binary and multi-class classification tasks, it has been demonstrated that batch selection algorithms preferring samples with higher uncertainty achieve better performance than difficulty-based methods. Although there are two batch selection methods tailored for multi-label data, none of them leverage important uncertainty information. Adapting the concept of uncertainty to multi-label data is not a trivial task, since there are two issues that should be tackled. First, traditional variance or entropy-based uncertainty measures ignore fluctuations of predictions within sliding windows and the importance of the current model state. Second, existing multi-label methods do not explicitly exploit the label correlations, particularly the uncertainty-based label correlations that evolve during the training process. In this paper, we propose an uncertainty-based multi-label batch selection algorithm. It assesses uncertainty for each label by considering differences between successive predictions and the confidence of current outputs, and further leverages dynamic uncertainty-based label correlations to emphasize instances whose uncertainty is synergistically expressed across multiple labels. Empirical studies demonstrate the effectiveness of our method in improving the performance and accelerating the convergence of various multi-label deep learning models.

EAAI Journal 2025 Journal Article

Compact-sparse prototype calibration network for few-shot continual fault diagnosis of rotating machinery

  • Shen Yan
  • Haidong Shao
  • Xinyi Wang
  • Haomiao Zhang
  • Yiming Xiao
  • Bin Liu

Rotating machinery inevitably generates only a few samples of new fault categories during long-term operation, which requires the fault diagnosis model to incrementally learn few new categories and retain the existing fault knowledge. Recent few-shot continual fault diagnosis (FSCFD) methods mainly rely on constructing prototype classifiers and generating virtual samples to address the challenges of catastrophic forgetting and overfitting. However, this would ignore the rich feature information in the base session, while the limited incremental data makes it difficult to accurately depict the new category feature information. Therefore, a compact-sparse prototype calibration network (CSPCN) is proposed to improve the diagnosis capacity for new category faults in the FSCFD scenario. First, a compact-sparse base loss (CSBL) is employed to reserve sufficient space for new fault categories by maximizing the variance distribution among base prototypes. Second, an incremental prototype calibration classifier (IPCC) is designed to improve the ability to distinguish new categories by integrating new prototypes with the weighted base prototypes in real time. Extensive experiments conducted on the subway train bogie and variable load gearbox dataset validate the proposed method's exceptional diagnostic performance. Through multidimensional comparisons with state-of-the-art FSCFD methods and rigorous ablation experiments, CSPCN demonstrates significant improvements in effectively identifying and distinguishing new fault categories.

TARK Conference 2025 Conference Paper

Distributed Knowing How

  • Bin Liu
  • Yanjing Wang 0001

Distributed knowledge is a key concept in the standard epistemic logic of knowledge-that. In this paper, we propose a corresponding notion of distributed knowledge-how and study its logic. Our framework generalizes two existing traditions in the logic of know-how: the individual-based multi-step framework and the coalition-based single-step framework. In particular, we assume a group can accomplish more than what its individuals can jointly do. The distributed knowledge-how is based on the distributed knowledge-that of a group whose multi-step strategies derive from distributed actions that subgroups can collectively perform. As the main result, we obtain a sound and strongly complete proof system for our logic of distributed knowledge-how, which closely resembles the logic of distributed knowledge-that in both the axioms and the proof method of completeness.

AIIM Journal 2025 Journal Article

Hybrid approach for drug-target interaction predictions in ischemic stroke models

  • Jing-Jie Peng
  • Yi-Yue Zhang
  • Rui-Feng Li
  • Wen-Jun Zhu
  • Hong-Rui Liu
  • Hui-Yin Li
  • Bin Liu
  • Dong-Sheng Cao

Multiple cell death mechanisms are triggered during ischemic stroke and they are interconnected in a complex network with extensive crosstalk, complicating the development of targeted therapies. We therefore propose a novel framework for identifying disease-specific drug-target interaction (DTI), named strokeDTI, to extract key nodes within an interconnected graph network of activated pathways via leveraging transcriptomic sequencing data. Our findings reveal that the drugs a model can predict are highly representative of the characteristics of the database the model is trained on. However, models with comparable performance yield diametrically opposite predictions in real testing scenarios. Our analysis reveals a correlation between the reported literature on drug-target pairs and their binding scores. Leveraging this correlation, we introduced an additional module to assess the predictive validity of our model for each unique target, thereby improving the reliability of the framework's predictions. Our framework identified Cerdulatinib as a potential anti-stroke drug via targeting multiple cell death pathways, particularly necroptosis and apoptosis. Experimental validation in in vitro and in vivo models demonstrated that Cerdulatinib significantly attenuated stroke-induced brain injury via inhibiting multiple cell death pathways, improving neurological function, and reducing infarct volume. This highlights strokeDTI's potential for disease-specific drug-target identification and Cerdulatinib's potential as a potent anti-stroke drug.

JBHI Journal 2025 Journal Article

MMFmiRLocEL: A Multi-Model Fusion and Ensemble Learning Approach for Identifying miRNA Subcellular Localization Using RNA Structure Language Model

  • Tao Bai
  • Junxi Xie
  • Bin Liu
  • Yumeng Liu

MiRNA subcellular localizations (MSLs) are essential for uncovering and understanding miRNA functions in various biological processes. Several computational methods have been proposed for measuring MSL. However, existing methods only rely on manually crafted features based on sequence without considering RNA 3D structure information, and most methods often rely on single-model approaches, which fail to capture the full complexity of biological systems, further hindering predictive accuracy and performance. In this study, we introduce a deep learning-based approach, MMFmiRLocEL, which integrates multi-model fusion and ensemble learning for MSL identification. To the best of our knowledge, MMFmiRLocEL is the first method to combine sequence, structure, and function three information for MSL prediction. Specifically, it employs RNA 3D structure generated by the predicted structural model to construct a structure-based approach for MSL prediction. It also develops a sequence-based prediction method using sequence features and convolutional neural networks, while constructing a function-based prediction method using miRNA-disease association networks and deep residual neural networks. Furthermore, a multi-model fusion approach, employing weighted ensemble strategies, integrates sequence, structure, and function models to enhance the robustness and accuracy of MSL identification. Experimental results demonstrate that MMFmiRLocEL outperforms existing state-of-the-art methods, and then ablation analysis confirmed the significant contribution of the multi-model fusion mechanism to improve the prediction performance.

EAAI Journal 2025 Journal Article

Residual generative adversarial network-driven Data enhancement for magnetic flux leakage-based defect recognition in oil and gas pipelines

  • Bin Liu
  • Liying Ding
  • Luyao He
  • Lijian Yang
  • Xiaobei Zhang
  • Ye Tian

Magnetic flux leakage (MFL) in-line inspection technology is widely regarded as one of the most effective techniques for assessing the health of long-distance oil and gas pipelines. However, the limited availability of MFL data for defects presents remarkable challenges in training defect recognition models based on deep learning. Consequently, this paper proposed a generative adversarial network (Res-CosGAN), integrating deep residual modules with cosine similarity loss. In comparison to existing data enhancement techniques for defect recognition networks, Res-CosGAN possesses three key advantages. Firstly, it directly utilizes the MFL data matrix as the input, thereby significantly reducing the time required to convert MFL data into images. Secondly, the method incorporates skip connections in the generator and introduces cosine similarity loss, mitigating the vanishing gradient problem. Lastly, utilizing the characteristics of MFL data, it integrates the physical information of defects into the generation loss, thereby minimizing the loss of data feature information and enhancing the quality of the generated data. Experimental results across multiple datasets indicated that, when applied to defect recognition networks, this data enhancement method could increase the average recognition accuracy by 25. 1 % compared with using raw MFL data and by 7. 4 % compared with the best results achieved with other enhancement networks. These findings demonstrate the effectiveness of the proposed method in pipeline MFL defect detection and its potential contribution to pipeline transportation safety analysis and assessment.

AAAI Conference 2025 Conference Paper

Rethinking Masked Data Reconstruction Pretraining for Strong 3D Action Representation Learning

  • Tao Gong
  • Qi Chu
  • Bin Liu
  • Nenghai Yu

In 3D human action recognition, limited supervised data makes it challenging to fully tap into the modeling potential of powerful networks such as transformers. As a result, researchers have been actively investigating effective self-supervised pre-training strategies. For example, MAMP shows that instead of following the prevalent masked joint reconstruction, explicit masked motion reconstruction is key to the success of learning effective feature representation for 3D action recognition. However, we find that if we make a simple and effective change to the reconstructed target of masked joint reconstruction, masked joint reconstruction can achieve the same results as masked motion reconstruction. The devil is in the special characteristic of 3D skeleton data and the normalization process of training targets. We need to dig for all effective information of targets during normalization. Besides, considering that mask data reconstruction focuses more on learning local relations in input data for fulfilling the reconstruction task, instead of modeling the relation among samples, we further employ contrastive learning to learn more discriminative 3D action representations. We show that contrastive learning can consistently boost the performance of model pre-trained by masked joint prediction under various settings, especially in the semi-supervised setting that has a very limited number of labeled samples. Extensive experiments on NTU-60, NTU-120, and PKU-MMD datasets show that the proposed pre-training strategy achieves state-of-the-art results without bells and whistles.

EAAI Journal 2025 Journal Article

Self-supervised contrastive learning for implicit collaborative filtering

  • Shipeng Song
  • Bin Liu
  • Fei Teng
  • Tianrui Li

Recommendation systems are a critical application of artificial intelligence (AI), driving personalized user experiences across various platforms. Recent advancements in contrastive learning-based recommendation algorithms have led to significant progress in self-supervised recommendation. A key method in this field is Bayesian Personalized Ranking (BPR), which has become a dominant approach for implicit collaborative filtering. However, the challenge of false-positive and false-negative examples in implicit feedback continues to hinder accurate preference learning. In this study, we introduce an efficient self-supervised contrastive learning framework that enhances the supervisory signal by incorporating positive feature augmentation and negative label augmentation. Our theoretical analysis reveals that this approach is equivalent to maximizing the likelihood estimation with latent variables representing user interest centers. Additionally, we present a novel negative label augmentation technique that selects unlabeled examples based on their relative ranking positions, enabling efficient augmentation with constant time complexity. Validation on the MovieLens-100k, MovieLens-1M, Yahoo! -R3, Yelp2018, and Gowalla datasets demonstrates that our method achieves over a 5% improvement in precision compared to the widely used BPR optimization objective, while maintaining comparable runtime efficiency.

NeurIPS Conference 2025 Conference Paper

Think before Recommendation: Autonomous Reasoning-enhanced Recommender

  • Xiaoyu Kong
  • Junguang Jiang
  • Bin Liu
  • Ziru Xu
  • Han Zhu
  • Jian Xu
  • Bo Zheng
  • Jiancan Wu

The core task of recommender systems is to learn user preferences from historical user-item interactions. With the rapid development of large language models (LLMs), recent research has explored leveraging the reasoning capabilities of LLMs to enhance rating prediction tasks. However, existing distillation-based methods suffer from limitations such as the teacher model's insufficient recommendation capability, costly and static supervision, and superficial transfer of reasoning ability. To address these issues, this paper proposes RecZero, a reinforcement learning (RL)-based recommendation paradigm that abandons the traditional multi-model and multi-stage distillation approach. Instead, RecZero trains a single LLM through pure RL to autonomously develop reasoning capabilities for rating prediction. RecZero consists of two key components: (1) "Think-before-Recommendation" prompt construction, which employs a structured reasoning template to guide the model in step-wise analysis of user interests, item features, and user-item compatibility; and (2) rule-based reward modeling, which adopts group relative policy optimization (GRPO) to compute rewards for reasoning trajectories and optimize the LLM. Additionally, the paper explores a hybrid paradigm, RecOne, which combines supervised fine-tuning with RL, initializing the model with cold-start reasoning samples and further optimizing it with RL. Experimental results demonstrate that RecZero and RecOne significantly outperform existing baseline methods on multiple benchmark datasets, validating the superiority of the RL paradigm in achieving autonomous reasoning-enhanced recommender systems.

TMLR Journal 2025 Journal Article

Thompson Sampling For Bandits With Cool-Down Periods

  • Jingxuan Zhu
  • Bin Liu

This paper investigates a variation of dynamic bandits, characterized by arms that follow a periodic availability pattern. Upon a "successful" selection, each arm transitions to an inactive state and requires a possibly unknown cool-down period before becoming active again. We devise Thompson Sampling algorithms specifically designed for this problem, guaranteeing logarithmic regrets. Notably, this work is the first to address scenarios in which the agent lacks knowledge of each arm's active state. Furthermore, the theoretical findings extend to the sleeping bandit framework, offering a notably superior regret bound compared to existing literature.

IJCAI Conference 2025 Conference Paper

Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification

  • Xulin Li
  • Yan Lu
  • Bin Liu
  • Jiaze Li
  • Qinhong Yang
  • Tao Gong
  • Qi Chu
  • Mang Ye

In real applications, person re-identification (ReID) expects to retrieve the target person at any time, including both daytime and nighttime, ranging from short-term to long-term. However, existing ReID tasks and datasets cannot meet this requirement, as they are constrained by available time and only provide training and evaluation for specific scenarios. Therefore, we investigate a new task called Anytime Person Re-identification (AT-ReID), which aims to achieve effective retrieval in multiple scenarios based on variations in time. To address the AT-ReID problem, we collect the first large-scale dataset, AT-USTC, which contains 135k images of individuals wearing multiple clothes captured by RGB and IR cameras. Our data collection spans over an entire year and 270 volunteers were photographed on average 29. 1 times across different dates or scenes, 4-15 times more than current datasets, providing conditions for follow-up investigations in AT-ReID. Further, to tackle the new challenge of multi-scenario retrieval, we propose a unified model named Uni-AT, which comprises a multi-scenario ReID (MS-ReID) framework for scenario-specific features learning, a Mixture-of-Attribute-Experts (MoAE) module to alleviate inter-scenario interference, and a Hierarchical Dynamic Weighting (HDW) strategy to ensure balanced training across all scenarios. Extensive experiments show that our model leads to satisfactory results and exhibits excellent generalization to all scenarios.

AAAI Conference 2025 Conference Paper

Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region Matching

  • Xuanpu Zhao
  • Dianmo Sheng
  • Zhentao Tan
  • Zhiwei Zhao
  • Tao Gong
  • Qi Chu
  • Bin Liu
  • Nenghai Yu

Open-vocabulary semantic segmentation (OVSS) aims to segment images of arbitrary categories specified by class labels. While previous approaches relied on extensive image-text pairs or dense semantic annotations, recent training-free methods attempted to overcome these limitations by constructing semantic prototypes in the construction stage and image-to-image matching (i.e., prototype matching) during testing. However, these methods often struggle to effectively capture the visual characteristics of categories and fail to utilize local features during prototype matching. To deal with these problems, we propose a novel training-free framework for OVSS that constructs diverse prototypes and performs fine-grained sub-region matching. Specifically, our method leverages Large Language Models (LLMs) to guide support image generation by descriptions of different attributes of categories and employs coarse-fine clustering to obtain diverse and robust part-level prototypes in the construction stage. During testing, we propose a sub-region matching method, which assigns part-level prototypes to sub-regions utilizing optimal transport, to fully utilize local image features among part-level prototypes. Extensive experiments demonstrate the effectiveness of our method and show that our method achieves state-of-the-art performance, outperforming previous methods across five datasets.

NeurIPS Conference 2025 Conference Paper

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection

  • Yang Zhao
  • Kai Xiong
  • Xiao Ding
  • Li Du
  • Yangou Ouyang
  • Zhouhao Sun
  • Jiannan Guan
  • Wenbin Zhang

A primary impediment to scaling reinforcement learning (RL) for large language model (LLM) training is the substantial computational cost, predominantly arising from the necessity of multi-sampling for policy optimization and evaluation. This underscores the critical yet challenging nature of efficient training data selection. Drawing inspiration from the Zone of Proximal Development (ZPD) theory, which posits that learners acquire knowledge more effectively from tasks of intermediate difficulty, we hypothesize that LLMs exhibit optimal learning from data they have not yet mastered but demonstrate the potential to comprehend. Conventional methodologies for assessing data difficulty or informativeness typically rely on computationally intensive multi-sampling or iterative procedures. To address this limitation, we introduce UFO-RL (**U**ncertainty-**F**ocused **O**ptimization for **R**einforcement **L**earning), a novel framework that employs a computationally efficient single-pass uncertainty estimation technique to identify informative training instances. This method, requiring only a single forward pass and obviating the need for iterative next-token computation, achieves a significant acceleration (up to 185$\times$) in data evaluation compared to multi-sampling approaches. UFO-RL leverages this efficient metric to select data within the model's estimated ZPD for training. Extensive experimentation across diverse LLMs and mathematical benchmarks demonstrates that training with a mere 10\% of the data, carefully selected by UFO-RL, yields performance comparable to or even surpassing that of full-data training. Furthermore, this targeted data selection results in up to a 16$\times$ reduction in overall training time, concurrently enhancing training stability and improving generalization capabilities. Thus, UFO-RL presents a practical and highly efficient strategy for scaling RL fine-tuning of LLMs by focusing learning efforts on the most informative and valuable data, thereby mitigating the computational bottlenecks associated with traditional RL training.

EAAI Journal 2025 Journal Article

Uncertainty-aware deep variational attention network: A trustworthy mechanical fault diagnostic model assisted by out-of-distribution detection

  • Yiming Xiao
  • Haidong Shao
  • Haomiao Zhang
  • Rongming Wei
  • Bin Liu

Most attention mechanisms used for fault diagnosis are deterministic attention with fixed attention weights and thus can only provide overconfident point estimate predictions. When handling out-of-distribution (OOD) inputs, they typically produce untrustworthy diagnostic decisions that incorrectly classify them into any seen fault classes. Unlike deterministic approaches, variational attention treats attention weights as latent random variables drawn from their posterior distribution, enabling uncertainty estimation to reject OOD samples. However, current implementations of variational attention typically require customized structural modifications to original deterministic attention, hindering deployment across different attention architectures. Moreover, existing variational attention primarily focuses on self-attention, with limited exploration in other types of attention. To address these limitations, this paper proposes: (1) an architecture-agnostic variational attention applicable to various popular deterministic attention; (2) a Monte Carlo shaping layer ensuring attention weights’ diversity for reliable uncertainty quantification; (3) a model uncertainty estimation scheme detecting distribution shifts. These components form an uncertainty-aware deep variational attention network for OOD detection-assisted trustworthy mechanical fault diagnosis. The experimental results show that the proposed method not only gives accurate predictions for in-distribution samples but also effectively detects OOD samples.

EAAI Journal 2024 Journal Article

A novel software defect prediction approach via weighted classification based on association rule mining

  • Wentao Wu
  • Shihai Wang
  • Bin Liu
  • Yuanxun Shao
  • Wandong Xie

Software defect prediction technology is used to assist software practitioners in effectively allocating test resources and identifying hidden defects in a timely manner. However, the prediction of defect-prone software using association rule mining algorithms is limited because of the unbalanced distribution of defect data. Furthermore, although the existing weighted association rule mining approach considers item strength, the weight calculation still relies on expert experience and lacks fine granularity. We propose a novel software defect prediction approach based on mutual information and correlation coefficient weighted class association rule mining (MCWCAR). The MCWCAR model employs a cost-sensitive strategy and generates frequent itemsets according to three mining objectives while maintaining the original item distribution: defective class rules, non-defective class rules, and feature association relationships. During the weighted frequent itemset mining process, it combines feature selection and itemset screening to determine the appropriate feature combination through mutual information weighted support. Meanwhile, the correlation coefficient is applied to accurately depict the correlation between feature items and defect classes, serving as the weight to mine class association rules. Additionally, to ensure that interestingness measures have asymmetry and effectively represent negative associations under the condition of class imbalance, the a d d e d v a l u e is adopted in the filtering association rules. We conducted experiments on 27 open-source datasets and evaluated the performance differences between MCWCAR and state-of-the-art baseline classifiers. Experimental results demonstrate that the proposed algorithm significantly outperforms other baselines in terms of B a l a n c e, G m e a n, M C C, and F - m e a s u r e.

TCS Journal 2024 Journal Article

An accelerated deterministic algorithm for maximizing monotone submodular minus modular function with cardinality constraint

  • Shufang Gong
  • Bin Liu
  • Qizhi Fang

Submodular optimization not only covers some classical combinatorial optimization problems, but also has a wide range of applications in fields such as machine learning and artificial intelligence. For submodular maximization problems with constraints, some work has been done including the design of approximation algorithms, the measurement of approximation algorithms in terms of quality and efficiency, etc. In this paper, we consider the problem of maximizing a non-negative monotone submodular function minus a non-negative modular function with the cardinality constraint. This model has been applied to many scenarios, such as team formation problem, influence maximization problem, recommender systems problem, etc. We propose a threshold algorithm that achieve a ( 1 / 2 − O ( ε ), 2 ) -bicriteria approximation ratio and query complexity O ( n log ⁡ n ). Our algorithm makes a small sacrifice in the approximation ratio but improves the best query complexity result of existing deterministic algorithms from O ( n 2 ) to O ( n log ⁡ n ) in the worst case.

YNICL Journal 2024 Journal Article

Chronic hypercortisolism disrupts the principal functional gradient in Cushing’s disease: A multi-scale connectomics and transcriptomics study

  • Guosong Shang
  • Tao Zhou
  • Xiaoteng Yu
  • Xinyuan Yan
  • Kunyu He
  • Bin Liu
  • Zhebin Feng
  • Junpeng Xu

Cushing's disease (CD) represents a state of cortisol excess, serving as a model to investigate the effects of prolonged hypercortisolism on functional brain. Potential alterations in the functional connectome of the brain may explain frequently reported cognitive deficits and affective disorders in CD patients. This study aims to elucidate the effects of chronic hypercortisolism on the principal functional gradient, which represents a hierarchical architecture with gradual transitions across cognitive processes, by integrating connectomics and transcriptomics approaches. Utilizing resting-state functional magnetic resonance imaging data from 140 participants (86 CD patients, 54 healthy controls) recruited at a single center, we explored the alterations in the principal gradient in CD patients. Further, we thoroughly explored the underlying associative mechanisms of the observed characteristic alterations with cognitive function domains, biological attributes, and neuropsychiatric representations, as well as gene expression profiles. Compared to healthy controls, CD patients demonstrated changes in connectome patterns in both primary and higher-order networks, exhibiting an overall converged trend along the principal gradient axis. The gradient values in CD patients' right prefrontal cortex and bilateral sensorimotor cortices exhibited a significant correlation with cortisol levels. Moreover, the cortical regions showing gradient alterations were principally associated with sensory information processing and higher-cognitive functions, as well as correlated with the gene expression patterns which involved synaptic components and function. The findings suggest that converged alterations in the principal gradient in CD patients may mediate the relationship between hypercortisolism and cognitive impairments, potentially involving genes regulating synaptic components and function.

AAAI Conference 2024 Conference Paper

Collaborative Consortium of Foundation Models for Open-World Few-Shot Learning

  • Shuai Shao
  • Yu Bai
  • Yan Wang
  • Baodi Liu
  • Bin Liu

Open-World Few-Shot Learning (OFSL) is a crucial research field dedicated to accurately identifying target samples in scenarios where data is limited and labels are unreliable. This research holds significant practical implications and is highly relevant to real-world applications. Recently, the advancements in foundation models like CLIP and DINO have showcased their robust representation capabilities even in resource-constrained settings with scarce data. This realization has brought about a transformative shift in focus, moving away from “building models from scratch” towards “effectively harnessing the potential of foundation models to extract pertinent prior knowledge suitable for OFSL and utilizing it sensibly”. Motivated by this perspective, we introduce the Collaborative Consortium of Foundation Models (CO3), which leverages CLIP, DINO, GPT-3, and DALL-E to collectively address the OFSL problem. CO3 comprises four key blocks: (1) the Label Correction Block (LC-Block) corrects unreliable labels, (2) the Data Augmentation Block (DA-Block) enhances available data, (3) the Feature Extraction Block (FE-Block) extracts multi-modal features, and (4) the Text-guided Fusion Adapter (TeFu-Adapter) integrates multiple features while mitigating the impact of noisy labels through semantic constraints. Only the adapter's parameters are adjustable, while the others remain frozen. Through collaboration among these foundation models, CO3 effectively unlocks their potential and unifies their capabilities to achieve state-of-the-art performance on multiple benchmark datasets. https://github.com/The-Shuai/CO3.

YNIMG Journal 2024 Journal Article

Disentangling sex-dependent effects of APOE on diverse trajectories of cognitive decline in Alzheimer's disease

  • Haixu Ma
  • Zhuoyu Shi
  • Minjeong Kim
  • Bin Liu
  • Patrick J. Smith
  • Yufeng Liu
  • Guorong Wu

Current diagnostic systems for Alzheimer's disease (AD) rely upon clinical signs and symptoms, despite the fact that the multiplicity of clinical symptoms renders various neuropsychological assessments inadequate to reflect the underlying pathophysiological mechanisms. Since putative neuroimaging biomarkers play a crucial role in understanding the etiology of AD, we sought to stratify the diverse relationships between AD biomarkers and cognitive decline in the aging population and uncover risk factors contributing to the diversities in AD. To do so, we capitalized on a large amount of neuroimaging data from the ADNI study to examine the inflection points along the dynamic relationship between cognitive decline trajectories and whole-brain neuroimaging biomarkers, using a state-of-the-art statistical model of change point detection. Our findings indicated that the temporal relationship between AD biomarkers and cognitive decline may differ depending on the synergistic effect of genetic risk and biological sex. Specifically, tauopathy-PET biomarkers exhibit a more dynamic and age-dependent association with Mini-Mental State Examination scores (p<0.05), with inflection points at 72, 78, and 83 years old, compared with amyloid-PET and neurodegeneration (cortical thickness from MRI) biomarkers. In the landscape of health disparities in AD, our analysis indicated that biological sex moderates the rate of cognitive decline associated with APOE4 genotype. Meanwhile, we found that higher education levels may moderate the effect of APOE4, acting as a marker of cognitive reserve.

RLJ Journal 2024 Journal Article

Enabling Intelligent Interactions between an Agent and an LLM: A Reinforcement Learning Approach

  • Bin Hu
  • Chenyang Zhao
  • Pu Zhang
  • Zihao Zhou
  • Yuanhang Yang
  • Zenglin Xu
  • Bin Liu

Large language models (LLMs) encode a vast amount of world knowledge acquired from massive text datasets. Recent studies have demonstrated that LLMs can assist an embodied agent in solving complex sequential decision making tasks by providing high-level instructions. However, interactions with LLMs can be time-consuming. In many practical scenarios, it requires a significant amount of storage space that can only be deployed on remote cloud servers. Additionally, using commercial LLMs can be costly since they may charge based on usage frequency. In this paper, we explore how to enable intelligent cost-effective interactions between a down stream task oriented agent and an LLM. We find that this problem can be naturally formulated by a Markov decision process (MDP), and propose When2Ask, a reinforcement learning based approach that learns when it is necessary to query LLMs for high-level instructions to accomplish a target task. On one side, When2Ask discourages unnecessary redundant interactions, while on the other side, it enables the agent to identify and follow useful instructions from the LLM. This enables the agent to halt an ongoing plan and transition to a more suitable one based on new environmental observations. Experiments on MiniGrid and Habitat environments that entail planning sub-goals demonstrate that When2Ask learns to solve target tasks with only a few necessary interactions with the LLM, significantly reducing interaction costs in testing environments compared with baseline methods. Our code is available at: https://github.com/ZJLAB-AMMI/LLM4RL.

RLC Conference 2024 Conference Paper

Enabling Intelligent Interactions between an Agent and an LLM: A Reinforcement Learning Approach

  • Bin Hu
  • Chenyang Zhao
  • Pu Zhang
  • Zihao Zhou
  • Yuanhang Yang
  • Zenglin Xu
  • Bin Liu

Large language models (LLMs) encode a vast amount of world knowledge acquired from massive text datasets. Recent studies have demonstrated that LLMs can assist an embodied agent in solving complex sequential decision making tasks by providing high-level instructions. However, interactions with LLMs can be time-consuming. In many practical scenarios, it requires a significant amount of storage space that can only be deployed on remote cloud servers. Additionally, using commercial LLMs can be costly since they may charge based on usage frequency. In this paper, we explore how to enable intelligent cost-effective interactions between a down stream task oriented agent and an LLM. We find that this problem can be naturally formulated by a Markov decision process (MDP), and propose When2Ask, a reinforcement learning based approach that learns when it is necessary to query LLMs for high-level instructions to accomplish a target task. On one side, When2Ask discourages unnecessary redundant interactions, while on the other side, it enables the agent to identify and follow useful instructions from the LLM. This enables the agent to halt an ongoing plan and transition to a more suitable one based on new environmental observations. Experiments on MiniGrid and Habitat environments that entail planning sub-goals demonstrate that When2Ask learns to solve target tasks with only a few necessary interactions with the LLM, significantly reducing interaction costs in testing environments compared with baseline methods. Our code is available at: https: //github. com/ZJLAB-AMMI/LLM4RL.

AAAI Conference 2024 Conference Paper

Enhancing Semi-supervised Domain Adaptation via Effective Target Labeling

  • Jiujun He
  • Bin Liu
  • Guosheng Yin

Existing semi-supervised domain adaptation (SSDA) models have exhibited impressive performance on the target domain by effectively utilizing few labeled target samples per class (e.g., 3 samples per class). To guarantee an equal number of labeled target samples for each class, however, they require domain experts to manually recognize a considerable amount of the unlabeled target data. Moreover, as the target samples are not equally informative for shaping the decision boundaries of the learning models, it is crucial to select the most informative target samples for labeling, which is, however, impossible for human selectors. As a remedy, we propose an EFfective Target Labeling (EFTL) framework that harnesses active learning and pseudo-labeling strategies to automatically select some informative target samples to annotate. Concretely, we introduce a novel sample query strategy, called non-maximal degree node suppression (NDNS), that iteratively performs maximal degree node query and non-maximal degree node removal to select representative and diverse target samples for labeling. To learn target-specific characteristics, we propose a novel pseudo-labeling strategy that attempts to label low-confidence target samples accurately via clustering consistency (CC), and then inject information of the model uncertainty into our query process. CC enhances the utilization of the annotation budget and increases the number of “labeled” target samples while requiring no additional manual effort. Our proposed EFTL framework can be easily coupled with existing SSDA models, showing significant improvements on three benchmarks

IJCAI Conference 2024 Conference Paper

Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents

  • Zihao Zhou
  • Bin Hu
  • Chenyang Zhao
  • Pu Zhang
  • Bin Liu

Recent studies have uncovered the potential of Large Language Models (LLMs) in addressing complex sequential decision-making tasks through the provision of high-level instructions. However, LLM-based agents lack specialization in tackling specific target problems, particularly in real-time dynamic environments. Additionally, deploying an LLM-based agent in practical scenarios can be both costly and time-consuming. On the other hand, reinforcement learning (RL) approaches train agents that specialize in the target task but often suffer from low sampling efficiency and high exploration costs. In this paper, we introduce a novel framework that addresses these challenges by training a smaller, specialized student RL agent using instructions from an LLM-based teacher agent. By incorporating the guidance from the teacher agent, the student agent can distill the prior knowledge of the LLM into its own model. Consequently, the student agent can be trained with significantly less data. Moreover, through further training with environment feedback, the student agent surpasses the capabilities of its teacher for completing the target task. We conducted experiments on challenging MiniGrid and Habitat environments, specifically designed for embodied AI research, to evaluate the effectiveness of our framework. The results clearly demonstrate that our approach achieves superior performance compared to strong baseline methods. Our code is available at https: //github. com/ZJLAB-AMMI/LLM4Teach.

JBHI Journal 2024 Journal Article

MMLmiRLocNet: miRNA Subcellular Localization Prediction based on Multi-view Multi-label Learning for Drug Design

  • Tao Bai
  • Junxi Xie
  • Yumeng Liu
  • Bin Liu

Identifying subcellular localization of microRNAs (miRNAs) is essential for comprehensive understanding of cellular function and has significant implications for drug design. In the past, several computational methods for miRNA subcellular localization is being used for uncovering multiple facets of RNA function to facilitate the biological applications. Unfortunately, most existing classification methods rely on a single sequencebased view, making the effective fusion of data from multiple heterogeneous networks a primary challenge. Inspired by multi-view multi-label learning strategy, we propose a computational method, named MMLmiRLocNet, for predicting the subcellular localizations of miRNAs. The MMLmiRLocNet predictor extracts multi-perspective sequence representations by analyzing lexical, syntactic, and semantic aspects of biological sequences. Specifically, it integrates lexical attributes derived from k-mer physicochemical profiles, syntactic characteristics obtained via word2vec embeddings, and semantic representations generated by pre-trained feature embeddings. Finally, module for extracting multi-view consensus-level features and specific-level features was constructed to capture consensus and specific features from various perspectives. The full connection networks are utilized as the output module to predict the miRNA subcellular localization. Experimental results suggest that MMLmiRLocNet outperforms existing methods in terms of F1, subACC, and Accuracy, and achieves best performance with the help of multi-view consensus features and specific features extract network.

AAAI Conference 2024 Conference Paper

MotionGPT: Finetuned LLMs Are General-Purpose Motion Generators

  • Yaqi Zhang
  • Di Huang
  • Bin Liu
  • Shixiang Tang
  • Yan Lu
  • Lu Chen
  • Lei Bai
  • Qi Chu

Generating realistic human motion from given action descriptions has experienced significant advancements because of the emerging requirement of digital humans. While recent works have achieved impressive results in generating motion directly from textual action descriptions, they often support only a single modality of the control signal, which limits their application in the real digital human industry. This paper presents a Motion General-Purpose generaTor (MotionGPT) that can use multimodal control signals, e.g., text and single-frame poses, for generating consecutive human motions by treating multimodal signals as special input tokens in large language models (LLMs). Specifically, we first quantize multimodal control signals into discrete codes and then formulate them in a unified prompt instruction to ask the LLMs to generate the motion answer. Our MotionGPT demonstrates a unified human motion generation model with multimodal control signals by tuning a mere 0.4% of LLM parameters. To the best of our knowledge, MotionGPT is the first method to generate human motion by multimodal control signals, which we hope can shed light on this new direction. Visit our webpage at https://qiqiapink.github.io/MotionGPT/.

EAAI Journal 2024 Journal Article

Multiscale information enhanced spatial-temporal graph convolutional network for multivariate traffic flow forecasting via magnifying perceptual scope

  • Xinyu zheng
  • Haidong Shao
  • Shen Yan
  • Yiming Xiao
  • Bin Liu

Graph Convolutional Networks (GCNs), which can model data in non-Euclidean space, have received extensive attention in multivariate traffic flow forecasting in recent years. However, most of the existing GCN forecasting methods rely on adjacency matrices defined by road network knowledge, which cannot accurately capture the inherent spatial-temporal dependencies of complex multivariate traffic flows. In this paper, a multiscale information enhanced spatial-temporal GCN is proposed, which utilizes an adaptive graph learning layer to automatically extract the dependencies between variables and constructs adaptive adjacency matrices for different datasets. In addition, an information enhancement module is introduced to magnify perceptual scope of the model and strengthen the connection between information, so that the spatial-temporal graph convolution module can better capture the spatial-temporal dependencies in the traffic flow. We conducted experiments on multivariate traffic flow data from three different real-world scenarios to evaluate the effectiveness of the proposed method. Comparative experimental results consistently show that the proposed method outperforms existing GCNs, achieving an average improvement of 21. 16% on Mean Absolute Error (MAE), 18. 87% on Root Mean Square Error (RMSE), and 8. 27% on R2 score (R2) in the three datasets, respectively, and the results demonstrate its excellent forecasting performance in the field of multivariate traffic flow forecasting.

IJCAI Conference 2024 Conference Paper

Predicting Housing Transaction with Common Covariance GNNs

  • Jinjin Li
  • Bin Liu
  • Chengyan Liu
  • Hongli Zhang

Urban migration is a significant aspect of a city's economy. The exploration of the underlying determinants of housing purchases among current residents contributes to the study of future trends in urban migration, enabling governments to formulate appropriate policies to guide future economic growth. This article employs a factor model to analyze data on residents' rentals, first-time home purchases, and subsequent housing upgrades. We decompose the factors influencing housing purchases into common drivers and specific drivers. Our hypothesis is that common drivers reflect universal social patterns, while personalized drivers represent stochastic elements. We construct a correlation matrix capturing the inter-resident relationships based on the common drivers of housing purchases. We then propose a graph neural network based on the correlation matrix to model housing predictions as a node classification problem. Our model addresses two critical questions. Firstly, we aim to identify which part of rental residents will engage in first-time home purchases in the future. Secondly, we seek to determine which group of residents, having completed rental and first-time home purchases, will opt for a second home purchase. The results of our testing on real-world datasets demonstrate that based solely on rental and home purchase records, we can achieve a sensitivity for housing predictions exceeding 80%.

AAAI Conference 2024 Conference Paper

TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target Detection

  • Tianxiang Chen
  • Zhentao Tan
  • Qi Chu
  • Yue Wu
  • Bin Liu
  • Nenghai Yu

Infrared small target detection (ISTD) is critical to national security and has been extensively applied in military areas. ISTD aims to segment small target pixels from background. Most ISTD networks focus on designing feature extraction blocks or feature fusion modules, but rarely describe the ISTD process from the feature map evolution perspective. In the ISTD process, the network attention gradually shifts towards target areas. We abstract this process as the directional movement of feature map pixels to target areas through convolution, pooling and interactions with surrounding pixels, which can be analogous to the movement of thermal particles constrained by surrounding variables and particles. In light of this analogy, we propose Thermal Conduction-Inspired Transformer (TCI-Former) based on the theoretical principles of thermal conduction. According to thermal conduction differential equation in heat dynamics, we derive the pixel movement differential equation (PMDE) in the image domain and further develop two modules: Thermal Conduction-Inspired Attention (TCIA) and Thermal Conduction Boundary Module (TCBM). TCIA incorporates finite difference method with PMDE to reach a numerical approximation so that target body features can be extracted. To further remove errors in boundary areas, TCBM is designed and supervised by boundary masks to refine target body features with fine boundary details. Experiments on IRSTD-1k and NUAA-SIRST demonstrate the superiority of our method.

AAAI Conference 2024 Conference Paper

Unifying Multi-Modal Uncertainty Modeling and Semantic Alignment for Text-to-Image Person Re-identification

  • Zhiwei Zhao
  • Bin Liu
  • Yan Lu
  • Qi Chu
  • Nenghai Yu

Text-to-Image person re-identification (TI-ReID) aims to retrieve the images of target identity according to the given textual description. The existing methods in TI-ReID focus on aligning the visual and textual modalities through contrastive feature alignment or reconstructive masked language modeling (MLM). However, these methods parameterize the image/text instances as deterministic embeddings and do not explicitly consider the inherent uncertainty in pedestrian images and their textual descriptions, leading to limited image-text relationship expression and semantic alignment. To address the above problem, in this paper, we propose a novel method that unifies multi-modal uncertainty modeling and semantic alignment for TI-ReID. Specifically, we model the image and textual feature vectors of pedestrian as Gaussian distributions, where the multi-granularity uncertainty of the distribution is estimated by incorporating batch-level and identity-level feature variances for each modality. The multi-modal uncertainty modeling acts as a feature augmentation and provides richer image-text semantic relationship. Then we present a bi-directional cross-modal circle loss to more effectively align the probabilistic features between image and text in a self-paced manner. To further promote more comprehensive image-text semantic alignment, we design a task that complements the masked language modeling, focusing on the cross-modality semantic recovery of global masked token after cross-modal interaction. Extensive experiments conducted on three TI-ReID datasets highlight the effectiveness and superiority of our method over state-of-the-arts.

EAAI Journal 2023 Journal Article

A novel dynamic distance coding identification method for oil–gas gathering and transportation process

  • Zijian Liu
  • Wende Tian
  • Bin Liu
  • Zhe Cui

Deep learning (DL) has become a mainstream method for fault identification in petrochemical processes. However, the high noise and nonlinear coupling of complex data samples have led to different degrees of low accuracy and robustness problems in the method. Meanwhile, the fault cause is difficult to capture due to the complex chemical process operation mechanism. To address this challenge, a novel dynamic distance coding method incorporating DL is proposed to identify anomalies in real time. First, the collected normal process data are smoothed by the Savitzky–Golay filter to build a normal sample set. Then, dynamic coding based on the distance metric is introduced to compute the distribution of normal and real-time samples for extracting the spatial domain features. By a sliding window, dynamic coded maps are generated and analyzed for fault causes. Finally, the time-domain information is extracted by long short-term memory (LSTM) to learn the deep features of the encoded graph for fault identification. The proposed method was applied to an oil–gas gathering and transportation process, which proves its feasibility and effectiveness. Compared with the conventional LSTM, the F1 score of the method is improved by 0. 193, reaching 0. 986. The obtained visualization information enables explaining the causes and supplements the fault database, providing a valuable reference for workers’ feedback operations.

NeurIPS Conference 2023 Conference Paper

A Reduction-based Framework for Sequential Decision Making with Delayed Feedback

  • Yunchang Yang
  • Han Zhong
  • Tianhao Wu
  • Bin Liu
  • Liwei Wang
  • Simon S. Du

We study stochastic delayed feedback in general single-agent and multi-agent sequential decision making, which includes bandits, single-agent Markov decision processes (MDPs), and Markov games (MGs). We propose a novel reduction-based framework, which turns any multi-batched algorithm for sequential decision making with instantaneous feedback into a sample-efficient algorithm that can handle stochastic delays in sequential decision making. By plugging different multi-batched algorithms into our framework, we provide several examples demonstrating that our framework not only matches or improves existing results for bandits, tabular MDPs, and tabular MGs, but also provides the first line of studies on delays in sequential decision making with function approximation. In summary, we provide a complete set of sharp results for single-agent and multi-agent sequential decision making with delayed feedback.

NeurIPS Conference 2023 Conference Paper

ALIM: Adjusting Label Importance Mechanism for Noisy Partial Label Learning

  • Mingyu Xu
  • Zheng Lian
  • Lei Feng
  • Bin Liu
  • Jianhua Tao

Noisy partial label learning (noisy PLL) is an important branch of weakly supervised learning. Unlike PLL where the ground-truth label must conceal in the candidate label set, noisy PLL relaxes this constraint and allows the ground-truth label may not be in the candidate label set. To address this challenging problem, most of the existing works attempt to detect noisy samples and estimate the ground-truth label for each noisy sample. However, detection errors are unavoidable. These errors can accumulate during training and continuously affect model optimization. To this end, we propose a novel framework for noisy PLL with theoretical interpretations, called ``Adjusting Label Importance Mechanism (ALIM)''. It aims to reduce the negative impact of detection errors by trading off the initial candidate set and model outputs. ALIM is a plug-in strategy that can be integrated with existing PLL approaches. Experimental results on multiple benchmark datasets demonstrate that our method can achieve state-of-the-art performance on noisy PLL. Our code is available at: https: //github. com/zeroQiaoba/ALIM.

EAAI Journal 2023 Journal Article

C-ECAFormer: A new lightweight fault diagnosis framework towards heavy noise and small samples

  • Jie Wang
  • Haidong Shao
  • Shen Yan
  • Bin Liu

In engineering practice, small-sample fault diagnosis of mechanical equipment towards heavy noise interference poses great challenges for the existing Transformer based intelligent models. To address these challenges, this paper proposes a new lightweight model called C-ECAFormer. Firstly, inverted residual block is used to establish signal correlations and inductive bias capability and to extract richer local feature information by varying the input channel dimensions. Secondly, ECAFormer module is designed to enhance the relationship awareness between different channels in the input signal features, thereby improving the model's attention to important channels. Finally, collaborative self-attention block is developed to facilitate spatial interaction between window local and grid global in vibration signals, reducing the number of parameters and computational complexity of the model. The results of two experiments demonstrate that the proposed approach accommodates advantages of lightweight and robustness in small-sample fault diagnosis tasks, compared to the existing mainstream Transformer and CNN fault diagnosis frameworks.

YNICL Journal 2023 Journal Article

Cerebellar gray matter alterations predict deep brain stimulation outcomes in Meige syndrome

  • Bin Liu
  • Zhiqi Mao
  • Zhiqiang Cui
  • Zhipei Ling
  • Xin Xu
  • Kunyu He
  • Mengchu Cui
  • Zhebin Feng

BACKGROUND: The physiopathologic mechanism of Meige syndrome (MS) has not been clarified, and neuroimaging studies centering on cerebellar changes in MS are scarce. Moreover, even though deep brain stimulation (DBS) of the subthalamic nucleus (STN) has been recognized as an effective surgical treatment for MS, there has been no reliable biomarker to predict its efficacy. OBJECTIVE: To characterize the volumetric alterations of gray matter (GM) in the cerebellum in MS and to identify GM measurements related to a good STN-DBS outcome. METHODS: We used voxel-based morphometry and lobule-based morphometry to compare the regional and lobular GM differences in the cerebellum between 47 MS patients and 52 normal human controls (HCs), as well as between 31 DBS responders and 10 DBS non-responders. Both volumetric analyses were achieved using the Spatially Unbiased Infratentorial Toolbox (SUIT). Further, we performed partial correlation analyses to probe the relationship between the cerebellar GM changes and clinical scores. Finally, we plotted the receiver operating characteristic (ROC) curve to select biomarkers for MS diagnosis and DBS outcomes prediction. RESULTS: Compared to HCs, MS patients had GM atrophy in lobule Crus I, lobule VI, lobule VIIb, lobule VIIIa, and lobule VIIIb. Compared to DBS responders, DBS non-responders had lower GM volume in the left lobule VIIIb. Moreover, partial correlation analyses revealed a positive relationship between the GM volume of the significant regions/lobules and the symptom improvement rate after DBS surgery. ROC analyses demonstrated that the GM volume of the significant cluster in the left lobule VIIIb could not only distinguish MS patients from HCs but also predict the outcomes of STN-DBS surgery with high accuracy. CONCLUSION: MS patients display bilateral GM shrinkage in the cerebellum relative to HCs. Regional GM volume of the left lobule VIIIb can be a reliable biomarker for MS diagnosis and DBS outcomes prediction.

IJCAI Conference 2023 Conference Paper

Fluid Dynamics-Inspired Network for Infrared Small Target Detection

  • Tianxiang Chen
  • Qi Chu
  • Bin Liu
  • Nenghai Yu

Most infrared small target detection (ISTD) networks focus on building effective neural blocks or feature fusion modules but none describes the ISTD process from the image evolution perspective. The directional evolution of image pixels influenced by convolution, pooling and surrounding pixels is analogous to the movement of fluid elements constrained by surrounding variables ang particles. Inspired by this, we explore a novel research routine by abstracting the movement of pixels in the ISTD process as the flow of fluid in fluid dynamics (FD). Specifically, a new Fluid Dynamics-Inspired Network (FDI-Net) is devised for ISTD. Based on Taylor Central Difference (TCD) method, the TCD feature extraction block is designed, where convolution and Transformer structures are combined for local and global information. The pixel motion equation during the ISTD process is derived from the Navier–Stokes (N-S) equation, constructing a N-S Refinement Module that refines extracted features with edge details. Thus, the TCD feature extraction block determines the primary movement direction of pixels during detection, while the N-S Refinement Module corrects some skewed directions of the pixel stream to supplement the edge details. Experiments on IRSTD-1k and SIRST demonstrate that our method achieves SOTA performance in terms of evaluation metrics.

JAIR Journal 2023 Journal Article

Graphmax for Text Generation

  • Bin Liu
  • Guosheng Yin

In text generation, a large language model (LM) makes a choice of each new word based only on the former selection of its context using the softmax function. Nevertheless, the link statistics information of concurrent words based on a scene-specific corpus is valuable in choosing the next word, which can help to ensure the topic of the generated text to be aligned with the current task. To fully explore the co-occurrence information, we propose a graphmax function for task-specific text generation. Using the graph-based regularization, graphmax enables the final word choice to be determined by both the global knowledge from the LM and the local knowledge from the scene-specific corpus. The traditional softmax function is regularized with a graph total variation (GTV) term, which incorporates the local knowledge into the LM and encourages the model to consider the statistical relationships between words in a scene-specific corpus. The proposed graphmax is versatile and can be readily plugged into any large pre-trained LM for text generation and machine translation. Through extensive experiments, we demonstrate that the new GTV-based regularization can improve performances in various natural language processing (NLP) tasks in comparison with existing methods. Moreover, through human experiments, we observe that participants can easily distinguish the text generated by graphmax or softmax.

IJCAI Conference 2023 Conference Paper

Interpret ESG Rating’s Impact on the Industrial Chain Using Graph Neural Networks

  • Bin Liu
  • Jiujun He
  • Ziyuan Li
  • Xiaoyang Huang
  • Xiang Zhang
  • Guosheng Yin

We conduct a quantitative analysis of the development of the industry chain from the environmental, social, and governance (ESG) perspective, which is an overall measure of sustainability. Factors that may impact the performance of the industrial chain have been studied in the literature, such as government regulation, monetary policy, etc. Our interest lies in how the sustainability change (i. e. , ESG shock) affects the performance of the industrial chain. To achieve this goal, we model the industrial chain with a graph neural network (GNN) and conduct node regression on two financial performance metrics, namely, the aggregated profitability ratios and operating margin. To quantify the effects of ESG, we propose to compute the interaction between ESG shocks and industrial chain features with a cross-attention module, and then filter the original node features in the graph regression. Experiments on two real datasets demonstrate that (i) there are significant effects of ESG shocks on the industrial chain, and (ii) model parameters including regression coefficients and the attention map can explain how ESG shocks affect the performance of the industrial chain.

IS Journal 2023 Journal Article

On the Persistence of Multilabel Learning, Its Recent Trends, and Its Open Issues

  • Nikolaos Mylonas
  • Ioannis Mollas
  • Bin Liu
  • Yannis Manolopoulos
  • Grigorios Tsoumakas

Multilabel data comprise instances associated with multiple binary target variables. The main learning task from such data is multilabel classification, where the goal is to output a bipartition of the target variables into relevant and irrelevant ones for a given instance. Other tasks involve ranking the target variables from the most to the least relevant one or even outputting a full joint distribution for every possible assignment of values to the binary targets.

TCS Journal 2023 Journal Article

Order based algorithms for the core maintenance problem on edge-weighted graphs

  • Feiteng Zhang
  • Bin Liu
  • Zhenming Liu
  • Qizhi Fang

The cohesive subgraph k-core is a maximal connected subgraph with the minimum degree δ ≥ k of a simple graph, where integer k ≥ 0. Define the core number of a vertex w as the maximum k such that w is contained in a k-core. The core decomposition problem which is calculating the core numbers of all vertices in static graphs, and the core maintenance problem which is updating the core numbers in dynamic graphs are our main concern. Although, core numbers can be updated by the core decomposition algorithms, only a small part of vertices' core numbers have changed after the change of a graph. Thus, it is necessary to update core numbers locally to reduce the cost. In this paper, we study the core maintenance problem on edge-weighted graphs by using the vertex sequence k-order which is ordered by the order that the core decomposition algorithm removes vertices. We design the core maintenance algorithms for inserting one edge at a time and the method of updating the k-order, which reduce the searching range and the time cost evidently. For the removing case, we use the existing subcore algorithm to do the core maintenance and modify it with the method of updating k-order we design. Finally, we do extensive experiments to evaluate the effectiveness and the efficiency of our algorithms, which shows that the order based algorithm has a better performance than the existing for the core maintenance on edge-weighted graphs.

EAAI Journal 2023 Journal Article

Soft-margin Ellipsoid generative adversarial networks

  • Zheng Jiang
  • Bin Liu
  • Weihua Huang

Generative adversarial networks (GANs) are of great significance for synthetizing realistic images. However, GANs are potentially unstable during the training process, posing some challenges for their development. By defining an integral probability metric (IPM) on the hypersphere, Sphere GAN enforces the discriminator to satisfy Lipschitz continuity and stabilizes the training process. Developed from Sphere GAN, a soft-margin Ellipsoid GAN is proposed for improving the quality of generated samples and the stability of training process. In the presented method, the geometric moment difference defined on the hypersphere is generalized to the hyperellipsoid. The hyperellipsoid is realized to relax the upper bound of the IPM by extending measurable functions space, thus the quality of generated samples can be improved. Furthermore, a nonlinear separating hyperellipsoid is designed to prevent the discriminator from gradient vanishing and exploding on the classification boundary. The proposed soft-margin Ellipsoid GAN is proved theoretically to have a global optimal solution, i. e. , the probability density of generated samples approaches to that of real samples infinitely when both the discriminator and the generator are optimal. The CIFAR10 and LSUN-bedrooms datasets are selected to evaluate the performance of the proposed methods. The quantitative results show that the proposed approach decreases the Fréchet inception distance (FID) on the two datasets by 8. 2% and 16. 0%, respectively. The qualitative results show that the proposed soft-margin mechanism improves the stability of the training process.

NeurIPS Conference 2023 Conference Paper

VRA: Variational Rectified Activation for Out-of-distribution Detection

  • Mingyu Xu
  • Zheng Lian
  • Bin Liu
  • Jianhua Tao

Out-of-distribution (OOD) detection is critical to building reliable machine learning systems in the open world. Researchers have proposed various strategies to reduce model overconfidence on OOD data. Among them, ReAct is a typical and effective technique to deal with model overconfidence, which truncates high activations to increase the gap between in-distribution and OOD. Despite its promising results, is this technique the best choice? To answer this question, we leverage the variational method to find the optimal operation and verify the necessity of suppressing abnormally low and high activations and amplifying intermediate activations in OOD detection, rather than focusing only on high activations like ReAct. This motivates us to propose a novel technique called ``Variational Rectified Activation (VRA)'', which simulates these suppression and amplification operations using piecewise functions. Experimental results on multiple benchmark datasets demonstrate that our method outperforms existing post-hoc strategies. Meanwhile, VRA is compatible with different scoring functions and network architectures. Our code is available at https: //github. com/zeroQiaoba/VRA.

JBHI Journal 2022 Journal Article

An Efficient Ciphertext-Policy Weighted Attribute-Based Encryption for the Internet of Health Things

  • Hang Li
  • Keping Yu
  • Bin Liu
  • Chaosheng Feng
  • Zhiguang Qin
  • Gautam Srivastava

The Internet of Health Things (IoHT) is a medical concept that describes uniquely identifiable devices connected to the Internet that can communicate with each other. As one of the most important components of smart health monitoring and improvement systems, the IoHT presents numerous challenges, among which cybersecurity is a priority. As a well-received security solution to achieve fine-grained access control, ciphertext-policy weighted attribute-based encryption (CP-WABE) has the potential to ensure data security in the IoHT. However, many issues remain, such as inflexibility, poor computational capability, and insufficient storage efficiency in attributes comparison. To address these issues, we propose a novel access policy expression method using 0-1 coding technology. Based on this method, a flexible and efficient CP-WABE is constructed for the IoHT. Our scheme supports not only weighted attributes but also any form of comparison of weighted attributes. Furthermore, we use offline/online encryption and outsourced decryption technology to ensure that the scheme can run on an inefficient IoT terminal. Both theoretical and experimental analyses show that our scheme is more efficient and feasible than other schemes. Moreover, security analysis indicates that our scheme achieves security against a chosen-plaintext attack.

TIST Journal 2021 Journal Article

Identifying Illicit Drug Dealers on Instagram with Large-scale Multimodal Data Fusion

  • Chuanbo Hu
  • Minglei Yin
  • Bin Liu
  • Xin Li
  • Yanfang Ye

Illicit drug trafficking via social media sites such as Instagram have become a severe problem, thus drawing a great deal of attention from law enforcement and public health agencies. How to identify illicit drug dealers from social media data has remained a technical challenge for the following reasons. On the one hand, the available data are limited because of privacy concerns with crawling social media sites; on the other hand, the diversity of drug dealing patterns makes it difficult to reliably distinguish drug dealers from common drug users. Unlike existing methods that focus on posting-based detection, we propose to tackle the problem of illicit drug dealer identification by constructing a large-scale multimodal dataset named Identifying Drug Dealers on Instagram (IDDIG). Nearly 4,000 user accounts, of which more than 1,400 are drug dealers, have been collected from Instagram with multiple data sources including post comments, post images, homepage bio, and homepage images. We then design a quadruple-based multimodal fusion method to combine the multiple data sources associated with each user account for drug dealer identification. Experimental results on the constructed IDDIG dataset demonstrate the effectiveness of the proposed method in identifying drug dealers (almost 95% accuracy). Moreover, we have developed a hashtag-based community detection technique for discovering evolving patterns, especially those related to geography and drug types.

AAAI Conference 2021 Conference Paper

Joint Color-irrelevant Consistency Learning and Identity-aware Modality Adaptation for Visible-infrared Cross Modality Person Re-identification

  • Zhiwei Zhao
  • Bin Liu
  • Qi Chu
  • Yan Lu
  • Nenghai Yu

Visible-infrared cross modality person re-identification (VI- ReID) is a core but challenging technology in the 24-hours intelligent surveillance system. How to eliminate the large modality gap lies in the heart of VI-ReID. Conventional methods mainly focus on directly aligning the heterogeneous modalities into the same space. However, due to the unbalanced color information between the visible and infrared images, the features of visible images tend to overfit the clothing color information, which would be harmful to the modality alignment. Besides, these methods mainly align the heterogeneous feature distributions in dataset-level while ignoring the valuable identity information, which may cause the feature misalignment of some identities and weaken the discrimination of features. To tackle above problems, we propose a novel approach for VI-ReID. It learns the colorirrelevant features through the color-irrelevant consistency learning (CICL) and aligns the identity-level feature distributions by the identity-aware modality adaptation (IAMA). The CICL and IAMA are integrated into a joint learning framework and can promote each other. Extensive experiments on two popular datasets SYSU-MM01 and RegDB demonstrate the superiority and effectiveness of our approach against the state-of-the-art methods.

JMLR Journal 2021 Journal Article

Simultaneous Change Point Inference and Structure Recovery for High Dimensional Gaussian Graphical Models

  • Bin Liu
  • Xinsheng Zhang
  • Yufeng Liu

In this article, we investigate the problem of simultaneous change point inference and structure recovery in the context of high dimensional Gaussian graphical models with possible abrupt changes. In particular, motivated by neighborhood selection, we incorporate a threshold variable and an unknown threshold parameter into a joint sparse regression model which combines p l1-regularized node-wise regression problems together. The change point estimator and the corresponding estimated coefficients of precision matrices are obtained together. Based on that, a classifier is introduced to distinguish whether a change point exists. To recover the graphical structure correctly, a data-driven thresholding procedure is proposed. In theory, under some sparsity conditions and regularity assumptions, our method can correctly choose a homogeneous or heterogeneous model with high accuracy. Furthermore, in the latter case with a change point, we establish estimation consistency of the change point estimator, by allowing the number of nodes being much larger than the sample size. Moreover, it is shown that, in terms of structure recovery of Gaussian graphical models, the proposed thresholding procedure achieves model selection consistency and controls the number of false positives. The validity of our proposed method is justified via extensive numerical studies. Finally, we apply our proposed method to the S&P 500 dataset to show its empirical usefulness. [abs] [ pdf ][ bib ] &copy JMLR 2021. ( edit, beta )

TCS Journal 2020 Journal Article

A random algorithm for profit maximization in online social networks

  • Tiantian Chen
  • Bin Liu
  • Wenjing Liu
  • Qizhi Fang
  • Jing Yuan
  • Weili Wu

Given a social network G and a positive integer k, the influence maximization problem seeks for k nodes in G that can influence the largest number of nodes. This problem has found important applications, and a large amount of works have been devoted to identifying the few most influential users. But most of existing works only focus on the diffusion of a single idea or product in social networks. However, in reality, one company may produce multiple kinds of products and one user may also have multiple adoptions. For multiple kinds of different products with different activation costs and profits, it is crucial for the company to distribute the limited budget among multiple products in order to achieve profit maximization. The Profit Maximization with Multiple Adoptions (PM2A) problem aims to seek for a seed set within the budget to maximize the overall profit. In this paper, a Randomized Modified Greedy (RMG) algorithm based on the Reverse Influence Sampling (RIS) technique is presented for the PM2A problem, which could achieve a ( 1 − 1 / e − ε ) -approximate solution with high probability and is also the best performance ratio of the PM2A problem. Comprehensive experiments on three real-world social networks are conducted, and the results demonstrate that our RMG algorithm outperforms the algorithm proposed in [16] and other heuristics in terms of profit maximization, and could better allocate the budget.

AAAI Conference 2020 Conference Paper

DASOT: A Unified Framework Integrating Data Association and Single Object Tracking for Online Multi-Object Tracking

  • Qi Chu
  • Wanli Ouyang
  • Bin Liu
  • Feng Zhu
  • Nenghai Yu

In this paper, we propose an online multi-object tracking (MOT) approach that integrates data association and single object tracking (SOT) with a unified convolutional network (ConvNet), named DASOTNet. The intuition behind integrating data association and SOT is that they can complement each other. Following Siamese network architecture, DASOT- Net consists of the shared feature ConvNet, the data association branch and the SOT branch. Data association is treated as a special re-identification task and solved by learning discriminative features for different targets in the data association branch. To handle the problem that the computational cost of SOT grows intolerably as the number of tracked objects increases, we propose an efficient two-stage tracking method in the SOT branch, which utilizes the merits of correlation features and can simultaneously track all the existing targets within one forward propagation. With feature sharing and the interaction between them, data association branch and the SOT branch learn to better complement each other. Using a multi-task objective, the whole network can be trained endto-end. Compared with state-of-the-art online MOT methods, our method is much faster while maintaining a comparable performance.

IJCAI Conference 2020 Conference Paper

GSM: Graph Similarity Model for Multi-Object Tracking

  • Qiankun Liu
  • Qi Chu
  • Bin Liu
  • Nenghai Yu

The popular tracking-by-detection paradigm for multi-object tracking (MOT) focuses on solving data association problem, of which a robust similarity model lies in the heart. Most previous works make effort to improve feature representation for individual object while leaving the relations among objects less explored, which may be problematic in some complex scenarios. In this paper, we focus on leveraging the relations among objects to improve robustness of the similarity model. To this end, we propose a novel graph representation that takes both the feature of individual object and the relations among objects into consideration. Besides, a graph matching module is specially designed for the proposed graph representation to alleviate the impact of unreliable relations. With the help of the graph representation and the graph matching module, the proposed graph similarity model, named GSM, is more robust to the occlusion and the targets sharing similar appearance. We conduct extensive experiments on challenging MOT benchmarks and the experimental results demonstrate the effectiveness of the proposed method.

NeurIPS Conference 2020 Conference Paper

Parametric Instance Classification for Unsupervised Visual Feature learning

  • Yue Cao
  • Zhenda Xie
  • Bin Liu
  • Yutong Lin
  • Zheng Zhang
  • Han Hu

This paper presents parametric instance classification (PIC) for unsupervised visual feature learning. Unlike the state-of-the-art approaches which do instance discrimination in a dual-branch non-parametric fashion, PIC directly performs a one-branch parametric instance classification, revealing a simple framework similar to supervised classification and without the need to address the information leakage issue. We show that the simple PIC framework can be as effective as the state-of-the-art approaches, i. e. SimCLR and MoCo v2, by adapting several common component settings used in the state-of-the-art approaches. We also propose two novel techniques to further improve effectiveness and practicality of PIC: 1) a sliding-window data scheduler, instead of the previous epoch-based data scheduler, which addresses the extremely infrequent instance visiting issue in PIC and improves the effectiveness; 2) a negative sampling and weight update correction approach to reduce the training time and GPU memory consumption, which also enables application of PIC to almost unlimited training images. We hope that the PIC framework can serve as a simple baseline to facilitate future study. The code and network configurations are available at \url{https: //github. com/bl0/PIC}.

TCS Journal 2020 Journal Article

Profit Maximization problem with Coupons in social networks

  • Bin Liu
  • Xiao Li
  • Huijuan Wang
  • Qizhi Fang
  • Junyu Dong
  • Weili Wu

Viral marketing has become one of the most effective marketing strategies. In the process of real commercialization, in order to let some seed individuals know the products, companies can provide free samples to them. However, for some companies, especially famous ones, they are more willing to offer coupons than give samples. In this paper, we consider the Profit Maximization problem with Coupons (PM-C) in our new diffusion model named the Independent Cascade Model with Coupons and Valuations (IC-CV). To solve this problem, we propose the PMCA algorithm which can return a ( 1 3 − ε ) -approximate solution with at least 1 − 2 n − l probability, and runs in O ( log ⁡ ( n p ) ⋅ m n 3 log ⁡ n ( l log ⁡ n + n log ⁡ 2 ) / ε 3 ) expected time. Furthermore, during the analysis we provide a method to estimate the non-monotone submodular function.

NeurIPS Conference 2019 Conference Paper

Dynamic Ensemble Modeling Approach to Nonstationary Neural Decoding in Brain-Computer Interfaces

  • Yu Qi
  • Bin Liu
  • Yueming Wang
  • Gang Pan

Brain-computer interfaces (BCIs) have enabled prosthetic device control by decoding motor movements from neural activities. Neural signals recorded from cortex exhibit nonstationary property due to abrupt noises and neuroplastic changes in brain activities during motor control. Current state-of-the-art neural signal decoders such as Kalman filter assume fixed relationship between neural activities and motor movements, thus will fail if this assumption is not satisfied. We propose a dynamic ensemble modeling (DyEnsemble) approach that is capable of adapting to changes in neural signals by employing a proper combination of decoding functions. The DyEnsemble method firstly learns a set of diverse candidate models. Then, it dynamically selects and combines these models online according to Bayesian updating mechanism. Our method can mitigate the effect of noises and cope with different task behaviors by automatic model switching, thus gives more accurate predictions. Experiments with neural data demonstrate that the DyEnsemble method outperforms Kalman filters remarkably, and its advantage is more obvious with noisy signals.

AAAI Conference 2018 Conference Paper

Early Prediction of Diabetes Complications from Electronic Health Records: A Multi-Task Survival Analysis Approach

  • Bin Liu
  • Ying Li
  • Zhaonan Sun
  • Soumya Ghosh
  • Kenney Ng

Type 2 diabetes mellitus (T2DM) is a chronic disease that usually results in multiple complications. Early identification of individuals at risk for complications after being diagnosed with T2DM is of significant clinical value. In this paper, we present a new data-driven predictive approach to predict when a patient will develop complications after the initial T2DM diagnosis. We propose a novel survival analysis method to model the time-to-event of T2DM complications designed to simultaneously achieve two important metrics: 1) accurate prediction of event times, and 2) good ranking of the relative risks of two patients. Moreover, to better capture the correlations of time-to-events of the multiple complications, we further develop a multi-task version of the survival model. To assess the performance of these approaches, we perform extensive experiments on patient level data extracted from a large electronic health record claims database. The results show that our new proposed survival analysis approach consistently outperforms traditional survival models and demonstrate the effectiveness of the multi-task framework over modeling each complication independently.

IJCAI Conference 2018 Conference Paper

Representing Urban Functions through Zone Embedding with Human Mobility Patterns

  • Zijun Yao
  • Yanjie Fu
  • Bin Liu
  • Wangsu Hu
  • Hui Xiong

Urban functions refer to the purposes of land use in cities where each zone plays a distinct role and cooperates with each other to serve people’s various life needs. Understanding zone functions helps to solve a variety of urban related problems, such as increasing traffic capacity and enhancing location-based service. Therefore, it is beneficial to investigate how to learn the representations of city zones in terms of urban functions, for better supporting urban analytic applications. To this end, in this paper, we propose a framework to learn the vector representation (embedding) of city zones by exploiting large-scale taxi trajectories. Specifically, we extract human mobility patterns from taxi trajectories, and use the co-occurrence of origin-destination zones to learn zone embeddings. To utilize the spatio-temporal characteristics of human mobility patterns, we incorporate mobility direction, departure/arrival time, destination attraction, and travel distance into the modeling of zone embeddings. We conduct extensive experiments with real-world urban datasets of New York City. Experimental results demonstrate the effectiveness of the proposed embedding model to represent urban functions of zones with human mobility data.

TIST Journal 2017 Journal Article

Personalized Air Travel Prediction

  • Jie Liu
  • Bin Liu
  • Yanchi Liu
  • Huipeng Chen
  • Lina Feng
  • Hui Xiong
  • Yalou Huang

Human mobility analysis is one of the most important research problems in the field of urban computing. Existing research mainly focuses on the intra-city ground travel behavior modeling, while the inter-city air travel behavior modeling has been largely ignored. Actually, the inter-city travel analysis can be of equivalent importance and complementary to the intra-city travel analysis. Understanding massive passenger-air-travel behavior delivers intelligence for airlines’ precision marketing and related socioeconomic activities, such as airport planning, emergency management, local transportation planning, and tourism-related businesses. Moreover, it provides opportunities to study the characteristics of cities and the mutual relationships between them. However, modeling and predicting air traveler behavior is challenging due to the complex factors of the market situation and individual characteristics of customers (e.g., airlines’ market share, customer membership, and travelers’ intrinsic interests on destinations). To this end, in this article, we present a systematic study on the personalized air travel prediction problem, namely where a customer will fly to and which airline carrier to fly with, by leveraging real-world anonymized Passenger Name Record (PNR) data. Specifically, we first propose a relational travel topic model, which combines the merits of latent factor model with a neighborhood-based method, to uncover the personal travel preferences of aviation customers and the latent travel topics of air routes and airline carriers simultaneously. Then we present a multi-factor travel prediction framework, which fuses complex factors of the market situation and individual characteristics of customers, to predict airline customers’ personalized travel demands. Experimental results on two real-world PNR datasets demonstrate the effectiveness of our approach on both travel topic discovery and customer travel prediction.

AIIM Journal 2017 Journal Article

Protein fold recognition based on sparse representation based classification

  • Ke Yan
  • Yong Xu
  • Xiaozhao Fang
  • Chunhou Zheng
  • Bin Liu

Knowledge of protein fold type is critical for determining the protein structure and function. Because of its importance, several computational methods for fold recognition have been proposed. Most of them are based on well-known machine learning techniques, such as Support Vector Machines (SVMs), Artificial Neural Network (ANN), etc. Although these machine learning methods play a role in stimulating the development of this important area, new techniques are still needed to further improve the predictive performance for fold recognition. Sparse Representation based Classification (SRC) has been widely used in image processing, and shows better performance than other related machine learning methods. In this study, we apply the SRC to solve the protein fold recognition problem. Experimental results on a widely used benchmark dataset show that the proposed method is able to improve the performance of some basic classifiers and three state-of-the-art methods to feature selection, including autocross-covariance (ACC) fold, D-D, and Bi-gram. Finally, we propose a novel computational predictor called MF-SRC for fold recognition by combining these three features into the framework of SRC to achieve further performance improvement. Compared with other computational methods in this field on DD dataset, EDD dataset and TG dataset, the proposed method achieves stable performance by reducing the influence of the noise in the dataset. It is anticipated that the proposed predictor may become a useful high throughput tool for large-scale fold recognition or at least, play a complementary role to the existing predictors in this regard.

AAAI Conference 2017 Conference Paper

Unsupervised Deep Learning for Optical Flow Estimation

  • Zhe Ren
  • Junchi Yan
  • Bingbing Ni
  • Bin Liu
  • Xiaokang Yang
  • Hongyuan Zha

Recent work has shown that optical flow estimation can be formulated as a supervised learning problem. Moreover, convolutional networks have been successfully applied to this task. However, supervised flow learning is obfuscated by the shortage of labeled training data. As a consequence, existing methods have to turn to large synthetic datasets for easily computer generated ground truth. In this work, we explore if a deep network for flow estimation can be trained without supervision. Using image warping by the estimated flow, we devise a simple yet effective unsupervised method for learning optical flow, by directly minimizing photometric consistency. We demonstrate that a flow network can be trained from endto-end using our unsupervised scheme. In some cases, our results come tantalizingly close to the performance of methods trained with full supervision.

TCS Journal 2014 Journal Article

Total coloring of embedded graphs with maximum degree at least seven

  • Huijuan Wang
  • Bin Liu
  • Jianliang Wu
  • Guizhen Liu

A k-total-coloring of a graph G is a coloring of V ( G ) ∪ E ( G ) using k colors such that no two adjacent or incident elements receive the same color. A graph G is k-total-colorable if it admits a k-total-coloring. In this paper, it is proved that any graph G which can be embedded in a surface Σ of Euler characteristic χ ( Σ ) ⩾ 0 is ( Δ ( G ) + 2 ) -total-colorable if Δ ( G ) ⩾ 7, where Δ ( G ) denotes the maximum degree of G.

TCS Journal 2009 Journal Article

Acyclic edge coloring of planar graphs with large girth

  • Dongxiao Yu
  • Jianfeng Hou
  • Guizhen Liu
  • Bin Liu
  • Lan Xu

Acyclic coloring problem is a specialized problem that arises in the efficient computation of Hessians. A proper edge coloring of a graph G is called acyclic if there is no 2 -colored cycle in G. The acyclic edge chromatic number χ a ′ ( G ) of G is the least number of colors in an acyclic edge coloring of G. Alon et al. conjectured that χ a ′ ( G ) ≤ Δ ( G ) + 2. In this paper, we consider the sufficient conditions for the planar graphs satisfying χ a ′ ( G ) ≤ Δ ( G ) + 1 and χ a ′ ( G ) = Δ ( G ).

v2026.09.13