Arrow Research search

Author name cluster

Liang Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

91 papers
2 author rows

Possible papers

91

AAAI Conference 2026 Conference Paper

Gait Transformer: End-to-End Transformer Backbone for Gait Recognition

  • Saihui Hou
  • Wenpeng Lang
  • Jilong Wang
  • Yan Huang
  • Liang Wang
  • Yongzhen Huang

Gait recognition has emerged as a promising biometric technique for long-distance and non-intrusive human identification. While Transformers have revolutionized vision tasks, their adaptation to gait recognition remains underexplored due to domain-specific challenges such as sparse silhouette modality, spatial-temporal dynamics, fine-grained motion cues, and limited training data. In this paper, we propose Gait Transformer (GaT), an end-to-end Transformer backbone specifically tailored for silhouette-based gait recognition. GaT introduces three key components: (1) a hybrid patch embedding module that combines convolutional stems with group-batch normalization to enhance structural preservation; (2) a decomposed token mixer that explicitly models both short-range and long-range dependencies across spatial-temporal dimensions; and (3) a hybrid positional encoding strategy that integrates absolute, relative, and rotary embeddings to support efficient training under data scarcity. Without relying on any pretraining, GaT achieves state-of-the-art performance on Gait3D, GREW, and CCGR-MINI.

NeurIPS Conference 2025 Conference Paper

3D-GSRD: 3D Molecular Graph Auto-Encoder with Selective Re-mask Decoding

  • Chang Wu
  • Zhiyuan Liu
  • Wen Shu
  • Liang Wang
  • Yanchen Luo
  • Wenqiang Lei
  • Yatao Bian
  • Junfeng Fang

Masked graph modeling (MGM) is a promising approach for molecular representation learning (MRL). However, extending the success of re-mask decoding from 2D to 3D MGM is non-trivial, primarily due to two conflicting challenges: avoiding 2D structure leakage to the decoder, while still providing sufficient 2D context for reconstructing re-masked atoms. To address these challenges, we propose 3D-GSRD: a 3D Molecular Graph Auto-Encoder with Selective Re-mask Decoding. The core innovation of 3D-GSRD lies in its Selective Re-mask Decoding (SRD), which re-masks only 3D-relevant information from encoder representations while preserving the 2D graph structures. This SRD is synergistically integrated with a 3D Relational-Transformer (3D-ReTrans) encoder alongside a structure-independent decoder. We analyze that SRD, combined with the structure-independent decoder, enhances the encoder's role in MRL. Extensive experiments show that 3D-GSRD achieves strong downstream performance, setting a new state-of-the-art on 7 out of 8 targets in the widely used MD17 molecular property prediction benchmark. The code is released at https: //github. com/WuChang0124/3D-GSRD.

EAAI Journal 2025 Journal Article

A data-driven method approach for prediction of coal seam gas content combining feature selection and machine learning

  • Sheng Su
  • Liang Wang
  • Songwei Wu
  • Yuechen Zhao
  • Chenghao Wang
  • Longyong Shu

Coal seam gas content (CSGC) is a key factor for mine safety and methane resource utilization. To address these limitations, an integrated feature selection and machine learning approach was developed for improved gas content prediction. In this study, a total of 27 sets of experimental data covering geologic structure, coal quality characteristics, and coal occurrence conditions were collected from the left flank area of Taoyuan Coal Mine in Huaibei, China. Missing values in the dataset were handled using the chain interpolation method. Subsequently, the Genetic Algorithm with Ant Colony Optimization (GA-ACO) was introduced and compared with several traditional feature selection algorithms, including All Subset Regression (ASR), Forward Selection Regression (FSR), Backward Selection Regression (BSR), Random Forest Algorithm (RFA), and Lasso regression to obtain the main controlling factors of gas content. The Particle Swarm Optimization Support Vector Regression (PSO-SVR) to build a regression prediction model. Various regression models such as Multivariate Linear Regression (MLR), Multi-Layer Perceptron (MLP), Gradient Boosting Regression Tree (GBRT) and SVR were then used to construct regression prediction models. The dataset was randomly divided into training and validation sets in an 8: 2 ratio, and prediction performance was evaluated using metrices. The results showed that the use of feature selection and machine learning effectively predicted the gas content of coal seams. The PSO-SVR model based on ASR achieved the best prediction performance with an R 2 of 0. 99 on the training set and 0. 978 on the test set, whit a maximum error of 0. 077. Finally, the constructed model was used to predict regional gas content and a regional gas content distribution map was plotted, providing a visual representation for analyzing gas content distribution in coal seams. This study demonstrates the potential of combining feature selection with machine learning for regression prediction in coal bed geology, offering a feasible and promising approach for research in this field.

EAAI Journal 2025 Journal Article

A novel adaptive spatial–temporal cross-graph convolutional fusion learning network for skeleton-based abnormal gait recognition

  • Liang Wang
  • Xiaoyan Wu
  • Bin Wu
  • Jianning Wu

Developing graph-based abnormal gait classification models with high generalization has been a challenging problem in gait analysis. In this study, a novel adaptive spatial–temporal cross-graph convolutional fusion learning network is proposed to accurately recognize skeleton-based abnormal gait patterns. In the proposed model, with an adaptive fusion adjacency matrix including self-adaptive adjacency matrices and cross-adaptive adjacency matrices, a joint–bone gait graph convolutional fusion learning algorithm is constructed to capture spatial gait abnormality features hidden in skeleton data. A temporal convolution network is then adopted to explore temporal dependencies of gait abnormality embedded in the spatial feature space. This could discover the most discriminative spatial–temporal gait abnormality representations containing richer information about interaction coupling across joints and bones for high-generalization. The skeleton data of mimic abnormal gait from 57 participants were collected to evaluate the feasibility of our model. The experimental results based on the leave-one-subject-out (LOSO) cross-validation scheme show that our proposed model reaches the optimal performance with the highest accuracy of 99. 43%, and significantly outcompetes several recent state-of-the-art models. Our model can feasibly take advantage of the adaptive fusion adjacency matrix to greatly enhance the aggregation degree of joints and bones. This helps to learn excellent gait abnormality representations containing richer interaction information for high generalization while keeping a low learning complexity. Our findings hopefully provide a powerful technical solution for abnormal gait recognition in practical clinical application.

EAAI Journal 2025 Journal Article

A real-time spatiotemporal error compensation framework for face gear grinding

  • Jialan Liu
  • Chi Ma
  • Mingming Li
  • Jialong He
  • Giovanni Totis
  • Chunlei Hua
  • Gangwei Cui
  • Liang Wang

Geometric and thermal errors critically affect the precision of face gear grinding, yet current modeling approaches are computationally intensive and lack real-time adaptability. This study proposes a real-time spatiotemporal error compensation framework for face gear grinding. A closed-loop feedback mechanism is introduced to adaptively update compensation intensity based on residual error feedback, ensuring robustness and efficiency under fluctuating machining conditions. Moreover, a novel spatial-temporal thermal error model is developed by integrating Taylor-graph convolutional network and modified-long short term memory network to capture both node-level spatial fusion and long-term temporal dependencies. High-order terms in geometric error modeling are eliminated using a vector decomposition and truncation-based approach, significantly reducing computational complexity. Furthermore, a high-efficiency multi-source error-tooth flank mapping model is developed based on vector decomposition and truncation function methods, enabling accurate prediction with reduced computational cost. To identify dominant error contributors, an improved Morris-based sensitivity analysis method is integrated, distinguishing geometric and thermal errors affecting tooth flank deviation. Experimental results demonstrate sub-65 ms real-time response, 24. 2 μm maximum error reduction, and robust adaptability under fluctuating machining conditions. Compared with recent gear-flank compensation studies, the proposed closed-loop framework achieves a 63. 4 % reduction in maximum normal flank error under real machining and <65 ms response latency. This level is comparable to reported reductions based on grid-aggregated metrics in spiral bevel gears (76. 82 % reduction of the sum of absolute grid errors), while additionally ensuring real-time, delay-aware execution. These findings validate the proposed system's potential for precision, real-time compensation in multi-axis manufacturing environments.

NeurIPS Conference 2025 Conference Paper

AuroRA: Breaking Low-Rank Bottleneck of LoRA with Nonlinear Mapping

  • Haonan Dong
  • Wenhao Zhu
  • Guojie Song
  • Liang Wang

Low-Rank Adaptation (LoRA) is a widely adopted parameter-efficient fine-tuning (PEFT) method validated across NLP and CV domains. However, LoRA faces an inherent low-rank bottleneck: narrowing its performance gap with full fine-tuning requires increasing the rank of its parameter matrix, resulting in significant parameter overhead. Recent linear LoRA variants have attempted to enhance expressiveness by introducing additional linear mappings; however, their composition remains inherently linear and fails to fundamentally improve LoRA’s representational capacity. To address this limitation, we propose \ourmethod, which incorporates an Adaptive Nonlinear Layer (ANL) between two linear projectors to capture \emph{fixed} and \emph{learnable} nonlinearities. This combination forms an {\fontfamily{lmtt}\selectfont \textbf{MLP-like structure}} with a compressed rank, enabling flexible and precise approximation of diverse target functions while theoretically guaranteeing lower approximation errors and bounded gradients. Extensive experiments on 22 datasets and 6 pretrained models demonstrate that \ourmethod: (\textbf{I}) not only matches or surpasses full fine-tuning performance with only $6. 18\%\sim25\%$ of LoRA’s parameters but also (\textbf{II}) outperforms state-of-the-art PEFT methods by up to $10. 88\%$ in both NLP and CV tasks, and \textbf{(III)} exhibits robust performance across various rank configurations.

TIST Journal 2025 Journal Article

Balancing Cooperation and Competition: Selfish Worker Coalition Formation in Spatial Crowdsourcing

  • Liang Wang
  • Shan Su
  • Rongchang Cheng
  • Dingqi Yang
  • Lianbo Ma
  • Fei Xiong
  • Bin Guo
  • Zhiwen Yu

Spatial Crowdsourcing (SC), which outsources location-dependent tasks to workers for physical completion, is gaining popularity. Recently, more complex tasks have emerged that require a group of workers collaborating in a coalition. Several pioneering studies have examined this issue using the server assigned tasks mode from an overall perspective, such as maximizing the total benefits of all workers. Unfortunately, maximizing the overall benefit does not necessarily align with maximizing individual benefits. In practice, crowd workers are often self-interested and autonomous, making decisions based on their personal perspectives. In this article, under the worker selected tasks mode, we investigate an important problem: Selfish Workers Coalition Formation (SWCF) problem in SC. Here, selfish workers autonomously form coalitions to accomplish tasks to maximize their individual benefits. Achieving a stable coalition formation for SWCF problem requires balancing cooperation and competition. First, we transform the SWCF problem into a hedonic coalition formation game using a devised exploited skills-based reward distribution model. Subsequently, we propose a distributed algorithm HCFTA and prove its Nash stability and performance bounds. Additionally, to enhance coalition formation efficiency, we propose a Markov blanket coloring parallel optimization algorithm MCPHCF. Extensive experiments demonstrate the superiority of the proposed methods on both synthetic and real-world datasets.

NeurIPS Conference 2025 Conference Paper

BridgeVLA: Input-Output Alignment for Efficient 3D Manipulation Learning with Vision-Language Models

  • Peiyan Li
  • Yixiang Chen
  • Hongtao Wu
  • Xiao Ma
  • Xiangnan Wu
  • Yan Huang
  • Liang Wang
  • Tao Kong

Recently, leveraging pre-trained vision-language models (VLMs) for building vision-language-action (VLA) models has emerged as a promising approach to effective robot manipulation learning. However, only few methods incorporate 3D signals into VLMs for action prediction, and they do not fully leverage the spatial structure inherent in 3D data, leading to low data efficiency. In this paper, we introduce a new paradigm for constructing 3D VLAs. Specifically, we first pre-train the VLM backbone to take 2D images as input and produce 2D heatmaps as output. Using this pre-trained VLM as the backbone, we then fine-tune the entire VLA model while maintaining alignment between inputs and outputs by: (1) projecting raw point cloud inputs into multi-view images, and (2) predicting heatmaps before generating the final action. Extensive experiments show that the resulting model, BridgeVLA, can learn 3D manipulation both efficiently and effectively. BridgeVLA outperforms state-of-the-art baselines across three simulation benchmarks. In RLBench, it improves the average success rate from 81. 4\% to 88. 2\%. In COLOSSEUM, it demonstrates significantly better performance in challenging generalization settings, boosting the average success rate from 56. 7\% to 64. 0\%. In GemBench, it surpasses all the comparing baseline methods in terms of average success rate. In real-robot experiments, BridgeVLA outperforms a state-of-the-art baseline method by 32\% on average. It generalizes robustly in multiple out-of-distribution settings, including visual disturbances and unseen instructions. Remarkably, it is able to achieve a success rate of 95. 4\% on 10+ tasks with only 3 trajectories per task, while other VLA methods such as $\pi_{0}$ fail completely. Project Website: https: //bridgevla. github. io/.

NeurIPS Conference 2025 Conference Paper

Chain-of-Retrieval Augmented Generation

  • Liang Wang
  • Haonan Chen
  • Nan Yang
  • Xiaolong Huang
  • Zhicheng Dou
  • Furu Wei

This paper introduces an approach for training o1-like RAG models that retrieve and reason over relevant information step by step before generating the final answer. Conventional RAG methods usually perform a single retrieval step before the generation process, which limits their effectiveness in addressing complex queries due to imperfect retrieval results. In contrast, our proposed method, CoRAG (Chain-of-Retrieval Augmented Generation), allows the model to dynamically reformulate the query based on the evolving state. To train CoRAG effectively, we utilize rejection sampling to automatically generate intermediate retrieval chains, thereby augmenting existing RAG datasets that only provide the correct final answer. At test time, we propose various decoding strategies to scale the model's test-time compute by controlling the length and number of sampled retrieval chains. Experimental results across multiple benchmarks validate the efficacy of CoRAG, particularly in multi-hop question answering tasks, where we observe more than $10$ points improvement in EM score compared to strong baselines. On the KILT benchmark, CoRAG establishes a new state-of-the-art performance across a diverse range of knowledge-intensive tasks. Furthermore, we offer comprehensive analyses to understand the scaling behavior of CoRAG, laying the groundwork for future research aimed at developing factual and grounded foundation models.

NeurIPS Conference 2025 Conference Paper

DAA: Amplifying Unknown Discrepancy for Test-Time Discovery

  • Tianle Liu
  • Fan Lyu
  • Chenggong Ni
  • Zhang Zhang
  • Fuyuan Hu
  • Liang Wang

Test-Time Discovery (TTD) addresses the critical challenge of identifying and adapting to novel classes during inference while maintaining performance on known classes, which is a capability essential for dynamic real-world environments such as healthcare and autonomous driving. Recent TTD methods adopt training-free, memory-based strategies but rely on frozen models and static representations, resulting in poor generalization. In this paper, we propose a Discrepancy-Amplifying Adapter (DAA), a trainable module that enables real-time adaptation by amplifying feature-level discrepancies between known and unknown classes. During training, DAA is optimized using simulated unknowns and a novel warm-up strategy to enhance its discriminative capacity. To ensure continual adaptation at test time, we introduce a Short-Term Memory Renewal (STMR) mechanism, which maintains a queue-based memory for unknown classes and selectively refreshes prototypes using recent, reliable samples. DAA is further updated through self-supervised learning, promoting knowledge retention for known classes while improving discrimination of emerging categories. Extensive experiments show that our method maintains high adaptability and stability, and significantly improves novel class discovery performance. Our code will be available.

AAAI Conference 2025 Conference Paper

Domain-Level Disentanglement Framework Based on Information Enhancement for Cross-Domain Cold-Start Recommendation

  • Nian Rong
  • Fei Xiong
  • Shirui Pan
  • Guixun Luo
  • Jia Wu
  • Liang Wang

Recommender systems in various applications often encounter the challenge of cold-start, which refers to how to provide recommendations for completely new users. Cross-domain recommendation offers a solution to address this cold-start issue by leveraging user interaction information from other domains and providing recommendations for users in the target domain. However, applying the classic two-tower model in cross-domain scenarios for pure cold-start users proves challenging, and most existing cross-domain cold-start recommendation models adopt an embedding-mapping framework that lacks end-to-end efficiency. The parallel training recommendation method lacks consideration of the domain-level intrinsic characteristics of cross-domain information. In this paper, we propose a generalized framework that Domain-level Disentanglement framework based on information enhancement for Cross-domain Cold-start Recommendation. On one hand, we achieve deep utilization of domain-level information through independent extraction of domain knowledge and fusion using heuristic strategies. On the other hand, our model is incorporated with an information enhancement network based on user attention and a user personalized adaptor. We introduce measures to assess user variability and immutability in cross-domain recommendation, aiming to eliminate inter-domain bias and highlight individual user preferences. Experimental results on widely used cross-domain recommendation datasets demonstrate that our proposed model outperforms state-of-the-art methods, validating its effectiveness.

NeurIPS Conference 2025 Conference Paper

EnzyControl: Adding Functional and Substrate-Specific Control for Enzyme Backbone Generation

  • Chao Song
  • Zhiyuan Liu
  • Han Huang
  • Liang Wang
  • Qiong Wang
  • Jian-Yu Shi
  • Hui Yu
  • Yihang Zhou

Designing enzyme backbones with substrate-specific functionality is a critical challenge in computational protein engineering. Current generative models excel in protein design but face limitations in binding data, substrate-specific control, and flexibility for de novo enzyme backbone generation. To address this, we introduce EnzyBind, a dataset with 11, 100 experimentally validated enzyme-substrate pairs specifically curated from PDBbind. Building on this, we propose EnzyControl, a method that enables functional and substrate-specific control in enzyme backbone generation. Our approach generates enzyme backbones conditioned on MSA-annotated catalytic sites and their corresponding substrates, which are automatically extracted from curated enzyme-substrate data. At the core of EnzyControl is EnzyAdapter, a lightweight, modular component integrated into a pretrained motif-scaffolding model, allowing it to become substrate-aware. A two-stage training paradigm further refines the model's ability to generate accurate and functional enzyme structures. Experiments show that our EnzyControl achieves the best performance across structural and functional metrics on EnzyBind and EnzyBench benchmarks, with particularly notable improvements of 13% in designability and 13% in catalytic efficiency compared to the baseline models. The code is released at https: //github. com/Vecteur-libre/EnzyControl.

NeurIPS Conference 2025 Conference Paper

GOOD: Training-Free Guided Diffusion Sampling for Out-of-Distribution Detection

  • Xin Gao
  • Jiyao Liu
  • Guanghao Li
  • Yueming LYU
  • Jianxiong Gao
  • Weichen Yu
  • Ningsheng Xu
  • Liang Wang

Recent advancements have explored text-to-image diffusion models for synthesizing out-of-distribution (OOD) samples, substantially enhancing the performance of OOD detection. However, existing approaches typically rely on perturbing text-conditioned embeddings, resulting in semantic instability and insufficient shift diversity, which limit generalization to realistic OOD. To address these challenges, we propose GOOD, a novel and flexible framework that directly guides diffusion sampling trajectories towards OOD regions using off-the-shelf in-distribution (ID) classifiers. GOOD incorporates dual-level guidance: (1) Image-level guidance based on the gradient of log partition to reduce input likelihood, drives samples toward low-density regions in pixel space. (2) Feature-level guidance, derived from k-NN distance in the classifier’s latent space, promotes sampling in feature-sparse regions. Hence, this dual-guidance design enables more controllable and diverse OOD sample generation. Additionally, we introduce a unified OOD score that adaptively combines image and feature discrepancies, enhancing detection robustness. We perform thorough quantitative and qualitative analyses to evaluate the effectiveness of GOOD, demonstrating that training with samples generated by GOOD can notably enhance OOD detection performance.

AAAI Conference 2025 Conference Paper

Learning Fine-Grained Alignment for Aerial Vision-Dialog Navigation

  • Yifei Su
  • Dong An
  • Kehan Chen
  • Weichen Yu
  • Baiyang Ning
  • Yonggen Ling
  • Yan Huang
  • Liang Wang

Aerial Vision-Dialog Navigation (AVDN) is a new task that requires drones to navigate to a target location based on human-robot dialog history. This paper focuses on the critical fine-grained cross-modal alignment problem in AVDN, requiring the drone to align language entities with visual landmarks in top-down views. To achieve this, we first construct a Fine-Grained AVDN (FG-AVDN) dataset via a semi-automatic annotation pipeline, providing diverse multimodal annotations at the entity-landmark level. Based on this, a novel Fine-grained Entity-Landmark Alignment (FELA) method is proposed to learn the cross-modal alignment explicitly. Concretely, FELA first boosts the drone's visual understanding with a precise semantic grid representation, which captures the environmental semantics and spatial structure simultaneously. Subsequently, to learn the entity-landmark alignment, we devise cross-modal auxiliary tasks from three perspectives, including grounding, captioning, and contrastive learning. Extensive experiments demonstrate that our explicit entity-landmark alignment learning is beneficial for AVDN. As a result, FELA achieves leading performance with 3.2% SR and 4.9% GP improvements over prior arts. Code and dataset will be publicly available.

NeurIPS Conference 2025 Conference Paper

MME-VideoOCR: Evaluating OCR-Based Capabilities of Multimodal LLMs in Video Scenarios

  • Yang Shi
  • Huanqian Wang
  • Xie Xie
  • Huanyao Zhang
  • Lijie Zhao
  • Yifan Zhang
  • Xinfeng Li
  • Chaoyou Fu

Multimodal Large Language Models (MLLMs) have achieved considerable accuracy in Optical Character Recognition (OCR) from static images. However, their efficacy in video OCR is significantly diminished due to factors such as motion blur, temporal variations, and visual effects inherent in video content. To provide clearer guidance for training practical MLLMs, we introduce MME-VideoOCR benchmark, which encompasses a comprehensive range of video OCR application scenarios. MME-VideoOCR features 10 task categories comprising 25 individual tasks and spans 44 diverse scenarios. These tasks extend beyond text recognition to incorporate deeper comprehension and reasoning of textual content within videos. The benchmark consists of 1, 464 videos with varying resolutions, aspect ratios, and durations, along with 2, 000 meticulously curated, manually annotated question-answer pairs. We evaluate 18 state-of-the-art MLLMs on MME-VideoOCR, revealing that even the best-performing model (Gemini-2. 5 Pro) achieves only an accuracy of 73. 7%. Fine-grained analysis indicates that while existing MLLMs demonstrate strong performance on tasks where relevant texts are contained within a single or few frames, they exhibit limited capability in effectively handling tasks that demand holistic video comprehension. These limitations are especially evident in scenarios that require spatio-temporal reasoning, cross-frame information integration, or resistance to language prior bias. Our findings also highlight the importance of high-resolution visual input and sufficient temporal coverage for reliable OCR in dynamic video scenarios.

NeurIPS Conference 2025 Conference Paper

Path-Enhanced Contrastive Learning for Recommendation

  • Haoran Sun
  • Fei Xiong
  • Yuanzhe Hu
  • Liang Wang

Collaborative filtering (CF) methods are now facing the challenge of data sparsity in recommender systems. In order to reduce the effect of data sparsity, researchers proposed contrastive learning methods to extract self-supervised signals from raw data. Contrastive learning methods address this problem by graph augmentation and maximizing the consistency of node representations between different augmented graphs. However, these methods tends to unintentionally distance the target node from its path nodes on the interaction path, thus limiting its effectiveness. In this regard, we propose a solution that uses paths as samples in the contrastive loss function. In order to obtain the path samples, we design a path sampling method. In addition to the contrast of the relationship between the target node and the nodes within the path (intra-path contrast), we also designed a method of contrasting the relationship between the paths (inter-path contrast) to better pull the target node and its path nodes closer to each other. We use Simplifying and Powering Graph Convolution Network (LightGCN) as the basis and combine with a new path-enhanced graph approach proposed for graph augmentation. It effectively improves the performance of recommendation models. Our proposed Path Enhanced Contrastive Loss (PECL) model replaces the common contrastive loss function with our novel loss function, showing significant performance improvement. Experiments on three real-world datasets demonstrate the effectiveness of our model.

NeurIPS Conference 2025 Conference Paper

Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

  • Junfei Wu
  • Jian Guan
  • Kaituo Feng
  • Qiang Liu
  • Shu Wu
  • Liang Wang
  • Wei Wu
  • Tieniu Tan

As textual reasoning with large language models (LLMs) has advanced significant, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods primarily approach multimodal reasoning in a straightforward, text-centric manner, where both reasoning and answer derivation are conducted purely through text, with the only difference being the presence of multimodal input. As a result, these methods often encounter fundamental limitations in spatial reasoning tasks that demand precise geometric understanding and continuous spatial tracking\textemdash capabilities that humans achieve through mental visualization and manipulation. To address the limitations, we propose drawing to reason in space, a novel paradigm that enables LVLMs to reason through elementary drawing operations in the visual space. By equipping models with basic drawing operations including annotating bounding boxes and drawing auxiliary lines, we empower them to express and analyze spatial relationships through direct visual manipulation, meanwhile avoiding the performance ceiling imposed by specialized perception tools in previous tool-integrated reasoning approaches. To cultivate this capability, we develop a three-stage training framework: cold-start training with synthetic data to establish basic drawing abilities, reflective rejection sampling to enhance self-reflection behaviors, and reinforcement learning to directly optimize for target rewards. Extensive experiments demonstrate that our model, named \textsc{Spark}, consistently outperforms existing methods across diverse spatial reasoning benchmarks involving maze navigation, static spatial reasoning, video-based reasoning and multi-view-based reasoning tasks, with an average improvement of 11. 5\%. Ablation studies reveal the critical role of each training stage, with reflective rejection sampling particularly enhancing the model's self-correction capabilities and reasoning potential.

AAAI Conference 2025 Conference Paper

Robust Graph Based Social Recommendation Through Contrastive Multi-View Learning

  • Fei Xiong
  • Tao Zhang
  • Shirui Pan
  • Guixun Luo
  • Liang Wang

Social recommendation leverages the social connections between users to mitigate the issue of data sparsity and enhance recommendation quality. Although existing related works show their effectiveness, there remain two critical questions: i) The patterns of preference interactions among users are varied and heterogeneous. Current models struggle to accurately capture preference shifts from user interactions in noisy social environments. ii) Existing methods handle the integration of auxiliary information coarsely, potentially introducing noise and leading to biases in user preferences. To address the limitations above, we introduce a novel framework named Robust Graph Based Social Recommendation through Contrastive Multi-view Learning (RGCML). This framework leverages denoised social relations and global intents as dual auxiliary information sources to provide comprehensive characterization of users. Firstly, RGCML employs the concept of opinion dynamics to simulate how user preferences evolve due to noisy social relations. Then, it utilizes a specifically designed information fusion module to extract critical contextual information from multiple semantic perspectives, thereby achieving efficient personalized information fusion. Finally, it adopts the designed global-local contrastive learning paradigm that untangles and discriminates user preferences from global intents, further addressing the noise problem and enhancing the quality of user representations. Extensive experiments conducted on three real-world datasets demonstrate the superior performance of RGCML compared to several state-of-the-art (SOTA) baselines.

JBHI Journal 2025 Journal Article

SliceMamba With Neural Architecture Search for Medical Image Segmentation

  • Chao Fan
  • Hongyuan Yu
  • Yan Huang
  • Liang Wang
  • Zhenghan Yang
  • Xibin Jia

Despite the progress made in Mamba-based medical image segmentation models, existing methods utilizing unidirectional or multi-directional feature scanning mechanisms struggle to effectively capture dependencies between neighboring positions, limiting the discriminant representation learning of local features. These local features are crucial for medical image segmentation as they provide critical structural information about lesions and organs. To address this limitation, we propose SliceMamba, a simple yet effective locally sensitive Mamba-based medical image segmentation model. SliceMamba features an efficient Bidirectional Slicing and Scanning (BSS) module, which performs bidirectional feature slicing and employs varied scanning mechanisms for sliced features with distinct shapes. This design keeps spatially adjacent features close in the scan sequence, preserving the local structure of the image and enhancing segmentation performance. Additionally, to fit the varying sizes and shapes of lesions and organs, we introduce an Adaptive Slicing Search method that automatically identifies the optimal feature slicing method based on the characteristics of the target data. Extensive experiments on two skin lesion datasets (ISIC2017 and ISIC2018), two polyp segmentation datasets (Kvasir and ClinicDB), one ultra-wide field retinal hemorrhage segmentation dataset (UWF-RHS), and one multi-organ segmentation dataset (Synapse) demonstrate the effectiveness of our method.

AAAI Conference 2025 Conference Paper

S²DN: Learning to Denoise Unconvincing Knowledge for Inductive Knowledge Graph Completion

  • Tengfei Ma
  • Yujie Chen
  • Liang Wang
  • Xuan Lin
  • Bosheng Song
  • Xiangxiang Zeng

Inductive Knowledge Graph Completion (KGC) aims to infer missing facts between newly emerged entities within knowledge graphs (KGs), posing a significant challenge. While recent studies have shown promising results in inferring such entities through knowledge subgraph reasoning, they suffer from (i) the semantic inconsistencies of similar relations, and (ii) noisy interactions inherent in KGs due to the presence of unconvincing knowledge for emerging entities. To address these challenges, we propose a Semantic Structure-aware Denoising Network (S2DN) for inductive KGC. Our goal is to learn adaptable general semantics and reliable structures to distill consistent semantic knowledge while preserving reliable interactions within KGs. Specifically, we introduce a semantic smoothing module over the enclosing subgraphs to retain the universal semantic knowledge of relations. We incorporate a structure refining module to filter out unreliable interactions and offer additional knowledge, retaining robust structure surrounding target links. Extensive experiments conducted on three benchmark KGs demonstrate that S2DN surpasses the performance of state-of-the-art models. These results demonstrate the effectiveness of S2DN in preserving semantic consistency and enhancing the robustness of filtering out unreliable interactions in contaminated KGs.

YNIMG Journal 2025 Journal Article

Theta oscillations between the ventromedial prefrontal cortex and amygdala support dynamic representations of threat and safety

  • Pingping Lu
  • Dong Chen
  • Wenran Xia
  • Si Chen
  • Zheng Tan
  • Wenjing Zhou
  • Liang Wang

The amygdala exhibits distinct different activity patterns to threat and safety stimuli. Animal studies have demonstrated that the fear (i.e., threat) and extinction (i.e., safety) memory are encoded by the amygdala and its interaction with the ventromedial prefrontal cortex (vmPFC). Recent studies in both animals and humans suggest that the inter-regional interaction between amygdala and vmPFC can be supported by theta oscillations during fear processing. However, the mechanism by which the human vmPFC-amygdala pathway dynamically supports neural representations of the same stimulus remains elusive, as it alternatively reflects threat and safety situations. To investigate this phenomenon, we conducted intracranial EEG recordings in drug-resistant epilepsy patients (n = 8) with implanted depth electrodes who performed a fear conditioning and extinction task. This task was designed with a fixed structure whereby specific CS+ stimulus could be either safe (never paired with US) or threatening (possibly paired with US) based on an implicit rule during fear acquisition. Our findings showed that the stimulus embodying potential threat information was accompanied by increased theta activities in amygdala during both fear acquisition and early extinction. Furthermore, the learning of safety information was associated with enhanced theta-related direction from the vmPFC to the amygdala. This study provided directly electrophysiological evidence supporting the dynamic oscillatory modulation of threat and safety representations in the human amygdala-vmPFC circuit, and suggests that amygdala safety processing depends on theta inputs from the vmPFC in both fear acquisition and extinction.

EAAI Journal 2024 Journal Article

A multi-domain adversarial transfer network for cross domain fault diagnosis under imbalanced data

  • Guofa Li
  • Shaoyang Liu
  • Jialong He
  • Liang Wang
  • Chenchen Wu
  • Chenhui Qian

In the intelligent fault diagnosis of rolling bearings, transfer learning methods extend the applicability of models to diverse working scenarios. However, real-world scenarios often suffer from data imbalance, which reduces diagnostic accuracy. To address this issue, this paper proposes a multi-domain adversarial transfer (MDAT) framework to enhance cross-domain fault diagnosis accuracy for rolling bearings under imbalanced data conditions. First, an enhanced information generation method is introduced to produce realistic and useable synthetic data to mitigate data imbalance. Subsequently, an adversarial multi-domain adaptation module is designed to learn invariant features across multiple domains. Finally, a domain reweighting method is proposed to improve domain alignment and enhance domain confusion. The effectiveness of the proposed method was validated through two case studies on rolling bearing fault diagnosis. The results demonstrated that MDAT achieved cross-domain diagnosis accuracies of 89. 2% and 99. 0% under imbalanced data conditions, confirming the effectiveness and superiority of the MDAT framework.

IJCAI Conference 2024 Conference Paper

AnchorGT: Efficient and Flexible Attention Architecture for Scalable Graph Transformers

  • Wenhao Zhu
  • Guojie Song
  • Liang Wang
  • Shaoguo Liu

Graph Transformers (GTs) have significantly advanced the field of graph representation learning by overcoming the limitations of message-passing graph neural networks (GNNs) and demonstrating promising performance and expressive power. However, the quadratic complexity of self-attention mechanism in GTs has limited their scalability, and previous approaches to address this issue often suffer from expressiveness degradation or lack of versatility. To address this issue, we propose AnchorGT, a novel attention architecture for GTs with global receptive field and almost linear complexity, which serves as a flexible building block to improve the scalability of a wide range of GT models. Inspired by anchor-based GNNs, we employ structurally important k-dominating node set as anchors and design an attention mechanism that focuses on the relationship between individual nodes and anchors, while retaining the global receptive field for all nodes. With its intuitive design, AnchorGT can easily replace the attention module in various GT models with different network architectures and structural encodings, resulting in reduced computational overhead without sacrificing performance. In addition, we theoretically prove that AnchorGT attention can be strictly more expressive than Weisfeiler-Lehman test, showing its superiority in representing graph structures. Our experiments on three state-of-the-art GT models demonstrate that their AnchorGT variants can achieve similar results while being faster and significantly more memory efficient.

NeurIPS Conference 2024 Conference Paper

Antigen-Specific Antibody Design via Direct Energy-based Preference Optimization

  • Xiangxin Zhou
  • Dongyu Xue
  • Ruizhe Chen
  • Zaixiang Zheng
  • Liang Wang
  • Quanquan Gu

Antibody design, a crucial task with significant implications across various disciplines such as therapeutics and biology, presents considerable challenges due to its intricate nature. In this paper, we tackle antigen-specific antibody sequence-structure co-design as an optimization problem towards specific preferences, considering both rationality and functionality. Leveraging a pre-trained conditional diffusion model that jointly models sequences and structures of antibodies with equivariant neural networks, we propose direct energy-based preference optimization to guide the generation of antibodies with both rational structures and considerable binding affinities to given antigens. Our method involves fine-tuning the pre-trained diffusion model using a residue-level decomposed energy preference. Additionally, we employ gradient surgery to address conflicts between various types of energy, such as attraction and repulsion. Experiments on RAbD benchmark show that our approach effectively optimizes the energy of generated antibodies and achieves state-of-the-art performance in designing high-quality antibodies with low total energy and high binding affinity simultaneously, demonstrating the superiority of our approach.

NeurIPS Conference 2024 Conference Paper

Beyond Efficiency: Molecular Data Pruning for Enhanced Generalization

  • Dingshuo Chen
  • Zhixun Li
  • Yuyan Ni
  • Guibin Zhang
  • Ding Wang
  • Qiang Liu
  • Shu Wu
  • Jeffrey X. Yu

With the emergence of various molecular tasks and massive datasets, how to perform efficient training has become an urgent yet under-explored issue in the area. Data pruning (DP), as an oft-stated approach to saving training burdens, filters out less influential samples to form a coreset for training. However, the increasing reliance on pretrained models for molecular tasks renders traditional in-domain DP methods incompatible. Therefore, we propose a Mol ecular data P runing framework for e nhanced G eneralization ( MolPeg ), which focuses on the source-free data pruning scenario, where data pruning is applied with pretrained models. By maintaining two models with different updating paces during training, we introduce a novel scoring function to measure the informativeness of samples based on the loss discrepancy. As a plug-and-play framework, MolPeg realizes the perception of both source and target domain and consistently outperforms existing DP methods across four downstream tasks. Remarkably, it can surpass the performance obtained from full-dataset training, even when pruning up to 60-70% of the data on HIV and PCBA dataset. Our work suggests that the discovery of effective data-pruning metrics could provide a viable path to both enhanced efficiency and superior generalization in transfer learning.

NeurIPS Conference 2024 Conference Paper

Everyday Object Meets Vision-and-Language Navigation Agent via Backdoor

  • Keji He
  • Kehan Chen
  • Jiawang Bai
  • Yan Huang
  • Qi Wu
  • Shu-Tao Xia
  • Liang Wang

Vision-and-Language Navigation (VLN) requires an agent to dynamically explore environments following natural language. The VLN agent, closely integrated into daily lives, poses a substantial threat to the security of privacy and property upon the occurrence of malicious behavior. However, this serious issue has long been overlooked. In this paper, we pioneer the exploration of an object-aware backdoored VLN, achieved by implanting object-aware backdoors during the training phase. Tailored to the unique VLN nature of cross-modality and continuous decision-making, we propose a novel backdoored VLN paradigm: IPR Backdoor. This enables the agent to act in abnormal behavior once encountering the object triggers during language-guided navigation in unseen environments, thereby executing an attack on the target scene. Experiments demonstrate the effectiveness of our method in both physical and digital spaces across different VLN agents, as well as its robustness to various visual and textual variations. Additionally, our method also well ensures navigation performance in normal scenarios with remarkable stealthiness.

IJCAI Conference 2024 Conference Paper

Graph Attention Network with High-Order Neighbor Information Propagation for Social Recommendation

  • Fei Xiong
  • Haoran Sun
  • Guixun Luo
  • Shirui Pan
  • Meikang Qiu
  • Liang Wang

In recommender systems, graph neural networks (GNN) can integrate interactions between users and items with their attributes, which makes GNN-based methods more powerful. However, directly stacking multiple layers in a graph neural network can easily lead to over-smoothing, hence recommendation systems based on graph neural networks typically underutilize higher-order neighborhoods in their learning. Although some heterogeneous graph random walk methods based on meta-paths can achieve higher-order aggregation, the focus is predominantly on the nodes at the ends of the paths. Moreover, these methods require manually defined meta-paths, which limits the model’s expressiveness and flexibility. Furthermore, path encoding in graph neural networks usually focuses only on the sequence leading to the target node. However, real-world interactions often do not follow this strict sequence, limiting the predictive performance of sequence-based network models. These problems prevent GNN-based methods from being fully effective. We propose a Graph Attention network with Information Propagation path aggregation for Social Recommendation (GAIPSRec). Firstly, we propose a universal heterogeneous graph sampling framework that does not require manually defining meta-paths for path sampling, thereby offering greater flexibility. Moreover, our method takes into account all nodes on the aggregation path and is capable of learning information from higher-order neighbors without succumbing to over-smoothing. Finally, our method utilizes a gate mechanism to fuse sequential and non-sequential dependence in encoding path instances, allowing a more holistic view of the data. Extensive experiments on real-world datasets show that our proposed GAIPSRec improves the performance significantly and outperforms state-of-the-art methods.

AAAI Conference 2024 Conference Paper

Heterogeneous Graph Reasoning for Fact Checking over Texts and Tables

  • Haisong Gong
  • Weizhi Xu
  • Shu Wu
  • Qiang Liu
  • Liang Wang

Fact checking aims to predict claim veracity by reasoning over multiple evidence pieces. It usually involves evidence retrieval and veracity reasoning. In this paper, we focus on the latter, reasoning over unstructured text and structured table information. Previous works have primarily relied on fine-tuning pretrained language models or training homogeneous-graph-based models. Despite their effectiveness, we argue that they fail to explore the rich semantic information underlying the evidence with different structures. To address this, we propose a novel word-level Heterogeneous-graph-based model for Fact Checking over unstructured and structured information, namely HeterFC. Our approach leverages a heterogeneous evidence graph, with words as nodes and thoughtfully designed edges representing different evidence properties. We perform information propagation via a relational graph neural network, facilitating interactions between claims and evidence. An attention-based method is utilized to integrate information, combined with a language model for generating predictions. We introduce a multitask loss function to account for potential inaccuracies in evidence retrieval. Comprehensive experiments on the large fact checking dataset FEVEROUS demonstrate the effectiveness of HeterFC. Code will be released at: https://github.com/Deno-V/HeterFC.

JBHI Journal 2024 Journal Article

Integrating Smart Computility for Subflow Orchestration in Remote Virtual Services

  • Liang Wang
  • Wei Su
  • Fei Song
  • IIsun You

The burgeoning domain of the metaverse has sparked significant interest from a diverse array of industries, including healthcare services. However, the metaverse and its associated applications present various challenges to existing networks. First, to meet the increasing demands of the metaverse, there is a need for enhanced bandwidth, reduced latency, and improved packet loss control. Furthermore, the transmission mechanism should exhibit flexibility to automatically adapt to the diverse hybrid needs of different healthcare services. In this article, a multipath transmission-based paradigm tailored for the metaverse-based healthcare services is developed. Significantly, we devise an orchestration framework to reconcile edge-side subflow management with diverse healthcare applications. Using machine learning techniques, the framework can produce near-optimal subflow adjustment strategies for client nodes and miscellaneous services. Comprehensive experiments are performed on applications with diverse requirements to validate the adaptability of the framework to the application needs. The experimental results demonstrate that the proposed method enables the network to autonomously adapt to changing network conditions and service requirements. This includes applications' preferences for high throughput, low delay, and high stability. Moreover, the test results show that the proposed approach can notably decrease the occurrences of network quality falling below the minimum requirement. Given its adaptability and impact on network quality, this work paves the way for future metaverse-based healthcare services.

AAAI Conference 2024 Conference Paper

Learning to Rank in Generative Retrieval

  • Yongqi Li
  • Nan Yang
  • Liang Wang
  • Furu Wei
  • Wenjie Li

Generative retrieval stands out as a promising new paradigm in text retrieval that aims to generate identifier strings of relevant passages as the retrieval target. This generative paradigm taps into powerful generative language models, distinct from traditional sparse or dense retrieval methods. However, only learning to generate is insufficient for generative retrieval. Generative retrieval learns to generate identifiers of relevant passages as an intermediate goal and then converts predicted identifiers into the final passage rank list. The disconnect between the learning objective of autoregressive models and the desired passage ranking target leads to a learning gap. To bridge this gap, we propose a learning-to-rank framework for generative retrieval, dubbed LTRGR. LTRGR enables generative retrieval to learn to rank passages directly, optimizing the autoregressive model toward the final passage ranking target via a rank loss. This framework only requires an additional learning-to-rank training phase to enhance current generative retrieval systems and does not add any burden to the inference stage. We conducted experiments on three public benchmarks, and the results demonstrate that LTRGR achieves state-of-the-art performance among generative retrieval methods. The code and checkpoints are released at https://github.com/liyongqi67/LTRGR.

AAAI Conference 2024 Conference Paper

Percentile Risk-Constrained Budget Pacing for Guaranteed Display Advertising in Online Optimization

  • Liang Dai
  • Kejie Lyu
  • Chengcheng Zhang
  • Guangming Zhao
  • Zhonglin Zu
  • Liang Wang
  • Bo Zheng

Guaranteed display (GD) advertising is a critical component of advertising since it provides publishers with stable revenue and enables advertisers to target specific audiences with guaranteed impressions. However, smooth pacing control for online ad delivery presents a challenge due to significant budget disparities, user arrival distribution drift, and dynamic change between supply and demand. This paper presents robust risk-constrained pacing (RCPacing) that utilizes Lagrangian dual multipliers to fine-tune probabilistic throttling through monotonic mapping functions within the percentile space of impression performance distribution. RCPacing combines distribution drift resilience and compatibility with guaranteed allocation mechanism, enabling us to provide near-optimal online services. We also show that RCPacing achieves O(sqrt(T)) dynamic regret where T is the length of the horizon. RCPacing's effectiveness is validated through offline evaluations and online A/B testing conducted on Taobao brand advertising platform.

NeurIPS Conference 2024 Conference Paper

Pin-Tuning: Parameter-Efficient In-Context Tuning for Few-Shot Molecular Property Prediction

  • Qiang Liu
  • Shaozhen Liu
  • Xin Sun
  • Shu Wu
  • Liang Wang

Molecular property prediction (MPP) is integral to drug discovery and material science, but often faces the challenge of data scarcity in real-world scenarios. Addressing this, few-shot molecular property prediction (FSMPP) has been developed. Unlike other few-shot tasks, FSMPP typically employs a pre-trained molecular encoder and a context-aware classifier, benefiting from molecular pre-training and molecular context information. Despite these advancements, existing methods struggle with the ineffective fine-tuning of pre-trained encoders. We attribute this issue to the imbalance between the abundance of tunable parameters and the scarcity of labeled molecules, and the lack of contextual perceptiveness in the encoders. To overcome this hurdle, we propose a parameter-efficient in-context tuning method, named Pin-Tuning. Specifically, we propose a lightweight adapter for pre-trained message passing layers (MP-Adapter) and Bayesian weight consolidation for pre-trained atom/bond embedding layers (Emb-BWC), to achieve parameter-efficient tuning while preventing over-fitting and catastrophic forgetting. Additionally, we enhance the MP-Adapters with contextual perceptiveness. This innovation allows for in-context tuning of the pre-trained encoder, thereby improving its adaptability for specific FSMPP tasks. When evaluated on public datasets, our method demonstrates superior tuning with fewer trainable parameters, improving few-shot predictive performance.

IROS Conference 2024 Conference Paper

Real-time terrain assessment and Bayesian-based path planning for off-road navigation

  • Tianwei Niu
  • Shuwei Yu
  • Liang Wang
  • Haoyu Yuan
  • Shoukun Wang
  • Junzheng Wang

In the context of unstructured and unknown environment, the autonomous navigation still faces many challenges, such as assessing rough terrain and deciding how to safely navigate complex terrain. In this work, we propose a robust and practical off-road navigation framework that has been successfully deployed on a vibroseis truck for land exploration. First, in degraded wild scenes, a tightly coupled lidar-GNSS-inertial fusion odometry and mapping framework is adopted to construct a local point cloud map around the vehicle in real-time and provide precise localization. Then, based on amplitude-frequency characteristic analysis and point cloud PCA, a multi-layer terrain assessment map containing terrain roughness, obstacles and slope information is obtained. Finally, combining Gaussian distribution based adaptive sampler and Bayesian sequentially updated proposal distribution, a local graph is efficiently built to obtain multiple path solutions under constrained conditions. Both simulations and field experiments show that the proposed navigation framework can decide how to travel on a flat road even in harsh terrain conditions, naturally suppressing frequent attitude angle changes and preventing vehicle accidents.

NeurIPS Conference 2024 Conference Paper

Reprogramming Pretrained Target-Specific Diffusion Models for Dual-Target Drug Design

  • Xiangxin Zhou
  • Jiaqi Guan
  • Yijia Zhang
  • Xingang Peng
  • Liang Wang
  • Jianzhu Ma

Dual-target therapeutic strategies have become a compelling approach and attracted significant attention due to various benefits, such as their potential in overcoming drug resistance in cancer therapy. Considering the tremendous success that deep generative models have achieved in structure-based drug design in recent years, we formulate dual-target drug design as a generative task and curate a novel dataset of potential target pairs based on synergistic drug combinations. We propose to design dual-target drugs with diffusion models that are trained on single-target protein-ligand complex pairs. Specifically, we align two pockets in 3D space with protein-ligand binding priors and build two complex graphs with shared ligand nodes for SE(3)-equivariant composed message passing, based on which we derive a composed drift in both 3D and categorical probability space in the generative process. Our algorithm can well transfer the knowledge gained in single-target pretraining to dual-target scenarios in a zero-shot manner. We also repurpose linker design methods as strong baselines for this task. Extensive experiments demonstrate the effectiveness of our method compared with various baselines.

AAAI Conference 2024 Conference Paper

Rethinking Graph Masked Autoencoders through Alignment and Uniformity

  • Xiang Tao
  • Qiang Liu
  • Shu Wu
  • Liang Wang

Self-supervised learning on graphs can be bifurcated into contrastive and generative methods. Contrastive methods, also known as graph contrastive learning (GCL), have dominated graph self-supervised learning in the past few years, but the recent advent of graph masked autoencoder (GraphMAE) rekindles the momentum behind generative methods. Despite the empirical success of GraphMAE, there is still a dearth of theoretical understanding regarding its efficacy. Moreover, while both generative and contrastive methods have been shown to be effective, their connections and differences have yet to be thoroughly investigated. Therefore, we theoretically build a bridge between GraphMAE and GCL, and prove that the node-level reconstruction objective in GraphMAE implicitly performs context-level GCL. Based on our theoretical analysis, we further identify the limitations of the GraphMAE from the perspectives of alignment and uniformity, which have been considered as two key properties of high-quality representations in GCL. We point out that GraphMAE's alignment performance is restricted by the masking strategy, and the uniformity is not strictly guaranteed. To remedy the aforementioned limitations, we propose an Alignment-Uniformity enhanced Graph Masked AutoEncoder, named AUG-MAE. Specifically, we propose an easy-to-hard adversarial masking strategy to provide hard-to-align samples, which improves the alignment performance. Meanwhile, we introduce an explicit uniformity regularizer to ensure the uniformity of the learned representations. Experimental results on benchmark datasets demonstrate the superiority of our model over existing state-of-the-art methods. The code is available at: https://github.com/AzureLeon1/AUG-MAE.

YNIMG Journal 2024 Journal Article

Temporal interference stimulation targets deep primate brain

  • Ruobing Liu
  • Guanyu Zhu
  • Zhengping Wu
  • Yifei Gan
  • Jianguo Zhang
  • Jiali Liu
  • Liang Wang

Temporal interference (TI) stimulation, a novel non-invasive stimulation strategy, has recently been shown to modulate neural activity in deep brain regions of living mice. Yet, it is uncertain if this method is applicable to larger brains and whether the electric field produced under traditional safety currents can penetrate deep regions as observed in mice. Despite recent model-based simulation studies offering positive evidence at both macro- and micro-scale levels, the absence of electrophysiological data from actual brains hinders comprehensive understanding and potential application of TI. This study aims to directly measure the spatiotemporal properties of the interfered electric field in the rhesus monkey brain and to validate the effects of TI on the human brain. Two monkeys were involved in the measurement, with implantation of several stereo-electroencephalography (SEEG) depth electrodes. TI stimulation was applied to anesthetized monkeys using two pairs of surface electrodes at differing stimulation parameters. Model-based simulations were also conducted and subsequently compared with actual recordings. Additionally, TI stimulation was administered to patients with motor disorders to validate its effects on motor symptoms. Through the integration of computational electric field simulation with empirical measurements, it was determined that the temporally interfering electric fields in the deep central regions are capable of attaining a magnitude sufficient to induce a subthreshold modulation effect on neural signals. Additionally, an improvement in movement disorders was observed as a result of TI stimulation. This study is the first to systematically measure the TI electric field in living non-human primates, offering empirical evidence that TI holds promise as a more focal and precise method for modulating neural activities in deep regions of a large brain. This advancement paves the way for future applications of TI in treating neuropsychiatric disorders.

AAAI Conference 2024 Conference Paper

Text-Guided Molecule Generation with Diffusion Language Model

  • Haisong Gong
  • Qiang Liu
  • Shu Wu
  • Liang Wang

Text-guided molecule generation is a task where molecules are generated to match specific textual descriptions. Recently, most existing SMILES-based molecule generation methods rely on an autoregressive architecture. In this work, we propose the Text-Guided Molecule Generation with Diffusion Language Model (TGM-DLM), a novel approach that leverages diffusion models to address the limitations of autoregressive methods. TGM-DLM updates token embeddings within the SMILES string collectively and iteratively, using a two-phase diffusion generation process. The first phase optimizes embeddings from random noise, guided by the text description, while the second phase corrects invalid SMILES strings to form valid molecular representations. We demonstrate that TGM-DLM outperforms MolT5-Base, an autoregressive model, without the need for additional data resources. Our findings underscore the remarkable effectiveness of TGM-DLM in generating coherent and precise molecules with specific properties, opening new avenues in drug discovery and related scientific domains. Code will be released at: https://github.com/Deno-V/tgm-dlm.

ICRA Conference 2024 Conference Paper

Trajectory-prediction-based Dynamic Tracking of a UGV to a Moving Target under Multi-disturbed Conditions

  • Jinge Si
  • Bin Li 0037
  • Yongkang Xu
  • Liang Wang
  • Chencheng Deng
  • Shoukun Wang
  • Junzheng Wang

Tracking dynamic targets poses a significant challenge for Unmanned Ground Vehicles (UGVs). Existing methods often lack research on multi-disturbed conditions. To address this issue, we propose a trajectory-prediction-based dynamic tracking scheme, which includes target localization, trajectory prediction, and UGV control. Firstly, an estimation algorithm based on the Extended Kalman Filter (EKF) is employed to mitigate noise and estimate the absolute states of the target accurately. To enhance robustness, we present an Adaptive Trajectory Prediction (ATP) algorithm based on prediction anchors. In this method, a quantization standard for trajectory disturbance is designed for adaptive control. Subsequently, we iteratively solve prediction anchor points based on two motion models to robustly predict the target trajectory even in the presence of unknown disturbances. Finally, the Linear Time-Varying Model Predictive Control (LTV-MPC) is utilized in the UGV controller for dynamic tracking. Experimental results demonstrate that the ATP exhibits superior prediction robustness and accuracy in perturbed environments compared to other prediction algorithms. In addition, the proposed scheme effectively achieves dynamic tracking of the Unmanned Aerial Vehicle (UAV) by the UGV under multi-disturbed conditions. Specifically, when the target moves at a speed of 1. 0 m/s, the UGV can maintain a tracking error within 0. 346 m.

NeurIPS Conference 2024 Conference Paper

VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark

  • Han Huang
  • Haitian Zhong
  • Tao Yu
  • Qiang Liu
  • Shu Wu
  • Liang Wang
  • Tieniu Tan

Recently, knowledge editing on large language models (LLMs) has received considerable attention. Compared to this, editing Large Vision-Language Models (LVLMs) faces extra challenges from diverse data modalities and complicated model components, and data for LVLMs editing are limited. The existing LVLM editing benchmark, which comprises three metrics (Reliability, Locality, and Generality), falls short in the quality of synthesized evaluation images and cannot assess whether models apply edited knowledge in relevant content. Therefore, we employ more reliable data collection methods to construct a new Large $\textbf{V}$ision-$\textbf{L}$anguage Model $\textbf{K}$nowledge $\textbf{E}$diting $\textbf{B}$enchmark, $\textbf{VLKEB}$, and extend the Portability metric for more comprehensive evaluation. Leveraging a multi-modal knowledge graph, our image data are bound with knowledge entities. This can be further used to extract entity-related knowledge, which constitutes the base of editing data. We conduct experiments of different editing methods on five LVLMs, and thoroughly analyze how do they impact the models. The results reveal strengths and deficiencies of these methods and hopefully provide insights for future research. The codes and dataset are available at: https: //github. com/VLKEB/VLKEB.

NeurIPS Conference 2023 Conference Paper

Combating Bilateral Edge Noise for Robust Link Prediction

  • Zhanke Zhou
  • Jiangchao Yao
  • Jiaxu Liu
  • Xiawei Guo
  • Quanming Yao
  • Li He
  • Liang Wang
  • Bo Zheng

Although link prediction on graphs has achieved great success with the development of graph neural networks (GNNs), the potential robustness under the edge noise is still less investigated. To close this gap, we first conduct an empirical study to disclose that the edge noise bilaterally perturbs both input topology and target label, yielding severe performance degradation and representation collapse. To address this dilemma, we propose an information-theory-guided principle, Robust Graph Information Bottleneck (RGIB), to extract reliable supervision signals and avoid representation collapse. Different from the basic information bottleneck, RGIB further decouples and balances the mutual dependence among graph topology, target labels, and representation, building new learning objectives for robust representation against the bilateral noise. Two instantiations, RGIB-SSL and RGIB-REP, are explored to leverage the merits of different methodologies, i. e. , self-supervised learning and data reparameterization, for implicit and explicit data denoising, respectively. Extensive experiments on six datasets and three GNNs with diverse noisy scenarios verify the effectiveness of our RGIB instantiations. The code is publicly available at: https: //github. com/tmlr-group/RGIB.

YNICL Journal 2023 Journal Article

Early detection of acute ischemic stroke using Contrast-enhanced electrical impedance tomography perfusion

  • Weirui Zhang
  • Yang Jiao
  • Tao Zhang
  • Xuechao Liu
  • Jianan Ye
  • Yuyan Zhang
  • Bin Yang
  • Meng Dai

A cerebral contrast-enhanced electrical impedance tomography perfusion method is developed for acute ischemic stroke during intravenous thrombolytic therapy. Several clinical contrast agents with stable impedance characteristics and high-conductivity contrast were screened experimentally as electrical impedance contrast agent candidates. The electrical impedance tomography perfusion method was tested on rabbits with focal cerebral infarction, and its capability for early detection was verified based on perfusion images. The experimental results showed that ioversol 350 performed significantly better as an electrical impedance contrast agent than other contrast agents (p < 0.01). Additionally, perfusion images of focal cerebral infarction in rabbits confirmed that the electrical impedance tomography perfusion method could accurately detect the location and area of different cerebral infarction lesions (p < 0.001). Therefore, the cerebral contrast-enhanced electrical impedance tomography perfusion method proposed herein combines traditional, dynamic continuous imaging with rapid detection and could be applied as an early, rapid-detection, auxiliary, bedside imaging method for patients after a suspected ischemic stroke in both prehospital and in-hospital settings.

EAAI Journal 2023 Journal Article

Ensemble deep random vector functional link for self-supervised direction-of-arrival estimation

  • Jiawen He
  • Xiaolei Li
  • Peishun Liu
  • Liang Wang
  • Hao Zhou
  • Jinyu Wang
  • Ruichun Tang

Direction-of-arrival (DOA) estimation is a key step in the passive target location. The primary issues with traditional DOA estimation methods are the huge computation and weak noise immunity in extreme noise environments. Random vector functional link (RVFL) and its variants (RVFL without direct links, RVFL v ) have demonstrated high learning efficiency and strong generalization ability in previous studies. However, due to the shallow network structure, they may not be effective for underwater acoustic array signals with complex features. Therefore, we propose a model-embedded self-supervised ensemble deep RVFL (ME-SedRVFL) network to estimate the DOA of underwater acoustic array signals. To prove the efficiency and generalization ability, ME-SedRVFL is compared with its variants (ME-SRVFL v ), as well as other well-known randomization-based networks. The results testify the noise immunity of ME-SedRVFL and ME-SRVFL v is 9. 62% and 9. 34% better than traditional signal model-based methods, 1. 68% and 1. 40% better than randomization-based parameter estimation methods (Signal-to-noise ratio is −20 dB, frequency is 200 Hz). The statistical box diagrams and statistical comparisons are performed to evaluate different methods, which indicate that the ME-SedRVFL obtains superior DOA estimation performance to ME-SRVFL v in most cases, due to direct input–output connections helping regularize the randomization. Hence, ME-SedRVFL is identified as the best-performing DOA estimation method through a comprehensive evaluation of real-world and simulated datasets.

EAAI Journal 2023 Journal Article

EORNet: An improved rotating box detection model for counting juvenile fish under occlusion and overlap

  • Pan Zhang
  • Liang Wang
  • Guangxu Wang
  • Daoliang Li

The juvenile fish cultivation stage is a very critical stage in the aquaculture process, and real-time monitoring and statistics of the number of juvenile fish is very important for the management of the aquaculture process. However, there are small targets, occlusion, and overlapping phenomena in cultivation stage, which seriously hinder the accurate detection and counting of juvenile fish. Conventional horizontal box detection has the phenomenon of feature reuse between targets in the face of severe occlusion or overlap. Therefore, this study tried to use the rotating box detection model to explore its ability to solve occlusion and overlap problems. First of all, this study constructed a data set of two kinds of juvenile fish in the incubation stage (including Brocarded Carp and Carp). Secondly, on the basis of Oriented RepPoints, the ECA attention mechanism is introduced to strengthen the feature map output at the end of the backbone feature extraction network, and an improved rotation box detection model EORNet is obtained. Finally, based on EORNet, the actual modeling effects (counting) of rotation box and horizontal box are compared to further explore the advantages of rotation box in mitigating occlusion and overlap problems. The results showed that the Recall and Precision of the rotation box detection model in the detection task reached 0. 978 and 0. 905 respectively, and the R2 of the counting model in the counting task reached 0. 937. The rotating frame detection method has been proved to be superior to the horizontal frame detection method in mitigating occlusion and overlap problems, and the rotating box detection is more conducive to capturing the key feature information of the target object, and then achieving accurate prediction. In addition to counting, the rotating box detection method is also expected to be used to accurately estimate the length, width and weight of aquaculture objects in the process of aquaculture.

NeurIPS Conference 2023 Conference Paper

Frequency-Enhanced Data Augmentation for Vision-and-Language Navigation

  • Keji He
  • Chenyang Si
  • Zhihe Lu
  • Yan Huang
  • Liang Wang
  • Xinchao Wang

Vision-and-Language Navigation (VLN) is a challenging task that requires an agent to navigate through complex environments based on natural language instructions. In contrast to conventional approaches, which primarily focus on the spatial domain exploration, we propose a paradigm shift toward the Fourier domain. This alternative perspective aims to enhance visual-textual matching, ultimately improving the agent's ability to understand and execute navigation tasks based on the given instructions. In this study, we first explore the significance of high-frequency information in VLN and provide evidence that it is instrumental in bolstering visual-textual matching processes. Building upon this insight, we further propose a sophisticated and versatile Frequency-enhanced Data Augmentation (FDA) technique to improve the VLN model's capability of capturing critical high-frequency information. Specifically, this approach requires the agent to navigate in environments where only a subset of high-frequency visual information corresponds with the provided textual instructions, ultimately fostering the agent's ability to selectively discern and capture pertinent high-frequency features according to the given instructions. Promising results on R2R, RxR, CVDN and REVERIE demonstrate that our FDA can be readily integrated with existing VLN approaches, improving performance without adding extra parameters, and keeping models simple and efficient. The code is available at https: //github. com/hekj/FDA.

NeurIPS Conference 2023 Conference Paper

GSLB: The Graph Structure Learning Benchmark

  • Zhixun Li
  • Xin Sun
  • Yifan Luo
  • Yanqiao Zhu
  • Dingshuo Chen
  • Yingtao Luo
  • Xiangxin Zhou
  • Qiang Liu

Graph Structure Learning (GSL) has recently garnered considerable attention due to its ability to optimize both the parameters of Graph Neural Networks (GNNs) and the computation graph structure simultaneously. Despite the proliferation of GSL methods developed in recent years, there is no standard experimental setting or fair comparison for performance evaluation, which creates a great obstacle to understanding the progress in this field. To fill this gap, we systematically analyze the performance of GSL in different scenarios and develop a comprehensive Graph Structure Learning Benchmark (GSLB) curated from 20 diverse graph datasets and 16 distinct GSL algorithms. Specifically, GSLB systematically investigates the characteristics of GSL in terms of three dimensions: effectiveness, robustness, and complexity. We comprehensively evaluate state-of-the-art GSL algorithms in node- and graph-level tasks, and analyze their performance in robust learning and model complexity. Further, to facilitate reproducible research, we have developed an easy-to-use library for training, evaluating, and visualizing different GSL methods. Empirical results of our extensive experiments demonstrate the ability of GSL and reveal its potential benefits on various downstream tasks, offering insights and opportunities for future research. The code of GSLB is available at: https: //github. com/GSL-Benchmark/GSLB.

IJCAI Conference 2023 Conference Paper

Hierarchical Transformer for Scalable Graph Learning

  • Wenhao Zhu
  • Tianyu Wen
  • Guojie Song
  • Xiaojun Ma
  • Liang Wang

Graph Transformer is gaining increasing attention in the field of machine learning and has demonstrated state-of-the-art performance on benchmarks for graph representation learning. However, as current implementations of Graph Transformer primarily focus on learning representations of small-scale graphs, the quadratic complexity of the global self-attention mechanism presents a challenge for full-batch training when applied to larger graphs. Additionally, conventional sampling-based methods fail to capture necessary high-level contextual information, resulting in a significant loss of performance. In this paper, we introduce the Hierarchical Scalable Graph Transformer (HSGT) as a solution to these challenges. HSGT successfully scales the Transformer architecture to node representation learning tasks on large-scale graphs, while maintaining high performance. By utilizing graph hierarchies constructed through coarsening techniques, HSGT efficiently updates and stores multi-scale information in node embeddings at different levels. Together with sampling-based training methods, HSGT effectively captures and aggregates multi-level information on the hierarchical graph using only Transformer blocks. Empirical evaluations demonstrate that HSGT achieves state-of-the-art performance on large-scale benchmarks with graphs containing millions of nodes with high efficiency.

IJCAI Conference 2023 Conference Paper

KDLGT: A Linear Graph Transformer Framework via Kernel Decomposition Approach

  • Yi Wu
  • Yanyang Xu
  • Wenhao Zhu
  • Guojie Song
  • Zhouchen Lin
  • Liang Wang
  • Shaoguo Liu

In recent years, graph Transformers (GTs) have been demonstrated as a robust architecture for a wide range of graph learning tasks. However, the quadratic complexity of GTs limits their scalability on large-scale data, in comparison to Graph Neural Networks (GNNs). In this work, we propose the Kernel Decomposition Linear Graph Transformer (KDLGT), an accelerating framework for building scalable and powerful GTs. KDLGT employs the kernel decomposition approach to rearrange the order of matrix multiplication, thereby reducing complexity to linear. Additionally, it categorizes GTs into three distinct types and provides tailored accelerating methods for each category to encompass all types of GTs. Furthermore, we provide a theoretical analysis of the performance gap between KDLGT and self-attention to ensure its effectiveness. Under this framework, we select two representative GTs to design our models. Experiments on both real-world and synthetic datasets indicate that KDLGT not only achieves state-of-the-art performance on various datasets but also reaches an acceleration ratio of approximately 10 on graphs of certain sizes.

JBHI Journal 2023 Journal Article

Multi-Level Adversarial Spatio-Temporal Learning for Footstep Pressure Based FoG Detection

  • Kun Hu
  • Shaohui Mei
  • Wei Wang
  • Kaylena A. Ehgoetz Martens
  • Liang Wang
  • Simon J. G. Lewis
  • David D. Feng
  • Zhiyong Wang

Freezing of gait (FoG) is one of the most common symptoms of Parkinson's disease, which is a neurodegenerative disorder of the central nervous system impacting millions of people around the world. To address the pressing need to improve the quality of treatment for FoG, devising a computer-aided detection and quantification tool for FoG has been increasingly important. As a non-invasive technique for collecting motion patterns, the footstep pressure sequences obtained from pressure sensitive gait mats provide a great opportunity for evaluating FoG in the clinic and potentially in the home environment. In this study, FoG detection is formulated as a sequential modelling task and a novel deep learning architecture, namely Adversarial Spatio-temporal Network (ASTN), is proposed to learn FoG patterns across multiple levels. ASTN introduces a novel adversarial training scheme with a multi-level subject discriminator to obtain subject-independent FoG representations, which helps to reduce the over-fitting risk due to the high inter-subject variance. As a result, robust FoG detection can be achieved for unseen subjects. The proposed scheme also sheds light on improving subject-level clinical studies from other scenarios as it can be integrated with many existing deep architectures. To the best of our knowledge, this is one of the first studies of footstep pressure-based FoG detection and the approach of utilizing ASTN is the first deep neural network architecture in pursuit of subject-independent representations. In our experiments on 393 trials collected from 21 subjects, the proposed ASTN achieved an AUC 0. 85, clearly outperforming conventional learning methods.

NeurIPS Conference 2023 Conference Paper

OneNet: Enhancing Time Series Forecasting Models under Concept Drift by Online Ensembling

  • Yifan Zhang
  • Qingsong Wen
  • Xue Wang
  • Weiqi Chen
  • Liang Sun
  • Zhang Zhang
  • Liang Wang
  • Rong Jin

Online updating of time series forecasting models aims to address the concept drifting problem by efficiently updating forecasting models based on streaming data. Many algorithms are designed for online time series forecasting, with some exploiting cross-variable dependency while others assume independence among variables. Given every data assumption has its own pros and cons in online time series modeling, we propose **On**line **e**nsembling **Net**work (**OneNet**). It dynamically updates and combines two models, with one focusing on modeling the dependency across the time dimension and the other on cross-variate dependency. Our method incorporates a reinforcement learning-based approach into the traditional online convex programming framework, allowing for the linear combination of the two models with dynamically adjusted weights. OneNet addresses the main shortcoming of classical online learning methods that tend to be slow in adapting to the concept drift. Empirical results show that OneNet reduces online forecasting error by more than $\mathbf{50}\\%$ compared to the State-Of-The-Art (SOTA) method.

IROS Conference 2023 Conference Paper

Relative Roughness Measurement Based Real-Time Speed Planning for Autonomous Vehicles on Rugged Road

  • Liang Wang
  • Tianwei Niu
  • Shoukun Wang
  • Shuai Wang
  • Junzheng Wang

In order to guarantee autonomous vehicles' autonomy, mobility, and ride quality in rugged environments, a real-time speed planning method based on the time-frequency transformation of terrain characteristics is designed to achieve adaptive speed planning of autonomous vehicles in rough ground. On the one hand, the vertical profile of the lidar's point cloud data is converted from the time domain to the frequency domain in real time, and the integrated area of the sub-frequency range in the frequency domain is chosen as the relative roughness quantification value to realize the roughness quantification under various terrains. On the other hand, to model the relationship between vehicle speed and relative roughness, iterative search is utilized to create a speed and roughness model, and sliding windows are employed to update the roughness to achieve continuous mapping between speed and roughness. Ultimately, a number of tests were conducted on various rough roads using the oil exploration vehicle EV-56 as the study object. The experimental results show that the proposed method can identify the terrain roughness changes under complex terrain and change their speed within 0. 2 m accuracy.

NeurIPS Conference 2023 Conference Paper

Uncovering Neural Scaling Laws in Molecular Representation Learning

  • Dingshuo Chen
  • Yanqiao Zhu
  • Jieyu Zhang
  • Yuanqi Du
  • Zhixun Li
  • Qiang Liu
  • Shu Wu
  • Liang Wang

Molecular Representation Learning (MRL) has emerged as a powerful tool for drug and materials discovery in a variety of tasks such as virtual screening and inverse design. While there has been a surge of interest in advancing model-centric techniques, the influence of both data quantity and quality on molecular representations is not yet clearly understood within this field. In this paper, we delve into the neural scaling behaviors of MRL from a data-centric viewpoint, examining four key dimensions: (1) data modalities, (2) dataset splitting, (3) the role of pre-training, and (4) model capacity. Our empirical studies confirm a consistent power-law relationship between data volume and MRL performance across these dimensions. Additionally, through detailed analysis, we identify potential avenues for improving learning efficiency. To challenge these scaling laws, we adapt seven popular data pruning strategies to molecular data and benchmark their performance. Our findings underline the importance of data-centric MRL and highlight possible directions for future research.

YNIMG Journal 2022 Journal Article

Advances in human intracranial electroencephalography research, guidelines and good practices

  • Manuel R. Mercier
  • Anne-Sophie Dubarry
  • François Tadel
  • Pietro Avanzini
  • Nikolai Axmacher
  • Dillan Cellier
  • Maria Del Vecchio
  • Liberty S. Hamilton

Since the second-half of the twentieth century, intracranial electroencephalography (iEEG), including both electrocorticography (ECoG) and stereo-electroencephalography (sEEG), has provided an intimate view into the human brain. At the interface between fundamental research and the clinic, iEEG provides both high temporal resolution and high spatial specificity but comes with constraints, such as the individual's tailored sparsity of electrode sampling. Over the years, researchers in neuroscience developed their practices to make the most of the iEEG approach. Here we offer a critical review of iEEG research practices in a didactic framework for newcomers, as well addressing issues encountered by proficient researchers. The scope is threefold: (i) review common practices in iEEG research, (ii) suggest potential guidelines for working with iEEG data and answer frequently asked questions based on the most widespread practices, and (iii) based on current neurophysiological knowledge and methodologies, pave the way to good practice standards in iEEG research. The organization of this paper follows the steps of iEEG data processing. The first section contextualizes iEEG data collection. The second section focuses on localization of intracranial electrodes. The third section highlights the main pre-processing steps. The fourth section presents iEEG signal analysis methods. The fifth section discusses statistical approaches. The sixth section draws some unique perspectives on iEEG research. Finally, to ensure a consistent nomenclature throughout the manuscript and to align with other guidelines, e.g., Brain Imaging Data Structure (BIDS) and the OHBM Committee on Best Practices in Data Analysis and Sharing (COBIDAS), we provide a glossary to disambiguate terms related to iEEG research.

TIST Journal 2022 Journal Article

Data-driven Targeted Advertising Recommendation System for Outdoor Billboard

  • Liang Wang
  • Zhiwen Yu
  • Bin Guo
  • Dingqi Yang
  • Lianbo Ma
  • Zhidan Liu
  • Fei Xiong

In this article, we propose and study a novel data-driven framework for Targeted Outdoor Advertising Recommendation (TOAR) with a special consideration of user profiles and advertisement topics. Given an advertisement query and a set of outdoor billboards with different spatial locations and rental prices, our goal is to find a subset of billboards, such that the total targeted influence is maximum under a limited budget constraint. To achieve this goal, we are facing two challenges: (1) it is difficult to estimate targeted advertising influence in physical world; (2) due to NP hardness, many common search techniques fail to provide a satisfied solution with an acceptable time, especially for large-scale problem settings. Taking into account the exposure strength, advertisement matching degree, and advertising repetition effect, we first build a targeted influence model that can characterize that the advertising influence spreads along with users mobility. Subsequently, based on a divide-and-conquer strategy, we develop two effective approaches, i.e., a master–slave-based sequential optimization method, TOAR-MSS, and a cooperative co-evolution-based optimization method, TOAR-CC, to solve our studied problem. Extensive experiments on two real-world datasets clearly validate the effectiveness and efficiency of our proposed approaches.

AAAI Conference 2022 Conference Paper

Generalizable Person Re-identification via Self-Supervised Batch Norm Test-Time Adaption

  • Ke Han
  • Chenyang Si
  • Yan Huang
  • Liang Wang
  • Tieniu Tan

In this paper, we investigate the generalization problem of person re-identification (re-id), whose major challenge is the distribution shift on an unseen domain. As an important tool of regularizing the distribution, batch normalization (BN) has been widely used in existing methods. However, they neglect that BN is severely biased to the training domain and inevitably suffers the performance drop if directly generalized without being updated. To tackle this issue, we propose Batch Norm Test-time Adaption (BNTA), a novel re-id framework that applies the self-supervised strategy to update BN parameters adaptively. Specifically, BNTA quickly explores the domain-aware information within unlabeled target data before inference, and accordingly modulates the feature distribution normalized by BN to adapt to the target domain. This is accomplished by two designed self-supervised auxiliary tasks, namely part positioning and part nearest neighbor matching, which help the model mine the domain-aware information with respect to the structure and identity of body parts, respectively. To demonstrate the effectiveness of our method, we conduct extensive experiments on three re-id datasets and confirm the superior performance to the stateof-the-art methods.

IJCAI Conference 2022 Conference Paper

GraphDIVE: Graph Classification by Mixture of Diverse Experts

  • Fenyu Hu
  • Liping Wang
  • Qiang Liu
  • Shu Wu
  • Liang Wang
  • Tieniu Tan

Graph classification is a challenging research task in many applications across a broad range of domains. Recently, Graph Neural Network (GNN) models have achieved superior performance on various real-world graph datasets. Despite their successes, most of current GNN models largely suffer from the ubiquitous class imbalance problem, which typically results in prediction bias towards majority classes. Although many imbalanced learning methods have been proposed, they mainly focus on regular Euclidean data and cannot well utilize topological structure of graph (non-Euclidean) data. To boost the performance of GNNs and investigate the relationship between topological structure and class imbalance, we propose GraphDIVE, which learns multi-view graph representations and combine multi-view experts (i. e. , classifiers). Specifically, multi-view graph representations correspond to the intrinsic diverse graph topological structure characteristics. Extensive experiments on molecular benchmark datasets demonstrate the effectiveness of the proposed approach.

NeurIPS Conference 2022 Conference Paper

Honor of Kings Arena: an Environment for Generalization in Competitive Reinforcement Learning

  • Hua Wei
  • Jingxiao Chen
  • Xiyang Ji
  • Hongyang Qin
  • Minwen Deng
  • Siqin Li
  • Liang Wang
  • Weinan Zhang

This paper introduces Honor of Kings Arena, a reinforcement learning (RL) environment based on the Honor of Kings, one of the world’s most popular games at present. Compared to other environments studied in most previous work, ours presents new generalization challenges for competitive reinforcement learning. It is a multi-agent problem with one agent competing against its opponent; and it requires the generalization ability as it has diverse targets to control and diverse opponents to compete with. We describe the observation, action, and reward specifications for the Honor of Kings domain and provide an open-source Python-based interface for communicating with the game engine. We provide twenty target heroes with a variety of tasks in Honor of Kings Arena and present initial baseline results for RL-based methods with feasible computing resources. Finally, we showcase the generalization challenges imposed by Honor of Kings Arena and possible remedies to the challenges. All of the software, including the environment-class, are publicly available.

NeurIPS Conference 2022 Conference Paper

MACK: Multimodal Aligned Conceptual Knowledge for Unpaired Image-text Matching

  • Yan Huang
  • Yuming Wang
  • Yunan Zeng
  • Liang Wang

Recently, the accuracy of image-text matching has been greatly improved by multimodal pretrained models, all of which are trained on millions or billions of paired images and texts. Different from them, this paper studies a new scenario as unpaired image-text matching, in which paired images and texts are assumed to be unavailable during model training. To deal with this, we propose a simple yet effective method namely Multimodal Aligned Conceptual Knowledge (MACK), which is inspired by the knowledge use in human brain. It can be directly used as general knowledge to correlate images and texts even without model training, or further fine-tuned based on unpaired images and texts to better generalize to certain datasets. In addition, we extend it as a re-ranking method, which can be easily combined with existing image-text matching models to substantially improve their performance.

IJCAI Conference 2022 Conference Paper

Regularized Graph Structure Learning with Semantic Knowledge for Multi-variates Time-Series Forecasting

  • Hongyuan Yu
  • Ting Li
  • Weichen Yu
  • Jianguo Li
  • Yan Huang
  • Liang Wang
  • Alex Liu

Multivariate time-series forecasting is a critical task for many applications, and graph time-series network is widely studied due to its capability to capture the spatial-temporal correlation simultaneously. However, most existing works focus more on learning with the explicit prior graph structure, while ignoring potential information from the implicit graph structure, yielding incomplete structure modeling. Some recent works attempts to learn the intrinsic or implicit graph structure directly, while lacking a way to combine explicit prior structure with implicit structure together. In this paper, we propose Regularized Graph Structure Learning (RGSL) model to incorporate both explicit prior structure and implicit structure together, and learn the forecasting deep networks along with the graph structure. RGSL consists of two innovative modules. First, we derive an implicit dense similarity matrix through node embedding, and learn the sparse graph structure using the Regularized Graph Generation (RGG) based on the Gumbel Softmax trick. Second, we propose a Laplacian Matrix Mixed-up Module (LM3) to fuse the explicit graph and implicit graph together. We conduct experiments on three real-word datasets. Results show that the proposed RGSL model outperforms existing graph forecasting algorithms with a notable margin, while learning meaningful graph structure simultaneously. Our code and models are made publicly available at https: //github. com/alipay/RGSL. git.

AAAI Conference 2021 Conference Paper

A Graph-based Relevance Matching Model for Ad-hoc Retrieval

  • Yufeng Zhang
  • Jinghao Zhang
  • Zeyu Cui
  • Shu Wu
  • Liang Wang

To retrieve more relevant, appropriate and useful documents given a query, finding clues about that query through the text is crucial. Recent deep learning models regard the task as a term-level matching problem, which seeks exact or similar query patterns in the document. However, we argue that they are inherently based on local interactions and do not generalise to ubiquitous, non-consecutive contextual relationships. In this work, we propose a novel relevance matching model based on graph neural networks to leverage the documentlevel word relationships for ad-hoc retrieval. In addition to the local interactions, we explicitly incorporate all contexts of a term through the graph-of-word text format. Matching patterns can be revealed accordingly to provide a more accurate relevance score. Our approach significantly outperforms strong baselines on two ad-hoc benchmarks. We also experimentally compare our model with BERT and show our advantages on long documents.

TIST Journal 2021 Journal Article

Disentangled Item Representation for Recommender Systems

  • Zeyu Cui
  • Feng Yu
  • Shu Wu
  • Qiang Liu
  • Liang Wang

Item representations in recommendation systems are expected to reveal the properties of items. Collaborative recommender methods usually represent an item as one single latent vector. Nowadays the e-commercial platforms provide various kinds of attribute information for items (e.g., category, price, and style of clothing). Utilizing this attribute information for better item representations is popular in recent years. Some studies use the given attribute information as side information, which is concatenated with the item latent vector to augment representations. However, the mixed item representations fail to fully exploit the rich attribute information or provide explanation in recommender systems. To this end, we propose a fine-grained Disentangled Item Representation (DIR) for recommender systems in this article, where the items are represented as several separated attribute vectors instead of a single latent vector. In this way, the items are represented at the attribute level, which can provide fine-grained information of items in recommendation. We introduce a learning strategy, LearnDIR, which can allocate the corresponding attribute vectors to items. We show how DIR can be applied to two typical models, Matrix Factorization (MF) and Recurrent Neural Network (RNN). Experimental results on two real-world datasets show that the models developed under the framework of DIR are effective and efficient. Even using fewer parameters, the proposed model can outperform the state-of-the-art methods, especially in the cold-start situation. In addition, we make visualizations to show that our proposition can provide explanation for users in real-world applications.

ICRA Conference 2021 Conference Paper

ENCODE: a dEep poiNt Cloud ODometry nEtwork

  • Yihuan Zhang
  • Liang Wang
  • Chen Fu
  • Yifan Dai
  • John M. Dolan

Ego-motion estimation is a key requirement for the simultaneous localization and mapping (SLAM) problem. The traditional pipeline goes through feature extraction, feature matching and pose estimation, whose performance depends on the manually designed features. In this paper, we are motivated by the strong performance of deep learning methods in other computer vision and robotics tasks. We replace hand-crafted features with a neural network and directly estimate the relative pose between two adjacent scans from a LiDAR sensor using ENCODE: a dEep poiNt Cloud ODometry nEtwork. Firstly, a spherical projection of the input point cloud is performed to acquire a multi-channel vertex map. Then a multi-layer network backbone is applied to learn the abstracted features and a fully connected layer is adopted to estimate the 6-DoF ego-motion. Additionally, a map-to-map optimization module is applied to update the local poses and output a smooth map. Experiments on multiple datasets demonstrate that the proposed method achieves the best performance in comparison to state-of-the-art methods and is capable of providing accurate poses with low drift in various kinds of scenarios.

IJCAI Conference 2021 Conference Paper

Few-Shot Learning with Part Discovery and Augmentation from Unlabeled Images

  • Wentao Chen
  • Chenyang Si
  • Wei Wang
  • Liang Wang
  • Zilei Wang
  • Tieniu Tan

Few-shot learning is a challenging task since only few instances are given for recognizing an unseen class. One way to alleviate this problem is to acquire a strong inductive bias via meta-learning on similar tasks. In this paper, we show that such inductive bias can be learned from a flat collection of unlabeled images, and instantiated as transferable representations among seen and unseen classes. Specifically, we propose a novel part-based self-supervised representation learning scheme to learn transferable representations by maximizing the similarity of an image to its discriminative part. To mitigate the overfitting in few-shot classification caused by data scarcity, we further propose a part augmentation strategy by retrieving extra images from a base dataset. We conduct systematic studies on miniImageNet and tieredImageNet benchmarks. Remarkably, our method yields impressive results, outperforming the previous best unsupervised methods by 7. 74% and 9. 24% under 5-way 1-shot and 5-way 5-shot settings, which are comparable with state-of-the-art supervised methods.

YNICL Journal 2021 Journal Article

Lack of association between acute stroke, post-stroke dementia, race, and β-amyloid status

  • Lauren N. Koenig
  • Lena M. McCue
  • Elizabeth Grant
  • Parinaz Massoumzadeh
  • Catherine M. Roe
  • Chengjie Xiong
  • Krista L. Moulder
  • Liang Wang

INTRODUCTION: Stroke and Alzheimer disease share risk factors and often co-occur, and both have been reported to have a higher prevalence in African Americans as compared to non-Hispanic whites. However, their interaction has not been established. The objective of this study was to determine if preclinical Alzheimer disease is a risk factor for stroke and post-stroke dementia and whether racial differences moderate this relationship. METHODS: This case-control study was analyzed in 2019 using retrospective data from 2007 to 2013. Participants were adults age 65 and older with and without acute ischemic stroke. Recruitment included word of mouth and referrals in Saint Louis, MO, with stroke participants recruited from acutely hospitalized patients and non-stroke participants from community living older adults who were research volunteers. Our assessment included radiologic reads of infarcts, microbleeds, and white matter hyperintensitites (WMH); a Pittsburgh Compound B PET measure of cortical β-amyloid binding; quantitative measures of hippocampal and WMH volume; longitudinal Mini Mental State Examination (MMSE) scores; and Clinical Dementia Rating (CDR) 1 year post-stroke. RESULTS: A total of 243 participants were enrolled, 81 of which had a recent ischemic stroke. Participants had a mean age of 75, 57% were women, and 52% were African American. Cortical amyloid did not differ significantly by race, stroke status, or CDR post-stroke. There were racial differences in MMSE scores at baseline (mean 26.8 for African Americans, 27.9 for non-Hispanic whites, p = 0.03), but not longitudinally. African Americans were more likely to have microbleeds (32.8% vs 22.6%, p = 0.04), and within the acute stroke group, African Americans were more likely to have small infarcts (75.6% vs 56.8%, p = 0.049). CONCLUSION: Preclinical Alzheimer disease did not show evidence of being a risk factor for stroke nor predictive of post-stroke dementia. We did not observe racial differences in β-amyloid levels. However, even after controlling for several vascular risk factors, African Americans with clinical stroke presentations had greater levels of vascular pathology on MRI.

NeurIPS Conference 2021 Conference Paper

Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision

  • Keji He
  • Yan Huang
  • Qi Wu
  • Jianhua Yang
  • Dong An
  • Shuanglin Sima
  • Liang Wang

In Vision-and-Language Navigation (VLN) task, an agent is asked to navigate inside 3D indoor environments following given instructions. Cross-modal alignment is one of the most critical challenges in VLN because the predicted trajectory needs to match the given instruction accurately. In this paper, we address the cross-modal alignment challenge from the perspective of fine-grain. Firstly, to alleviate weak cross-modal alignment supervision from coarse-grained data, we introduce a human-annotated fine-grained VLN dataset, namely Landmark-RxR. Secondly, to further enhance local cross-modal alignment under fine-grained supervision, we investigate the focal-oriented rewards with soft and hard forms, by focusing on the critical points sampled from fine-grained Landmark-RxR. Moreover, to fully evaluate the navigation process, we also propose a re-initialization mechanism that makes metrics insensitive to difficult points, which can cause the agent to deviate from the correct trajectories. Experimental results show that our agent has superior navigation performance on Landmark-RxR, en-RxR and R2R. Our dataset and code are available at https: //github. com/hekj/Landmark-RxR.

NeurIPS Conference 2021 Conference Paper

Learning Diverse Policies in MOBA Games via Macro-Goals

  • Yiming Gao
  • Bei Shi
  • Xueying Du
  • Liang Wang
  • Guangwei Chen
  • Zhenjie Lian
  • Fuhao Qiu
  • GUOAN HAN

Recently, many researchers have made successful progress in building the AI systems for MOBA-game-playing with deep reinforcement learning, such as on Dota 2 and Honor of Kings. Even though these AI systems have achieved or even exceeded human-level performance, they still suffer from the lack of policy diversity. In this paper, we propose a novel Macro-Goals Guided framework, called MGG, to learn diverse policies in MOBA games. MGG abstracts strategies as macro-goals from human demonstrations and trains a Meta-Controller to predict these macro-goals. To enhance policy diversity, MGG samples macro-goals from the Meta-Controller prediction and guides the training process towards these goals. Experimental results on the typical MOBA game Honor of Kings demonstrate that MGG can execute diverse policies in different matches and lineups, and also outperform the state-of-the-art methods over 102 heroes.

NeurIPS Conference 2021 Conference Paper

The Tufts fNIRS Mental Workload Dataset & Benchmark for Brain-Computer Interfaces that Generalize

  • Zhe Huang
  • Liang Wang
  • Giles Blaney
  • Christopher Slaughter
  • Devon McKeon
  • Ziyu Zhou
  • Robert Jacob
  • Michael Hughes

Functional near-infrared spectroscopy (fNIRS) promises a non-intrusive way to measure real-time brain activity and build responsive brain-computer interfaces. A primary barrier to realizing this technology's potential has been that observed fNIRS signals vary significantly across human users. Building models that generalize well to never-before-seen users has been difficult; a large amount of subject-specific data has been needed to train effective models. To help overcome this barrier, we introduce the largest open-access dataset of its kind, containing multivariate fNIRS recordings from 68 participants, each with labeled segments indicating four possible mental workload intensity levels. Labels were collected via a controlled setting in which subjects performed standard n-back tasks to induce desired working memory levels. We propose a benchmark analysis of this dataset with a standardized training and evaluation protocol, which allows future researchers to report comparable numbers and fairly assess generalization potential while avoiding any overlap or leakage between train and test data. Using this dataset and benchmark, we show how models trained using abundant fNIRS data from many other participants can effectively classify a new target subject's data, thus reducing calibration and setup time for new subjects. We further show how performance improves as the size of the available dataset grows, while also analyzing error rates across key subpopulations to audit equity concerns. We share our open-access Tufts fNIRS to Mental Workload (fNIRS2MW) dataset and open-source code as a step toward advancing brain computer interfaces.

AAAI Conference 2020 Conference Paper

Mastering Complex Control in MOBA Games with Deep Reinforcement Learning

  • Deheng Ye
  • Zhao Liu
  • Mingfei Sun
  • Bei Shi
  • Peilin Zhao
  • Hao Wu
  • Hongsheng Yu
  • Shaojie Yang

We study the reinforcement learning problem of complex action control in the Multi-player Online Battle Arena (MOBA) 1v1 games. This problem involves far more complicated state and action spaces than those of traditional 1v1 games, such as Go and Atari series, which makes it very difficult to search any policies with human-level performance. In this paper, we present a deep reinforcement learning framework to tackle this problem from the perspectives of both system and algorithm. Our system is of low coupling and high scalability, which enables efficient explorations at large scale. Our algorithm includes several novel strategies, including control dependency decoupling, action mask, target attention, and dualclip PPO, with which our proposed actor-critic network can be effectively trained in our system. Tested on the MOBA game Honor of Kings, the trained AI agents can defeat top professional human players in full 1v1 games.

AAAI Conference 2020 Conference Paper

Part-Level Graph Convolutional Network for Skeleton-Based Action Recognition

  • Linjiang Huang
  • Yan Huang
  • Wanli Ouyang
  • Liang Wang

Recently, graph convolutional networks have achieved remarkable performance for skeleton-based action recognition. In this work, we identify a problem posed by the GCNs for skeleton-based action recognition, namely part-level action modeling. To address this problem, a novel Part-Level Graph Convolutional Network (PL-GCN) is proposed to capture part-level information of skeletons. Different from previous methods, the partition of body parts is learnable rather than manually defined. We propose two part-level blocks, namely Part Relation block (PR block) and Part Attention block (PA block), which are achieved by two differentiable operations, namely graph pooling operation and graph unpooling operation. The PR block aims at learning high-level relations between body parts while the PA block aims at highlighting the important body parts in the action. Integrating the original GCN with the two blocks, the PL-GCN can learn both part-level and joint-level information of the action. Extensive experiments on two benchmark datasets show the state-ofthe-art performance on skeleton-based action recognition and demonstrate the effectiveness of the proposed method.

AAAI Conference 2020 Conference Paper

Pose-Guided Multi-Granularity Attention Network for Text-Based Person Search

  • Ya Jing
  • Chenyang Si
  • Junbo Wang
  • Wei Wang
  • Liang Wang
  • Tieniu Tan

Text-based person search aims to retrieve the corresponding person images in an image database by virtue of a describing sentence about the person, which poses great potential for various applications such as video surveillance. Extracting visual contents corresponding to the human description is the key to this cross-modal matching problem. Moreover, correlated images and descriptions involve different granularities of semantic relevance, which is usually ignored in previous methods. To exploit the multilevel corresponding visual contents, we propose a pose-guided multi-granularity attention network (PMA). Firstly, we propose a coarse alignment network (CA) to select the related image regions to the global description by a similarity-based attention. To further capture the phrase-related visual body part, a fine-grained alignment network (FA) is proposed, which employs pose information to learn latent semantic alignment between visual body part and textual noun phrase. To verify the effectiveness of our model, we perform extensive experiments on the CUHK Person Description Dataset (CUHK-PEDES) which is currently the only available dataset for text-based person search. Experimental results show that our approach outperforms the state-of-the-art methods by 15 % in terms of the top-1 metric.

AAAI Conference 2020 Conference Paper

Relational Prototypical Network for Weakly Supervised Temporal Action Localization

  • Linjiang Huang
  • Yan Huang
  • Wanli Ouyang
  • Liang Wang

In this paper, we propose a weakly supervised temporal action localization method on untrimmed videos based on prototypical networks. We observe two challenges posed by weakly supervision, namely action-background separation and action relation construction. Unlike the previous method, we propose to achieve action-background separation only by the original videos. To achieve this, a clustering loss is adopted to separate actions from backgrounds and learn intra-compact features, which helps in detecting complete action instances. Besides, a similarity weighting module is devised to further separate actions from backgrounds. To effectively identify actions, we propose to construct relations among actions for prototype learning. A GCN-based prototype embedding module is introduced to generate relational prototypes. Experiments on THUMOS14 and ActivityNet1. 2 datasets show that our method outperforms the state-of-the-art methods.

NeurIPS Conference 2020 Conference Paper

Towards Playing Full MOBA Games with Deep Reinforcement Learning

  • Deheng Ye
  • Guibin Chen
  • Wen Zhang
  • Sheng Chen
  • Bo Yuan
  • Bo Liu
  • Jia Chen
  • Zhao Liu

MOBA games, e. g. , Honor of Kings, League of Legends, and Dota 2, pose grand challenges to AI systems such as multi-agent, enormous state-action space, complex action control, etc. Developing AI for playing MOBA games has raised much attention accordingly. However, existing work falls short in handling the raw game complexity caused by the explosion of agent combinations, i. e. , lineups, when expanding the hero pool in case that OpenAI's Dota AI limits the play to a pool of only 17 heroes. As a result, full MOBA games without restrictions are far from being mastered by any existing AI system. In this paper, we propose a MOBA AI learning paradigm that methodologically enables playing full MOBA games with deep reinforcement learning. Specifically, we develop a combination of novel and existing learning techniques, including off-policy adaption, multi-head value estimation, curriculum self-play learning, policy distillation, and Monte-Carlo tree-search, in training and playing a large pool of heroes, meanwhile addressing the scalability issue skillfully. Tested on Honor of Kings, a popular MOBA game, we show how to build superhuman AI agents that can defeat top esports players. The superiority of our AI is demonstrated by the first large-scale performance test of MOBA AI agent in the literature.

NeurIPS Conference 2020 Conference Paper

Unfolding the Alternating Optimization for Blind Super Resolution

  • zhengxiong luo
  • Yan Huang
  • Shang Li
  • Liang Wang
  • Tieniu Tan

Previous methods decompose blind super resolution (SR) problem into two sequential steps: \textit{i}) estimating blur kernel from given low-resolution (LR) image and \textit{ii}) restoring SR image based on estimated kernel. This two-step solution involves two independently trained models, which may not well compatible with each other. Small estimation error of the first step could cause severe performance drop of the second one. While on the other hand, the first step can only utilize limited information from LR image, which makes it difficult to predict highly accurate blur kernel. Towards these issues, instead of considering these two steps separately, we adopt an alternating optimization algorithm, which can estimate blur kernel and restore SR image in a single model. Specifically, we design two convolutional neural modules, namely \textit{Restorer} and \textit{Estimator}. \textit{Restorer} restores SR image based on predicted kernel, and \textit{Estimator} estimates blur kernel with the help of restored SR image. We alternate these two modules repeatedly and unfold this process to form an end-to-end trainable network. In this way, \textit{Estimator} utilizes information from both LR and SR images, which makes the estimation of blur kernel easier. More importantly, \textit{Restorer} is trained with the kernel estimated by \textit{Estimator}, instead of ground-truth kernel, thus \textit{Restorer} could be more tolerant to the estimation error of \textit{Estimator}. Extensive experiments on synthetic datasets and real-world images show that our model can largely outperform state-of-the-art methods and produce more visually favorable results at much higher speed. The source code will be publicly available.

AAAI Conference 2019 Conference Paper

Few-Shot Image and Sentence Matching via Gated Visual-Semantic Embedding

  • Yan Huang
  • Yang Long
  • Liang Wang

Although image and sentence matching has been widely studied, its intrinsic few-shot problem is commonly ignored, which has become a bottleneck for further performance improvement. In this work, we focus on this challenging problem of few-shot image and sentence matching, and propose a Gated Visual-Semantic Embedding (GVSE) model to deal with it. The model consists of three corporative modules in terms of uncommon VSE, common VSE, and gated metric fusion. The uncommon VSE exploits external auxiliary resources to extract generic features for representing uncommon instances and words in images and sentences, and then integrates them by modeling their semantic relation to obtain global representations for association analysis. To better model other common instances and words in rest content of images and sentences, the common VSE learns their discriminative representations directly from scratch. After obtaining two similarity metrics from the two VSE modules with different advantages, the gated metric fusion module adaptively fuses them by automatically balancing their relative importance. Based on the fused metric, we perform extensive experiments in terms of few-shot and conventional image and sentence matching, and demonstrate the effectiveness of the proposed model by achieving the state-of-the-art results on two public benchmark datasets.

AAAI Conference 2019 Conference Paper

Session-Based Recommendation with Graph Neural Networks

  • Shu Wu
  • Yuyuan Tang
  • Yanqiao Zhu
  • Liang Wang
  • Xing Xie
  • Tieniu Tan

The problem of session-based recommendation aims to predict user actions based on anonymous sessions. Previous methods model a session as a sequence and estimate user representations besides item representations to make recommendations. Though achieved promising results, they are insufficient to obtain accurate user vectors in sessions and neglect complex transitions of items. To obtain accurate item embedding and take complex transitions of items into account, we propose a novel method, i. e. Session-based Recommendation with Graph Neural Networks, SR-GNN for brevity. In the proposed method, session sequences are modeled as graphstructured data. Based on the session graph, GNN can capture complex transitions of items, which are difficult to be revealed by previous conventional sequential methods. Each session is then represented as the composition of the global preference and the current interest of that session using an attention network. Extensive experiments conducted on two real datasets show that SR-GNN evidently outperforms the state-of-the-art session-based recommendation methods consistently.

AAAI Conference 2018 Conference Paper

Lateral Inhibition-Inspired Convolutional Neural Network for Visual Attention and Saliency Detection

  • Chunshui Cao
  • Yongzhen Huang
  • Zilei Wang
  • Liang Wang
  • Ninglong Xu
  • Tieniu Tan

Lateral inhibition in top-down feedback is widely existing in visual neurobiology, but such an important mechanism has not be well explored yet in computer vision. In our recent research, we find that modeling lateral inhibition in convolutional neural network (LICNN) is very useful for visual attention and saliency detection. In this paper, we propose to formulate lateral inhibition inspired by the related studies from neurobiology, and embed it into the top-down gradient computation of a general CNN for classification, i. e. only category-level information is used. After this operation (only conducted once), the network has the ability to generate accurate category-specific attention maps. Further, we apply LICNN for weakly-supervised salient object detection. Extensive experimental studies on a set of databases, e. g. , EC- SSD, HKU-IS, PASCAL-S and DUT-OMRON, demonstrate the great advantage of LICNN which achieves the state-ofthe-art performance. It is especially impressive that LICNN with only category-level supervised information even outperforms some recent methods with segmentation-level supervised learning.

TIST Journal 2018 Journal Article

Mining Significant Microblogs for Misinformation Identification

  • Qiang Liu
  • Feng Yu
  • Shu Wu
  • Liang Wang

With the rapid growth of social media, massive misinformation is also spreading widely on social media, e.g., Weibo and Twitter, and brings negative effects to human life. Today, automatic misinformation identification has drawn attention from academic and industrial communities. Whereas an event on social media usually consists of multiple microblogs, current methods are mainly constructed based on global statistical features. However, information on social media is full of noise, which should be alleviated. Moreover, most of the microblogs about an event have little contribution to the identification of misinformation, where useful information can be easily overwhelmed by useless information. Thus, it is important to mine significant microblogs for constructing a reliable misinformation identification method. In this article, we propose an attention-based approach for identification of misinformation (AIM). Based on the attention mechanism, AIM can select microblogs with the largest attention values for misinformation identification. The attention mechanism in AIM contains two parts: content attention and dynamic attention. Content attention is the calculated-based textual features of each microblog. Dynamic attention is related to the time interval between the posting time of a microblog and the beginning of the event. To evaluate AIM, we conduct a series of experiments on the Weibo and Twitter datasets, and the experimental results show that the proposed AIM model outperforms the state-of-the-art methods.

IJCAI Conference 2017 Conference Paper

A Convolutional Approach for Misinformation Identification

  • Feng Yu
  • Qiang Liu
  • Shu Wu
  • Liang Wang
  • Tieniu Tan

The fast expanding of social media fuels the spreading of misinformation which disrupts people's normal lives. It is urgent to achieve goals of misinformation identification and early detection in social media. In dynamic and complicated social media scenarios, some conventional methods mainly concentrate on feature engineering which fail to cover potential features in new scenarios and have difficulty in shaping elaborate high-level interactions among significant features. Moreover, a recent Recurrent Neural Network (RNN) based method suffers from deficiencies that it is not qualified for practical early detection of misinformation and poses a bias to the latest input. In this paper, we propose a novel method, Convolutional Approach for Misinformation Identification (CAMI) based on Convolutional Neural Network (CNN). CAMI can flexibly extract key features scattered among an input sequence and shape high-level interactions among significant features, which help effectively identify misinformation and achieve practical early detection. Experiment results on two large-scale datasets validate the effectiveness of CAMI model on both misinformation identification and early detection tasks.

AAAI Conference 2017 Conference Paper

How to Train a Compact Binary Neural Network with High Accuracy?

  • Wei Tang
  • Gang Hua
  • Liang Wang

How to train a binary neural network (BinaryNet) with both high compression rate and high accuracy on large scale datasets? We answer this question through a careful analysis of previous work on BinaryNets, in terms of training strategies, regularization, and activation approximation. Our findings first reveal that a low learning rate is highly preferred to avoid frequent sign changes of the weights, which often makes the learning of BinaryNets unstable. Secondly, we propose to use PReLU instead of ReLU in a BinaryNet to conveniently absorb the scale factor for weights to the activation function, which enjoys high computation efficiency for binarized layers while maintains high approximation accuracy. Thirdly, we reveal that instead of imposing L2 regularization, driving all weights to zero which contradicts with the setting of BinaryNets, we introduce a regularization term that encourages the weights to be bipolar. Fourthly, we discover that the failure of binarizing the last layer, which is essential for high compression rate, is due to the improper output range. We propose to use a scale layer to bring it to normal. Last but not least, we propose multiple binarizations to improve the approximation of the activations. The composition of all these enables us to train BinaryNets with both high compression rate and high accuracy, which is strongly supported by our extensive empirical study.

AAAI Conference 2016 Conference Paper

Information Credibility Evaluation on Social Media

  • Shu Wu
  • Qiang Liu
  • Yong Liu
  • Liang Wang
  • Tieniu Tan

With the growth of social media, rumors are spread fast and viewed by more and more people on the Internet. Rumors bring significant harm to daily life and public security. It is crucial to evaluate the credibility of information and detect the rumors on social media automatically. In this work, we establish a Network Information Credibility Evaluation (NICE) platform, which collects a database of rumors that have been verified on Sina Weibo and automatically evaluates the information which is generated by users on social media but has not been verified. Users can use a query to search related information. If the according information appears in our database, users can identify it is a rumor immediately. Otherwise, NICE will show users with realtime results crawled automatically from social media and can calculate credibility of a specific result with our algorithm. Our algorithm learns dynamic representations for information on social media based on behavior information, dynamic information, user information and comment information. Then, we use an ordinary logistic regression to classify information into rumors and non-rumors. Based on our algorithm, NICE system achieves satisfactory performance on evaluating information credibility and detecting rumors on social media.

AAAI Conference 2016 Conference Paper

Predicting the Next Location: A Recurrent Model with Spatial and Temporal Contexts

  • Qiang Liu
  • Shu Wu
  • Liang Wang
  • Tieniu Tan

Spatial and temporal contextual information plays a key role for analyzing user behaviors, and is helpful for predicting where he or she will go next. With the growing ability of collecting information, more and more temporal and spatial contextual information is collected in systems, and the location prediction problem becomes crucial and feasible. Some works have been proposed to address this problem, but they all have their limitations. Factorizing Personalized Markov Chain (FPMC) is constructed based on a strong independence assumption among different factors, which limits its performance. Tensor Factorization (TF) faces the cold start problem in predicting future actions. Recurrent Neural Networks (RNN) model shows promising performance comparing with PFMC and TF, but all these methods have problem in modeling continuous time interval and geographical distance. In this paper, we extend RNN and propose a novel method called Spatial Temporal Recurrent Neural Networks (ST-RNN). ST-RNN can model local temporal and spatial contexts in each layer with time-specific transition matrices for different time intervals and distance-specific transition matrices for different geographical distances. Experimental results show that the proposed ST-RNN model yields significant improvements over the competitive compared methods on two typical datasets, i. e. , Global Terrorism Database (GTD) and Gowalla dataset.

AAAI Conference 2016 Conference Paper

SAPE: A System for Situation-Aware Public Security Evaluation

  • Shu Wu
  • Qiang Liu
  • Ping Bai
  • Liang Wang
  • Tieniu Tan

Public security events are occurring all over the world, bringing threat to personal and property safety, and homeland security. It is vital to construct an effective model to evaluate and predict the public security. In this work, we establish a Situation-Aware Public Security Evaluation (SAPE) platform. Based on conventional Recurrent Neural Networks (RNN), we develop a new variant for temporal contexts in public security event datasets. This model can achieve better performance than the compared state-of-the-art methods. SAPE has two demonstrations, i. e. , global public security evaluation and China public security evaluation. In the global part, based on Global Terrorism Database from UMD, for each country, SAPE can predict risk level and top-n potential terrorist organizations which might attack the country. Users can also view the actual attacking organizations and predicted results. For each province in China, SAPE can predict the risk level and the probability scores of different types of events in the next month. Users can also view the actual numbers of events and predicted risk levels of the past one year.

NeurIPS Conference 2015 Conference Paper

Bidirectional Recurrent Convolutional Networks for Multi-Frame Super-Resolution

  • Yan Huang
  • Wei Wang
  • Liang Wang

Super resolving a low-resolution video is usually handled by either single-image super-resolution (SR) or multi-frame SR. Single-Image SR deals with each video frame independently, and ignores intrinsic temporal dependency of video frames which actually plays a very important role in video super-resolution. Multi-Frame SR generally extracts motion information, e. g. optical flow, to model the temporal dependency, which often shows high computational cost. Considering that recurrent neural network (RNN) can model long-term contextual information of temporal sequences well, we propose a bidirectional recurrent convolutional network for efficient multi-frame SR. Different from vanilla RNN, 1) the commonly-used recurrent full connections are replaced with weight-sharing convolutional connections and 2) conditional convolutional connections from previous input layers to current hidden layer are added for enhancing visual-temporal dependency modelling. With the powerful temporal dependency modelling, our model can super resolve videos with complex motions and achieve state-of-the-art performance. Due to the cheap convolution operations, our model has a low computational complexity and runs orders of magnitude faster than other multi-frame methods.

AAAI Conference 2015 Conference Paper

COT: Contextual Operating Tensor for Context-Aware Recommender Systems

  • Qiang Liu
  • Shu Wu
  • Liang Wang

With rapid growth of information on the internet, recommender systems become fundamental for helping users alleviate the problem of information overload. Since contextual information can be used as a significant factor in modeling user behavior, various contextaware recommendation methods are proposed. However, the state-of-the-art context modeling methods treat contexts as other dimensions similar to the dimensions of users and items, and cannot capture the special semantic operation of contexts. On the other hand, some works on multi-domain relation prediction can be used for the context-aware recommendation, but they have problems in generating recommendation under a large amount of contextual information. In this work, we propose Contextual Operating Tensor (COT) model, which represents the common semantic effects of contexts as a contextual operating tensor and represents a context as a latent vector. Then, to model the semantic operation of a context combination, we generate contextual operating matrix from the contextual operating tensor and latent vectors of contexts. Thus latent vectors of users and items can be operated by the contextual operating matrices. Experimental results show that the proposed COT model yields significant improvements over the competitive compared methods on three typical datasets, i. e. , Food, Adom and Movielens-1M datasets.

JBHI Journal 2013 Journal Article

An Online One Class Support Vector Machine-Based Person-Specific Fall Detection System for Monitoring an Elderly Individual in a Room Environment

  • Miao Yu
  • Yuanzhang Yu
  • Adel Rhuma
  • Syed Mohsen Raza Naqvi
  • Liang Wang
  • Jonathon A. Chambers

In this paper, we propose a novel computer vision-based fall detection system for monitoring an elderly person in a home care, assistive living application. Initially, a single camera covering the full view of the room environment is used for the video recording of an elderly person's daily activities for a certain time period. The recorded video is then manually segmented into short video clips containing normal postures, which are used to compose the normal dataset. We use the codebook background subtraction technique to extract the human body silhouettes from the video clips in the normal dataset and information from ellipse fitting and shape description, together with position information, is used to provide features to describe the extracted posture silhouettes. The features are collected and an online one class support vector machine (OCSVM) method is applied to find the region in feature space to distinguish normal daily postures and abnormal postures such as falls. The resultant OCSVM model can also be updated by using the online scheme to adapt to new emerging normal postures and certain rules are added to reduce false alarm rate and thereby improve fall detection performance. From the comprehensive experimental evaluations on datasets for 12 people, we confirm that our proposed person-specific fall detection system can achieve excellent fall detection performance with 100% fall detection rate and only 3% false detection rate with the optimally tuned parameters. This work is a semiunsupervised fall detection system from a system perspective because although an unsupervised-type algorithm (OCSVM) is applied, human intervention is needed for segmenting and selecting of video clips containing normal postures. As such, our research represents a step toward a complete unsupervised fall detection system.

NeurIPS Conference 2013 Conference Paper

Relevance Topic Model for Unstructured Social Group Activity Recognition

  • Fang Zhao
  • Yongzhen Huang
  • Liang Wang
  • Tieniu Tan

Unstructured social group activity recognition in web videos is a challenging task due to 1) the semantic gap between class labels and low-level visual features and 2) the lack of labeled training data. To tackle this problem, we propose a relevance topic model" for jointly learning meaningful mid-level representations upon bag-of-words (BoW) video representations and a classifier with sparse weights. In our approach, sparse Bayesian learning is incorporated into an undirected topic model (i. e. , Replicated Softmax) to discover topics which are relevant to video classes and suitable for prediction. Rectified linear units are utilized to increase the expressive power of topics so as to explain better video data containing complex contents and make variational inference tractable for the proposed model. An efficient variational EM algorithm is presented for model parameter estimation and inference. Experimental results on the Unstructured Social Activity Attribute dataset show that our model achieves state of the art performance and outperforms other supervised topic model in terms of classification accuracy, particularly in the case of a very small number of labeled training videos. "

YNIMG Journal 2011 Journal Article

Characterizing dynamic functional connectivity in the resting brain using variable parameter regression and Kalman filtering approaches

  • Jin Kang
  • Liang Wang
  • Chaogan Yan
  • Jinhui Wang
  • Xia Liang
  • Yong He

The cognitive activity of the human brain benefits from the functional connectivity of multiple brain regions that form specific, functional brain networks. Recent studies have indicated that the relationship between brain regions can be investigated by examining the temporal interaction (known as functional connectivity) of spontaneous blood oxygen level-dependent (BOLD) signals derived from resting-state functional MRI. Most of these studies plausibly assumed that inter-regional interactions were temporally stationary. However, little is known about the dynamic characteristics of resting-state functional connectivity (RSFC). In this study, we thoroughly examined this question within and between multiple functional brain networks. Twenty-two healthy subjects were scanned in a resting state. Several of the RSFC networks observed, including the default-mode, motor, attention, memory, auditory, visual, language and subcortical networks, were first identified using a conventional voxel-wise correlation analysis with predefined region of interests (ROIs). Then, a variable parameter regression model combined with the Kalman filtering method was employed to detect the dynamic interactions between each ROI and all other brain voxels within each of the RSFC maps extracted above. Experimental results revealed that the functional interactions within each RSFC map showed time-varying properties, and that approximately 10–20% of the voxels within each RSFC map showed significant functional connectivity to each ROI during the scanning session. This dynamic pattern was also observed for the interactions between different functional networks. In addition, the spatial pattern of dynamic connectivity maps obtained from neighboring time points had a high similarity. Overall, this study provides insights into the dynamic properties of resting-state functional networks.

YNIMG Journal 2011 Journal Article

Effective connectivity of brain networks during self-initiated movement in Parkinson's disease

  • Tao Wu
  • Liang Wang
  • Mark Hallett
  • Yi Chen
  • Kuncheng Li
  • Piu Chan

Patients with Parkinson's disease (PD) have difficulty in performing self-initiated movements. The neural mechanism of this deficiency remains unclear. In the current study, we used functional MRI (fMRI) and psychophysiological interaction (PPI) methods to investigate the changes in effective connectivity of the brain networks during performance of self-initiated movement in PD patients. Effective connectivity is defined as the influence one neuronal system exerts over another. fMRIs were acquired in 18 PD patients and in 18 age- and sex-matched healthy controls, when performing a self-initiated right hand tapping task. We chose the left primary motor cortex (M1), rostral supplementary motor area (pre-SMA), left premotor cortex (PMC), left putamen, and right cerebellum as index areas for PPI analysis. During the performance of self-initiated movement, connectivity between the putamen and M1, PMC, SMA, and cerebellum was decreased in PD patients compared to controls. In contrast, connections between the M1, pre-SMA, PMC, parietal cortex, and cerebellum were increased in PD patients compared to controls. In addition, the M1, pre-SMA, PMC, and cerebellum also had less connectivity with the dorsal lateral prefrontal cortex in PD. In PD patients, the effective connectivity between the putamen and M1, PMC, SMA, and cerebellum negatively correlated with the Unified Parkinson's Disease Rating Scale (UPDRS) motor scores; whereas the connectivity between the M1, pre-SMA, PMC, and cerebellum positively correlated with the UPDRS motor scores. Our findings demonstrate that the pattern of interactions of brain networks is disrupted in PD during performance of self-initiated movements. The striatum-cortical and striatum-cerebellar connections are weakened. In contrast, the connections between cortico-cerebellar motor regions are strengthened and may compensate for basal ganglia dysfunction. These altered interregional connections are more deviant when the disorder is more severe, and, therefore, our results give further insight into the explanation for the difficulty in performing self-initiated movements in PD.

YNIMG Journal 2010 Journal Article

Age-related changes in topological patterns of large-scale brain functional networks during memory encoding and recognition

  • Liang Wang
  • Yanfang Li
  • Paul Metzak
  • Yong He
  • Todd S. Woodward

In this study we used functional magnetic resonance imaging to investigate age-related changes in large-scale brain functional networks during memory encoding and recognition in 12 younger and 16 older adults. For each participant, functional brain networks were constructed by computing temporal correlation matrices of 90 brain regions and analyzed using graph theoretical approaches. We found the age-related changes mainly in the long-range connections with widespread reductions associated with aging in the fronto-temporal and temporo-parietal regions, and a few age-related increases in the posterior parietal regions. Graph theoretical analysis revealed that the older adults had longer path lengths linking different regions in the functional brain networks as compared to the younger adults. Further analysis indicated that the increases in shortest path length in the networks were combined with the loss of long-range connections. Finally, we showed that for older adults, frontal areas played reduced roles in the network (reduced regional centrality), whereas several default-mode regions played increased roles relative to younger subjects (increased regional centrality). Together, our results suggest that normal aging is associated with disruption of large-scale brain systems during the performance of memory tasks, which provides novel insights into the understanding of age-related decline in multiple cognitive functions.

YNIMG Journal 2010 Journal Article

Intrinsic connectivity between the hippocampus and posteromedial cortex predicts memory performance in cognitively intact older individuals

  • Liang Wang
  • Peter LaViolette
  • Kelly O'Keefe
  • Deepti Putcha
  • Akram Bakkour
  • Koene R.A. Van Dijk
  • Maija Pihlajamäki
  • Bradford C. Dickerson

Coherent fluctuations of spontaneous brain activity are present in distinct functional-anatomic brain systems during undirected wakefulness. However, the behavioral significance of this spontaneous activity has only begun to be investigated. Our previous studies have demonstrated that successful memory formation requires coordinated neural activity in a distributed memory network including the hippocampus and posteromedial cortices, specifically the precuneus and posterior cingulate (PPC), thought to be integral nodes of the default network. In this study, we examined whether intrinsic connectivity during the resting state between the hippocampus and PPC can predict individual differences in the performance of an associative memory task among cognitively intact older individuals. The intrinsic connectivity, between regions within the hippocampus and PPC that were maximally engaged during a subsequent memory fMRI task, was measured during a period of rest prior to the performance of the memory paradigm. Stronger connectivity between the hippocampal and posteromedial regions during rest predicted better performance on the memory task. Furthermore, hippocampal-PPC intrinsic connectivity was also significantly correlated with episodic memory measures on neuropsychological tests, but not with performance in non-memory domains. Whole-brain exploratory analyses further confirmed the spatial specificity of the relationship between hippocampal-default network posteromedial cortical connectivity and memory performance in older subjects. Our findings provide support for the hypothesis that one of the functions of this large-scale brain network is to subserve episodic memory processes. Research is ongoing to determine if impaired connectivity between these regions may serve as a predictor of memory decline related to early Alzheimer's disease.

YNIMG Journal 2007 Journal Article

Regional coherence changes in the early stages of Alzheimer’s disease: A combined structural and resting-state functional MRI study

  • Yong He
  • Liang Wang
  • Yufeng Zang
  • Lixia Tian
  • Xinqing Zhang
  • Kuncheng Li
  • Tianzi Jiang

Recent functional imaging studies have indicated that the pathophysiology of Alzheimer’s disease (AD) can be associated with the changes in spontaneous low-frequency (<0. 08 Hz) blood oxygenation level-dependent fluctuations (LFBF) measured during a resting state. The purpose of this study was to examine regional LFBF coherence patterns in early AD and the impact of regional brain atrophy on the functional results. Both structural MRI and resting-state functional MRI scans were collected from 14 AD subjects and 14 age-matched normal controls. We found significant regional coherence decreases in the posterior cingulate cortex/precuneus (PCC/PCu) in the AD patients when compared with the normal controls. Moreover, the decrease in the PCC/PCu coherence was correlated with the disease progression measured by the Mini-Mental State Exam scores. The changes in LFBF in the PCC/PCu may be related to the resting hypometabolism in this region commonly detected in previous positron emission tomography studies of early AD. When the regional PCC/PCu atrophy was controlled, these results still remained significant but with a decrease in the statistical power, suggesting that the LFBF results are at least partly explained by the regional atrophy. In addition, we also found increased LFBF coherence in the bilateral cuneus, right lingual gyrus and left fusiform gyrus in the AD patients. These regions are consistent with previous findings of AD-related increased activation during cognitive tasks explained in terms of a compensatory-recruitment hypothesis. Finally, our study indicated that regional brain atrophy could be an important consideration in functional imaging studies of neurodegenerative diseases.

YNIMG Journal 2006 Journal Article

Changes in hippocampal connectivity in the early stages of Alzheimer's disease: Evidence from resting state fMRI

  • Liang Wang
  • Yufeng Zang
  • Yong He
  • Meng Liang
  • Xinqing Zhang
  • Lixia Tian
  • Tao Wu
  • Tianzi Jiang

A selective distribution of Alzheimer's disease (AD) pathological lesions in specific cortical layers isolates the hippocampus from the rest of the brain. However, functional connectivity between the hippocampus and other brain regions remains unclear in AD. Here, we employ a resting state functional MRI (fMRI) to examine changes in hippocampal connectivity comparing 13 patients with mild AD versus 13 healthy age-matched controls. Hippocampal connectivity was investigated by examination of the correlation between low frequency fMRI signal fluctuations in the hippocampus and those in all other brain regions. We found that functional connectivity between the right hippocampus and a set of regions was disrupted in AD; these regions are: medial prefrontal cortex (MPFC), ventral anterior cingulate cortex (vACC), right inferotemporal cortex, right cuneus extending into precuneus, left cuneus, right superior and middle temporal gyrus and posterior cingulate cortex (PCC). We also found increased functional connectivity between the left hippocampus and the right lateral prefrontal cortex in AD. In addition, rightward asymmetry of hippocampal connectivity observed in elderly controls was diminished in AD patients. The disrupted hippocampal connectivity to the MPFC, vACC and PCC provides further support for decreased activity in “default mode network” previously shown in AD. The decreased connectivity between the hippocampus and the visual cortices might indicate reduced integrity of hippocampus-related cortical networks in AD. Moreover, these findings suggest that resting-state fMRI might be an appropriate approach for studying pathophysiological changes in early AD.

v2026.09.13