Arrow Research search

Author name cluster

Wei Liang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

30 papers
2 author rows

Possible papers

30

EAAI Journal 2026 Journal Article

A joint economic synergy planning framework for coupled energy–carbon systems with cooperative cost-sharing

  • Hongxiang Ge
  • Wei Liang

This paper develops a cost-aware synergy cooperative -planning framework for the coordinated expansion of multi-energy systems and carbon capture, utilization, and storage (CCUS) infrastructure. To address key limitations in existing studies, including the lack of regional carbon transport representation and simplified cost-sharing structures, an integrated investment and operational planning approach is proposed based on convex relaxation and cooperative coordination principles. In addition to physics-informed thermochemical and electrochemical constraints, an artificial intelligence (AI)–assisted decision support layer is incorporated to enhance computational efficiency and scalability. This layer employs regression-based surrogate learning to approximate nonlinear carbon capture and conversion behavior and applies data-driven scenario screening to reduce high-dimensional uncertainty prior to full-scale optimization. Cooperative interactions among stakeholders are represented through a carbon-oriented cost allocation mechanism that accounts for marginal economic contributions and emission intensity differences. The proposed framework is evaluated on benchmark and extended large-scale test systems, demonstrating up to 10% total cost reduction, full carbon neutrality, and improved infrastructure deployment under coordinated planning. Sensitivity analyses under varying carbon price signals confirm economic resilience, while scalability assessments verify applicability to large regional networks. The results indicate that cooperative energy–carbon planning, supported by AI–based decision assistance, can effectively balance environmental objectives with economic performance in emerging decentralized carbon markets.

EAAI Journal 2026 Journal Article

CGMAE: Self-supervised Masked Auto-Encoder with Cross-Graph node alignment for node classification

  • Ruoxian Song
  • Peng Cao
  • Guangqi Wen
  • Lanting Li
  • Wei Liang
  • Weiping Li
  • Jinzhu Yang
  • Osmar R. Zaiane

Masked Auto-Encoder (MAE) is widely adopted for node classification by recovering the randomly masked graph structure or node attributes. However, traditional MAE methods face two critical challenges: (1) features learned for reconstruction may not align with the downstream classification task, and (2) masking edges risks distorting inherent semantic relationships, degrading representation quality. To overcome these limitations, we propose a simple yet effective self-supervised M asked A uto- E ncoder with C ross- G raph node alignment (CGMAE) for node classification. It leverages labeled nodes from an auxiliary graph to enhance discriminative feature learning in an unlabeled target graph, bridging the task gap between reconstruction and classification. CGMAE introduces a node-level alignment mechanism to address distribution shifts across graphs. This design jointly learns structural patterns and node attributes through specific encoders, enabling multi-view feature matching to refine node representations. Furthermore, CGMAE innovatively predicts masked target edges using aligned nodes from the auxiliary graph, preserving the semantic relationships during reconstruction. Extensive experiments on six diverse networks (standard, complex, sparse, and large-scale graphs) verify the effectiveness and robustness of the proposed method in self-supervised/unsupervised node classification tasks, with accuracy improvements ranging from 1. 5%/1. 1% to 3. 5%/7. 0% over state-of-the-art methods. Code is available at https: //github. com/songruoxian/CGMAE.

JBHI Journal 2026 Journal Article

Chemistry-Structure Dual-Perception Large Language Models: Advancing Molecular Property Prediction for Precise Disease Treatment

  • Li Peng
  • Jun Kai Gao
  • Zong Yi Yang
  • Xin Yi Ai
  • Wei Liang

Accurate prediction of drug molecular properties is crucial for precision drug discovery, which is closely related to precise disease diagnosis. Understanding the physicochemical properties, biological activities, and mechanisms of action of molecules in biological systems can support early disease diagnosis and personalized treatment. Machine learning (ML) and deep learning (DL) technologies have significantly enhanced the accuracy of predicting these properties. However, current methods face challenges: heavy reliance on substantial computational resources and limited ability to incorporate chemists' perspectives. We propose CSLLM, a novel method that uses instructions to guide large language models (LLMs) to generate drug molecular representations embedded with chemical knowledge. CSLLM introduces a three-dimensional instruction framework: (1) task guidance, focusing LLMs on key information for specific prediction tasks; (2) chemical perception, enabling LLMs to reason like chemists; and (3) structural perception, improving LLMs' understanding of drug molecular structures. Evaluation on nine datasets shows CSLLM outperforms existing models. In addition, we demonstrate through visualization that CSLLM is capable of reasoning from a chemist's perspective. In summary, CSLLM generates chemically knowledge-rich drug molecular representations with limited computational resources, illuminating molecules' potential applications in disease diagnosis.

JBHI Journal 2026 Journal Article

CPFTransGAN: A Cross Perception Fusion Transformer-based Generative Adversarial Network for Head and Neck Cancer Dose Prediction in Radiotherapy

  • Miao Liao
  • Enyu Zhou
  • Xiong Li
  • Wei Liang
  • Yuqian Zhao
  • Shuanhu Di
  • Victor Chang

Radiation therapy is one of the primary treatment modalities for head and neck (H&N) cancer in clinical practice, aiming to deliver sufficient dose to Planning Target Volume (PTV) while protecting surrounding Organs at Risk (OAR) from or minimizing exposure to radiation. Quantitative dose prediction of various tissues and organs is a prerequisite for implementing intelligent precision radiotherapy. In order to improve dose prediction accuracy, we propose a generative adversarial network CPFTrans- GAN based on Cross Perception Fusion Transformer (CPF Transformer). Specifically, we design a CPF Transformer module through deeply integrating CNN and Transformer. Using the CPF Transformer as basic unit, we constructed a generator with four-stage encoding-decoding structure called CPFTransGenerator. An adaptive weight loss is used to train the discriminator to alleviate the issues of imbalance training in adversarial learning. To further improve the prediction accuracy, a multiscale cross-window encoding network is designed, which can constrain the differences between predicted dose and the reference one at different granularity levels by calculating feature losses between them at different scales. The proposed method is evaluated on two public head and neck cancer datasets and a local clinical dataset. Extensive experiments demonstrate the superior performance of our method compared with the state-of-the-art ones.

EAAI Journal 2026 Journal Article

MAFSA: A multi-layer asynchronous federated learning with staleness-awareness in edge computing

  • Shiwen Zhang
  • Shuang Chen
  • Yujia Zhang
  • Wei Liang
  • Kuanching Li
  • Ling-Huey Li
  • Keqin Li

Federated Learning (FL) has emerged as a promising privacy-preserving scheme in edge computing. However, traditional cloud-based FL architectures still suffer from high communication overhead, which motivates the development of hierarchical and asynchronous variants to improve communication efficiency. However, traditional cloud-side-end architecture must wait for the results of all devices to complete the update, resulting in inefficient training. Furthermore, the phenomenon of end-device dropouts can lead to the waste of resources, thereby compromising the system’s fairness. In this work, we propose MAFSA, a multi-layer asynchronous federated learning scheme with staleness-awareness, which aims to enhance the system’s efficiency and resource utilization while maintaining accuracy and ensuring fairness. MAFSA proposes a composite asynchronous aggregation strategy to address the inefficiency issue caused by device heterogeneity and enhance the system’s communication efficiency. In addition, MAFSA proposes a staleness-awareness mechanism to cope with end device dropping, improve resource utilization, and ensure the system’s fairness. Extensive experiments demonstrate that the proposed scheme effectively utilizes stale information, significantly benefiting the federated learning system and proving the method’s effectiveness.

AAAI Conference 2026 Conference Paper

RadarMP: Motion Perception for 4D mmWave Radar in Autonomous Driving

  • Ruiqi Cheng
  • Huijun Di
  • Jian Li
  • Feng Liu
  • Wei Liang

Accurate 3D scene motion perception significantly enhances the safety and reliability of an autonomous driving system. Benefiting from its all-weather operational capability and unique perceptual properties, 4D mmWave radar has emerged as an essential component in advanced autonomous driving. However, sparse and noisy radar points often lead to imprecise motion perception, leaving autonomous vehicles with limited sensing capabilities when optical sensors degrade under adverse weather conditions. In this paper, we propose RadarMP, a novel method for precise 3D scene motion perception using low-level radar echo signals from two consecutive frames. Unlike existing methods that separate radar target detection and motion estimation, RadarMP jointly models both tasks in a unified architecture, enabling consistent radar point cloud generation and pointwise 3D scene flow prediction. Tailored to radar characteristics, we design specialized self-supervised loss functions guided by Doppler shifts and echo intensity, effectively supervising spatial and motion consistency without explicit annotations. Extensive experiments on the public dataset demonstrate that RadarMP achieves reliable motion perception across diverse weather and illumination conditions, outperforming radar-based decoupled motion perception pipelines and enhancing perception capabilities for full-scenario autonomous driving systems.

JBHI Journal 2025 Journal Article

Drug Repositioning via Multi-View Representation Learning With Heterogeneous Graph Neural Network

  • Li Peng
  • Cheng Yang
  • Jiahuai Yang
  • Yuan Tu
  • Qingchun Yu
  • Zejun Li
  • Min Chen
  • Wei Liang

Exploring simple and efficient computational methods for drug repositioning has emerged as a popular and compelling topic in the realm of comprehensive drug development. The crux of this technology lies in identifying potential drug-disease associations, which can effectively mitigate the burdens caused by the exorbitant costs and lengthy periods of conventional drugs development. However, existing computational drug repositioning methods continue to encounter challenges in accurately predicting associations between drugs and diseases. In this paper, we propose a Multi-view Representation Learning method (MRLHGNN) with Heterogeneous Graph Neural Network for drug repositioning. This method is based on a collection of data from multiple biological entities associated with drugs or diseases. It consists of a view-specific feature aggregation module with meta-paths and auto multi-view fusion encoder. To better utilize local structural and semantic information from specific views in heterogeneous graph, MRLHGNN employs a feature aggregation model with variable-length meta-paths to expand the local receptive field. Additionally, it utilizes a transformer based semantic aggregation module to aggregate semantic features across different view-specific graphs. Finally, potential drug-disease associations are obtained through a multi-view fusion decoder with an attention mechanism. Cross-validation experiments demonstrate the effectiveness and interpretability of the MRLHGNN in comparison to nine state-of-the-art approaches. Case studies further reveal that MRLHGNN can serve as a powerful tool for drug repositioning.

IROS Conference 2025 Conference Paper

Env-Mani: Quadrupedal Robot Loco-Manipulation with Environment-in-the-Loop

  • Yixuan Li
  • Zan Wang
  • Wei Liang

Dogs can climb onto tables using their front legs for support, enabling them to retrieve objects and significantly expand their workspace by leveraging the external environment. However, the ability of quadrupedal robots to perform similar skills remains largely unexplored. In this work, we introduce a unified, learning-based loco-manipulation framework for quadrupedal robots, allowing them to utilize the external environment as support to extend their workspace and enhance their manipulation capabilities. Specifically, our method proposes a unified policy that takes limited onboard sensors and proprioception as input, generating whole-body actions that enable the robot to manipulate objects. To guide the policy learning for environment-in-the-loop manipulation, we design a set of rewards that address challenges such as imprecise perception and center-of-mass shifts. Additionally, we employ curriculum learning to train both teacher and student policies, ensuring effective skill transfer in complex tasks. We train the policy in simulation and conduct extensive experiments, demonstrating that our approach allows robots to manipulate previously inaccessible objects, opening up new possibilities for enhancing quadrupedal robot capabilities without the need for hardware modifications or additional costs. The project page is available at https://sites.google.com/view/env-mani.

YNIMG Journal 2025 Journal Article

Exploring the impact of APOE ɛ4 on functional connectivity in Alzheimer’s disease across cognitive impairment levels

  • Kangli Dong
  • Wei Liang
  • Ting Hou
  • Zhijie Lu
  • Yixuan Hao
  • Chenrui Li
  • Yue Qiu
  • Nan Kong

The apolipoprotein E (APOE) ɛ4 allele is a recognized genetic risk factor for Alzheimer's Disease (AD). Studies have shown that APOE ɛ4 mediates the modulation of intrinsic functional brain networks in cognitively normal individuals and significantly disrupts the whole-brain topological structure in AD patients. However, how APOE ɛ4 regulates brain functional connectivity (FC) and consequently affects the levels of cognitive impairment in AD patients remains unknown. In this study, we systematically analyzed functional magnetic resonance imaging (fMRI) data from two distinct cohorts: an In-house dataset includes 59 AD patients (73.37 ± 6.42 years), and the ADNI dataset includes 117 AD patients (74.91 ± 7.91 years). Experimental comparisons were conducted by grouping AD patients based on both APOE ɛ4 status and cognitive impairment levels of AD. Network-Based Statistic (NBS) method and the Graph Neural Network Explainer (GNN-Explainer) were combined to identify significant FC changes across different comparisons. Importantly, the GNN-Explainer method was introduced as an enhancement over the NBS method to better model complex high-order nonlinear characteristics for discovering FC features that significantly contribute to classification tasks. The results showed that APOE ɛ4 primarily influenced temporal lobe FCs, while it influenced different cognitive impairment levels of AD by adjusting prefrontal-parietal FCs. These findings were validated by p-values < 0.05 from NBS method, and 5-fold cross-validation along with ablation studies from the GNN-Explainer method. In conclusion, our findings provide new insights into the role of APOE ɛ4 in altering FC dynamics during the progression of AD, highlighting potential targets for early intervention.

AAAI Conference 2025 Conference Paper

FloNa: Floor Plan Guided Embodied Visual Navigation

  • Jiaxin Li
  • Weiqi Huang
  • Zan Wang
  • Wei Liang
  • Huijun Di
  • Feng Liu

Humans naturally rely on floor plans to navigate in unfamiliar environments, as they are readily available, reliable, and provide rich geometrical guidance. However, existing visual navigation settings overlook this valuable prior knowledge, leading to limited efficiency and accuracy. To eliminate this gap, we introduce a novel navigation task: Floor Plan Visual Navigation (FloNa), the first attempt to incorporate floor plans into embodied visual navigation. While the floor plan offers significant advantages, two key challenges emerge: (1) handling the spatial inconsistency between the floor plan and the actual scene layout for collision-free navigation, and (2) aligning observed images with the floor plan sketch despite their distinct modalities. To address these challenges, we propose FloDiff, a novel diffusion policy framework incorporating a localization module to facilitate alignment between the current observation and the floor plan. We further collect 20k navigation episodes across 117 scenes in the iGibson simulator to support the training and evaluation. Extensive experiments demonstrate the effectiveness and efficiency of our framework in unfamiliar scenes using floor plan knowledge.

NeurIPS Conference 2025 Conference Paper

Heterogeneous Adversarial Play in Interactive Environments

  • Manjie Xu
  • Xinyi Yang
  • Jiayu Zhan
  • Wei Liang
  • Chi Zhang
  • Yixin Zhu

Self-play constitutes a fundamental paradigm for autonomous skill acquisition, whereby agents iteratively enhance their capabilities through self-directed environmental exploration. Conventional self-play frameworks exploit agent symmetry within zero-sum competitive settings, yet this approach proves inadequate for open-ended learning scenarios characterized by inherent asymmetry. Human pedagogical systems exemplify asymmetric instructional frameworks wherein educators systematically construct challenges calibrated to individual learners' developmental trajectories. The principal challenge resides in operationalizing these asymmetric, adaptive pedagogical mechanisms within artificial systems capable of autonomously synthesizing appropriate curricula without predetermined task hierarchies. Here we present Heterogeneous Adversarial Play (HAP), an adversarial Automatic Curriculum Learning framework that formalizes teacher-student interactions as a minimax optimization wherein task-generating instructor and problem-solving learner co-evolve through adversarial dynamics. In contrast to prevailing automatic curriculum learning methodologies that employ static curricula or unidirectional task selection mechanisms, HAP establishes a bidirectional feedback system wherein instructors continuously recalibrate task complexity in response to real-time learner performance metrics. Experimental validation across multi-task learning domains demonstrates that our framework achieves performance parity with SOTA baselines while generating curricula that enhance learning efficacy in both artificial agents and human subjects.

JBHI Journal 2025 Journal Article

Human Visual Transduction Mechanism-Inspired Adaptive Enhancement Network during Endoscopy

  • Xinzhen Ren
  • Wenju Zhou
  • Maoyu Jin
  • Wei Liang
  • Yiyou Gao
  • Shengqiang Cai

Disposable endoscopes are increasingly popular due to the lower risk of cross-contamination. Clearer endoscopes have shown the ability to improve adenoma detection rate. However, due to cost constraints, disposable endoscopes are equipped with smaller sensor arrays than reusable ones, resulting in lower resolution and degraded performance in both subjective and automated diagnostics. In this paper, an adaptive enhancement network (AEN) is proposed to recover the details of the lesion region, thereby assisting endoscopists during endoscopy. It combines the physiological mechanism of the human eyes with artificial intelligence. The framework of the AEN is analogous to the human visual transduction mechanism, which contains photoreceptor outer segments & outer nuclear layer, outer plexiform layer, inner nuclear layer, inner plexiform layer, and ganglion cell layer. The first layer includes the rod cell module (RCM) and the cone cell module (CCM). Rod cells are shape-sensitive and fast, while cone cells are sensitive to color and detail. Therefore, the RCM is designed to segment lesion regions and the CCM is designed to perform enhancement on them. The outer plexiform layer transmits the information from RCM to CCM. The inner nuclear layer refines boundaries according to the poisson equation with the information from the inner plexiform layer. Ultimately, a clearer image of the lesion regions is generated in the ganglion cell layer and transmitted to the brain (display). Experiments demonstrate that the AEN can effectively enhance lesion regions with much sharper edges, finer texture details, fewer blurs and artifacts. Comprehensive evaluations on four datasets further illustrate its generalizability. The AEN offers advantages in real-time processing.

JBHI Journal 2025 Journal Article

Multiscale Spatial-Temporal Feature Fusion Neural Network for Motor Imagery Brain-Computer Interfaces

  • Jing Jin
  • Weijie Chen
  • Ren Xu
  • Wei Liang
  • Xiao Wu
  • Xinjie He
  • Xingyu Wang
  • Andrzej Cichocki

Motor imagery, one of the main brain-computer interface (BCI) paradigms, has been extensively utilized in numerous BCI applications, such as the interaction between disabled people and external devices. Precise decoding, one of the most significant aspects of realizing efficient and stable interaction, has received a great deal of intensive research. However, the current decoding methods based on deep learning are still dominated by single-scale serial convolution, which leads to insufficient extraction of abundant information from motor imagery signals. To overcome such challenges, we propose a new end-to-end convolutional neural network based on multiscale spatial-temporal feature fusion (MSTFNet) for EEG classification of motor imagery. The architecture of MSTFNet consists of four distinct modules: feature enhancement module, multiscale temporal feature extraction module, spatial feature extraction module and feature fusion module, with the latter being further divided into the depthwise separable convolution block and efficient channel attention block. Moreover, we implement a straightforward yet potent data augmentation strategy to bolster the performance of MSTFNet significantly. To validate the performance of MSTFNet, we conduct cross-session experiments and leave-one-subject-out experiments. The cross-session experiment is conducted across two public datasets and one laboratory dataset. On the public datasets of BCI Competition IV 2a and BCI Competition IV 2b, MSTFNet achieves classification accuracies of 83. 62% and 89. 26%, respectively. On the laboratory dataset, MSTFNet achieves 86. 68% classification accuracy. Besides, the leave-one-subject-out experiment is performed on the BCI Competition IV 2a dataset, and MSTFNet achieves 66. 31% classification accuracy. These experimental results outperform several state-of-the-art methodologies, indicate the proposed MSTFNet's robust capability in decoding EEG signals associated with motor imagery.

JBHI Journal 2024 Journal Article

A Computational Framework for Predicting Novel Drug Indications Using Graph Convolutional Network With Contrastive Learning

  • Yuxun Luo
  • Wenyu Shan
  • Li Peng
  • Lingyun Luo
  • Pingjian Ding
  • Wei Liang

Inferring potential drug indications plays a vital role in the drug discovery process. It can be time-consuming and costly to discover novel drug indications through biological experiments. Recently, graph learning-based methods have gained popularity for this task. These methods typically treat the prediction task as a binary classification problem, focusing on modeling associations between drugs and diseases within a graph. However, labeled data for drug indication prediction is often limited and expensive to acquire. Contrastive learning addresses this challenge by aligning similar drug-disease pairs and separating dissimilar pairs in the embedding space. Thus, we developed a model called DrIGCL for drug indication prediction, which utilizes graph convolutional networks and contrastive learning. DrIGCL incorporates drug structure, disease comorbidities, and known drug indications to extract representations of drugs and diseases. By combining contrastive and classification losses, DrIGCL predicts drug indications effectively. In multiple runs of hold-out validation experiments, DrIGCL consistently outperformed existing computational methods for drug indication prediction, particularly in terms of top-k. Furthermore, our ablation study has demonstrated a significant improvement in the predictive capabilities of our model when utilizing contrastive learning. Finally, we validated the practical usefulness of DrIGCL by examining the predicted novel indications of Aspirin.

AIIM Journal 2024 Journal Article

FA-Net: A hierarchical feature fusion and interactive attention-based network for dose prediction in liver cancer patients

  • Miao Liao
  • Shuanhu Di
  • Yuqian Zhao
  • Wei Liang
  • Zhen Yang

Dose prediction is a crucial step in automated radiotherapy planning for liver cancer. Several deep learning-based approaches for dose prediction have been proposed to enhance the design efficiency and quality of radiotherapy plan. However, these approaches usually take CT images and contours of organs at risk (OARs) and planning target volume (PTV) as a multi-channel input and is thus difficult to extract sufficient feature information from each input, which results in unsatisfactory dose distribution. In this paper, we propose a novel dose prediction network for liver cancer based on hierarchical feature fusion and interactive attention. A feature extraction module is first constructed to extract multi-scale features from different inputs, and a hierarchical feature fusion module is then built to fuse these multi-scale features hierarchically. A decoder based on attention mechanism is designed to gradually reconstruct the fused features into dose distribution. Additionally, we design an autoencoder network to generate a perceptual loss during training stage, which is used to improve the accuracy of dose prediction. The proposed method is tested on private clinical dataset and obtains HI and CI of 0. 31 and 0. 87, respectively. The experimental results are better than those by several existing methods, indicating that the dose distribution generated by the proposed method is close to that approved in clinics. The codes are available at https: //github. com/hired-ld/FA-Net.

EAAI Journal 2024 Journal Article

Integrated optimization and engineering application for disassembly line balancing problem with preventive maintenance

  • Yanqing Zeng
  • Zeqiang Zhang
  • Tengfei Wu
  • Wei Liang

The current research on the disassembly line balancing (DLB) problem has not yet considered the preventive maintenance of workstation equipment. The lack of preventive maintenance can easily lead to sudden failure of the disassembly line equipment, which will affect the disassembly efficiency. Preventive maintenance is planned maintenance, which can effectively reduce the probability of equipment failure in actual disassembly line engineering applications. This study integrates preventive maintenance into the DLB. The objectives focus on optimizing both cycle times under the regular DLB scenario and the preventive maintenance DLB scenario, and the number of task adjustments between the two scenarios. A mixed integer linear programming model integrating preventive maintenance is established. Then an improved genetic simulated annealing (IGSA) algorithm based on the disassembly characteristics is designed. The exact solver and the proposed IGSA are used to solve a small-scale case simultaneously to verify the correctness of the model and algorithm. The performance of the proposed IGSA is verified by comparing the solution results of 21 benchmarks. Finally, the proposed model and algorithm are applied to a microwave disassembly line as the engineering application case. The results show that the integrated optimization can ensure the disassembly efficiency of the regular disassembly scenario, and can fully utilize workstations of unperformed maintenance during the preventive maintenance scenario. This can also effectively improve workstation utilization and reduce time costs.

ICML Conference 2024 Conference Paper

Prompt-guided Precise Audio Editing with Diffusion Models

  • Manjie Xu
  • Chenxing Li
  • Duzhen Zhang
  • Dan Su 0002
  • Wei Liang
  • Dong Yu 0001

Audio editing involves the arbitrary manipulation of audio content through precise control. Although text-guided diffusion models have made significant advancements in text-to-audio generation, they still face challenges in finding a flexible and precise way to modify target events within an audio track. We present a novel approach, referred to as PPAE, which serves as a general module for diffusion models and enables precise audio editing. The editing is based on the input textual prompt only and is entirely training-free. We exploit the cross-attention maps of diffusion models to facilitate accurate local editing and employ a hierarchical local-global pipeline to ensure a smoother editing process. Experimental results highlight the effectiveness of our method in various editing tasks.

NeurIPS Conference 2023 Conference Paper

Active Reasoning in an Open-World Environment

  • Manjie Xu
  • Guangyuan Jiang
  • Wei Liang
  • Chi Zhang
  • Yixin Zhu

Recent advances in vision-language learning have achieved notable success on complete-information question-answering datasets through the integration of extensive world knowledge. Yet, most models operate passively, responding to questions based on pre-stored knowledge. In stark contrast, humans possess the ability to actively explore, accumulate, and reason using both newfound and existing information to tackle incomplete-information questions. In response to this gap, we introduce Conan, an interactive open-world environment devised for the assessment of active reasoning. Conan facilitates active exploration and promotes multi-round abductive inference, reminiscent of rich, open-world settings like Minecraft. Diverging from previous works that lean primarily on single-round deduction via instruction following, Conan compels agents to actively interact with their surroundings, amalgamating new evidence with prior knowledge to elucidate events from incomplete observations. Our analysis on \bench underscores the shortcomings of contemporary state-of-the-art models in active exploration and understanding complex scenarios. Additionally, we explore Abduction from Deduction, where agents harness Bayesian rules to recast the challenge of abduction as a deductive process. Through Conan, we aim to galvanize advancements in active reasoning and set the stage for the next generation of artificial intelligence agents adept at dynamically engaging in environments.

NeurIPS Conference 2023 Conference Paper

Interactive Visual Reasoning under Uncertainty

  • Manjie Xu
  • Guangyuan Jiang
  • Wei Liang
  • Chi Zhang
  • Yixin Zhu

One of the fundamental cognitive abilities of humans is to quickly resolve uncertainty by generating hypotheses and testing them via active trials. Encountering a novel phenomenon accompanied by ambiguous cause-effect relationships, humans make hypotheses against data, conduct inferences from observation, test their theory via experimentation, and correct the proposition if inconsistency arises. These iterative processes persist until the underlying mechanism becomes clear. In this work, we devise the IVRE (pronounced as "ivory" ) environment for evaluating artificial agents' reasoning ability under uncertainty. IVRE is an interactive environment featuring rich scenarios centered around Blicket detection. Agents in IVRE are placed into environments with various ambiguous action-effect pairs and asked to determine each object's role. They are encouraged to propose effective and efficient experiments to validate their hypotheses based on observations and actively gather new information. The game ends when all uncertainties are resolved or the maximum number of trials is consumed. By evaluating modern artificial agents in IVRE, we notice a clear failure of today's learning methods compared to humans. Such inefficacy in interactive reasoning ability under uncertainty calls for future research in building human-like intelligence.

EAAI Journal 2023 Journal Article

LACN: A lightweight attention-guided ConvNeXt network for low-light image enhancement

  • Saijie Fan
  • Wei Liang
  • Derui Ding
  • Hui Yu

Images captured under low-light conditions usually have poor visual quality, and hence greatly reduce the accuracy of subsequent tasks such as image segmentation and detection. In the low-light image enhancement task, noises in the dark areas are generally amplified while the images’ brightness is enhanced. It should be pointed out that many deep learning methods cannot effectively suppress the noise at this stage and capture important feature information. To address the above problem, this paper proposes a Lightweight Attention-guided ConvNeXt Network (LACN) for low-light image enhancement. A novel Attention ConvNeXt Module (ACM) is first proposed by introducing a parameter-free attention module (i. e. SimAM) into the ConvNeXt backbone network. Then, a nontrivial lightweight network LACN based on a multi-attention mechanism is established through stacking two ACMs and fusing their features. In what follows, an improved hybrid attention mechanism, Selective Kernel Attention Module (SKAM), is adopted to effectively extract both global and local information. Such a module realizes the evaluation of lighting conditions for the whole image and the adaptive adjustment of the receptive field. Finally, through the feature fusion module, the features of different stages are aggregated to improve the ability of network to retain color information. Numerous experiments on low-light image enhancement are implemented via comparison with other state-of-the-art methods. Experiments show that the proposed method significantly improves the brightness and contrast of low-illumination images, preserves color information, and suppresses the generation of noises after image brightening.

TCS Journal 2022 Journal Article

Approximation algorithm for prize-collecting sweep cover with base stations

  • Wei Liang
  • Zhao Zhang

In a sweep cover problem, positions of interest (PoIs) are required to be visited periodically by mobile sensors. In this paper, we propose a new sweep cover problem: the prize-collecting sweep cover problem (PCSC), in which penalty is incurred by those PoIs which are not sweep-covered, and the goal is to minimize the covering cost plus the penalty. Assuming that every mobile sensor has to be linked to some base station, and the number of base stations is upper bounded by a constant, we present a 5-LMP (Lagrangian Multiplier Preserving) algorithm. As a step stone, we propose the prize-collecting forest with k components problem (PCF k ), which might be interesting in its own sense, and presented a 2-LMP for rooted PCF k.

NeurIPS Conference 2022 Conference Paper

HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes

  • Zan Wang
  • Yixin Chen
  • Tengyu Liu
  • Yixin Zhu
  • Wei Liang
  • Siyuan Huang

Learning to generate diverse scene-aware and goal-oriented human motions in 3D scenes remains challenging due to the mediocre characters of the existing datasets on Human-Scene Interaction (HSI); they only have limited scale/quality and lack semantics. To fill in the gap, we propose a large-scale and semantic-rich synthetic HSI dataset, denoted as HUMANISE, by aligning the captured human motion sequences with various 3D indoor scenes. We automatically annotate the aligned motions with language descriptions that depict the action and the individual interacting objects; e. g. , sit on the armchair near the desk. HUMANIZE thus enables a new generation task, language-conditioned human motion generation in 3D scenes. The proposed task is challenging as it requires joint modeling of the 3D scene, human motion, and natural language. To tackle this task, we present a novel scene-and-language conditioned generative model that can produce 3D human motions of the desirable action interacting with the specified objects. Our experiments demonstrate that our model generates diverse and semantically consistent human motions in 3D scenes.

NeurIPS Conference 2022 Conference Paper

Towards Versatile Embodied Navigation

  • Hanqing Wang
  • Wei Liang
  • Luc V Gool
  • Wenguan Wang

With the emergence of varied visual navigation tasks (e. g. , image-/object-/audio-goal and vision-language navigation) that specify the target in different ways, the community has made appealing advances in training specialized agents capable of handling individual navigation tasks well. Given plenty of embodied navigation tasks and task-specific solutions, we address a more fundamental question: can we learn a single powerful agent that masters not one but multiple navigation tasks concurrently? First, we propose VXN, a large-scale 3D dataset that instantiates~four classic navigation tasks in standardized, continuous, and audiovisual-rich environments. Second, we propose Vienna, a versatile embodied navigation agent that simultaneously learns to perform the four navigation tasks with one model. Building upon a full-attentive architecture, Vienna formulates various navigation tasks as a unified, parse-and-query procedure: the target description, augmented with four task embeddings, is comprehensively interpreted into a set of diversified goal vectors, which are refined as the navigation progresses, and used as queries to retrieve supportive context from episodic history for decision making. This enables the reuse of knowledge across navigation tasks with varying input domains/modalities. We empirically demonstrate that, compared with learning each visual navigation task individually, our multitask agent achieves comparable or even better performance with reduced complexity.

TCS Journal 2021 Journal Article

Minimum power partial multi-cover on a line

  • Wei Liang
  • Menghong Li
  • Zhao Zhang
  • Xiaohui Huang

This paper studies the minimum power partial multi-cover problem on a line (MinPowPMC-Line), the goal of which is to find an assignment of powers to sensors such that at least a required number of points are covered up to their covering requirements. We first present an LP method to show that the minimum power multi-cover problem on a line (without partial covering requirement) is solvable in polynomial time. But this method no longer works when facing partial covering requirement. We turn to dynamic programming method to find an optimal solution for MinPowPMC-Line in time O ( n 4 m 1 + 2 ( c r m a x ) ), where n, m are the number of points and the number of sensors, respectively, and c r m a x denotes the maximum covering requirement of elements. So, this problem is polynomial-time solvable when c r max is upper bounded by a constant.

AAAI Conference 2019 Conference Paper

3D Face Synthesis Driven by Personality Impression

  • Yining Lang
  • Wei Liang
  • Yujia Wang
  • Lap-Fai Yu

Synthesizing 3D faces that give certain personality impressions is commonly needed in computer games, animations, and virtual world applications for producing realistic virtual characters. In this paper, we propose a novel approach to synthesize 3D faces based on personality impression for creating virtual characters. Our approach consists of two major steps. In the first step, we train classifiers using deep convolutional neural networks on a dataset of images with personality impression annotations, which are capable of predicting the personality impression of a face. In the second step, given a 3D face and a desired personality impression type as user inputs, our approach optimizes the facial details against the trained classifiers, so as to synthesize a face which gives the desired personality impression. We demonstrate our approach for synthesizing 3D faces giving desired personality impressions on a variety of 3D face models. Perceptual studies show that the perceived personality impressions of the synthesized faces agree with the target personality impressions specified for synthesizing the faces.

AAAI Conference 2019 Conference Paper

Deep Single-View 3D Object Reconstruction with Visual Hull Embedding

  • Hanqing Wang
  • Jiaolong Yang
  • Wei Liang
  • Xin Tong

3D object reconstruction is a fundamental task of many robotics and AI problems. With the aid of deep convolutional neural networks (CNNs), 3D object reconstruction has witnessed a significant progress in recent years. However, possibly due to the prohibitively high dimension of the 3D object space, the results from deep CNNs are often prone to missing some shape details. In this paper, we present an approach which aims to preserve more shape details and improve the reconstruction quality. The key idea of our method is to leverage object mask and pose estimation from CNNs to assist the 3D shape learning by constructing a probabilistic singleview visual hull inside of the network. Our method works by first predicting a coarse shape as well as the object pose and silhouette using CNNs, followed by a novel 3D refinement CNN which refines the coarse shapes using the constructed probabilistic visual hulls. Experiment on both synthetic data and real images show that embedding a single-view visual hull for shape refinement can significantly improve the reconstruction quality by recovering more shapes details and improving shape consistency with the input image.

AAAI Conference 2018 Conference Paper

Tracking Occluded Objects and Recovering Incomplete Trajectories by Reasoning About Containment Relations and Human Actions

  • Wei Liang
  • Yixin Zhu
  • Song-Chun Zhu

This paper studies a challenging problem of tracking severely occluded objects in long video sequences. The proposed method reasons about the containment relations and human actions, thus infers and recovers occluded objects identities while contained or blocked by others. There are two conditions that lead to incomplete trajectories: i) Contained. The occlusion is caused by a containment relation formed between two objects, e. g. , an unobserved laptop inside a backpack forms containment relation between the laptop and the backpack. ii) Blocked. The occlusion is caused by other objects blocking the view from certain locations, during which the containment relation does not change. By explicitly distinguishing these two causes of occlusions, the proposed algorithm formulates tracking problem as a network flow representation encoding containment relations and their changes. By assuming all the occlusions are not spontaneously happened but only triggered by human actions, an MAP inference is applied to jointly interpret the trajectory of an object by detection in space and human actions in time. To quantitatively evaluate our algorithm, we collect a new occluded object dataset captured by Kinect sensor, including a set of RGB-D videos and human skeletons with multiple actors, various objects, and different changes of containment relations. In the experiments, we show that the proposed method demonstrates better performance on tracking occluded objects compared with baseline methods.

JBHI Journal 2018 Journal Article

Wearable Heading Estimation for Motion Tracking in Health Care by Adaptive Fusion of Visual–Inertial Measurements

  • Yinlong Zhang
  • Wei Liang
  • Hongsheng He
  • Jindong Tan

The increasing demand for health informatics has become a far-reaching trend in the aging society. The utilization of wearable sensors enables monitoring senior people daily activities in free-living environments, conveniently and effectively. Among the primary health-care sensing categories, the wearable visual–inertial modality for human motion tracking, gradually exerts promising potentials. In this paper, we present a novel wearable heading estimation strategy to track the movements of human limbs. It adaptively fuses inertial measurements with visual features following locality constraints. Body movements are classified into two types: general motion (which consists of both rotation and translation) or degenerate motion (which consists of only rotation). A specific number of feature correspondences between camera frames are adaptively chosen to satisfy both the feature descriptor similarity constraint and the locality constraint. The selected feature correspondences and inertial quaternions are employed to calculate the initial pose, followed by the coarse-to-fine procedure to iteratively remove visual outliers. Eventually, the ultimate heading is optimized using the correct feature matches. The proposed method has been thoroughly evaluated on the straight-line, rotatory, and ambulatory movement scenarios. As the system is lightweight and requires small computational resources, it enables effective and unobtrusive human motion monitoring, especially for the senior citizens in the long-term rehabilitation.

IJCAI Conference 2016 Conference Paper

What Is Where: Inferring Containment Relations from Videos

  • Wei Liang
  • Yibiao Zhao
  • Yixin Zhu
  • Song-Chun Zhu

In this paper, we present a probabilistic approach to explicitly infer containment relations between objects in 3D scenes. Given an input RGB-D video, our algorithm quantizes the perceptual space of a 3D scene by reasoning about containment relations over time. At each frame, we represent the containment relations in space by a containment graph, where each vertex represents an object and each edge represents a containment relation. We assume that human actions are the only cause that leads to containment relation changes over time, and classify human actions into four types of events: move-in, move-out, no change and paranormal-change. Here, paranormal-change refers to the events that are physically infeasible, and thus are ruled out through reasoning. A dynamic programming algorithm is adopted to finding both the optimal sequence of containment relations across the video, and the containment relation changes between adjacent frames. We evaluate the proposed method on our dataset with 1326 video clips taken in 9 indoor scenes, including some challenging cases, such as heavy occlusions and diverse changes of containment relations. The experimental results demonstrate good performance on the dataset.

EAAI Journal 2012 Journal Article

Assessing and classifying risk of pipeline third-party interference based on fault tree and SOM

  • Wei Liang
  • Jinqiu Hu
  • Laibin Zhang
  • Cunjie Guo
  • Weipeng Lin

Accidents to pipelines because of the third-party interference have been recorded and they often result in catastrophic consequences for environment and society with a great deal of economic loss. The third-party interference resulting from complicated origins occurs randomly, and is hard to be forecasted or controlled in advance, so it becomes a serious threat to the safe operation of long transmission pipeline. This paper focuses on the application of self-organizing maps (SOMs) to assess the risk of third-party interference and classify their risk patterns. In this work, fault tree is used first to establish the risk assessment index system, and then SOM is used in multi-parameter risk pattern classification approach, which is proposed to present various risk maps, incorporating the factors of pipeline laying conditions, historical damage records, safety-related actions, management measures and the environment around the underling pipeline. A field case study of Shaanxi–Beijing gas pipeline in China is undertaken so that the effectiveness of the proposed approach could be verified. By taking the classification results into consideration, the decision maker may well get precious and differentiated information about the pipeline risk distribution of third-party interference and make appropriate safety-related actions to prevent the damage.

v2026.09.13