Arrow Research search

Author name cluster

Rui Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

124 papers
2 author rows

Possible papers

124

EAAI Journal 2026 Journal Article

A multi-view collaborative heterogeneous graph neural network with semantic- and relation-aware for drug-disease association prediction

  • Jinzhou Wu
  • Donglin He
  • Xin Li
  • Rui Wang
  • Yujuan Zhang

Drug repositioning (DR) is crucial for accelerating drug development and reducing costs; computational methods offer efficient alternatives to costly traditional methods. However, existing methods rely on static linear operations for integrating multi-source similarity networks, causing noise accumulation and redundancy. Additionally, meta-path aggregation often fails to dynamically balance intra-path interactions and cross-path heterogeneity, limiting multi-granularity information fusion. To address these, we propose a Multi-view Collaborative Heterogeneous Graph Neural Network with Semantic and Relation-Aware (MSRHGNN) for Drug-Disease Association (DDA) prediction. MSRHGNN jointly encodes similarity and heterogeneous biological network features. It first uses an adaptive dynamic fusion mechanism to integrate multi-source similarity data, leveraging a Graph Transformer to capture richer structural features. Second, within the heterogeneous biological network, the low-order relational view aggregates first-order neighborhood information to capture local topology, while the high-order relational view designs a dual attention collaborative aggregation mechanism: node-level attention driven by central anchor points highlights key interactions within paths, and semantic and relation-aware mechanisms at the cross-path level quantify the consistency within paths and heterogeneity between paths, achieving complementary dynamic fusion. Additionally, MSRHGNN aligns node- and graph-level representations through a multi-view contrastive learning strategy and employs a multi-task balance strategy to alleviate gradient conflicts between the main and auxiliary tasks. Experimental results demonstrate that our method outperforms other baseline models across multiple evaluation metrics on several public datasets, with its stability and robustness validated under cold-start, imbalanced, and noisy conditions. Furthermore, case studies on specific diseases and molecular docking experiments highlight its potential value in practical drug discovery.

EAAI Journal 2026 Journal Article

A novel cognitive diagnostic network with color and spatial cues for skin disease recognition

  • Ming Ju
  • Fei Wang
  • Yao Huang
  • Rui Wang
  • Haiquan Wang
  • Chunhua Qian

Early recognition and diagnosis of skin diseases are crucial for subsequent treatment. However, current methods struggle to precisely recognize skin disease images with high interclass similarity and intraclass variability, and overlook the occult nature of skin lesions and the limitations of discriminative information mining. To tackle the aforementioned issues, we present a novel cognitive diagnostic network with color and spatial texture cues. This network aims to emulate the visual perception and decision-making processes involved in medical diagnosis. Specifically, we developed a color semantic diagnosis cue module to extract deep color semantic information from the Horizontal–Vertical-Intensity (HVI) domain. The proposed spatial-texture cue extraction process includes a local texture-shape perception enhancement module and a global perception module. These modules enhance local feature perception and extract global information from spatial and channel dimensions to improve global spatial distribution feature discrimination. Furthermore, we designed a color-spatial dynamic fusion module to effectively condense and synthesize the two types of cues. Employing the concept of multi-expert joint diagnosis, different cues are independently classified, and decision-making fusion is performed. Meanwhile, interclass separation and intraclass consistency loss functions are designed to prevent the diversity and similarity of skin diseases from confusing cognitive decisions. Comparative experiments on public and real-world clinical datasets validate the superiority of our method, offering valuable support for clinical diagnosis.

EAAI Journal 2026 Journal Article

A novel U-Net-physical informed neural network model for load forecasting of hydrogen power boat power system combined multi-source information spatiotemporal matrix construction

  • Xingdou Liu
  • Liang Zou
  • Zhiyun Han
  • Jundao Jiang
  • Yawei Wang
  • Rui Wang

The intelligent energy management of hydrogen powered boat (HPB) highly relies on the accurate prediction ability of power load, which directly determines the endurance performance and operational economy of the vessel. However, the power system load prediction (PSLP) method for roadbeds often fails to improve prediction accuracy when dealing with sudden changes in navigation conditions and hydrological disturbances. This is because it does not integrate the nonlinear interaction mechanisms of multi-source heterogeneous data and does not sufficiently consider the physical characteristics of the time-varying dynamics of boat power systems. This paper proposes a novel U-Net-physical informed neural network (U-Net-PINN) for multi-step PSLP in HPB. The encoder-decoder architecture of U-Net innovatively employs a spatial transformer network (STN) to integrate the spatial attention (SA) mechanism into the joint learning of spatio-temporal features. First, the improved complete ensemble empirical mode decomposition with adaptive noise (ICEEMDAN) and high-dimensional phase space reconstruction (HDPSR) are used to process HPB operational data and hydrometeorological data, constructing a 2 dimensions spatial-temporal feature matrix from multi-source heterogeneous time series data. Subsequently, a chaotic system was constructed using the Multidimensional Continuous Lorentz Equations (MDCLE) to model the interaction between electrical parameters and hydrometeorological environmental parameters, serving as the physical constraint term for the loss function of the U-Net-PINN architecture. Finally, the proposed U-Net-PINN was used for multi-step PSLP of HPB navigation processes. The proposed model addresses the challenge of integrating multi-source data information during HPB operation. The article uses mean absolute percentage error (MAPE), normalized root mean square error (NRMSE), and coefficient of determination (R2) as three basic average indicators, which improve the comparison model by 8 %–45 % and 3 %–25 % under cruising and working conditions, respectively. Simultaneously using hit rate (HR) to evaluate the probability of large errors occurring, and increasing the performance gap between the Diebold Mariano test comprehensive evaluation benchmark model and the proposed model. The effectiveness of the proposed method has been demonstrated by the HPB dataset operating on the Yangtze River channel. The results show that the proposed U-Net-PINN is significantly superior to the other nine advanced comparison models, and the physical constraint term accelerates the convergence of the loss function while significantly reducing the risk of large errors.

AAMAS Conference 2026 Conference Paper

Better Goals, Better Policies: LLM-Driven Relabeling for Offline Goal-Conditioned Reinforcement Learning

  • Xule Gao
  • Chuxiong Sun
  • Rui Wang
  • Changwen Zheng

Offline goal-conditioned reinforcement learning (offline GCRL) learns goal-conditioned policies from fixed, reward-free datasets. Existing methods often rely on hindsight experience replay(HER), which treats all future states uniformly, leading to many uninformative goals. We propose LLM-Driven Relabeling, an adaptive and trustworthy framework that uses LLM-generated semantic rules to identify task-relevant key states and prioritize them as informative relabeling goals. We theoretically demonstrate that such goals lead to larger TD errors, thereby reducing sample complexity. Empirical results on offline GCRL benchmarks demonstrate that LLM-Driven Relabeling significantly improves learning efficiency, particularly under reduced-data conditions.

JBHI Journal 2026 Journal Article

Can Information Representations Inspired by the Human Auditory Perception Benefit Computer Audition-Based Disease Detection? An Interpretable Comparative Study

  • Zhihua Wang
  • Haojie Zhang
  • Yang Tan
  • Rui Wang
  • Kun Qian
  • Bin Hu
  • Yoshiharu Yamamoto
  • Björn W. Schuller

Computer audition-based methods have attracted a great deal of attention in the field of disease detection due to their significant advantages, e. g. , non-invasive and convenient operation. Among them, the introduction of information representations inspired by human auditory perception, e. g. , Mel-frequency transformation, gives it great potential to approach and even exceed the limits of the human auditory system. However, according to previous research, it remains challenging to fairly assess whether information representations inspired by human auditory perception have a significant positive effect on disease detection. Moreover, performance differences among various information representations and their underlying causes are yet to be thoroughly investigated and analyzed. To this end, we propose an interpretable comparative study on information representations inspired by human auditory perception for disease detection. First, the detection accuracy of different information representations are investigated on two sound datasets (a psychological and a physiological disease) based on the classical model and the proposed Temporal-Spatial Multi-Scale Perception Network. Then, the noise robustness of these information representations are compared by introducing Gaussian noise with varying signal-to-noise ratios (SNRs). Finally, by combining the human auditory perception mechanism and explainable AI techniques, we analyze the reasons for performance differences among various information representations from qualitative and quantitative perspectives. Experimental results demonstrate that information representations inspired by human auditory perception can improve the performance of disease detection with statistical significance. Furthermore, Gammatone Frequency Cepstral Coefficients (GFCCs) outperform other information representations by achieving the highest accuracy, particularly under noisy conditions. The interpretable results further reveal the underlying reasons for GFCC's superior performance, highlighting its ability to capture critical auditory features robustly across varying noise levels. These findings emphasize the potential of auditory perception-inspired representations in advancing computer audition-based disease detection systems and provide a solid foundation for future research in this domain.

AAAI Conference 2026 Conference Paper

Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models

  • Fei Song
  • Yi Li
  • Rui Wang
  • Jiahuan Zhou
  • Changwen Zheng
  • Jiangmeng Li

Test-time prompt tuning for vision-language models has demonstrated impressive generalization capabilities under zero-shot settings. However, tuning the learnable prompts solely based on unlabeled test data may induce prompt optimization bias, ultimately leading to suboptimal performance on downstream tasks. In this work, we analyze the underlying causes of prompt optimization bias from both the model and data perspectives. In terms of the model, the entropy minimization objective typically focuses on reducing the entropy of model predictions while overlooking their correctness. This can result in overconfident yet incorrect outputs, thereby compromising the quality of prompt optimization. On the data side, prompts affected by optimization bias can introduce misalignment between visual and textual modalities, which further aggravates the prompt optimization bias. To this end, we propose a Doubly Debiased Test-Time Prompt Tuning method, abbreviated as D2TPT. Specifically, we first introduce a dynamic retrieval-augmented modulation module that retrieves high-confidence knowledge from a dynamic knowledge base using the test image feature as a query, and uses the retrieved knowledge to modulate the predictions. Guided by the refined predictions, we further develop a reliability-aware prompt optimization module that incorporates a confidence-based weighted ensemble and cross-modal consistency distillation to impose regularization constraints during prompt tuning. Extensive experiments across 15 benchmark datasets involving both natural distribution shifts and cross-datasets generalization demonstrate that D2TPT outperforms baselines, validating its effectiveness in mitigating prompt optimization bias.

AAMAS Conference 2026 Conference Paper

DR2: Revisiting Visual Reinforcement Learning from the Dimensional Analysis Perspective

  • Chuxiong Sun
  • Jinli Chen
  • Zehua Zang
  • Jiangmeng Li
  • Rui Wang
  • Changwen Zheng

Despite impressive progress on visual control challenges, visual reinforcement learning (VRL) remains sample-inefficient. Existing work commonly leverages auxiliary objectives and data augmentation to learn discriminative representations from observation space containing redundant and task-irrelevant information. However, our analysis shows that the learned representation space still contain dimensional redundancy and dimensional confounders, impeding policy learning. To address these problems, we introduce DR2, a simple plug-and-play module that first constructs a redundancy-reduced representation space and then identifies dimensions most critical for decision making. Concretely, DR2 first applies a redundancy-reduction regularizer to decorrelate latent dimensions, then learns a dimensional mask that models each dimension’s gradient contribution to policy learning, dynamically downweighting task-irrelevant confounders during training. Across diversevisualcontrolbenchmarks, DR2consistentlyimprovessample efficiency and generalization over state-of-the-art baselines. These results indicate that addressing redundancy and confounding at the representation level provides a complementary—rather than substitutive—benefit to existing augmentation and self-supervised strategies.

AAAI Conference 2026 Conference Paper

False Positives Matter: Multidimensional Localization Evaluation and Training-Free Explainable Adversarial Patch Defense

  • Lihua Jing
  • Rui Wang
  • Jinwen Zhong
  • Runbo Li
  • Zixuan Zhu

Adversarial patch attacks pose a significant threat to visual systems. While current patch purification-based defense methods enhance core metrics of visual perception models, they overlook the critical issue of false positive patches, severely compromising image usability. This paper reveals the inadequacy of existing evaluations for adversarial patch defenses, and pioneers a multidimensional adversarial patch localization evaluation framework, which comprehensively quantifies false positives, recall capability, and overall localization accuracy, providing a novel perspective for comparative analysis within the field. Furthermore, building upon the observation that false positives stem from a lack of semantic understanding, we propose a Semantic-Aware Training-free Explainable Defense method (SATED). SATED achieves zero-shot patch localization, false detection correction, and decision explanation by constructing a patch reasoning chain, while simultaneously performing integrated text-guided patch inpainting. Extensive experiments across digital and physical scenarios, detection and segmentation tasks, and diverse adversarial patches, demonstrate that our method significantly reduces false positives and doubles the overall patch localization accuracy, boosting both the generalizability and explainability of the defense.

EAAI Journal 2026 Journal Article

How recycling technologies play stage-specific roles in renewable energy and energy storage systems? Insights from a patent claim analysis

  • Xuefeng Zhao
  • Qianwen Hao
  • Wei Zhang
  • Rui Wang
  • Chengjiang Li

The continued expansion of Renewable Energy and Energy Storage Systems (REESS) has led to increasing challenges related to material consumption and environmental sustainability. Recycling technologies (RT) have emerged as key enablers of resource efficiency and circularity. However, the stage-specific contributions of different RT types within REESS remain poorly understood. This study employs a large language model (LLM) to generate search terms for constructing corpora and classifying patents, subdivides patent subsets, builds a Type & Dependency mechanism for claim analysis, and uses three analytical approaches to elucidate the evolving role of RTs in REESS development. The results reveal three main findings: (1) RTs display stage-specific application patterns, with Chemical RTs dominating the control stage, while Biological RTs remain limited in the generation stage; (2) Each RT type plays a distinct role across stages. Chemical RTs show the strongest impact, Physical RTs offer stable support, and Biological RTs remain emerging with limited but growing potential; (3) RTs exhibit increasing interconnectivity across REESS stages, indicating a shift toward more integrated and circular technological development. These findings contribute to a deeper understanding of RT integration in REESS and offer valuable implications for advancing sustainable energy systems.

AAAI Conference 2026 Conference Paper

LPPG-RL: Lexicographically Projected Policy Gradient Reinforcement Learning with Subproblem Exploration

  • Ruiyu Qiu
  • Rui Wang
  • Guanghui Yang
  • Xiang Li
  • Zhijiang Shao

Lexicographic multi-objective problems, which consist of multiple conflicting subtasks with explicit priorities, are common in real-world applications. Despite the advantages of Reinforcement Learning (RL) in single tasks, extending conventional RL methods to prioritized multiple objectives remains challenging. In particular, traditional Safe RL and Multi-Objective RL (MORL) methods have difficulty enforcing priority orderings efficiently. Therefore, Lexicographic Multi-Objective RL (LMORL) methods have been developed to address these challenges. However, existing LMORL methods either rely on heuristic threshold tuning with prior knowledge or are restricted to discrete domains. To overcome these limitations, we propose Lexicographically Projected Policy Gradient RL (LPPG-RL), a novel LMORL framework which leverages sequential gradient projections to identify feasible policy update directions, thereby enabling LPPG-RL broadly compatible with all policy gradient algorithms in continuous spaces. LPPG-RL reformulates the projection step as an optimization problem, and utilizes Dykstra's projection rather than generic solvers to deliver great speedups, especially for small- to medium-scale instances. In addition, LPPG-RL introduces Subproblem Exploration (SE) to prevent gradient vanishing, accelerate convergence and enhance stability. We provide theoretical guarantees for convergence and establish a lower bound on policy improvement. Finally, through extensive experiments in a 2D navigation environment, we demonstrate the effectiveness of LPPG-RL, showing that it outperforms existing state-of-the-art continuous LMORL methods.

AAAI Conference 2026 Conference Paper

M2I2: Learning Efficient Multi-Agent Communication via Masked State Modeling and Intention Inference

  • Chuxiong Sun
  • Peng He
  • Qirui Ji
  • Zehua Zang
  • Jiangmeng Li
  • Rui Wang
  • Wei Wang

Communication is essential in coordinating the behaviors of multiple agents. However, existing methods primarily emphasize content, timing, and partners for information sharing, often neglecting the critical aspect of integrating shared information. This gap can significantly impact agents' ability to understand and respond to complex, uncertain interactions, thus affecting overall communication efficiency. To address this issue, we introduce M2I2, a novel framework designed to enhance the agents' capabilities to assimilate and utilize received information effectively. M2I2 equips agents with advanced capabilities for masked state modeling and joint-action prediction, enriching their perception of environmental uncertainties and facilitating the anticipation of teammates' intentions. This approach ensures that agents are furnished with both comprehensive and relevant information, bolstering more informed and synergistic behaviors. Moreover, we propose a Dimensional Rational Network, innovatively trained via a meta-learning paradigm, to identify the importance of dimensional pieces of information, evaluating their contributions to decision-making and auxiliary tasks. Then, we implement an importance-based heuristic for selective information masking and sharing. This strategy optimizes the efficiency of masked state modeling and the rationale behind information sharing. We evaluate M2I2 across diverse multi-agent tasks, the results demonstrate its superior performance, efficiency, and generalization capabilities, over existing state-of-the-art methods in various complex scenarios.

AAAI Conference 2026 Conference Paper

Neural Graph Navigation for Intelligent Subgraph Matching

  • Yuchen Ying
  • Yiyang Dai
  • Wenda Li
  • Wenjie Huang
  • Rui Wang
  • Tongya Zheng
  • Yu Wang
  • Hanyang Yuan

Subgraph matching, a cornerstone of relational pattern detection in domains ranging from biochemical systems to social network analysis, faces significant computational challenges due to the dramatically growing search space. Existing methods address this problem within a filtering-ordering-enumeration framework, in which the enumeration stage recursively matches the query graph against the candidate subgraphs of the data graph. However, the lack of awareness of subgraph structural patterns leads to a costly brute-force enumeration, thereby critically motivating the need for intelligent navigation in subgraph matching. To address this challenge, we propose Neural Graph Navigation (NeuGN), a neuro-heuristic framework that transforms brute-force enumeration into neural-guided search by integrating neural navigation mechanisms into the core enumeration process. By preserving heuristic-based completeness guarantees while incorporating neural intelligence, NeuGN significantly reduces the First Match Steps by up to 98.2% compared to state-of-the-art methods across six real-world datasets.

AAAI Conference 2026 Conference Paper

OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding

  • Dianbing Xi
  • Jiepeng Wang
  • Yuanzhi Liang
  • Xi Qiu
  • Yuchi Huo
  • Rui Wang
  • Chi Zhang
  • Xuelong Li

In this paper, we propose a novel framework for controllable video diffusion, OmniVDiff, aiming to synthesize and comprehend multiple video visual content in a single diffusion model. To achieve this, OmniVDiff treats all video visual modalities in the color space to learn a joint distribution, while employing an adaptive control strategy that dynamically adjusts the role of each visual modality during the diffusion process, either as a generation modality or a conditioning modality. Our framework supports three key capabilities: (1) Text-conditioned video generation, where all modalities are jointly synthesized from a textual prompt; (2) Video understanding, where structural modalities are predicted from rgb inputs in a coherent manner; and (3) X-conditioned video generation, where video synthesis is guided by finegrained inputs such as depth, canny and segmentation. Extensive experiments demonstrate that OmniVDiff achieves state-of-the-art performance in video generation tasks and competitive results in video understanding. Its flexibility and scalability make it well-suited for downstream applications such as video-to-video translation, modality adaptation for visual tasks, and scene reconstruction.

AAAI Conference 2026 Conference Paper

PFAvatar: Pose-Fusion 3D Personalized Avatar Reconstruction from Real-World Outfit-of-the-Day Photos

  • Dianbing Xi
  • Guoyuan An
  • Jingsen Zhu
  • Zhijian Liu
  • Yuan Liu
  • Ruiyuan Zhang
  • Jiayuan Lu
  • Yuchi Huo

We propose PFAvatar (Pose-Fusion Avatar), a new method that reconstructs high-quality 3D avatars from Outfit of the Day (OOTD) photos, which exhibit diverse poses, occlusions, and complex backgrounds. Our method consists of two stages: (1) fine-tuning a pose-aware diffusion model from few-shot OOTD examples and (2) distilling a 3D avatar represented by a neural radiance field (NeRF). In the first stage, unlike previous methods that segment images into assets (e.g. garments, accessories) for 3D assembly, which is prone to inconsistency, we avoid decomposition and directly model the full-body appearance. By integrating a pre-trained ControlNet for pose estimation and a novel Condition Prior Preservation Loss (CPPL), our method enables end-to-end learning of fine details while mitigating language drift in few-shot training. Our method completes personalization in just 5 minutes, achieving a 48x speed-up compared to previous approaches. In the second stage, we introduce a NeRF-based avatar representation optimized by canonical SMPL-X space sampling and Multi-Resolution 3D-SDS. Compared to mesh-based representations that suffer from resolution-dependent discretization and erroneous occluded geometry, our continuous radiance field can preserve high-frequency textures (e.g., hair) and handle occlusions correctly through transmittance. Experiments demonstrate that PFAvatar outperforms state-of-the-art methods in terms of reconstruction fidelity, detail preservation, and robustness to occlusions/truncations, advancing practical 3D avatar generation from real-world OOTD albums. In addition, the reconstructed 3D avatars support downstream applications such as virtual try-on, animation, and human video reenactment, further demonstrating the versatility and practical value of our approach.

AAMAS Conference 2026 Conference Paper

RBC: Retroactive Belief State Compensation for Multi-Agent Collaboration Under Information Delay

  • Dongkun Huo
  • Hongbo Liu
  • Shu Yin
  • Yixue Hao
  • Long Hu
  • Rui Wang
  • Min Chen

Real-time information is usually not satisfied in real world due to communication or observation delay. Although existing works address individual delay, they do not fully consider the complex effects of composite delay, denoted as “Information Delay”, which severely reduce the efficiency of these methods. To address information delay, we propose Retroactive Belief state Compensation (RBC), a multi-agent framework with enhanced robustness and collaboration. Specifically, we design a multi-step reconstruction model that retroactively rebuilds agents’ belief states starting from the generation time of the information. This process corrects the accumulated deviation in the current belief state caused by information delay. Moreover, to enhance proactive collaboration, we introduce an intent inference module. This module enables agents to generate intents, which represent short-term action plans, as content of communication. By aggregating intents from teammates, agents will choose more coherent and synchronized joint actions. To evaluate the performance of RBC, we design scenarios with multiple levels of observation, communication, and composite delays. Experimental results demonstrate that RBC outperforms the baselines in all scenarios with delays. Yixue Hao is corresponding author. Email: yixuehao@hust. edu. cn. This work is licensed under a Creative Commons Attribution International 4. 0 License. Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), C. Amato, L. Dennis, V. Mascardi, J. Thangarajah (eds.), May 25 – 29, 2026, Paphos, Cyprus. © 2026 International Foundation for Autonomous Agents and Multiagent Systems (www. ifaamas. org). https: //doi. org/10. 65109/MFLP4403

AAAI Conference 2026 Conference Paper

Resource Efficient Sleep Staging via Multi-Level Masking and Prompt Learning

  • Lejun Ai
  • Yulong Li
  • Haodong Yi
  • Jixuan Xie
  • Yue Wang
  • Jia Liu
  • Min Chen
  • Rui Wang

Automatic sleep staging plays a vital role in assessing sleep quality and diagnosing sleep disorders. Most existing methods rely heavily on long and continuous EEG recordings, which poses significant challenges for data acquisition in resource-constrained systems, such as wearable or home-based monitoring systems. In this paper, we propose the task of resource-efficient sleep staging, which aims to reduce the amount of signal collected per sleep epoch while maintaining reliable classification performance. To solve this task, we adopt the masking and prompt learning strategy and propose a novel framework called Mask-Aware Sleep Staging (MASS). Specifically, we design a multi-level masking strategy to promote effective feature modeling under partial and irregular observations. To mitigate the loss of contextual information introduced by masking, we further propose a hierarchical prompt learning mechanism that aggregates unmasked data into a global prompt, serving as a semantic anchor for guiding both patch-level and epoch-level feature modeling. MASS is evalutaed on four datasets, demonstrating state-of-the-art performance, especially when the amount of data is very limited. This result highlights its potential for efficient and scalable deployment in real-world low-resource sleep monitoring environments.

AAAI Conference 2026 Conference Paper

Tensorized Label Learning via Balanced Tensor Regression

  • Guangyu Yang
  • Yuzhuo Feng
  • Qin Li
  • Quanxue Gao
  • Ming Yang
  • Rui Wang

The multi-view clustering methods based on tensor regression can make full use of the potential structural information between views and achieve data-level fusion. However, existing tensor regression-based approaches for anchor graph often overlook the probabilistic nature of anchor graph, focusing solely on sample labels while ignoring the influence of anchor labels on clustering results. To overcome these limitations, we introduce Tensorized Label Learning via Balanced Tensor Regression (TLL-BTR). Our key idea is to exploit the probabilistic nature of the anchor graph by regarding the sample labels as a projection tensor that maps the anchor graph into the label space, thereby producing anchor labels. By enforcing constraints on these anchor labels, we guide the concurrent learning of sample labels and achieve co-label learning between anchors and samples. To prevent trivial solutions, we maximize the nuclear norm to promote an even distribution of samples across clusters. Extensive experiments on benchmark datasets demonstrate that TLL-BTR consistently outperforms state-of-the-art methods.

AAAI Conference 2026 Conference Paper

TMAE:Learning Targeted Multi-Agent Exploration via Causal Inference

  • Chuxiong Sun
  • Dunqi Yao
  • Rui Wang
  • Wenwen Qiang
  • Changwen Zheng
  • Jiangmeng Li

Exploration in sparse-reward tasks remains a fundamental challenge in multi-agent reinforcement learning (MARL) due to complex inter-agent interactions and the expansive exploration space. To address this issue, we propose Targeted Multi-Agent Exploration (TMAE), a novel framework that uncovers the causal relationships between the state space and the reward function, thereby reducing the exploration space and enabling more targeted exploration. Specifically, we construct a structural causal model (SCM) to model the causality between sub-state variables and sparse rewards, providing a robust analytical foundation for subsequent causal inference. Through counterfactual causal intervention, TMAE identifies the most critical subspaces for discovering rare but pivotal events while filtering out confounders. By incorporating these causal insights into the exploration process, TMAE prioritizes subspaces with stronger causal effects on sparse rewards, significantly enhancing exploration efficiency. We evaluate TMAE on a range of MARL benchmarks featuring sparse rewards, consistently demonstrating superior exploration efficiency compared to state-of-the-art methods. Furthermore, visualized causal insights derived from TMAE reveal its ability to effectively capture intricate dependencies and priorities in targeted exploration, showcasing strong alignment with prior domain knowledge.

AAAI Conference 2026 Conference Paper

Wasserstein-Aligned Hyperbolic Multi-View Clustering

  • Rui Wang
  • Yuting Jiang
  • Xiaoqing Luo
  • Xiao-Jun Wu
  • Nicu Sebe
  • Ziheng Chen

Multi-view clustering (MVC) aims to uncover the latent structure of multi-view data by learning view-common and view-specific information. Although recent studies have explored hyperbolic representations for better tackling the representation gap between different views, they focus primarily on instance-level alignment and neglect global semantic consistency, rendering them vulnerable to view-specific information (e.g., noise and cross-view discrepancies). To this end, this paper proposes a novel Wasserstein-Aligned Hyperbolic (WAH) framework for multi-view clustering. Specifically, our method exploits a view-specific hyperbolic encoder for each view to embed features into the Lorentz manifold for hierarchical semantic modeling. Whereafter, a global semantic loss based on the hyperbolic sliced-Wasserstein distance is introduced to align manifold distributions across views. This is followed by soft cluster assignments to encourage cross-view semantic consistency. Extensive experiments on multiple benchmarking datasets show that our method can achieve SOTA clustering performance.

IJCAI Conference 2025 Conference Paper

A Correlation Manifold Self-Attention Network for EEG Decoding

  • Chen Hu
  • Rui Wang
  • Xiaoning Song
  • Tao Zhou
  • Xiao-Jun Wu
  • Nicu Sebe
  • Ziheng Chen

Riemannian neural networks, which generalize the deep learning paradigm to non-Euclidean geometries, have garnered widespread attention across diverse applications in artificial intelligence. Among these, the representative attention models have been studied on various non-Euclidean spaces to geometrically capture the spatiotemporal dependencies inherent in time series data, e. g. , electroencephalography (EEG). Recent studies have highlighted the full-rank correlation matrix as an advantageous alternative to the covariance matrix for data representation, owing to its invariance to the scale of variables. Motivated by these advancements, we propose the Correlation Attention Network (CorAtt) tailored for full-rank correlation matrices and implement it under the permutation-invariant and computationally efficient Off-Log and Log-Scaled geometries, respectively. Extensive evaluations on three benchmarking EEG datasets provide substantial evidence for the effectiveness of our introduced CorAtt. The code and supplementary material can be found at https: //github. com/ChenHu-ML/CorAtt.

TIST Journal 2025 Journal Article

AEKG4APT: An AI-Enhanced Knowledge Graph for Advanced Persistent Threats with Large Language Model Analysis

  • Yinghai Zhou
  • Ziyu Wang
  • Yunxin Jiang
  • Bingqi Ma
  • Rui Wang
  • Yuan Liu
  • Yue Zhao
  • Zhihong Tian

This paper introduces AEKG4APT, an APT Knowledge Graph (KG) enhanced by Large Language Models (LLMs), as a way to deal with the cybersecurity problems caused by Advanced Persistent Threats (APTs). The core of AEKG4APT lies in the combined application of LLMs, Cyber Threat Intelligence (CTI), and KG. The first part of the paper goes into great detail about how the AEKG4APT was constructed, including its ontology schema, data sources, and dataset features. There are also statistics on the AEKG4APT’s nodes, relationships, and key attributes. Secondly, it was shown how to utilize LLMs and public sandboxes for the collection and analysis of CTI Additionally, tests that compare traditional deep learning models to LLM methods show that LLM is both more efficient and more accurate at extracting information. Subsequently, the Decision Making Trial and Evaluation Laboratory - Interpretive Structural Modeling (DEMATEL-ISM) analytical method was introduced to identify and analyse the factors and their interrelationships within the AEKG4APT data, thereby revealing the key dependencies and influence paths within the data structure. Experiments were designed to demonstrate its applications in modeling, computing, and obtaining interpretable computational results on AEKG4APT. In addition, this paper also explores the dynamic expansion capabilities of AEKG4APT, including data expansion, schema expansion, and permanent maintenance strategies, to address the evolving APT threats. Finally, this paper summarizes the competitiveness and application value of AEKG4APT by comparing it with other CTI KGs and platforms in academia and industry, demonstrating its extensive application potential in the field of cybersecurity.

ICRA Conference 2025 Conference Paper

An End-to-End Learning-Based Multi-Sensor Fusion for Autonomous Vehicle Localization

  • Changhong Lin
  • Jiarong Lin
  • Zhiqiang Sui
  • XiaoZhi Qu
  • Rui Wang
  • Kehua Sheng
  • Bo Zhang 0106

Multi-sensor fusion is essential for autonomous vehicle localization, as it is capable of integrating data from various sources for enhanced accuracy and reliability. The accuracy of the integrated location and orientation depends on the precision of the uncertainty modeling. Traditional methods of uncertainty modeling typically assume a Gaussian distribution and involve manual heuristic parameter tuning. However, these methods struggle to scale effectively and address long-tail scenarios. To address these challenges, we propose a learning-based method that encodes sensor information using higher-order neural network features, thereby eliminating the need for uncertainty estimation. This method significantly eliminates the need for parameter fine-tuning by developing an end-to-end neural network that is specifically designed for multi-sensor fusion. In our experiments, we demonstrate the effectiveness of our approach in real-world autonomous driving scenarios. Results show that the proposed method outperforms existing multi-sensor fusion methods in terms of both accuracy and robustness. A video of the results can be viewed at https://youtu.be/q4iuobMbjME.

AAAI Conference 2025 Conference Paper

Attention-Imperceptible Backdoor Attacks on Vision Transformers

  • Zhishen Wang
  • Rui Wang
  • Lihua Jing

With the successful transition of Transformers from natural language processing (NLP) to computer vision (CV) domains, Vision Transformers (ViTs) have achieved state-of-the-art performance in many CV tasks. However, backdoor attacks, a significant threat in deep learning, also pose a risk to the security of ViT models. Recently, several backdoor attack methods targeting the patch-level self-attention mechanism in ViTs have been proposed, but they are relatively naive in terms of stealthiness and robustness against defensive measures, lacking in-depth investigation. In this paper, we explore the crucial role of attention-level imperceptibility in backdoor attacks for ViTs and propose an Attention-Imperceptible Backdoor Attacks on Vision Transformers (AIBA). In AIBA, a constrained adversarial perturbation is used as the trigger to achieve visual imperceptibility. Additionally, the trigger is designed to seamlessly implant into the focal areas of the image, ensuring that the trigger receives enough attention from the model without causing anomalies at the attention level. During the backdoor learning process, we designed an efficient constrained bi-level optimization training strategy at the mini-batch level to implant an effective backdoor in the victim model using the imperceptible trigger. We evaluated the effectiveness of the proposed AIBA across multiple datasets and ViT benchmarks and explored the robustness of AIBA against current ViT-specific defense methods. The experimental results demonstrate that our backdoor attack method can successfully implant a powerful and stealthy backdoor into ViTs.

NeurIPS Conference 2025 Conference Paper

Code Graph Model (CGM): A Graph-Integrated Large Language Model for Repository-Level Software Engineering Tasks

  • Hongyuan Tao
  • Ying Zhang
  • Zhenhao Tang
  • Hongen Peng
  • Xukun Zhu
  • Bingchang Liu
  • Yingguang Yang
  • Ziyin Zhang

Recent advances in Large Language Models (LLMs) have shown promise in function-level code generation, yet repository-level software engineering tasks remain challenging. Current solutions predominantly rely on proprietary LLM agents, which introduce unpredictability and limit accessibility, raising concerns about data privacy and model customization. This paper investigates whether open-source LLMs can effectively address repository-level tasks without requiring agent-based approaches. We demonstrate this is possible by enabling LLMs to comprehend functions and files within codebases through their semantic information and structural dependencies. To this end, we introduce Code Graph Models (CGMs), which integrate repository code graph structures into the LLM's attention mechanism and map node attributes to the LLM's input space using a specialized adapter. When combined with an agentless graph RAG framework, our approach achieves a 43. 00% resolution rate on the SWE-bench Lite benchmark using the open-source Qwen2. 5-72B model. This performance ranks first among open weight models, second among methods with open-source systems, and eighth overall, surpassing the previous best open-source model-based method by 12. 33%.

EAAI Journal 2025 Journal Article

Dynamic health prediction of plain reservoirs based on deep learning algorithms

  • Zhaohui Zhu
  • Hao Wu
  • Zhicheng Zhang
  • Rui Wang
  • Qiang Yue

The healthy operation of reservoirs is crucial for fulfilling their functions and preventing harm to human populations and riverine ecosystems. This study addresses the issues of poor model generalization and low diagnostic accuracy resulting from imbalanced sample distribution in the health prediction of plain reservoirs. We innovatively propose a sample augmentation model (VAE-CGAN) that combines Variational Autoencoder (VAE) and Conditional Generative Adversarial Network (CGAN). Additionally, we present a classification model (CNN-BiLSTM) based on Convolutional Neural Network (CNN) and Bidirectional Long Short-Term Memory Network (BiLSTM). The VAE-CGAN model learns the data distribution of authentic samples through alternating training of the encoder, generator, and discriminator, thereby achieving augmentation of fault samples and effectively addressing the issue of sample imbalance. The CNN-BiLSTM model utilizes CNN to capture global key features and BiLSTM to capture bi-directional features of time series, efficiently classifying the augmented and balanced data and accurately identifying the health status of the reservoir. In practical applications at two plain reservoirs in China, our method demonstrated superior robustness compared to three other mainstream deep learning models when handling data with varying degrees of imbalance, achieving an accuracy rate of 0. 94, which is significantly higher than that of other models. Even under extreme imbalance ratios of 1: 20, the accuracy rate improved from 0. 80 to 0. 93 through sample augmentation. This study not only enhances the dynamic perception and understanding of the operational health of reservoirs but also significantly bolsters the reliability of risk mitigation decisions, offering a novel technical approach for reservoir health management.

IJCAI Conference 2025 Conference Paper

Efficient Dynamic Graphs Learning with Refined Batch Parallel Training

  • Zhengzhao Feng
  • Rui Wang
  • Longjiao Zhang
  • Tongya Zheng
  • Ziqi Huang
  • Mingli Song

Memory-based temporal graph neural networks (MTGNN) use node memory to store historical information, enabling efficient processing of large dynamic graphs through batch parallel training, with larger batch sizes leading to increased training efficiency. However, this approach overlooks the interdependency among edges within the same batch, leading to outdated memory states and reduced training accuracy. Previous studies have attempted to mitigate this issue through methods such as measuring memory loss, overlap training, and additional compensation modules. Despite these efforts, challenges persist, including imprecise coarse-grained memory loss measurement and ineffective compensation modules. To address these challenges, we propose the Refined Batch parallel Training (RBT) framework, which accurately evaluates intra-batch information loss and optimizes batch partitioning to minimize loss, enhancing the training process's effectiveness and efficiency. RBT also includes a precise and efficient memory compensation algorithm. Experimental results demonstrate RBT's superior performance compared to existing MTGNN frameworks like TGL, ETC, and PRES in terms of training efficiency and accuracy across various dynamic graph datasets. Our code is made publicly available at https: //github. com/fengwudi/RBT.

ICML Conference 2025 Conference Paper

Federated In-Context Learning: Iterative Refinement for Improved Answer Quality

  • Ruhan Wang
  • Zhiyong Wang
  • Chengkai Huang
  • Rui Wang
  • Tong Yu 0001
  • Lina Yao 0001
  • John C. S. Lui
  • Dongruo Zhou

For question-answering (QA) tasks, in-context learning (ICL) enables language models (LMs) to generate responses without modifying their parameters by leveraging examples provided in the input. However, the effectiveness of ICL heavily depends on the availability of high-quality examples, which are often scarce due to data privacy constraints, annotation costs, and distribution disparities. A natural solution is to utilize examples stored on client devices, but existing approaches either require transmitting model parameters—incurring significant communication overhead—or fail to fully exploit local datasets, limiting their effectiveness. To address these challenges, we propose Federated In-Context Learning (Fed-ICL), a general framework that enhances ICL through an iterative, collaborative process. Fed-ICL progressively refines responses by leveraging multi-round interactions between clients and a central server, improving answer quality without the need to transmit model parameters. We establish theoretical guarantees for the convergence of Fed-ICL and conduct extensive experiments on standard QA benchmarks, demonstrating that our proposed approach achieves strong performance while maintaining low communication costs.

EAAI Journal 2025 Journal Article

Federated Reinforcement Learning for smart and privacy-preserving energy management of residential microgrids clusters

  • Mao Tan
  • Jie Zhao
  • Xiao Liu
  • Yongxin Su
  • Ling Wang
  • Rui Wang
  • Zhuocen Dai

Real-time energy management optimizes energy utilization and manages electrical loads, which is crucial for improving the operational efficiency of residential microgrids. However, existing management methods suffer from model complexity and slow training speed. To solve this problem, we introduce Federated Reinforcement Learning to manage residential microgrids by training a control strategy in a decentralized and privacy-preserving manner. Specifically, a residential microgrid energy optimization management model is first established based on the Proximal Policy Optimization (PPO) method. Then, we propose a cooperative training strategy for multiple Residential microgrids based on Federated Reinforcement Learning (RFRL). The proposed method improves the training speed of residential microgrid models by sharing parameter information, such as network weights, while protects users’ usage data. Finally, clustering analysis is introduced in the case of heterogeneous residential microgrid data. Extensive experimental evaluation shows that our method outperforms the alternative residential microgrid management methods in terms of cost efficiency.

IJCAI Conference 2025 Conference Paper

Find and Perceive: Tell Visual Change with Fine-Grained Comparison

  • Feixiao Lv
  • Rui Wang
  • Lihua Jing
  • Lijun Liu

The goal of the image change captioning task is to capture the differences between two similar images and describe them in natural language. In this paper, we decompose this task into two sub-problems, i. e. , fine-grained change feature learning and discrimination of changed regions. Compared with existing methods which only focus on change feature learning, we propose a novel change captioning learning paradigm, Find and Perceive (F&P). Our proposed F&P consists of two main ideas, i. e. , the Fine-Grained Semantic Change Perception (FGSCP) module for improving the model's perception ability of subtle changes and the Weakly-Supervised Discriminator (WSD) of changed regions for improving the model's sensitivity of localising the important regions. Specifically, the FGSCP deploys a two-step manner, firstly introducing the fine-grained categorisation and then enhancing the interaction of the two paired images. And the WSD adopts the contributions of each image region for final generated captions, accurately indicating which regions are important for change captions without any extra annotations. Finally, we conduct extensive experiments on four change captioning datasets, and experimental results show that our proposed method F&P outperforms existing change caption methods and achieves new state-of-the-art performance.

NeurIPS Conference 2025 Conference Paper

Flexible Realignment of Language Models

  • Wenhong Zhu
  • Ruobing Xie
  • Weinan Zhang
  • Rui Wang

Realignment becomes necessary when a language model (LM) fails to meet expected performance. We propose a flexible realignment framework that supports quantitative control of alignment degree during training and inference. This framework incorporates Training-time Realignment (TrRa), which efficiently realigns the reference model by leveraging the controllable fusion of logits from both the reference and already aligned models. For example, TrRa reduces token usage by 54. 63% on DeepSeek-R1-Distill-Qwen-1. 5B without any performance degradation, outperforming DeepScaleR-1. 5B’s 33. 86%. To complement TrRa during inference, we introduce a layer adapter that enables smooth Inference-time Realignment (InRa). This adapter is initialized to perform an identity transformation at the bottom layer and is inserted preceding the original layers. During inference, input embeddings are simultaneously processed by the adapter and the original layer, followed by the remaining layers, and then controllably interpolated at the logit level. We upgraded DeepSeek-R1-Distill-Qwen-7B from a slow-thinking model to one that supports both fast and slow thinking, allowing flexible alignment control even during inference. By encouraging deeper reasoning, it even surpassed its original performance.

NeurIPS Conference 2025 Conference Paper

LIFEBENCH: Evaluating Length Instruction Following in Large Language Models

  • Wei Zhang
  • Zhenhong Zhou
  • Kun Wang
  • Junfeng Fang
  • Rongwu Xu
  • Yuanhe Zhang
  • Rui Wang
  • Ge Zhang

While large language models (LLMs) can solve PhD-level reasoning problems over long context inputs, they still struggle with a seemingly simpler task: following explicit length instructions —e. g. , write a 10, 000-word novel. Additionally, models often generate far too short outputs, terminate prematurely, or even refuse the request. Existing benchmarks focus primarily on evaluating generations quality, but often overlook whether the generations meet length constraints. To this end, we introduce Length Instruction Following Evaluation Benchmark (LIFEBench) to comprehensively evaluate LLMs' ability to follow length instructions across diverse tasks and a wide range of specified lengths. LIFEBench consists of 10, 800 instances across 4 task categories in both English and Chinese, covering length constraints ranging from 16 to 8192 words. We evaluate 26 widely-used LLMs and find that most models reasonably follow short-length instructions but deteriorate sharply beyond a certain threshold. Surprisingly, almost all models fail to reach the vendor-claimed maximum output lengths in practice, as further confirmed by our evaluations extending up to 32K words. Even long-context LLMs, despite their extended input-output windows, counterintuitively fail to improve length-instructions following. Notably, Reasoning LLMs outperform even specialized long-text generation models, achieving state-of-the-art length following. Overall, LIFEBench uncovers fundamental limitations in current LLMs' length instructions following ability, offering critical insights for future progress.

AAAI Conference 2025 Conference Paper

LiON: Learning Point-Wise Abstaining Penalty for LiDAR Outlier DetectioN Using Diverse Synthetic Data

  • Shaocong Xu
  • Pengfei Li
  • Qianpu Sun
  • Xinyu Liu
  • Yang Li
  • Shihui Guo
  • Zhen Wang
  • Bo Jiang

LiDAR-based semantic scene understanding is an important module in the modern autonomous driving perception stack. However, identifying outlier points in a LiDAR point cloud is challenging as LiDAR point clouds lack semantically-rich information. While former SOTA methods adopt heuristic architectures, we revisit this problem from the perspective of Selective Classification, which introduces a selective function into the standard closed-set classification setup. Our solution is built upon the basic idea of abstaining from choosing any inlier categories but learns a point-wise abstaining penalty with a margin-based loss. Apart from learning paradigms, synthesizing outliers to approximate unlimited real outliers is also critical, so we propose a strong synthesis pipeline that generates outliers originated from various factors: object categories, sampling patterns and sizes. We demonstrate that learning different abstaining penalties, apart from point-wise penalty, for different types of (synthesized) outliers can further improve the performance. We benchmark our method on SemanticKITTI and nuScenes and achieve SOTA results.

ICLR Conference 2025 Conference Paper

Longhorn: State Space Models are Amortized Online Learners

  • Bo Liu 0042
  • Rui Wang
  • Lemeng Wu
  • Yihao Feng
  • Peter Stone 0001
  • Qiang Liu 0001

The most fundamental capability of modern AI methods such as Large Language Models (LLMs) is the ability to predict the next token in a long sequence of tokens, known as “sequence modeling.” Although the Transformers model is the current dominant approach to sequence modeling, its quadratic computational cost with respect to sequence length is a significant drawback. State-space models (SSMs) offer a promising alternative due to their linear decoding efficiency and high parallelizability during training. However, existing SSMs often rely on seemingly ad hoc linear recurrence designs. In this work, we explore SSM design through the lens of online learning, conceptualizing SSMs as meta-modules for specific online learning problems. This approach links SSM design to formulating precise online learning objectives, with state transition rules derived from optimizing these objectives. Based on this insight, we introduce a novel deep SSM architecture based on the implicit update for optimizing an online regression objective. Our experimental results show that our models outperform state-of-the-art SSMs, including the Mamba model, on standard sequence modeling benchmarks and language modeling tasks.

IJCAI Conference 2025 Conference Paper

M4Bench: A Benchmark of Multi-domain Multi-granularity Multi-image Understanding for Multi-modal Large Language Models

  • Xiaojun Ye
  • Guanbao Liang
  • Chun Wang
  • Liangcheng Li
  • Pengfei Ke
  • Rui Wang
  • Bingxin Jia
  • Gang Huang

The increasing demands in analyzing complex associated scenes pose necessities to researching multi-image understanding abilities. Compared with understanding individual images, both the alignments and differences between images are essential aspects of understanding the intricate relationships for multi-image inference tasks. However, existing benchmarks face difficulties in addressing both of these aspects simultaneously, resulting in obstacles to modeling relationships under various granularities and domains of images. In this paper, we introduce M4Bench to enhance the capability of aligning and distinguishing multi-images with multi-domain multi-granularity comparison. We carefully design five comparison tasks related to coarse and fine-grained granularities in single and multiple domains of images and evaluate them on 13 state-of-the-art multi-modal large language models with various sizes. Besides, we analyze the evaluation results and provide several observations and viewpoints for the multi-image understanding research. The data and evaluation code are available at https: //github. com/eaglelab-zju/M4Bench.

JBHI Journal 2025 Journal Article

MSMTSeg: Multi-Stained Multi-Tissue Segmentation of Kidney Histology Images via Generative Self-Supervised Meta-Learning Framework

  • Xueyu Liu
  • Rui Wang
  • Yexin Lai
  • Yongfei Wu
  • Hangbei Cheng
  • Yuanyue Lu
  • Jianan Zhang
  • Ning Hao

Accurately diagnosing chronic kidney disease requires pathologists to assess the structure of multiple tissues under different stains, a process that is time-consuming and labor-intensive. Current AI-based methods for automatic structure assessment, like segmentation, often demand extensive manual annotation and focus on single stain domain. To address these challenges, we introduce MSMTSeg, a generative self-supervised meta-learning framework for multi-stained multi-tissue segmentation in renal biopsy whole slide images (WSIs). MSMTSeg incorporates multiple stain transform models for style translation of inter-stain domains, a self-supervision module for obtaining pre-trained models with the domain-specific feature representation, and a meta-learning strategy that leverages generated virtual data and pre-trained models to learn the domain-invariant feature representation across multiple stains, thereby enhancing segmentation performance. Experimental results demonstrate that MSMTSeg achieves superior and robust performance, with mDSC of 0. 836 and mIoU of 0. 718 for multiple tissues under different stains, using only one annotated training sample for each stain. Our ablation study confirms the effectiveness of each component, positioning MSMTSeg ahead of classic advanced segmentation networks, recent few-shot segmentation methods, and unsupervised domain adaptation methods. In conclusion, our proposed few-shot cross-domain technology offers a feasible and cost-effective solution for multi-stained renal histology segmentation, providing convenient assistance to pathologists in clinical practice.

NeurIPS Conference 2025 Conference Paper

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

  • Jingjing Chang
  • Yixiao Fang
  • Peng Xing
  • Shuhan Wu
  • Wei Cheng
  • Rui Wang
  • Xianfang Zeng
  • Gang Yu

Text-to-image (T2I) models have garnered significant attention for generating high-quality images aligned with text prompts. However, rapid T2I model advancements reveal limitations in early benchmarks, lacking comprehensive evaluations, especially for text rendering and style. Notably, recent state-of-the-art models, with their rich knowledge modeling capabilities, show potential in reasoning-driven image generation, yet existing evaluation systems have not adequately addressed this frontier. To systematically address these gaps, we introduce $\textbf{OneIG-Bench}$, a meticulously designed comprehensive benchmark framework for fine-grained evaluation of T2I models across multiple dimensions, including subject-element alignment, text rendering precision, reasoning-generated content, stylization, and diversity. By structuring the evaluation, this benchmark enables in-depth analysis of model performance, helping researchers and practitioners pinpoint strengths and bottlenecks in the full pipeline of image generation. Our codebase and dataset are now publicly available to facilitate reproducible evaluation studies and cross-model comparisons within the T2I research community.

NeurIPS Conference 2025 Conference Paper

PolyMath: Evaluating Mathematical Reasoning in Multilingual Contexts

  • Yiming Wang
  • Pei Zhang
  • Jialong Tang
  • Hao-Ran Wei
  • Baosong Yang
  • Rui Wang
  • Chenshu Sun
  • Feitong Sun

In this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs. We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2. 5-pro, achieve only 54. 6 and 52. 2 benchmark scores, with about 40% accuracy under the highest level. From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning: (1) Reasoning performance varies widely across languages for current LLMs; (2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance; (3) The thinking length differs significantly by language for current LLMs. Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs.

IJCAI Conference 2025 Conference Paper

Preventing Latent Diffusion Model-Based Image Mimicry via Angle Shifting and Ensemble Learning

  • Minghao Li
  • Rui Wang
  • Ming Sun
  • Lihua Jing

The remarkable progress of Latent Diffusion Models (LDMs) in image generation has raised concerns about the potential for unauthorized image mimicry. To address these concerns, studies on adversarial attacks against LDMs have gained increasing attention in recent years. However, existing methods face bottlenecks when attacking the denoising module. In this work, we reveal that the robustness of the denoising module stems from two key factors: the cancellation effect between adversarial perturbations and estimated noise, and unstable gradients caused by randomly sampled timesteps and Gaussian noise. Based on these insights, we introduce a cosine similarity adversarial loss to prevent the generation of perturbations that are easily impaired and develop a more stable optimization strategy by ensembling gradients and fixing the noise in the latent space. Additionally, we propose an alternating iterative framework to reduce memory usage by mathematically dividing the optimization process into two spaces: latent space and pixel space. Compared to previous strategies, our proposed framework reduces video memory demands without sacrificing attack effectiveness. Extensive experiments demonstrate that the alternating iterative framework and the stable optimization strategy on cosine similarity loss are more efficient and more effective. Code is available at https: //github. com/MinghaoLi01/cosattack.

AAAI Conference 2025 Conference Paper

R-DTI: Drug Target Interaction Prediction Based on Second-Order Relevance Exploration

  • Yang Hua
  • Tianyang Xu
  • Xiaoning Song
  • Zhenhua Feng
  • Rui Wang
  • Wenjie Zhang
  • Xiaojun Wu

Drug Target Interaction (DTI) prediction has witnessed promising performance boosts accompanied by advanced multimodal feature extraction. However, existing approaches suffer from two main difficulties. First, the complex protein structures cannot be well represented by current protein-sequence-based feature extractors. Second, the gap between protein and drug features increases the vulnerability of the obtained classifier thus degrading the prediction robustness. To address these issues, we propose a novel R-DTI method by exploring the second-order relevance in both protein structural feature extraction and DTI prediction phases. Specifically, we construct a pre-trained structural feature extractor that mines the atomic relevance of each amino acid. Then, an inter-feature structure-preserved Riemannian network is designed to expand the existing protein extraction patterns. To improve the prediction robustness, we also develop a Riemannian classifier that uses the second-order protein-drug relevance with a unified feature space. Extensive experimental results demonstrate the merits and superiority of our R-DTI against the state-of-the-art, achieving 1.4% and 1.9% higher AUC-ROC on the BindingDB and DrugBank datasets, respectively.

AAAI Conference 2025 Conference Paper

Revisiting Change Captioning from Self-supervised Global-Part Alignment

  • Feixiao Lv
  • Rui Wang
  • Lihua Jing

The goal of image change captioning is to capture the content differences between two images and describe them in natural language. The key is how to learn stable content changes from noise such as viewpoint and image structure. However, current work mostly focuses on identifying changes, and the influence of global noise leads to unstable recognition of global features. In order to tackle this problem, we propose a Self-supervised Global-Part Alignment (SSGPA) network and revisit the image change captioning task by enhancing the construction process of overall image global features, enabling the model to integrate global changes such as viewpoint into local changes, and to detect and describe changes in the image through alignment. Concretely, we first design a Global-Part Transport Alignment mechanism to enhance global features and learn stable content changes through a self-supervised method of optimal transport. Further, we design a Change Fusion Adapter with pre-trained vision-language model to enhance the similar parts features of paired images, thereby enhancing global features, and expanding content changes. Extensive experiments show our method achieves the state-of-the-art results on four datasets.

AAMAS Conference 2025 Conference Paper

Revisiting Communication Efficiency in Multi-Agent Reinforcement Learning from the Dimensional Analysis Perspective

  • Chuxiong Sun
  • Peng He
  • Rui Wang
  • Changwen Zheng

In this work, we introduce a novel perspective—dimensional analysis—to address the challenge of communication efficiency in Multi- Agent Reinforcement Learning (MARL). Our findings reveal that simply optimizing the content and timing of communication at sending end is insufficient to fully resolve communication efficiency issues. Even after applying optimized and gated messages, dimensional redundancy and confounders still persist in the integrated message embeddings at receiving end, which negatively impact communication quality and decision-making. To address these challenges, we propose Dimensional Rational Multi-Agent Communication (DRMAC), designed to mitigate both dimensional redundancy and confounders in MARL. DRMAC incorporates a redundancy-reduction regularization term to encourage the decoupling of information across dimensions within the learned representations of integrated messages. Additionally, we introduce a dimensional mask that dynamically adjusts gradient weights during training to eliminate the influence of decision-irrelevant dimensions. We evaluate DRMAC across a diverse set of multi-agent tasks, demonstrating its superior performance over existing state-of-theart methods in complex scenarios. Furthermore, the plug-and-play nature of DRMAC’s key modules highlights its generalizable performance, serving as a valuable complement rather than a replacement for existing multi-agent communication strategies.

NeurIPS Conference 2025 Conference Paper

Revisiting Generative Infrared and Visible Image Fusion Based on Human Cognitive Laws

  • Lin Guo
  • Xiaoqing Luo
  • Wei Xie
  • Zhancheng Zhang
  • Hui Li
  • Rui Wang
  • Zhenhua Feng
  • Xiaoning Song

Existing infrared and visible image fusion methods often face the dilemma of balancing modal information. Generative fusion methods reconstruct fused images by learning from data distributions, but their generative capabilities remain limited. Moreover, the lack of interpretability in modal information selection further affects the reliability and consistency of fusion results in complex scenarios. This manuscript revisits the essence of generative image fusion under the inspiration of human cognitive laws and proposes a novel infrared and visible image fusion method, termed HCLFuse. First, HCLFuse investigates the quantification theory of information mapping in unsupervised fusion networks, which leads to the design of a multi-scale mask-regulated variational bottleneck encoder. This encoder applies posterior probability modeling and information decomposition to extract accurate and concise low-level modal information, thereby supporting the generation of high-fidelity structural details. Furthermore, the probabilistic generative capability of the diffusion model is integrated with physical laws, forming a time-varying physical guidance mechanism that adaptively regulates the generation process at different stages, thereby enhancing the ability of the model to perceive the intrinsic structure of data and reducing dependence on data quality. Experimental results show that the proposed method achieves state-of-the-art fusion performance in qualitative and quantitative evaluations across multiple datasets and significantly improves semantic segmentation metrics. This fully demonstrates the advantages of this generative image fusion method, drawing inspiration from human cognition, in enhancing structural consistency and detail quality.

NeurIPS Conference 2025 Conference Paper

SALoM: Structure Aware Temporal Graph Networks with Long-Short Memory Updater

  • Hanwen Liu
  • Longjiao Zhang
  • Rui Wang
  • Tongya Zheng
  • Sai Wu
  • Chang Yao
  • Mingli Song

Dynamic graph learning is crucial for accurately modeling complex systems by integrating topological structure and temporal information within graphs. While memory-based methods are commonly used and excel at capturing short-range temporal correlations, they struggle with modeling long-range dependencies, harmonizing long-range and short-range correlations, and integrating structural information effectively. To address these challenges, we present SALoM: Structure Aware Temporal Graph Networks with Long-Short Memory Updater. SALoM features a memory module that addresses gradient vanishing and information forgetting, enabling the capture of long-term dependencies across various time scales. Additionally, SALoM utilizes a long-short memory updater (LSMU) to dynamically balance long-range and short-range temporal correlations, preventing over-generalization. By integrating co-occurrence encoding and LSMU through information bottleneck-based fusion, SALoM effectively captures both the structural and temporal information within graphs. Experimental results across various graph datasets demonstrate SALoM's superior performance, achieving state-of-the-art results in dynamic graph link prediction. Our code is openly accessible at https: //github. com/wave5418/SALoM.

NeurIPS Conference 2025 Conference Paper

Sampling-Efficient Test-Time Scaling: Self-Estimating the Best-of-N Sampling in Early Decoding

  • Yiming Wang
  • Pei Zhang
  • Siyuan Huang
  • Baosong Yang
  • Zhuosheng Zhang
  • Fei Huang
  • Rui Wang

Test-time scaling enhances large language model performance by allocating additional compute resources during decoding. Best-of-$N$ (BoN) sampling serves as a common sampling-based scaling technique, broadening the search space in parallel to find better solutions from the model distribution. However, its cost–performance trade-off is still underexplored. Two main challenges limit the efficiency of BoN sampling: (1) Generating $N$ full samples consumes substantial GPU memory, reducing inference capacity under limited resources. (2) Reward models add extra memory and latency overhead, and training strong reward models introduces potential training data costs. Although some studies have explored efficiency improvements, none have addressed both challenges at once. To address this gap, we propose **Self-Truncation Best-of-$N$ (ST-BoN)**, a decoding method that avoids fully generating all $N$ samples and eliminates the need for reward models. It leverages early sampling consistency in the model’s internal states to identify the most promising path and truncate suboptimal ones. In terms of cost, ST-BoN reduces dynamic GPU memory usage by over 80% and inference latency by 50%. In terms of cost–performance trade-off, ST-BoN achieves the same performance as Full-BoN while saving computational cost by 70%–80%, and under the same cost, it can improve accuracy by 3–4 points.

YNIMG Journal 2025 Journal Article

Spatiospectral dynamics of electroencephalography patterns during propofol-induced alterations of consciousness states

  • Xuan Li
  • Dezhao Liu
  • Zheng Li
  • Rui Wang
  • Xiaoli Li
  • Tianyi Zhou

Altered consciousness induced by anesthetics is characterized by distinct spatial and spectral neural dynamics that are readily apparent in the human electroencephalogram. Despite considerable study, we remain uncertain which brain regions and neural oscillations are involved, as well as how they are impacted when consciousness is disrupted. The experimental data was obtained from the open-access dataset, which contains pre-processed EEG data recorded from 20 healthy participants during propofol sedation. Using unsupervised machine learning methods (i.e., non-negative matrix factorization, NMF), we investigated the spatiospectral dynamic evolution of brain activity from awake to sedation and back induced by propofol in healthy research volunteers. Our methods yielded six dynamical patterns that continuously reflect the neural activity changes in specific brain regions and frequency bands under propofol sedation. Temporal dynamic analyses showed that differences in alpha oscillation patterns were less pronounced in response group than drowsy group, with hemispheric asymmetry in posterior occipital lobe over the course of the sedation procedure. We designed an index 'hemispheric lateralization modulation of alpha [HLM(α)]' to measure asymmetry during awake state and predicting individual variability in propofol-induced alterations of consciousness states, obtaining prediction AUC of 0.8462. We present an alpha modulation index which characterizes how these patterns track the transition from awake to sedation as a function of increasing dosage. Our study reveals dynamics indices that track the evolution of neurophysiological of propofol on brain circuits. Analyzing the spatiospectral dynamics influenced by propofol provides valuable understanding of the mechanisms of these agents and strategies for monitoring and precisely controlling the level of consciousness in patients under sedation and general anesthesia.

NeurIPS Conference 2025 Conference Paper

Thoughts Are All Over the Place: On the Underthinking of Long Reasoning Models

  • Yue Wang
  • Qiuzhi Liu
  • Jiahao Xu
  • Tian Liang
  • Xingyu Chen
  • Zhiwei He
  • Linfeng Song
  • Dian Yu

Long reasoning models (LRMs) such as OpenAI's o1 and DeepSeek's R1 have demonstrated remarkable abilities in complex reasoning tasks by scaling test-time compute and exhibiting human-like deep thinking. However, we identify a phenomenon we term underthinking, where LRMs frequently switch between different reasoning thoughts without sufficiently exploring promising paths to reach a correct solution. This behavior leads to inadequate depth of reasoning and decreased performance, particularly on challenging mathematical problems. To systematically analyze this issue, we conduct experiments on three challenging test sets and two representative open-source LRMs, revealing that frequent thought switching correlates with incorrect responses. We introduce a novel metric to quantify underthinking by measuring token efficiency in incorrect answers. To address underthinking, we propose a decoding strategy with thought switching penalty (Tip) that discourages premature transitions between thoughts, encouraging deeper exploration of each reasoning path. Experimental results demonstrate that our approach improves accuracy across challenging datasets without requiring model fine-tuning. Our findings contribute to understanding reasoning inefficiencies in LRMs and offer a practical solution to enhance their problem-solving capabilities. Our code is open-source and available at https: //github. com/wangyuenlp/underthinking.

ICLR Conference 2025 Conference Paper

TLDR: Token-Level Detective Reward Model for Large Vision Language Models

  • Deqing Fu
  • Tong Xiao 0003
  • Rui Wang
  • Wang Zhu 0001
  • Pengchuan Zhang
  • Guan Pang
  • Robin Jia
  • Lawrence Chen 0002

Although reward models have been successful in improving multimodal large language models, the reward models themselves remain brutal and contain minimal information. Notably, existing reward models only mimic human annotations by assigning only one feedback to any text, no matter how long the text is. In the realm of multimodal language models, where models are required to process both images and texts, a naive reward model may learn implicit biases toward texts and become less grounded in images. In this paper, we propose a **T**oken-**L**evel **D**etective **R**eward Model (**TLDR**) to provide fine-grained annotations to each text token. We first introduce a perturbation-based method to generate synthetic hard negatives and their token-level labels to train TLDR models. Then we show the rich usefulness of TLDR models both in assisting off-the-shelf models to self-correct their generations, and in serving as a hallucination evaluation tool. We show that TLDR automatically trains a token-level likelihood optimization, and can improve the base model's performance significantly. Finally, we show that TLDR models can significantly speed up human annotation by 3 times to acquire a broader range of high-quality vision language data.

NeurIPS Conference 2025 Conference Paper

Towards a General Attention Framework on Gyrovector Spaces for Matrix Manifolds

  • Rui Wang
  • Chen Hu
  • Xiaoning Song
  • Xiaojun Wu
  • Nicu Sebe
  • Ziheng Chen

Deep neural networks operating on non-Euclidean geometries have recently demonstrated impressive performance across various machine-learning applications. Several studies have extended the attention mechanism to different manifolds. However, most existing non-Euclidean attention models are tailored to specific geometries, limiting their applicability. On the other hand, recent studies show that several matrix manifolds, such as Symmetric Positive Definite (SPD), Symmetric Positive Semi-Definite (SPSD), and Grassmannian manifolds, admit gyrovector structures, which extend vector addition and scalar product into manifolds. Leveraging these properties, we propose a Gyro Attention (GyroAtt) framework over general gyrovector spaces, applicable to various matrix geometries. Empirically, we manifest GyroAtt on three gyro structures on the SPD manifold, three on the SPSD manifold, and one on the Grassmannian manifold. Extensive experiments on four electroencephalography (EEG) datasets demonstrate the effectiveness of our framework.

IJCAI Conference 2025 Conference Paper

Towards Robust Incremental Learning Under Ambiguous Supervision

  • Rui Wang
  • Mingxuan Xia
  • Haobo Wang
  • Lei Feng
  • Junbo Zhao
  • Gang Chen
  • Chang Yao

Traditional Incremental Learning (IL) targets to handle sequential fully-supervised learning problems where novel classes emerge from time to time. However, due to inherent annotation uncertainty and ambiguity, collecting high-quality annotated data in a dynamic learning system can be extremely expensive. To mitigate this problem, we propose a novel weakly-supervised learning paradigm called Incremental Partial Label Learning (IPLL), where the sequentially arrived data relate to a set of candidate labels rather than the ground truth. Technically, we develop the Prototype-Guided Disambiguation and Replay Algorithm (PGDR) which leverages the class prototypes as a proxy to mitigate two intertwined challenges in IPLL, i. e. , label ambiguity and catastrophic forgetting. To handle the former, PGDR encapsulates a momentum-based pseudo-labeling algorithm along with prototype-guided initialization, resulting in a balanced perception of classes. To alleviate forgetting, we develop a memory replay technique that collects well-disambiguated samples while maintaining representativeness and diversity. By jointly distilling knowledge from curated memory data, our framework exhibits a great disambiguation ability for samples of new tasks and achieves less forgetting of knowledge. Extensive experiments demonstrate that PGDR achieves superior performance over the baselines in the IPLL task.

JBHI Journal 2025 Journal Article

TriCvT-DTI: Predicting Drug–Target Interactions Using Trimodal Representations and Convolutional Vision Transformers

  • Azouz Maroua
  • Gang Tian
  • Rui Wang
  • Jiehan Zhou

Predicting interactions between drugs and their targets is vital for drug discovery and repositioning. Conventional techniques are slow and labor-intensive, while deep learning algorithms offer efficient solutions. However, deep learning often focus on single drug representations or simplistic combinations, leading to suboptimal feature representation. Moreover, the prevalent use of convolutional neural networks (CNNs) in drug image representation neglects the necessity for both local and global drug information in Drug-Target Interaction (DTI) tasks. To address these challenges, we propose TriCvT-DTI, a novel approach that combines molecular images, chemical sequence features, and graph representations of drugs to comprehensively capture structural, spatial, and functional aspects. TriCvT-DTI introduces a bidirectional multi-head attention mechanism for interactive feature learning between drugs and targets, enhancing performance by modeling complex relationships. By using Convolutional Vision Transformers (CvTs), TriCvT-DTI can effectively extract structural and spatial features from drug images. We evaluate our model on three datasets: Human, C. elegans, and Davis, and we compare it with state-of-the-art methods. Then we train TriCvT-DTI with uni-modality and bi-modality to compare then extract the impact of each modality on TriCvT-DTI. Experimental results demonstrate that TriCvT-DTI outperforms existing methods on both balanced and unbalanced datasets. Moreover, it presents impressive generalization capabilities on the Drug-Target Interaction (DTI) task.

AAAI Conference 2025 Conference Paper

VIoTGPT: Learning to Schedule Vision Tools Towards Intelligent Video Internet of Things

  • Yaoyao Zhong
  • Mengshi Qi
  • Rui Wang
  • Yuhan Qiu
  • Yang Zhang
  • Huadong Ma

Video Internet of Things (VIoT) has shown full potential in collecting an unprecedented volume of video data. How to schedule the domain-specific perceiving models and analyze the collected videos uniformly, efficiently, and especially intelligently to accomplish complicated tasks is challenging. To address the challenge, we build VIoTGPT, the framework based on LLMs to correctly interact with humans, query knowledge videos, and invoke vision models to analyze multimedia data collaboratively. To support VIoTGPT and related future works, we meticulously crafted the VIoT-Tool dataset, including the training dataset and the benchmark involving 11 representative vision models across three categories based on semi-automatic annotations. To guide LLM to act as the intelligent agent towards intelligent VIoT, we resort to ReAct instruction tuning method based on VIoT-Tool to learn the tool capability. Quantitative and qualitative experiments and analyses demonstrate the effectiveness of VIoTGPT. We believe VIoTGPT contributes to improving human-centered experiences in VIoT applications.

TMLR Journal 2024 Journal Article

A general framework for formulating structured variable selection

  • Guanbo Wang
  • Mireille Schnitzer
  • Tom Chen
  • Rui Wang
  • Robert W Platt

In variable selection, a selection rule that prescribes the permissible sets of selected variables (called a "selection dictionary") is desirable due to the inherent structural constraints among the candidate variables. Such selection rules can be complex in real-world data analyses, and failing to incorporate such restrictions could not only compromise the interpretability of the model but also lead to decreased prediction accuracy. However, no general framework has been proposed to formalize selection rules and their applications, which poses a significant challenge for practitioners seeking to integrate these rules into their analyses. In this work, we establish a framework for structured variable selection that can incorporate universal structural constraints. We develop a mathematical language for constructing arbitrary selection rules, where the selection dictionary is formally defined. We demonstrate that all selection rules can be expressed as combinations of operations on constructs, facilitating the identification of the corresponding selection dictionary. We use a detailed and complex example to illustrate the developed framework. Once this selection dictionary is derived, practitioners can apply their own user-defined criteria to select the optimal model. Additionally, our framework enhances existing penalized regression methods for variable selection by providing guidance on how to appropriately group variables to achieve the desired selection rule. Furthermore, our innovative framework opens the door to establishing new $\ell_0$-based penalized regression techniques that can be tailored to respect arbitrary selection rules, thereby expanding the possibilities for more robust and tailored model development.

IJCAI Conference 2024 Conference Paper

A Grassmannian Manifold Self-Attention Network for Signal Classification

  • Rui Wang
  • Chen Hu
  • Ziheng Chen
  • Xiao-Jun Wu
  • Xiaoning Song

In the community of artificial intelligence, significant progress has been made in encoding sequential data using deep learning techniques. Nevertheless, how to effectively mine useful information from channel dimensions remains a major challenge, as these features have a submanifold structure. Linear subspace, the basic element of the Grassmannian manifold, has proven to be an effective manifold-valued feature descriptor in statistical representation. Besides, the Euclidean self-attention mechanism has shown great success in capturing long-range relationships of data. Inspired by these facts, we extend the self-attention mechanism to the Grassmannian manifold. Our framework can effectively characterize the spatiotemporal fluctuations of sequential data encoded in the Grassmannian manifold. Extensive experimental results on three benchmarking datasets (a drone recognition dataset and two EEG signal classification datasets) demonstrate the superiority of our method over the state-of-the-art. The code and supplementary material for this work can be found at https: //github. com/ChenHu-ML/GDLNet.

NeurIPS Conference 2024 Conference Paper

A Recipe for Charge Density Prediction

  • Xiang Fu
  • Andrew Rosen
  • Kyle Bystrom
  • Rui Wang
  • Albert Musaelian
  • Boris Kozinsky
  • Tess Smidt
  • Tommi Jaakkola

In density functional theory, charge density is the core attribute of atomic systems from which all chemical properties can be derived. Machine learning methods are promising in significantly accelerating charge density prediction, yet existing approaches either lack accuracy or scalability. We propose a recipe that can achieve both. In particular, we identify three key ingredients: (1) representing the charge density with atomic and virtual orbitals (spherical fields centered at atom/virtual coordinates); (2) using expressive and learnable orbital basis sets (basis function for the spherical fields); and (3) using high-capacity equivariant neural network architecture. Our method achieves state-of-the-art accuracy while being more than an order of magnitude faster than existing methods. Furthermore, our method enables flexible efficiency-accuracy trade-offs by adjusting the model/basis sizes.

NeurIPS Conference 2024 Conference Paper

DART-Math: Difficulty-Aware Rejection Tuning for Mathematical Problem-Solving

  • Yuxuan Tong
  • Xiwen Zhang
  • Rui Wang
  • Ruidong Wu
  • Junxian He

Solving mathematical problems requires advanced reasoning abilities and presents notable challenges for large language models. Previous works usually synthesize data from proprietary models to augment existing datasets, followed by instruction tuning to achieve top-tier results. However, our analysis of these datasets reveals severe biases towards easy queries, with frequent failures to generate any correct response for the most challenging queries. Hypothesizing that difficult queries are crucial to learning complex reasoning, we propose Difficulty-Aware Rejection Tuning ( DART ), a method that allocates difficult queries more trials during the synthesis phase, enabling more extensive training on difficult samples. Utilizing DART, we have created new datasets for mathematical problem-solving that focus more on difficult queries and are substantially smaller than previous ones. Remarkably, our synthesis process solely relies on a 7B-sized open-weight model, without reliance on the commonly used proprietary GPT-4. We fine-tune various base models on our datasets ranging from 7B to 70B in size, resulting in a series of strong models called DART-Math. In comprehensive in-domain and out-of-domain evaluation on 6 mathematical benchmarks, DART-Math outperforms vanilla rejection tuning significantly, being superior or comparable to previous arts, despite using much smaller datasets and no proprietary models. Furthermore, our results position our synthetic datasets as the most effective and cost-efficient publicly available resources for advancing mathematical problem-solving. Our datasets, models and code are publicly available at https: //github. com/hkust-nlp/dart-math.

NeurIPS Conference 2024 Conference Paper

DiffPhyCon: A Generative Approach to Control Complex Physical Systems

  • Long Wei
  • Peiyan Hu
  • Ruiqi Feng
  • Haodong Feng
  • Yixuan Du
  • Tao Zhang
  • Rui Wang
  • Yue Wang

Controlling the evolution of complex physical systems is a fundamental task across science and engineering. Classical techniques suffer from limited applicability or huge computational costs. On the other hand, recent deep learning and reinforcement learning-based approaches often struggle to optimize long-term control sequences under the constraints of system dynamics. In this work, we introduce Diffusion Physical systems Control (DiffPhyCon), a new class of method to address the physical systems control problem. DiffPhyCon excels by simultaneously minimizing both the learned generative energy function and the predefined control objectives across the entire trajectory and control sequence. Thus, it can explore globally and plan near-optimal control sequences. Moreover, we enhance DiffPhyCon with prior reweighting, enabling the discovery of control sequences that significantly deviate from the training distribution. We test our method on three tasks: 1D Burgers' equation, 2D jellyfish movement control, and 2D high-dimensional smoke control, where our generated jellyfish dataset is released as a benchmark for complex physical system control research. Our method outperforms widely applied classical approaches and state-of-the-art deep learning and reinforcement learning methods. Notably, DiffPhyCon unveils an intriguing fast-close-slow-open pattern observed in the jellyfish, aligning with established findings in the field of fluid dynamics. The project website, jellyfish dataset, and code can be found at https: //github. com/AI4Science-WestlakeU/diffphycon.

NeurIPS Conference 2024 Conference Paper

Embedding Trajectory for Out-of-Distribution Detection in Mathematical Reasoning

  • Yiming Wang
  • Pei Zhang
  • Baosong Yang
  • Derek F. Wong
  • Zhuosheng Zhang
  • Rui Wang

Real-world data deviating from the independent and identically distributed (\textit{i. i. d. }) assumption of in-distribution training data poses security threats to deep networks, thus advancing out-of-distribution (OOD) detection algorithms. Detection methods in generative language models (GLMs) mainly focus on uncertainty estimation and embedding distance measurement, with the latter proven to be most effective in traditional linguistic tasks like summarization and translation. However, another complex generative scenario mathematical reasoning poses significant challenges to embedding-based methods due to its high-density feature of output spaces, but this feature causes larger discrepancies in the embedding shift trajectory between different samples in latent spaces. Hence, we propose a trajectory-based method TV score, which uses trajectory volatility for OOD detection in mathematical reasoning. Experiments show that our method outperforms all traditional algorithms on GLMs under mathematical reasoning scenarios and can be extended to more applications with high-density features in output spaces, such as multiple-choice questions.

NeurIPS Conference 2024 Conference Paper

Enhancing Protein Mutation Effect Prediction through a Retrieval-Augmented Framework

  • Ruihan Guo
  • Rui Wang
  • Ruidong Wu
  • Zhizhou Ren
  • Jiahan Li
  • Shitong Luo
  • Zuofan Wu
  • Qiang Liu

Predicting the effects of protein mutations is crucial for analyzing protein functions and understanding genetic diseases. However, existing models struggle to effectively extract mutation-related local structure motifs from protein databases, which hinders their predictive accuracy and robustness. To tackle this problem, we design a novel retrieval-augmented framework for incorporating similar structure information in known protein structures. We create a vector database consisting of local structure motif embeddings from a pre-trained protein structure encoder, which allows for efficient retrieval of similar local structure motifs during mutation effect prediction. Our findings demonstrate that leveraging this method results in the SOTA performance across multiple protein mutation prediction datasets, and offers a scalable solution for studying mutation effects.

NeurIPS Conference 2024 Conference Paper

EnsIR: An Ensemble Algorithm for Image Restoration via Gaussian Mixture Models

  • Shangquan Sun
  • Wenqi Ren
  • Zikun Liu
  • Hyunhee Park
  • Rui Wang
  • Xiaochun Cao

Image restoration has experienced significant advancements due to the development of deep learning. Nevertheless, it encounters challenges related to ill-posed problems, resulting in deviations between single model predictions and ground-truths. Ensemble learning, as a powerful machine learning technique, aims to address these deviations by combining the predictions of multiple base models. Most existing works adopt ensemble learning during the design of restoration models, while only limited research focuses on the inference-stage ensemble of pre-trained restoration models. Regression-based methods fail to enable efficient inference, leading researchers in academia and industry to prefer averaging as their choice for post-training ensemble. To address this, we reformulate the ensemble problem of image restoration into Gaussian mixture models (GMMs) and employ an expectation maximization (EM)-based algorithm to estimate ensemble weights for aggregating prediction candidates. We estimate the range-wise ensemble weights on a reference set and store them in a lookup table (LUT) for efficient ensemble inference on the test set. Our algorithm is model-agnostic and training-free, allowing seamless integration and enhancement of various pre-trained image restoration models. It consistently outperforms regression-based methods and averaging ensemble approaches on 14 benchmarks across 3 image restoration tasks, including super-resolution, deblurring and deraining. The codes and all estimated weights have been released in Github.

IJCAI Conference 2024 Conference Paper

Error-aware Sampling in Adaptive Shells for Neural Surface Reconstruction

  • Qi Wang
  • Yuchi Huo
  • Qi Ye
  • Rui Wang
  • Hujun Bao

Neural implicit surfaces with signed distance functions (SDFs) achieve superior quality in 3D geometry reconstruction. However, training SDFs is time-consuming because it requires a great number of samples to calculate accurate weight distributions and a considerable amount of samples sampled from the distribution for integrating the rendering results. Some existing sampling strategies focus on this problem. During the training, they assume a spatially-consistent convergence speed of kernel size, thus still suffering from low convergence or errors. Instead, we introduce an error-aware sampling method based on thin intervals of valid weight distributions, dubbed adaptive shells, to reduce the number of samples while still maintaining the reconstruction accuracy. To this end, we first extend Laplace-based neural implicit surfaces with learned spatially-varying kernel sizes which indicates the range of valid weight distributions. Then, the adaptive shell for each ray is determined by an efficient double-clipping strategy with spatially-varying SDF values and kernel sizes, fitting larger kernel sizes to wider shells. Finally, we calculate the error-bounded cumulative distribution functions (CDFs) of shells to conduct efficient importance sampling, achieving low-variance rendering with fewer calculations. Extensive results in various scenes demonstrate the superiority of our sampling technique, including significantly reducing sample counts and training time, even improving the reconstruction quality. The code is available at https: //github. com/erernan/ESampling.

ICML Conference 2024 Conference Paper

FAFE: Immune Complex Modeling with Geodesic Distance Loss on Noisy Group Frames

  • Ruidong Wu
  • Ruihan Guo
  • Rui Wang
  • Shitong Luo
  • Yue Xu
  • Jiahan Li
  • Jianzhu Ma
  • Qiang Liu 0001

Despite the striking success of general protein folding models such as AlphaFold2 (AF2), the accurate computational modeling of antibody-antigen complexes remains a challenging task. In this paper, we first analyze AF2’s primary loss function, known as the Frame Aligned Point Error (FAPE), and raise a previously overlooked issue that FAPE tends to face gradient vanishing problem on high-rotational-error targets. To address this fundamental limitation, we propose a novel geodesic loss called Frame Aligned Frame Error (FAFE, denoted as F2E to distinguish from FAPE), which enables the model to better optimize both the rotational and translational errors between two frames. We then prove that F2E can be reformulated as a group-aware geodesic loss, which translates the optimization of the residue-to-residue error to optimizing group-to-group geodesic frame distance. By fine-tuning AF2 with our proposed new loss function, we attain a correct rate of 52. 3% (DockQ $>$ 0. 23) on an evaluation set and 43. 8% correct rate on a subset with low homology, with improvement over AF2 by 182% and 100% respectively.

AAAI Conference 2024 Conference Paper

Frequency Shuffling and Enhancement for Open Set Recognition

  • Lijun Liu
  • Rui Wang
  • Yuan Wang
  • Lihua Jing
  • Chuan Wang

Open-Set Recognition (OSR) aims to accurately identify known classes while effectively rejecting unknown classes to guarantee reliability. Most existing OSR methods focus on learning in the spatial domain, where subtle texture and global structure are potentially intertwined. Empirical studies have shown that DNNs trained in the original spatial domain are inclined to over-perceive subtle texture. The biased semantic perception could lead to catastrophic over-confidence when predicting both known and unknown classes. To this end, we propose an innovative approach by decomposing the spatial domain to the frequency domain to separately consider global (low-frequency) and subtle (high-frequency) information, named Frequency Shuffling and Enhancement (FreSH). To alleviate the overfitting of subtle texture, we introduce the High-Frequency Shuffling (HFS) strategy that generates diverse high-frequency information and promotes the capture of low-frequency invariance. Moreover, to enhance the perception of global structure, we propose the Low-Frequency Residual (LFR) learning procedure that constructs a composite feature space, integrating low-frequency and original spatial features. Experiments on various benchmarks demonstrate that the proposed FreSH consistently trumps the state-of-the-arts by a considerable margin.

ICRA Conference 2024 Conference Paper

Leveraging Neural Radiance Fields for Uncertainty-Aware Visual Localization

  • Le Chen
  • Weirong Chen
  • Rui Wang
  • Marc Pollefeys

As a promising fashion for visual localization, scene coordinate regression (SCR) has seen tremendous progress in the past decade. Most recent methods usually adopt neural networks to learn the mapping from image pixels to 3D scene coordinates, which requires a vast amount of annotated training data. We propose to leverage Neural Radiance Fields (NeRF) to generate training samples for SCR. Despite NeRF’s efficiency in rendering, many of the rendered data are polluted by artifacts or only contain minimal information gain, which can hinder the regression accuracy or bring unnecessary computational costs with redundant data. These challenges are addressed in three folds in this paper: (1) A NeRF is designed to separately predict uncertainties for the rendered color and depth images, which reveal data reliability at the pixel level. (2) SCR is formulated as deep evidential learning with epistemic uncertainty, which is used to evaluate information gain and scene coordinate quality. (3) Based on the three arts of uncertainties, a novel view selection policy is formed that significantly improves data efficiency. Experiments on public datasets demonstrate that our method could select the samples that bring the most information gain and promote the performance with the highest efficiency.

YNIMG Journal 2024 Journal Article

Noninvasive focused ultrasound-mediated delivery of rAAV9-EGFP vectors for neuronal targeting in rats

  • Rui Wang
  • Jiayi Li

OBJECTIVE: To evaluate the synergistic potential of Focused Ultrasound (FUS) in conjunction with microbubbles (MB) and recombinant adeno-associated virus serotype 9 (rAAV9) vectors for targeted gene delivery to neuronal cells in rats, optimizing gene expression conditions and assessing any adverse effects. METHODS: The parameters for permeability enhancement of the rat's blood-brain barrier (BBB) were established using FUS+MB, with MRI scans and Evans Blue (EB) dye assisting in the evaluation. Rats underwent FUS-mediated transfection using rAAV9-Syn-EGFP vectors produced via a triple-transfection in HEK293T cells. Following this, the uptake and expression of GFP in targeted brain regions were evaluated using confocal fluorescence microscopy at various time intervals. Inflammatory responses post-FUS treatment were tracked by observing levels of GFAP, a marker for astrocytic activation, and TNF-α, a pro-inflammatory cytokine. Motor behavior effects post-intervention were gauged using the Rotarod test across multiple groups over a span of four weeks. RESULTS: FUS+MB affected BBB permeability, with optimal results at 4 W for 200 s showing 85 % permeability and evident Gd-DTPA leakage. Settings beyond these resulted in tissue damage. Control groups exhibited a basal GFP expression of 2 % ± 0.5 %, whereas FUS+MB with rAAV-EGFP injections substantially increased GFP expression to about 67 % ± 6 % in targeted neurons. This GFP expression peaked at three weeks post-treatment and remained evident six months later. Following FUS treatment, both GFAP and TNF-α levels underwent fluctuations before eventually nearing their baseline values. The Rotarod test revealed no significant behavioral differences post-treatments among the groups. CONCLUSIONS: Combining FUS+MB with rAAV offers an innovative approach to enhance therapeutic delivery to the central nervous system (CNS) by transiently adjusting BBB permeability.

NeurIPS Conference 2024 Conference Paper

Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning

  • Bei Li
  • Tong Zheng
  • Rui Wang
  • Jiahao Liu
  • Qingyan Guo
  • Junliang Guo
  • Xu Tan
  • Tong Xiao

Residual networks, as discrete approximations of Ordinary Differential Equations (ODEs), have inspired significant advancements in neural network design, including multistep methods, high-order methods, and multi-particle dynamical systems. The precision of the solution to ODEs significantly affects parameter optimization, thereby impacting model performance. In this work, we present a series of advanced explorations of Transformer architecture design to minimize the error compared to the true ``solution. '' First, we introduce a predictor-corrector learning framework to minimize truncation errors, which consists of a high-order predictor and a multistep corrector. Second, we propose an exponential moving average-based coefficient learning method to strengthen our higher-order predictor. Extensive experiments on large-scale machine translation, abstractive summarization, language modeling, and natural language understanding benchmarks demonstrate the superiority of our approach. On the WMT'14 English-German and English-French tasks, our model achieved BLEU scores of 30. 95 and 44. 27, respectively. Furthermore, on the OPUS multilingual machine translation task, our model surpasses a robust 3. 8B DeepNet by an average of 2. 9 SacreBLEU, using only 1/3 parameters. Notably, it also beats LLama models by 5. 7 accuracy points on the LM Harness Evaluation.

AAAI Conference 2024 Conference Paper

Rethinking Dimensional Rationale in Graph Contrastive Learning from Causal Perspective

  • Qirui Ji
  • Jiangmeng Li
  • Jie Hu
  • Rui Wang
  • Changwen Zheng
  • Fanjiang Xu

Graph contrastive learning is a general learning paradigm excelling at capturing invariant information from diverse perturbations in graphs. Recent works focus on exploring the structural rationale from graphs, thereby increasing the discriminability of the invariant information. However, such methods may incur in the mis-learning of graph models towards the interpretability of graphs, and thus the learned noisy and task-agnostic information interferes with the prediction of graphs. To this end, with the purpose of exploring the intrinsic rationale of graphs, we accordingly propose to capture the dimensional rationale from graphs, which has not received sufficient attention in the literature. The conducted exploratory experiments attest to the feasibility of the aforementioned roadmap. To elucidate the innate mechanism behind the performance improvement arising from the dimensional rationale, we rethink the dimensional rationale in graph contrastive learning from a causal perspective and further formalize the causality among the variables in the pre-training stage to build the corresponding structural causal model. On the basis of the understanding of the structural causal model, we propose the dimensional rationale-aware graph contrastive learning approach, which introduces a learnable dimensional rationale acquiring network and a redundancy reduction constraint. The learnable dimensional rationale acquiring network is updated by leveraging a bi-level meta-learning technique, and the redundancy reduction constraint disentangles the redundant features through a decorrelation process during learning. Empirically, compared with state-of-the-art methods, our method can yield significant performance boosts on various benchmarks with respect to discriminability and transferability. The code implementation of our method is available at https://github.com/ByronJi/DRGCL.

NeurIPS Conference 2024 Conference Paper

RMLR: Extending Multinomial Logistic Regression into General Geometries

  • Ziheng Chen
  • Yue Song
  • Rui Wang
  • Xiao-Jun Wu
  • Nicu Sebe

Riemannian neural networks, which extend deep learning techniques to Riemannian spaces, have gained significant attention in machine learning. To better classify the manifold-valued features, researchers have started extending Euclidean multinomial logistic regression (MLR) into Riemannian manifolds. However, existing approaches suffer from limited applicability due to their strong reliance on specific geometric properties. This paper proposes a framework for designing Riemannian MLR over general geometries, referred to as RMLR. Our framework only requires minimal geometric properties, thus exhibiting broad applicability and enabling its use with a wide range of geometries. Specifically, we showcase our framework on the Symmetric Positive Definite (SPD) manifold and special orthogonal group, i. e. , the set of rotation matrices. On the SPD manifold, we develop five families of SPD MLRs under five types of power-deformed metrics. On rotation matrices we propose Lie MLR based on the popular bi-invariant metric. Extensive experiments on different Riemannian backbone networks validate the effectiveness of our framework.

JMLR Journal 2024 Journal Article

Sparse Representer Theorems for Learning in Reproducing Kernel Banach Spaces

  • Rui Wang
  • Yuesheng Xu
  • Mingsong Yan

Sparsity of a learning solution is a desirable feature in machine learning. Certain reproducing kernel Banach spaces (RKBSs) are appropriate hypothesis spaces for sparse learning methods. The goal of this paper is to understand what kind of RKBSs can promote sparsity for learning solutions. We consider two typical learning models in an RKBS: the minimum norm interpolation (MNI) problem and the regularization problem. We first establish an explicit representer theorem for solutions of these problems, which represents the extreme points of the solution set by a linear combination of the extreme points of the subdifferential set, of the norm function, which is data-dependent. We then propose sufficient conditions on the RKBS that can transform the explicit representation of the solutions to a sparse kernel representation having fewer terms than the number of the observed data. Under the proposed sufficient conditions, we investigate the role of the regularization parameter on sparsity of the regularized solutions. We further show that two specific RKBSs, the sequence space $\ell_1(\mathbb{N})$ and the measure space, can have sparse representer theorems for both MNI and regularization models. [abs] [ pdf ][ bib ] &copy JMLR 2024. ( edit, beta )

AAAI Conference 2024 Conference Paper

T2MAC: Targeted and Trusted Multi-Agent Communication through Selective Engagement and Evidence-Driven Integration

  • Chuxiong Sun
  • Zehua Zang
  • Jiabao Li
  • Jiangmeng Li
  • Xiao Xu
  • Rui Wang
  • Changwen Zheng

Communication stands as a potent mechanism to harmonize the behaviors of multiple agents. However, existing work primarily concentrates on broadcast communication, which not only lacks practicality, but also leads to information redundancy. This surplus, one-fits-all information could adversely impact the communication efficiency. Furthermore, existing works often resort to basic mechanisms to integrate observed and received information, impairing the learning process. To tackle these difficulties, we propose Targeted and Trusted Multi-Agent Communication (T2MAC), a straightforward yet effective method that enables agents to learn selective engagement and evidence-driven integration. With T2MAC, agents have the capability to craft individualized messages, pinpoint ideal communication windows, and engage with reliable partners, thereby refining communication efficiency. Following the reception of messages, the agents integrate information observed and received from different sources at an evidence level. This process enables agents to collectively use evidence garnered from multiple perspectives, fostering trusted and cooperative behaviors. We evaluate our method on a diverse set of cooperative multi-agent tasks, with varying difficulties, involving different scales and ranging from Hallway, MPE to SMAC. The experiments indicate that the proposed model not only surpasses the state-of-the-art methods in terms of cooperative performance and communication efficiency, but also exhibits impressive generalization.

TMLR Journal 2024 Journal Article

Unifying the Perspectives of NLP and Software Engineering: A Survey on Language Models for Code

  • Ziyin Zhang
  • Chaoyu Chen
  • Bingchang Liu
  • Cong Liao
  • Zi Gong
  • Hang Yu
  • Jianguo Li
  • Rui Wang

In this work we systematically review the recent advancements in software engineering with language models, covering 70+ models, 40+ evaluation tasks, 180+ datasets, and 900 related works. Unlike previous works, we integrate software engineering (SE) with natural language processing (NLP) by discussing the perspectives of both sides: SE applies language models for development automation, while NLP adopts SE tasks for language model evaluation. We break down code processing models into general language models represented by the GPT family and specialized models that are specifically pretrained on code, often with tailored objectives. We discuss the relations and differences between these models, and highlight the historical transition of code modeling from statistical models and RNNs to pretrained Transformers and LLMs, which is exactly the same course that had been taken by NLP. We also go beyond programming and review LLMs' application in other software engineering activities including requirement engineering, testing, deployment, and operations in an endeavor to provide a global view of NLP in SE, and identify key challenges and potential future directions in this domain.

NeurIPS Conference 2023 Conference Paper

A Fast and Accurate Estimator for Large Scale Linear Model via Data Averaging

  • Rui Wang
  • Yanyan Ouyang
  • Yu Panpan
  • Wangli Xu

This work is concerned with the estimation problem of linear model when thesample size is extremely large and the data dimension can vary with the samplesize. In this setting, the least square estimator based on the full data is not feasiblewith limited computational resources. Many existing methods for this problem arebased on the sketching technique which uses the sketched data to perform leastsquare estimation. We derive fine-grained lower bounds of the conditional meansquared error for sketching methods. For sampling methods, our lower boundprovides an attainable optimal convergence rate. Our result implies that when thedimension is large, there is hardly a sampling method can have a faster convergencerate than the uniform sampling method. To achieve a better statistical performance, we propose a new sketching method based on data averaging. The proposedmethod reduces the original data to a few averaged observations. These averagedobservations still satisfy the linear model and are used to estimate the regressioncoefficients. The asymptotic behavior of the proposed estimation procedure isstudied. Our theoretical results show that the proposed method can achieve afaster convergence rate than the optimal convergence rate for sampling methods. Theoretical and numerical results show that the proposed estimator has goodstatistical performance as well as low computational cost.

EAAI Journal 2023 Journal Article

AdaBoost-driven multi-parameter real-time warning of rock burst risk in coal mines

  • Rui Wang
  • Shaojie Chen
  • Xuelong Li
  • Gang Tian
  • Tongbin Zhao

The stope dynamic disaster, which occurs around the mining space and is represented by stress-type rock burst and fracture-type rock burst, seriously affects the safety production of coal mine. How to effectively forewarn rock burst risk to reduce the disaster caused by rock burst is an urgent problem to be solved in stope. In this paper, the spatiotemporal parameters that affect the occurrence of stope dynamic disaster are collected, and the big data of stope state is established. Due to the complex temporal and spatial parameters of rock bursts with different intensities occurrence and their varying degrees, it is difficult for some machine learning methods to excavate their internal relationships and make accurate warning. In this paper, (1) a big data platform is used to record and integrate the multiple parameters of stope dynamic disaster, and data preprocessing technology is used to de-noise and standardize them. (2) Based on the fused historical data, a strong classifier is iteratively found by the classification algorithm AdaBoost to identify the existence of the rock burst risk, so as to achieve the purpose of accurately and timely warning of rock burst risk. The research of this paper uses big data mining technology and machine learning method to carry out intelligent perception warning of the rock burst risk in real time. The experimental results show that the proposed method has a good effect and is of great significance to the prevention and control of rock burst disaster in stope.

EAAI Journal 2023 Journal Article

Deep reinforcement learning-PID based supervisor control method for indirect-contact heat transfer processes in energy systems

  • Xuan Wang
  • Jinwen Cai
  • Rui Wang
  • Gequn Shu
  • Hua Tian
  • Mingtao Wang
  • Bowen Yan

Indirect-contact heat exchangers have been widely used in various energy systems, and the precise tracking control of important heat transfer parameters, such as temperature, is vital for safe and efficient operation. However, the high nonlinearity of heat transfer and large disturbance brings difficulty to optimal control. Considering the strong perception and decision-making capabilities of deep reinforcement learning (DRL), this study proposed a supervisor control method combined DRL and proportional–integral–derivative (PID). A set of the fewest conveniently measurable variables was derived as agent observations to describe the heat transfer process effectively and thereby improve the control efficiency under large disturbances. In addition, the local heat transfer process was used as a training environment to reduce training costs significantly. Finally, superheat temperature control in a complex organic Rankine cycle was simulated with SIMULINK to evaluate the effectiveness of the proposed observation variables and the training and control methods. The results showed that the proposed control method achieved satisfactory performance. The average absolute tracking error was only 0. 246 K under trained and untrained disturbances, whereas that of the PID control was 4. 645 K. Compared with the model predictive control, the DRL-PID-based supervisory control evidently performed better under a large disturbance; the average absolute tracking errors under DRL-PID control and MPC were 0. 288 K and 0. 509 K, respectively.

AAAI Conference 2023 Conference Paper

Few-Shot Composition Learning for Image Retrieval with Prompt Tuning

  • Junda Wu
  • Rui Wang
  • Handong Zhao
  • Ruiyi Zhang
  • Chaochao Lu
  • Shuai Li
  • Ricardo Henao

We study the problem of composition learning for image retrieval, for which we learn to retrieve target images with search queries in the form of a composition of a reference image and a modification text that describes desired modifications of the image. Existing models of composition learning for image retrieval are generally built with large-scale datasets, demanding extensive training samples, i.e., query-target pairs, as supervision, which restricts their application for the scenario of few-shot learning with only few query-target pairs available. Recently, prompt tuning with frozen pretrained language models has shown remarkable performance when the amount of training data is limited. Inspired by this, we propose a prompt tuning mechanism with the pretrained CLIP model for the task of few-shot composition learning for image retrieval. Specifically, we regard the representation of the reference image as a trainable visual prompt, prefixed to the embedding of the text sequence. One challenge is to efficiently train visual prompt with few-shot samples. To deal with this issue, we further propose a self-upervised auxiliary task via ensuring that the reference image can retrieve itself when no modification information is given from the text, which facilitates training for the visual prompt, while not requiring additional annotations for query-target pairs. Experiments on multiple benchmarks show that our proposed model can yield superior performance when trained with only few query-target pairs.

NeurIPS Conference 2023 Conference Paper

InfoPrompt: Information-Theoretic Soft Prompt Tuning for Natural Language Understanding

  • Junda Wu
  • Tong Yu
  • Rui Wang
  • Zhao Song
  • Ruiyi Zhang
  • Handong Zhao
  • Chaochao Lu
  • Shuai Li

Soft prompt tuning achieves superior performances across a wide range of few-shot tasks. However, the performances of prompt tuning can be highly sensitive to the initialization of the prompts. We have also empirically observed that conventional prompt tuning methods cannot encode and learn sufficient task-relevant information from prompt tokens. In this work, we develop an information-theoretic framework that formulates soft prompt tuning as maximizing the mutual information between prompts and other model parameters (or encoded representations). This novel view helps us to develop a more efficient, accurate and robust soft prompt tuning method, InfoPrompt. With this framework, we develop two novel mutual information based loss functions, to (i) explore proper prompt initialization for the downstream tasks and learn sufficient task-relevant information from prompt tokens and (ii) encourage the output representation from the pretrained language model to be more aware of the task-relevant information captured in the learnt prompts. Extensive experiments validate that InfoPrompt can significantly accelerate the convergence of the prompt tuning and outperform traditional prompt tuning methods. Finally, we provide a formal theoretical result to show that a gradient descent type algorithm can be used to train our mutual information loss.

NeurIPS Conference 2023 Conference Paper

Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification

  • Rui Wang
  • Peipei Li
  • Huaibo Huang
  • Chunshui Cao
  • Ran He
  • Zhaofeng He

We present a novel language-driven ordering alignment method for ordinal classification. The labels in ordinal classification contain additional ordering relations, making them prone to overfitting when relying solely on training data. Recent developments in pre-trained vision-language models inspire us to leverage the rich ordinal priors in human language by converting the original task into a vision-language alignment task. Consequently, we propose L2RCLIP, which fully utilizes the language priors from two perspectives. First, we introduce a complementary prompt tuning technique called RankFormer, designed to enhance the ordering relation of original rank prompts. It employs token-level attention with residual-style prompt blending in the word embedding space. Second, to further incorporate language priors, we revisit the approximate bound optimization of vanilla cross-entropy loss and restructure it within the cross-modal embedding space. Consequently, we propose a cross-modal ordinal pairwise loss to refine the CLIP feature space, where texts and images maintain both semantic alignment and ordering alignment. Extensive experiments on three ordinal classification tasks, including facial age estimation, historical color image (HCI) classification, and aesthetic assessment demonstrate its promising performance.

NeurIPS Conference 2023 Conference Paper

Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation

  • Zibo Zhao
  • Wen Liu
  • Xin Chen
  • Xianfang Zeng
  • Rui Wang
  • Pei Cheng
  • Bin Fu
  • Tao Chen

We present a novel alignment-before-generation approach to tackle the challenging task of generating general 3D shapes based on 2D images or texts. Directly learning a conditional generative model from images or texts to 3D shapes is prone to producing inconsistent results with the conditions because 3D shapes have an additional dimension whose distribution significantly differs from that of 2D images and texts. To bridge the domain gap among the three modalities and facilitate multi-modal-conditioned 3D shape generation, we explore representing 3D shapes in a shape-image-text-aligned space. Our framework comprises two models: a Shape-Image-Text-Aligned Variational Auto-Encoder (SITA-VAE) and a conditional Aligned Shape Latent Diffusion Model (ASLDM). The former model encodes the 3D shapes into the shape latent space aligned to the image and text and reconstructs the fine-grained 3D neural fields corresponding to given shape embeddings via the transformer-based decoder. The latter model learns a probabilistic mapping function from the image or text space to the latent shape space. Our extensive experiments demonstrate that our proposed approach can generate higher-quality and more diverse 3D shapes that better semantically conform to the visual or textural conditional inputs, validating the effectiveness of the shape-image-text-aligned space for cross-modality 3D shape generation.

EAAI Journal 2023 Journal Article

Multi-stage distribution correction: A promising data augmentation method for few-shot fault diagnosis

  • Xiao Zhang
  • Weiguo Huang
  • Rui Wang
  • Yi Liao
  • Chuancang Ding
  • Jun Wang
  • Juanjuan Shi

Benefiting from the excellent capability of data processing, deep learning-based methods have been well applied in fault diagnosis. However, these methods may perform poorly due to lack of labeled data for training in real-world applications In few-shot learning settings, these methods may be trapped in overfitting since the data distribution estimated from a small number of labeled samples may be overly biased, which cannot cover the ground-truth data distribution. Given this problem, we propose a data augmentation method named Multi-Stage Distribution Correction (MSDC) for few-shot fault diagnosis. The proposed method can be divided into four training stages. In the first stage, unlabeled query features are clustered into several groups via an unsupervised Fuzzy C-means algorithm. Secondly, according to cosine similarity, the few labeled support features are assigned to the cluster with the highest similarity. Then, the Gaussian statistics of each cluster are extracted to correct the distributions of support features. Specifically, Our proposed method assumes that each dimension of the feature representation from the same class follows Gaussian distribution. New labeled features are generated by sampling from the corrected distributions. These generated features, along with the labeled support features, are used to train the task-specific classifier in the last stage. By joint training of the four stages, the scale of the labeled training dataset is effectively expanded, and the classification performance of the proposed method can be enhanced. The experimental results reveal that the proposed method outperforms the selected comparative methods in two case studies.

YNIMG Journal 2023 Journal Article

Multimodal and multiscale evidence for network-based cortical thinning in major depressive disorder

  • Junle Li
  • Rui Wang
  • Ning Mao
  • Manli Huang
  • Shijun Qiu
  • Jinhui Wang

BACKGROUND: Major depressive disorder (MDD) is associated with widespread, irregular cortical thickness (CT) reductions across the brain. However, little is known regarding mechanisms that govern spatial distribution of the reductions. METHODS: We combined multimodal MRI and genetic, cytoarchitectonic and chemoarchitectonic data to examine structural covariance, functional synchronization, gene co-expression, cytoarchitectonic similarity and chemoarchitectonic covariance between regions atrophied in MDD. RESULTS: Regions atrophied in MDD were associated with significantly higher structural covariance, functional synchronization, gene co-expression and chemoarchitectonic covariance. These results were robust against methodological variations in brain parcellation and null model, reproducible in patients and controls, and independent of age at onset of MDD. Despite no significant differences in the cytoarchitectonic similarity, MDD-related CT reductions were susceptible to specific cytoarchitectonic class of association cortex. Further, we found that nodal shortest path lengths to disease epicenters derived from structural (right supramarginal gyrus) and chemoarchitectonic covariance (right sulcus intermedius primus) networks of healthy brains were correlated with the extent to which a region was atrophied in MDD, supporting the transneuronal spread hypothesis that regions closer to the epicenters are more susceptible to MDD. Finally, we showed that structural covariance and functional synchronization among regions atrophied in MDD were mainly related to genes enriched in metabolic and membrane-related processes, driven by genes in excitatory neurons, and associated with specific neurotransmitter transporters and receptors. CONCLUSIONS: Altogether, our findings provide empirical evidence for and genetic and molecular insights into connectivity-constrained CT thinning in MDD.

EAAI Journal 2023 Journal Article

Revisiting the consistency improvement and consensus reaching processes of intuitionistic multiplicative preference relations

  • Rui Wang
  • Zhen-Song Chen
  • Bin Shuai
  • Luis Martínez
  • Wen-Tao Kong
  • Witold Pedrycz

Intuitionistic multiplicative preference relation (IMPR) has been successfully and widely used to model the decision makers’ preferences elicited by pairwise comparisons. As the critical links of group decision-making (GDM) procedure, the consistency improvement process (CIP) and consensus reaching process (CRP) have been investigated intensively; nevertheless, the extant methods explored the CIP and CRP of IMPRs without considering the situations of whether decision makers participate in the reassessment task. To overcome the limitation, this paper proposes two novel IMPR-based GDM approaches with CIP and CRP under different situations. First, when decision makers only express their original preferences, the novel consistency index and group consensus measure of IMPRs are defined to construct the CIP and CRP without feedback mechanism, in which the original preferences can be updated automatically. Second, when decision makers are required to modify their opinions, a series of consistency, consensus, and proximity degrees are put forwards to generate the identification and direction rules for improving the inconsistent and conflict preferences. Subsequently, the induced intuitionistic multiplicative ordered geometric averaging operator is constructed to fuse the individual preferences for determining the GDM results, in which the collective IMPR can maintain the acceptable consistency. Finally, a numerical example is applied to illustrate the feasibility and advantages of the proposed methods, which can improve the consistency and consensus levels of IMPRs according to the degrees of decision makers’ involvement in the re-evaluation process.

AAAI Conference 2023 Conference Paper

Riemannian Local Mechanism for SPD Neural Networks

  • Ziheng Chen
  • Tianyang Xu
  • Xiao-Jun Wu
  • Rui Wang
  • Zhiwu Huang
  • Josef Kittler

The Symmetric Positive Definite (SPD) matrices have received wide attention for data representation in many scientific areas. Although there are many different attempts to develop effective deep architectures for data processing on the Riemannian manifold of SPD matrices, very few solutions explicitly mine the local geometrical information in deep SPD feature representations. Given the great success of local mechanisms in Euclidean methods, we argue that it is of utmost importance to ensure the preservation of local geometric information in the SPD networks. We first analyse the convolution operator commonly used for capturing local information in Euclidean deep networks from the perspective of a higher level of abstraction afforded by category theory. Based on this analysis, we define the local information in the SPD manifold and design a multi-scale submanifold block for mining local geometry. Experiments involving multiple visual tasks validate the effectiveness of our approach.

AAAI Conference 2023 Conference Paper

SoftCorrect: Error Correction with Soft Detection for Automatic Speech Recognition

  • Yichong Leng
  • Xu Tan
  • Wenjie Liu
  • Kaitao Song
  • Rui Wang
  • Xiang-Yang Li
  • Tao Qin
  • Ed Lin

Error correction in automatic speech recognition (ASR) aims to correct those incorrect words in sentences generated by ASR models. Since recent ASR models usually have low word error rate (WER), to avoid affecting originally correct tokens, error correction models should only modify incorrect words, and therefore detecting incorrect words is important for error correction. Previous works on error correction either implicitly detect error words through target-source attention or CTC (connectionist temporal classification) loss, or explicitly locate specific deletion/substitution/insertion errors. However, implicit error detection does not provide clear signal about which tokens are incorrect and explicit error detection suffers from low detection accuracy. In this paper, we propose SoftCorrect with a soft error detection mechanism to avoid the limitations of both explicit and implicit error detection. Specifically, we first detect whether a token is correct or not through a probability produced by a dedicatedly designed language model, and then design a constrained CTC loss that only duplicates the detected incorrect tokens to let the decoder focus on the correction of error tokens. Compared with implicit error detection with CTC loss, SoftCorrect provides explicit signal about which words are incorrect and thus does not need to duplicate every token but only incorrect tokens; compared with explicit error detection, SoftCorrect does not detect specific deletion/substitution/insertion errors but just leaves it to CTC loss. Experiments on AISHELL-1 and Aidatatang datasets show that SoftCorrect achieves 26.1% and 9.4% CER reduction respectively, outperforming previous works by a large margin, while still enjoying fast speed of parallel generation.

AAAI Conference 2023 Conference Paper

Trafformer: Unify Time and Space in Traffic Prediction

  • Di Jin
  • Jiayi Shi
  • Rui Wang
  • Yawen Li
  • Yuxiao Huang
  • Yu-Bin Yang

Traffic prediction is an important component of the intelligent transportation system. Existing deep learning methods encode temporal information and spatial information separately or iteratively. However, the spatial and temporal information is highly correlated in a traffic network, so existing methods may not learn the complex spatial-temporal dependencies hidden in the traffic network due to the decomposed model design. To overcome this limitation, we propose a new model named Trafformer, which unifies spatial and temporal information in one transformer-style model. Trafformer enables every node at every timestamp interact with every other node in every other timestamp in just one step in the spatial-temporal correlation matrix. This design enables Trafformer to catch complex spatial-temporal dependencies. Following the same design principle, we use the generative style decoder to predict multiple timestamps in only one forward operation instead of the iterative style decoder in Transformer. Furthermore, to reduce the complexity brought about by the huge spatial-temporal self-attention matrix, we also propose two variants of Trafformer to further improve the training and inference speed without losing much effectivity. Extensive experiments on two traffic datasets demonstrate that Trafformer outperforms existing methods and provides a promising future direction for the spatial-temporal traffic prediction problem.

IJCAI Conference 2022 Conference Paper

Effective Graph Context Representation for Document-level Machine Translation

  • Kehai Chen
  • Muyun Yang
  • Masao Utiyama
  • Eiichiro Sumita
  • Rui Wang
  • Min Zhang

Document-level neural machine translation (DocNMT) universally encodes several local sentences or the entire document. Thus, DocNMT does not consider the relevance of document-level contextual information, for example, some context (i. e. , content words, logical order, and co-occurrence relation) is more effective than another auxiliary context (i. e. , functional and auxiliary words). To address this issue, we first utilize the word frequency information to recognize content words in the input document, and then use heuristical relations to summarize content words and sentences as a graph structure without relying on external syntactic knowledge. Furthermore, we apply graph attention networks to this graph structure to learn its feature representation, which allows DocNMT to more effectively capture the document-level context. Experimental results on several widely-used document-level benchmarks demonstrated the effectiveness of the proposed approach.

NeurIPS Conference 2022 Conference Paper

Few-Shot Fast-Adaptive Anomaly Detection

  • Ze Wang
  • Yipin Zhou
  • Rui Wang
  • Tsung-Yu Lin
  • Ashish Shah
  • Ser Nam Lim

The ability to detect anomaly has long been recognized as an inherent human ability, yet to date, practical AI solutions to mimic such capability have been lacking. This lack of progress can be attributed to several factors. To begin with, the distribution of ``abnormalities'' is intractable. Anything outside of a given normal population is by definition an anomaly. This explains why a large volume of work in this area has been dedicated to modeling the normal distribution of a given task followed by detecting deviations from it. This direction is however unsatisfying as it would require modeling the normal distribution of every task that comes along, which includes tedious data collection. In this paper, we report our work aiming to handle these issues. To deal with the intractability of abnormal distribution, we leverage Energy Based Model (EBM). EBMs learn to associates low energies to correct values and higher energies to incorrect values. At its core, the EBM employs Langevin Dynamics (LD) in generating these incorrect samples based on an iterative optimization procedure, alleviating the intractable problem of modeling the world of anomalies. Then, in order to avoid training an anomaly detector for every task, we utilize an adaptive sparse coding layer. Our intention is to design a plug and play feature that can be used to quickly update what is normal during inference time. Lastly, to avoid tedious data collection, this mentioned update of the sparse coding layer needs to be achievable with just a few shots. Here, we employ a meta learning scheme that simulates such a few shot setting during training. We support our findings with strong empirical evidence.

JBHI Journal 2022 Journal Article

GNN-Based Depression Recognition Using Spatio-Temporal Information: A fNIRS Study

  • Qiao Yu
  • Rui Wang
  • Jia Liu
  • Long Hu
  • Min Chen
  • Zhongchun Liu

In recent years, depression has become an increasingly serious problem globally. Previous studies of automatic depression recognition based on functional near-Infrared spectroscopy (fNIRS) or other brain imaging techniques have shown potential to serve as auxiliary diagnosis methods that provide assistance to medical professionals. Recently, some studies have found that, besides directly using the data themselves (temporal data), the use of functional connectivity among channels (spatial data) also can be effective. In this paper, we propose a method based on Graph Neural Network (GNN) that combines both temporal and spatial features of fNIRS data for automatic depression recognition. Specifically, fNIRS data of 96 subjects were collected and pre-processed. Basic statistical metrics of each channel were extracted as temporal features, and channel connectivity (coherence and correlation) were calculated as spatial features. Point-biserial analysis was conducted on these features and depression labels as a data-driven motivation. For classification, we considered data of each subject as a graph, with temporal features as node features and spatial features as edge weights. The graphs were fed into GNNs for training and testing. Experimental results showed that our GNN-based methods realized the best depression recognition performance compared with classical machine-learning methods regarding accuracy, F1 score, and precision, especially in F1 score for over 10%.

ICML Conference 2022 Conference Paper

Iterative Double Sketching for Faster Least-Squares Optimization

  • Rui Wang
  • Yanyan Ouyang
  • Wangli Xu

This work is concerned with the overdetermined linear least-squares problem for large scale data. We generalize the iterative Hessian sketching (IHS) algorithm and propose a new sketching framework named iterative double sketching (IDS) which uses approximations for both the gradient and the Hessian in each iteration. To understand the behavior of the IDS algorithm and choose the optimal hyperparameters, we derive the exact limit of the conditional prediction error of the IDS algorithm in the setting of Gaussian sketching. Guided by this theoretical result, we propose an efficient IDS algorithm via a new class of sequentially related sketching matrices. We give a non-asymptotic analysis of this efficient IDS algorithm which shows that the proposed algorithm achieves the state-of-the-art trade-off between accuracy and efficiency.

NeurIPS Conference 2022 Conference Paper

Meta-Learning Dynamics Forecasting Using Task Inference

  • Rui Wang
  • Robin Walters
  • Rose Yu

Current deep learning models for dynamics forecasting struggle with generalization. They can only forecast in a specific domain and fail when applied to systems with different parameters, external forces, or boundary conditions. We propose a model-based meta-learning method called DyAd which can generalize across heterogeneous domains by partitioning them into different tasks. DyAd has two parts: an encoder that infers the time-invariant hidden features of the task with weak supervision, and a forecaster which learns the shared dynamics of the entire domain. The encoder adapts and controls the forecaster during inference using adaptive instance normalization and adaptive padding. Theoretically, we prove that the generalization error of such a procedure is related to the task relatedness in the source domain, as well as the domain differences between source and target. Experimentally, we demonstrate that our model outperforms state-of-the-art approaches on forecasting complex physical dynamics including turbulent flow, real-world sea surface temperature, and ocean currents.

NeurIPS Conference 2022 Conference Paper

Museformer: Transformer with Fine- and Coarse-Grained Attention for Music Generation

  • Botao Yu
  • Peiling Lu
  • Rui Wang
  • Wei Hu
  • Xu Tan
  • Wei Ye
  • Shikun Zhang
  • Tao Qin

Symbolic music generation aims to generate music scores automatically. A recent trend is to use Transformer or its variants in music generation, which is, however, suboptimal, because the full attention cannot efficiently model the typically long music sequences (e. g. , over 10, 000 tokens), and the existing models have shortcomings in generating musical repetition structures. In this paper, we propose Museformer, a Transformer with a novel fine- and coarse-grained attention for music generation. Specifically, with the fine-grained attention, a token of a specific bar directly attends to all the tokens of the bars that are most relevant to music structures (e. g. , the previous 1st, 2nd, 4th and 8th bars, selected via similarity statistics); with the coarse-grained attention, a token only attends to the summarization of the other bars rather than each token of them so as to reduce the computational cost. The advantages are two-fold. First, it can capture both music structure-related correlations via the fine-grained attention, and other contextual information via the coarse-grained attention. Second, it is efficient and can model over 3X longer music sequences compared to its full-attention counterpart. Both objective and subjective experimental results demonstrate its ability to generate long music sequences with high quality and better structures.

AAAI Conference 2022 Conference Paper

Powerful Graph Convolutional Networks with Adaptive Propagation Mechanism for Homophily and Heterophily

  • Tao Wang
  • Di Jin
  • Rui Wang
  • Dongxiao He
  • Yuxiao Huang

Graph Convolutional Networks (GCNs) have been widely applied in various fields due to their significant power on processing graph-structured data. Typical GCN and its variants work under a homophily assumption (i. e. , nodes with same class are prone to connect to each other), while ignoring the heterophily which exists in many real-world networks (i. e. , nodes with different classes tend to form edges). Existing methods deal with heterophily by mainly aggregating higher-order neighborhoods or combing the immediate representations, which leads to noise and irrelevant information in the result. But these methods did not change the propagation mechanism which works under homophily assumption (that is a fundamental part of GCNs). This makes it difficult to distinguish the representation of nodes from different classes. To address this problem, in this paper we design a novel propagation mechanism, which can automatically change the propagation and aggregation process according to homophily or heterophily between node pairs. To adaptively learn the propagation process, we introduce two measurements of homophily degree between node pairs, which is learned based on topological and attribute information, respectively. Then we incorporate the learnable homophily degree into the graph convolution framework, which is trained in an end-to-end schema, enabling it to go beyond the assumption of homophily. More importantly, we theoretically prove that our model can constrain the similarity of representations between nodes according to their homophily degree. Experiments on seven real-world datasets demonstrate that this new approach outperforms the state-of-the-art methods under heterophily or low homophily, and gains competitive performance under homophily.

IJCAI Conference 2022 Conference Paper

RAW-GNN: RAndom Walk Aggregation based Graph Neural Network

  • Di Jin
  • Rui Wang
  • Meng Ge
  • Dongxiao He
  • Xiang Li
  • Wei Lin
  • Weixiong Zhang

Graph-Convolution-based methods have been successfully applied to representation learning on homophily graphs where nodes with the same label or similar attributes tend to connect with one another. Due to the homophily assumption of Graph Convolutional Networks (GCNs) that these methods use, they are not suitable for heterophily graphs where nodes with different labels or dissimilar attributes tend to be adjacent. Several methods have attempted to address this heterophily problem, but they do not change the fundamental aggregation mechanism of GCNs because they rely on summation operators to aggregate information from neighboring nodes, which is implicitly subject to the homophily assumption. Here, we introduce a novel aggregation mechanism and develop a RAndom Walk Aggregation-based Graph Neural Network (called RAW-GNN) method. The proposed approach integrates the random walk strategy with graph neural networks. The new method utilizes breadth-first random walk search to capture homophily information and depth-first search to collect heterophily information. It replaces the conventional neighborhoods with path-based neighborhoods and introduces a new path-based aggregator based on Recurrent Neural Networks. These designs make RAW-GNN suitable for both homophily and heterophily graphs. Extensive experimental results showed that the new method achieved state-of-the-art performance on a variety of homophily and heterophily graphs.

IJCAI Conference 2021 Conference Paper

A Survey on Low-Resource Neural Machine Translation

  • Rui Wang
  • Xu Tan
  • Renqian Luo
  • Tao Qin
  • Tie-Yan Liu

Neural approaches have achieved state-of-the-art accuracy on machine translation but suffer from the high cost of collecting large scale parallel data. Thus, a lot of research has been conducted for neural machine translation (NMT) with very limited parallel data, i. e. , the low-resource setting. In this paper, we provide a survey for low-resource NMT and classify related works into three categories according to the auxiliary data they used: (1) exploiting monolingual data of source and/or target languages, (2) exploiting data from auxiliary languages, and (3) exploiting multi-modal data. We hope that our survey can help researchers to better understand this field and inspire them to design better algorithms, and help industry practitioners to choose appropriate algorithms for their applications.

NeurIPS Conference 2021 Conference Paper

Curriculum Disentangled Recommendation with Noisy Multi-feedback

  • Hong Chen
  • Yudong Chen
  • Xin Wang
  • Ruobing Xie
  • Rui Wang
  • Feng Xia
  • Wenwu Zhu

Learning disentangled representations for user intentions from multi-feedback (i. e. , positive and negative feedback) can enhance the accuracy and explainability of recommendation algorithms. However, learning such disentangled representations from multi-feedback data is challenging because i) multi-feedback is complex: there exist complex relations among different types of feedback (e. g. , click, unclick, and dislike, etc) as well as various user intentions, and ii) multi-feedback is noisy: there exists noisy (useless) information both in features and labels, which may deteriorate the recommendation performance. Existing works on disentangled representation learning only focus on positive feedback, failing to handle the complex relations and noise hidden in multi-feedback data. To solve this problem, in this work we propose a Curriculum Disentangled Recommendation (CDR) model that is capable of efficiently learning disentangled representations from complex and noisy multi-feedback for better recommendation. Concretely, we design a co-filtering dynamic routing mechanism that simultaneously captures the complex relations among different behavioral feedback and user intentions as well as denoise the representations in the feature level. We then present an adjustable self-evaluating curriculum that is able to evaluate sample difficulties for better model training and conduct denoising in the label level via disregarding useless information. Our extensive experiments on several real-world datasets demonstrate that the proposed CDR model can significantly outperform several state-of-the-art methods in terms of recommendation accuracy.

JBHI Journal 2021 Journal Article

Depression Analysis and Recognition Based on Functional Near-Infrared Spectroscopy

  • Rui Wang
  • Yixue Hao
  • Qiao Yu
  • Min Chen
  • Iztok Humar
  • Giancarlo Fortino

Depression is the result of a complex interaction of social, psychological and physiological elements. Research into the brain disorders of patients suffering from depression can help doctors to understand the pathogenesis of depression and facilitate its diagnosis and treatment. Functional near-infrared spectroscopy (fNIRS) is a non-invasive approach to the detection of brain functions and activities. In this paper, a comprehensive fNIRS-based depression-processing architecture, including the layers of source, feature and model, is first established to guide the deep modeling for fNIRS. In view of the complexity of depression, we propose a methodology in the time and frequency domains for feature extraction and deep neural networks for depression recognition combined with current research. It is found that compared to non-depression people, patients with depression have a weaker encephalic area connectivity and lower level of activation in the prefrontal lobe during brain activity. Finally, based on raw data, manual features and channel correlations, the AlexNet model shows the best performance, especially in terms of the correlation features and presents an accuracy rate of 0. 90 and a precision rate of 0. 91, which is higher than ResNet18 and machine-learning algorithms on other data. Therefore, the correlation of brain regions can effectively recognize depression (from cases of non-depression), making it significant for the recognition of brain functions in the clinical diagnosis and treatment of depression.

IROS Conference 2021 Conference Paper

Exploring Imitation Learning for Autonomous Driving with Feedback Synthesizer and Differentiable Rasterization

  • Jinyun Zhou
  • Rui Wang
  • Xu Liu
  • Yifei Jiang
  • Shu Jiang
  • Jiaming Tao
  • Jinghao Miao
  • Shiyu Song

We present a learning-based planner that aims to robustly drive a vehicle by mimicking human drivers’ driving behavior. We leverage a mid-to-mid approach that allows us to manipulate the input to our imitation learning network freely. With that in mind, we propose a novel feedback synthesizer for data augmentation. It allows our agent to gain more driving experience in various previously unseen environments that are likely to encounter, thus improving overall performance. This is in contrast to prior works that rely purely on random synthesizers. Furthermore, rather than completely commit to imitating, we introduce task losses that penalize undesirable behaviors, such as collision, off-road, and so on. Unlike prior works, this is done by introducing a differentiable vehicle rasterizer that directly converts the waypoints output by the network into images. This effectively avoids the usage of heavyweight ConvLSTM networks, therefore, yields a faster model inference time. About the network architecture, we exploit an attention mechanism that allows the network to reason critical objects in the scene and produce better interpretable attention heatmaps. To further enhance the safety and robustness of the network, we add an optional optimization-based post-processing planner improving the driving comfort. We comprehensively validate our method’s effectiveness in different scenarios that are specifically created for evaluating self-driving vehicles. Results demonstrate that our learning-based planner achieves high intelligence and can handle complex situations. Detailed ablation and visualization analysis are included to further demonstrate each of our proposed modules’ effectiveness in our method.

IROS Conference 2021 Conference Paper

Generalization Through Hand-Eye Coordination: An Action Space for Learning Spatially-Invariant Visuomotor Control

  • Chen Wang 0053
  • Rui Wang
  • Ajay Mandlekar
  • Li Fei-Fei 0001
  • Silvio Savarese
  • Danfei Xu

Imitation Learning (IL) is an effective framework to learn visuomotor skills from offline demonstration data. However, IL methods often fail to generalize to new scene configurations not covered by training data. On the other hand, humans can manipulate objects in varying conditions. Key to such capability is hand-eye coordination, a cognitive ability that enables humans to adaptively direct their movements at task-relevant objects and be invariant to the objects’ absolute spatial location. In this work, we present a learnable action space, Hand-eye Action Networks (HAN) that learns coordinated hand-eye movements from human teleoperated demonstrations. Through a set of challenging multi-stage manipulation tasks, we show that a visuomotor policy equipped with HAN is able to inherit the key spatial invariance property of handeye coordination and achieve generalization to new scene configurations. Additional materials available at https://sites.google.com/stanford.edu/han

AAAI Conference 2021 Conference Paper

Hierarchical Reinforcement Learning for Integrated Recommendation

  • Ruobing Xie
  • Shaoliang Zhang
  • Rui Wang
  • Feng Xia
  • Leyu Lin

Integrated recommendation aims to jointly recommend heterogeneous items in the main feed from different sources via multiple channels, which needs to capture user preferences on both item and channel levels. It has been widely used in practical systems by billions of users, while few works concentrate on the integrated recommendation systematically. In this work, we propose a novel Hierarchical reinforcement learning framework for integrated recommendation (HRL-Rec), which divides the integrated recommendation into two tasks to recommend channels and items sequentially. The low-level agent is a channel selector, which generates a personalized channel list. The high-level agent is an item recommender, which recommends specific items from heterogeneous channels under the channel constraints. We design various rewards for both recommendation accuracy and diversity, and propose four losses for fast and stable model convergence. We also conduct an online exploration for sufficient training. In experiments, we conduct extensive offline and online experiments on a billion-level real-world dataset to show the effectiveness of HRL-Rec. HRL-Rec has also been deployed on WeChat Top Stories, affecting millions of users. The source codes are released in https: //github. com/modriczhang/HRL-Rec.

AAMAS Conference 2021 Conference Paper

Intrinsic Motivated Multi-Agent Communication

  • Chuxiong Sun
  • Bo Wu
  • Rui Wang
  • Xiaohui Hu
  • Xiaoya Yang
  • Cong Cong

Efficient communication is a promising way to achieve cooperation among agents in many real-world scenarios. However, aimless and motiveless information sharing may not work or even degrade the cooperative performance. Typically, the multi-agent communication behaviors are motivated by extrinsic rewards from environment. We conclude the mechanism as ’Communicate what rewards you’. In this work, we present a novel communication mechanism called Intrinsic Motivated Multi-Agent Communication (IMMAC). Our key insight can be summarized as ’Communicate what surprises you’. Concretely, we use an observation-dependent intrinsic value to represent the importance of observed information. Then a gating mechanism and an attentional mechanism based on intrinsic values are designed to control communication. By encouraging agent to communicate and focus on the observations with uncertain and important information, our algorithm achieves superior communication efficiency and cooperative performance. We evaluate IMMAC on a variety of challenging tasks, and demonstrate that intrinsic values are sufficient to drive efficient communication behaviors. Moreover, we found that the combination of intrinsic values and extrinsic values can further improve the communication efficiency. Consequently, intrinsic motivation is a promising way to control communication and it is capable of being a good complement to the existing extrinsic motivated communication methods.

JMLR Journal 2021 Journal Article

Representer Theorems in Banach Spaces: Minimum Norm Interpolation, Regularized Learning and Semi-Discrete Inverse Problems

  • Rui Wang
  • Yuesheng Xu

Learning a function from a finite number of sampled data points (measurements) is a fundamental problem in science and engineering. This is often formulated as a minimum norm interpolation (MNI) problem, a regularized learning problem or, in general, a semi-discrete inverse problem (SDIP), in either Hilbert spaces or Banach spaces. The goal of this paper is to systematically study solutions of these problems in Banach spaces. We aim at obtaining explicit representer theorems for their solutions, on which convenient solution methods can then be developed. For the MNI problem, the explicit representer theorems enable us to express the infimum in terms of the norm of the linear combination of the interpolation functionals. For the purpose of developing efficient computational algorithms, we establish the fixed-point equation formulation of solutions of these problems. We reveal that unlike in a Hilbert space, in general, solutions of these problems in a Banach space may not be able to be reduced to truly finite dimensional problems (with certain infinite dimensional components hidden). We demonstrate how this obstacle can be removed, reducing the original problem to a truly finite dimensional one, in the special case when the Banach space is $\ell_1(\mathbb{N})$. [abs] [ pdf ][ bib ] &copy JMLR 2021. ( edit, beta )

ICLR Conference 2021 Conference Paper

Shapley Explanation Networks

  • Rui Wang
  • Xiaoqian Wang 0001
  • David I. Inouye

Shapley values have become one of the most popular feature attribution explanation methods. However, most prior work has focused on post-hoc Shapley explanations, which can be computationally demanding due to its exponential time complexity and preclude model regularization based on Shapley explanations during training. Thus, we propose to incorporate Shapley values themselves as latent representations in deep models thereby making Shapley explanations first-class citizens in the modeling paradigm. This intrinsic explanation approach enables layer-wise explanations, explanation regularization of the model during training, and fast explanation computation at test time. We define the Shapley transform that transforms the input into a Shapley representation given a specific function. We operationalize the Shapley transform as a neural network module and construct both shallow and deep networks, called ShapNets, by composing Shapley modules. We prove that our Shallow ShapNets compute the exact Shapley values and our Deep ShapNets maintain the missingness and accuracy properties of Shapley values. We demonstrate on synthetic and real-world datasets that our ShapNets enable layer-wise Shapley explanations, novel Shapley regularizations during training, and fast computation while maintaining reasonable performance. Code is available at https://github.com/inouye-lab/ShapleyExplanationNetworks.

NeurIPS Conference 2021 Conference Paper

Universal Graph Convolutional Networks

  • Di Jin
  • Zhizhi Yu
  • Cuiying Huo
  • Rui Wang
  • Xiao Wang
  • Dongxiao He
  • Jiawei Han

Graph Convolutional Networks (GCNs), aiming to obtain the representation of a node by aggregating its neighbors, have demonstrated great power in tackling various analytics tasks on graph (network) data. The remarkable performance of GCNs typically relies on the homophily assumption of networks, while such assumption cannot always be satisfied, since the heterophily or randomness are also widespread in real-world. This gives rise to one fundamental question: whether networks with different structural properties should adopt different propagation mechanisms? In this paper, we first conduct an experimental investigation. Surprisingly, we discover that there are actually segmentation rules for the propagation mechanism, i. e. , 1-hop, 2-hop and $k$-nearest neighbor ($k$NN) neighbors are more suitable as neighborhoods of network with complete homophily, complete heterophily and randomness, respectively. However, the real-world networks are complex, and may present diverse structural properties, e. g. , the network dominated by homophily may contain a small amount of randomness. So can we reasonably utilize these segmentation rules to design a universal propagation mechanism independent of the network structural assumption? To tackle this challenge, we develop a new universal GCN framework, namely U-GCN. It first introduces a multi-type convolution to extract information from 1-hop, 2-hop and $k$NN networks simultaneously, and then designs a discriminative aggregation to sufficiently fuse them aiming to given learning objectives. Extensive experiments demonstrate the superiority of U-GCN over state-of-the-arts. The code and data are available at https: //github. com/jindi-tju.

AAAI Conference 2020 Conference Paper

Boundary Enhanced Neural Span Classification for Nested Named Entity Recognition

  • Chuanqi Tan
  • Wei Qiu
  • Mosha Chen
  • Rui Wang
  • Fei Huang

Named entity recognition (NER) is a well-studied task in natural language processing. However, the widely-used sequence labeling framework is usually difficult to detect entities with nested structures. The span-based method that can easily detect nested entities in different subsequences is naturally suitable for the nested NER problem. However, previous span-based methods have two main issues. First, classifying all subsequences is computationally expensive and very inefficient at inference. Second, the span-based methods mainly focus on learning span representations but lack of explicit boundary supervision. To tackle the above two issues, we propose a boundary enhanced neural span classification model. In addition to classifying the span, we propose incorporating an additional boundary detection task to predict those words that are boundaries of entities. The two tasks are jointly trained under a multitask learning framework, which enhances the span representation with additional boundary supervision. In addition, the boundary detection model has the ability to generate high-quality candidate spans, which greatly reduces the time complexity during inference. Experiments show that our approach outperforms all existing methods and achieves 85. 3, 83. 9, and 78. 3 scores in terms of F1 on the ACE2004, ACE2005, and GENIA datasets, respectively.

IJCAI Conference 2020 Conference Paper

Deep Feedback Network for Recommendation

  • Ruobing Xie
  • Cheng Ling
  • Yalong Wang
  • Rui Wang
  • Feng Xia
  • Leyu Lin

Both explicit and implicit feedbacks can reflect user opinions on items, which are essential for learning user preferences in recommendation. However, most current recommendation algorithms merely focus on implicit positive feedbacks (e. g. , click), ignoring other informative user behaviors. In this paper, we aim to jointly consider explicit/implicit and positive/negative feedbacks to learn user unbiased preferences for recommendation. Specifically, we propose a novel Deep feedback network (DFN) modeling click, unclick and dislike behaviors. DFN has an internal feedback interaction component that captures fine-grained interactions between individual behaviors, and an external feedback interaction component that uses precise but relatively rare feedbacks (click/dislike) to extract useful information from rich but noisy feedbacks (unclick). In experiments, we conduct both offline and online evaluations on a real-world recommendation system WeChat Top Stories used by millions of users. The significant improvements verify the effectiveness and robustness of DFN. The source code is in https: //github. com/qqxiaochongqq/DFN.

EAAI Journal 2020 Journal Article

Ensemble framework by using nature inspired algorithms for the early-stage forest fire rescue — A case study of dynamic optimization problems

  • HongGuang Zhang
  • ZiHan Liang
  • HuaJian Liu
  • Rui Wang
  • YuanAn Liu

In this paper, we propose rescue ensemble to simulate the dynamic rescue process between forest fire spread and forest fire rescue, while simultaneously formulating this process as a dynamic optimization problem. However, there is still little research about simulating this kind of the dynamic rescue process, even when many new unmanned monitoring systems and large-scale firefighting aircraft emerge in the forest-fire-rescue field. Our rescue ensemble that consists of rescue simulator and rescue algorithm is characterized by supporting the offline simulation of the dynamic rescue process between forest fire spread (like offensive forces) and forest fire rescue (like defensive forces). Based on modifying the cellular automaton model of forest fire spread, rescue simulator is able to simulate forest fire spread and aircraft firefighting, simultaneously. Besides, the main goal of rescue algorithm is to realize the aircraft task allocation. Firefighting particle swarm optimization is proposed by us as our rescue algorithm, which is characterized by considering fire edge suppression, the burning-cell continuity, and wind direction. We construct our test problems based on real forest maps and aircraft firefighting capability. Comparing with four compared rescue algorithms, we test the different capabilities of firefighting particle swarm optimization, such as searching dynamic optimal solution, shortening the rescue time, controlling the spread speed of fire edge, and minimizing the burned cost. Experimental results demonstrate that the framework of rescue ensemble is feasible. Meanwhile, the results of firefighting particle swarm optimization are satisfactory in most cases.

AAAI Conference 2020 Conference Paper

Explicit Sentence Compression for Neural Machine Translation

  • Zuchao Li
  • Rui Wang
  • Kehai Chen
  • Masao Utiyama
  • Eiichiro Sumita
  • Zhuosheng Zhang
  • Hai Zhao

State-of-the-art Transformer-based neural machine translation (NMT) systems still follow a standard encoder-decoder framework, in which source sentence representation can be well done by an encoder with self-attention mechanism. Though Transformer-based encoder may effectively capture general information in its resulting source sentence representation, the backbone information, which stands for the gist of a sentence, is not specifically focused on. In this paper, we propose an explicit sentence compression method to enhance the source sentence representation for NMT. In practice, an explicit sentence compression goal used to learn the backbone information in a sentence. We propose three ways, including backbone source-side fusion, targetside fusion, and both-side fusion, to integrate the compressed sentence into NMT. Our empirical tests on the WMT Englishto-French and English-to-German translation tasks show that the proposed sentence compression method significantly improves the translation performances over strong baselines.

NeurIPS Conference 2020 Conference Paper

Semi-Supervised Neural Architecture Search

  • Renqian Luo
  • Xu Tan
  • Rui Wang
  • Tao Qin
  • Enhong Chen
  • Tie-Yan Liu

Neural architecture search (NAS) relies on a good controller to generate better architectures or predict the accuracy of given architectures. However, training the controller requires both abundant and high-quality pairs of architectures and their accuracy, while it is costly to evaluate an architecture and obtain its accuracy. In this paper, we propose SemiNAS, a semi-supervised NAS approach that leverages numerous unlabeled architectures (without evaluation and thus nearly no cost). Specifically, SemiNAS 1) trains an initial accuracy predictor with a small set of architecture-accuracy data pairs; 2) uses the trained accuracy predictor to predict the accuracy of large amount of architectures (without evaluation); and 3) adds the generated data pairs to the original data to further improve the predictor. The trained accuracy predictor can be applied to various NAS algorithms by predicting the accuracy of candidate architectures for them. SemiNAS has two advantages: 1) It reduces the computational cost under the same accuracy guarantee. On NASBench-101 benchmark dataset, it achieves comparable accuracy with gradient-based method while using only 1/7 architecture-accuracy pairs. 2) It achieves higher accuracy under the same computational cost. It achieves 94. 02% test accuracy on NASBench-101, outperforming all the baselines when using the same number of architectures. On ImageNet, it achieves 23. 5% top-1 error rate (under 600M FLOPS constraint) using 4 GPU-days for search. We further apply it to LJSpeech text to speech task and it achieves 97% intelligibility rate in the low-resource setting and 15% test error rate in the robustness setting, with 9%, 7% improvements over the baseline respectively.

AAAI Conference 2020 Conference Paper

SG-Net: Syntax-Guided Machine Reading Comprehension

  • Zhuosheng Zhang
  • Yuwei Wu
  • Junru Zhou
  • Sufeng Duan
  • Hai Zhao
  • Rui Wang

For machine reading comprehension, the capacity of effectively modeling the linguistic knowledge from the detailriddled and lengthy passages and getting ride of the noises is essential to improve its performance. Traditional attentive models attend to all words without explicit constraint, which results in inaccurate concentration on some dispensable words. In this work, we propose using syntax to guide the text modeling by incorporating explicit syntactic constraints into attention mechanism for better linguistically motivated word representations. In detail, for self-attention network (SAN) sponsored Transformer-based encoder, we introduce syntactic dependency of interest (SDOI) design into the SAN to form an SDOI-SAN with syntax-guided selfattention. Syntax-guided network (SG-Net) is then composed of this extra SDOI-SAN and the SAN from the original Transformer encoder through a dual contextual architecture for better linguistics inspired representation. To verify its effectiveness, the proposed SG-Net is applied to typical pre-trained language model BERT which is right based on a Transformer encoder. Extensive experiments on popular benchmarks including SQuAD 2. 0 and RACE show that the proposed SG-Net design helps achieve substantial performance improvement over strong baselines.

AAAI Conference 2019 Conference Paper

A Deep Cascade Model for Multi-Document Reading Comprehension

  • Ming Yan
  • Jiangnan Xia
  • Chen Wu
  • Bin Bi
  • Zhongzhou Zhao
  • Ji Zhang
  • Luo Si
  • Rui Wang

A fundamental trade-off between effectiveness and efficiency needs to be balanced when designing an online question answering system. Effectiveness comes from sophisticated functions such as extractive machine reading comprehension (MRC), while efficiency is obtained from improvements in preliminary retrieval components such as candidate document selection and paragraph ranking. Given the complexity of the real-world multi-document MRC scenario, it is difficult to jointly optimize both in an end-to-end system. To address this problem, we develop a novel deep cascade learning model, which progressively evolves from the documentlevel and paragraph-level ranking of candidate texts to more precise answer extraction with machine reading comprehension. Specifically, irrelevant documents and paragraphs are first filtered out with simple functions for efficiency consideration. Then we jointly train three modules on the remaining texts for better tracking the answer: the document extraction, the paragraph extraction and the answer extraction. Experiment results show that the proposed method outperforms the previous state-of-the-art methods on two large-scale multidocument benchmark datasets, i. e. , TriviaQA and DuReader. In addition, our online system can stably serve typical scenarios with millions of daily requests in less than 50ms.

YNICL Journal 2019 Journal Article

A quantitative SVM approach potentially improves the accuracy of magnetic resonance spectroscopy in the preoperative evaluation of the grades of diffuse gliomas

  • Chong Qi
  • Yiming Li
  • Xing Fan
  • Yin Jiang
  • Rui Wang
  • Song Yang
  • Lanxi Meng
  • Tao Jiang

OBJECTIVES: H-MRS) metabolic features and the grade of gliomas, and to establish a machine-learning model to predict the glioma grade. METHODS: H-MRS image. The Student's t-test was conducted to screen for differentially expressed features between low- and high-grade gliomas (WHO grades II and III/IV, respectively). Next, the minimum Redundancy Maximum Relevance (mRMR) algorithm was performed to further select features for a support vector machine (SVM) classifier building. Performance of the predictive model was evaluated both in the training and validation sets using ROC curve analysis. RESULTS: H-MRS metabolic features, thirteen features were differentially expressed. Four features were further selected as grade-predictive imaging signatures using the mRMR algorithm. The predictive performance of the machine-learning model measured by the AUC was 0.825 and 0.820 in the training and validation sets, respectively. This was better than the predictive performances of individual metabolic features, the best of which was 0.812. CONCLUSIONS: H-MRS metabolic features could help in predicting the grade of gliomas. The machine-learning model achieved a better prediction performance in grading gliomas than individual features, indicating that it could complement the traditionally used metabolic features.

IJCAI Conference 2019 Conference Paper

An Atari Model Zoo for Analyzing, Visualizing, and Comparing Deep Reinforcement Learning Agents

  • Felipe Petroski Such
  • Vashisht Madhavan
  • Rosanne Liu
  • Rui Wang
  • Pablo Samuel Castro
  • Yulun Li
  • Jiale Zhi
  • Ludwig Schubert

Much human and computational effort has aimed to improve how deep reinforcement learning (DRL) algorithms perform on benchmarks such as the Atari Learning Environment. Comparatively less effort has focused on understanding what has been learned by such methods, and investigating and comparing the representations learned by different families of DRL algorithms. Sources of friction include the onerous computational requirements, and general logistical and architectural complications for running DRL algorithms at scale. We lessen this friction, by (1) training several algorithms at scale and releasing trained models, (2) integrating with a previous DRL model release, and (3) releasing code that makes it easy for anyone to load, visualize, and analyze such models. This paper introduces the Atari Zoo framework, which contains models trained across benchmark Atari games, in an easy-to-use format, as well as code that implements common modes of analysis and connects such models to a popular neural network visualization library. Further, to demonstrate the potential of this dataset and software package, we show initial quantitative and qualitative comparisons between the performance and representations of several DRL algorithms, highlighting interesting and previously unknown distinctions between them.

AAAI Conference 2019 Conference Paper

Chinese NER with Height-Limited Constituent Parsing

  • Rui Wang
  • Xin Xin
  • Wei Chang
  • Kun Ming
  • Biao Li
  • Xin Fan

In this paper, we investigate how to improve Chinese named entity recognition (NER) by jointly modeling NER and constituent parsing, in the framework of neural conditional random fields (CRF). We reformulate the parsing task to heightlimited constituent parsing, by which the computational complexity can be significantly reduced, and the majority of phrase-level grammars are retained. Specifically, an unified model of neural semi-CRF and neural tree-CRF is proposed, which simultaneously conducts word segmentation, part-ofspeech (POS) tagging, NER, and parsing. The challenge comes from how to train and infer the joint model, which has not been solved previously. We design a dynamic programming algorithm for both training and inference, whose complexity is O(n·4h ), where n is the sentence length and h is the height limit. In addition, we derive a pruning algorithm for the joint model, which further prunes 99. 9% of the search space with 2% loss of the ground truth data. Experimental results on the OntoNotes 4. 0 dataset have demonstrated that the proposed model outperforms the state-of-the-art method by 2. 79 points in the F1-measure.

NeurIPS Conference 2019 Conference Paper

Hierarchical Reinforcement Learning with Advantage-Based Auxiliary Rewards

  • Siyuan Li
  • Rui Wang
  • Minxue Tang
  • Chongjie Zhang

Hierarchical Reinforcement Learning (HRL) is a promising approach to solving long-horizon problems with sparse and delayed rewards. Many existing HRL algorithms either use pre-trained low-level skills that are unadaptable, or require domain-specific information to define low-level rewards. In this paper, we aim to adapt low-level skills to downstream tasks while maintaining the generality of reward design. We propose an HRL framework which sets auxiliary rewards for low-level skill training based on the advantage function of the high-level policy. This auxiliary reward enables efficient, simultaneous learning of the high-level policy and low-level skills without using task-specific knowledge. In addition, we also theoretically prove that optimizing low-level skills with this auxiliary reward will increase the task return for the joint policy. Experimental results show that our algorithm dramatically outperforms other state-of-the-art HRL methods in Mujoco domains. We also find both low-level and high-level policies trained by our algorithm transferable.

AAAI Conference 2019 Conference Paper

PCGAN: Partition-Controlled Human Image Generation

  • Dong Liang
  • Rui Wang
  • Xiaowei Tian
  • Cong Zou

Human image generation is a very challenging task since it is affected by many factors. Many human image generation methods focus on generating human images conditioned on a given pose, while the generated backgrounds are often blurred. In this paper, we propose a novel Partition-Controlled GAN to generate human images according to target pose and background. Firstly, human poses in the given images are extracted, and foreground/background are partitioned for further use. Secondly, we extract and fuse appearance features, pose features and background features to generate the desired images. Experiments on Market-1501 and DeepFashion datasets show that our model not only generates realistic human images but also produce the human pose and background as we want. Extensive experiments on COCO and LIP datasets indicate the potential of our method.

IJCAI Conference 2019 Conference Paper

Self-attentive Biaffine Dependency Parsing

  • Ying Li
  • Zhenghua Li
  • Min Zhang
  • Rui Wang
  • Sheng Li
  • Luo Si

The current state-of-the-art dependency parsing approaches employ BiLSTMs to encode input sentences. Motivated by the success of the transformer-based machine translation, this work for the first time applies the self-attention mechanism to dependency parsing as the replacement of the BiLSTM-based encoders, leading to competitive performance on both English and Chinese benchmark data. Based on the detailed error analysis, we then combine the power of both BiLSTM and self-attention via model ensembles, demonstrating their complementary capability of capturing contextual information. Finally, we explore the recently proposed contextualized word representations as extra input features, and further improve the parsing performance.

AAAI Conference 2019 Conference Paper

Syntax-Aware Neural Semantic Role Labeling

  • Qingrong Xia
  • Zhenghua Li
  • Min Zhang
  • Meishan Zhang
  • Guohong Fu
  • Rui Wang
  • Luo Si

Semantic role labeling (SRL), also known as shallow semantic parsing, is an important yet challenging task in NLP. Motivated by the close correlation between syntactic and semantic structures, traditional discrete-feature-based SRL approaches make heavy use of syntactic features. In contrast, deep-neural-network-based approaches usually encode the input sentence as a word sequence without considering the syntactic structures. In this work, we investigate several previous approaches for encoding syntactic trees, and make a thorough study on whether extra syntax-aware representations are beneficial for neural SRL models. Experiments on the benchmark CoNLL-2005 dataset show that syntax-aware SRL approaches can effectively improve performance over a strong baseline with external word representations from ELMo. With the extra syntax-aware representations, our approaches achieve new state-of-the-art 85. 6 F1 (single model) and 86. 6 F1 (ensemble) on the test data, outperforming the corresponding strong baselines with ELMo by 0. 8 and 1. 0, respectively. Detailed error analysis are conducted to gain more insights on the investigated approaches.

AAAI Conference 2018 Conference Paper

Audio Visual Attribute Discovery for Fine-Grained Object Recognition

  • Hua Zhang
  • Xiaochun Cao
  • Rui Wang

Current progresses on fine-grained recognition are mainly focus on learning the discriminative feature representation via introducing the visual supervisions e. g. part labels. However, it is time-consuming and needs the professional knowledge to obtain the accuracy annotations. Different from these existing methods based on the visual supervisions, in this paper, we introduce a novel feature named audio visual attributes via discovering the correlations between the visual and audio representations. Specifically, our unified framework is training with video-level category label, which consists of two important modules, the encoder module and the attribute discovery module, to encode the image and audio into vectors and learn the correlations between audio and images, respectively. On the encoder module, we present two types of feed forward convolutional neural network for the image and audio modalities. While an attention driven framework based on recurrent neural network is developed to generate the audio visual attribute representation. Thus, our proposed architecture can be implemented end-to-end in the step of inference. We exploit our models for the problem of fine-grained bird recognition on the CUB200-211 benchmark. The experimental results demonstrate that with the help of audio visual attribute, we achieve the superior or comparable performance to that of strongly supervised approaches on the bird recognition.

EAAI Journal 2018 Journal Article

Modeling and optimization of a road–rail intermodal transport system under uncertain information

  • Rui Wang
  • Kai Yang
  • Lixing Yang
  • Ziyou Gao

A realistic road–rail intermodal transport system can be suitably modeled as a hub-and-spoke (H&S) network for which the parameters are subject to fuzzy uncertainty: demand, cost and time. For modeling uncertainty, we present a bi-objective optimization formulation for the hub-and-spoke based road–rail intermodal transportation (HS-RRIT) network design problem by taking into account the expected value criterion and the critical value criterion. Using the weighted sum method, we reformulate a single-objective mixed-integer linear programming (MILP) model to solve the equivalent HS-RRIT network design problem. Given the inherent complexity for solving this problem, we develop a memetic algorithm (MA) to obtain high quality solutions. This algorithm utilizes a genetic search method to explore the search space and two different local search strategies called shift and exchange to exploit information in the search region. Finally, we conduct computational analysis over the Turkish network data set to demonstrate the applicability of proposed model and the effectiveness of solution method.

AAAI Conference 2018 Conference Paper

Syntax-Directed Attention for Neural Machine Translation

  • Kehai Chen
  • Rui Wang
  • Masao Utiyama
  • Eiichiro Sumita
  • Tiejun Zhao

Attention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window source words. However, alignment weights for the current target word often decrease to the left and right by linear distance centering on the aligned source position and neglect syntax distance constraints. In this paper, we extend the local attention with syntax-distance constraint, which focuses on syntactically related source words with the predicted target word to learning a more effective context vector for predicting translation. Moreover, we further propose a double context NMT architecture, which consists of a global context vector and a syntax-directed context vector from the global attention, to provide more translation performance for NMT from source representation. The experiments on the largescale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves a substantial and significant improvement over the baseline system.

IJCAI Conference 2016 Conference Paper

A Bilingual Graph-Based Semantic Model for Statistical Machine Translation

  • Rui Wang
  • Hai Zhao
  • Sabine Ploux
  • Bao-Liang Lu
  • Masao Utiyama

Most existing bilingual embedding methods for Statistical Machine Translation (SMT) suffer from two obvious drawbacks. First, they only focus on simple context such as word count and co-occurrence in document or sliding window to build word embedding, ignoring latent useful information from selected context. Second, word sense but not word form is supposed to be the minimal semantic unit while most existing works are still for word representation. This paper presents Bilingual Graph-based Semantic Model (BGSM) to alleviate such shortcomings. By means of maximum complete sub-graph (clique) for context selection, BGSM is capable of effectively modeling word sense representation instead of the word form itself. The proposed model is applied to phrase pair translation probability estimation and generation for SMT. The empirical results show that BGSM can enhance SMT both in performance (up to +1. 3 BLEU) and efficiency in comparison against existing methods.

TAAS Journal 2016 Journal Article

Managing Server Clusters on Renewable Energy Mix

  • Chao Li
  • Rui Wang
  • Depei Qian
  • Tao Li

As climate change has become a global concern and server energy demand continues to soar, many IT companies have started to explore server clusters running on various renewable energy sources. Existing green data center designs often yield suboptimal performance as they only look at a certain specific type of energy source. This article explores data centers powered by hybrid renewable energy systems. We propose GreenWorks, a framework for HPC data centers running on a renewable energy mix. Specifically, GreenWorks features a cross-layer power management scheme tailored to the timing behaviors and capacity constraints of different energy sources. Using realistic workload traces and renewable energy data, we show that GreenWorks could provide a near-optimal workload performance (within 3% difference) on average. It can also reduce the worst-case performance degradation by 43% compared to the state-of-the-art design. Moreover, the performance improvements are based on carbon-neutral operations and are not at the cost of significant efficiency degradation and reduced battery lifecycle. Our technique becomes more efficient when servers become more energy proportional and can effectively handle the ever-increasing depth of renewable power penetration in green data centers.

AAAI Conference 2011 Conference Paper

Towards Maximizing the Area Under the ROC Curve for Multi-Class Classification Problems

  • Ke Tang
  • Rui Wang
  • Tianshi Chen

The Area Under the ROC Curve (AUC) metric has achieved a big success in binary classification problems since they measure the performance of classifiers without making any specific assumptions about the class distribution and misclassification costs. This is desirable because the class distribution and misclassification costs may be unknown during training process or even change in environment. MAUC, the extension of AUC to multi-class problems, has also attracted a lot of attention. However, despite the emergence of approaches for training classifiers with large AUC, little has been done for MAUC. This paper analyzes MAUC in-depth, and reveals that the maximization of MAUC can be achieved by decomposing the multi-class problem into a number of independent sub-problems. These sub-problems are formulated in the form of a “learning to rank” problem, for which well-established methods already exist. Based on the analysis, a method that employs RankBoost algorithm as the sub-problem solver is proposed to achieve classification systems with maximum MAUC. Empirical studies have shown the advantages of the proposed method over other eight relevant methods. Due to the importance of MAUC to multi-class cost-sensitive learning and class imbalanced learning problems, the proposed method is a general technique for both problems. It can also be generalized to accommodate other learning algorithms as the subproblem solvers.

TCS Journal 2007 Journal Article

On the hardness of minimizing space for all-shortest-path interval routing schemes

  • Rui Wang
  • Francis C.M. Lau
  • Yan Yan Liu

k -Interval Routing Scheme ( k -IRS) is a compact routing method that allows up to k interval labels to be assigned to an arc; and global k -IRS allows not more than a total of k interval labels in the whole network. A fundamental problem is to characterize the networks that admit k -IRS (or global k -IRS). Many of the problems related to single-shortest-path k -IRS have already been shown to be NP-complete. For all-shortest-path k -IRS, the characterization problem remains open for k ⩾ 1. In this paper, we study the time complexity of devising minimal-space all-shortest-path k -IRSs and show that it is NP-complete to decide whether a graph admits an all-shortest-path k -IRS, for every integer k ⩾ 3, and so is that of deciding whether a graph admits an all-shortest-path k -strict IRS, for every integer k ⩾ 4. These are the first NP-completeness results for all-shortest-path k -IRS where k is a constant and the graph is unweighted. The NP-completeness holds also for the linear case. We also prove that it is NP-complete to decide whether an unweighted graph admits an all-shortest-path IRS with global compactness of at most k, which also holds for the linear and strict cases.

TCS Journal 2007 Journal Article

Optimal gossiping in square 2D meshes

  • Rui Wang
  • Francis C.M. Lau

Gossiping is the communication problem in which each node has a unique message to be transmitted to every other node. The nodes exchange their message by packets. A solution to the problem is judged by how many rounds of packet sending it requires. In this paper, we consider the version of the problem in which small-size packets each carrying exactly one message are used. The nodes of the target meshes are assumed to be all-port (a node’s incident edges can all be active at the same time); and their edges are either half-duplex or full-duplex, which are also known as the H* model and the F* model, respectively. We study the class of 2D square meshes. Soch and Tvrdik (SIROCCO’97, pp. 253–265; Tech. rep. DC-97-04, Dept. of CS&E, Czech Technical University) have obtained optimal algorithms for the F* model (for square or nonsquare meshes). Lau and Zhang (IEEE Trans. Parallel Distribut. Syst. 13 (4) (2002) 349–358) have obtained fast algorithms for the H* model. We present optimal algorithms for both models, with the interesting property that they route their messages along the same paths and in the same order, i. e. for any edge { u, v }, the i -th message from u to v under either model is the same message.

AAAI Conference 2007 Conference Paper

Recognizing Textual Entailment Using a Subsequence Kernel Method

  • Rui Wang

We present a novel approach to recognizing Textual Entailment. Structural features are constructed from abstract tree descriptions, which are automatically extracted from syntactic dependency trees. These features are then applied in a subsequence-kernel-based classifier to learn whether an entailment relation holds between two texts. Our method makes use of machine learning techniques using a limited data set, no external knowledge bases (e. g. WordNet), and no handcrafted inference rules. We achieve an accuracy of 74. 5% for text pairs in the Information Extraction and Question Answering task, 63. 6% for the RTE-2 test data, and 66. 9% for the RET-3 test data.

v2026.09.13