Arrow Research search

Author name cluster

Yuan Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

37 papers
2 author rows

Possible papers

37

AAAI Conference 2026 Conference Paper

Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation

  • Yuan Wang
  • Shujian Gao
  • Jiaxiang Liu
  • Songtao Jiang
  • Xia Haoxiang
  • Xiaotian Zhang
  • Zhaolu Kang
  • Yemin Wang

Automatic medical report generation can greatly reduce the workload of doctors, but it is often unreliable for real-world deployment. Current methods can write formally fluent sentences but may be factually flawed, introducing serious medical errors known as clinical hallucinations, which make them untrustworthy for diagnosis. To bridge this gap, we introduce HiMed-RL, a Hierarchical Medical Reward Learning Framework designed to explicitly prioritize clinical quality. HiMed-RL moves beyond simple text matching by deconstructing reward learning into three synergistic levels: it first ensures linguistic fluency at the token-level, then enforces factual grounding at the concept-level by aligning key medical terms with expert knowledge, and finally assesses high-level diagnostic consistency at the semantic-level using a specialized LLM verifier. This hierarchical reward is implemented via a Human-inspired Dynamic Reward Adjustment, a strategy which first teaches the model to learn basic facts before progressing to more complex diagnostic reasoning. Experimentally, HiMed-3B achieves state-of-the-art performance on both in-domain and out-of-domain benchmarks, particularly on the latter, with an improvement of 10.8% over the second-best baseline. Our work provides a robust paradigm for generating reports that not only improve fluency but clinical fine-grained quality.

AAMAS Conference 2026 Conference Paper

Heterogeneity in Multi-Agent Reinforcement Learning

  • Tianyi Hu
  • Zhiqiang Pu
  • Yuan Wang
  • Tenghai Qiu
  • Min Chen
  • Xin Yu

Heterogeneity is a fundamental property in multi-agent reinforcement learning (MARL), which is closely related not only to the functional differences of agents, but also to policy diversity and environmentalinteractions. However, theMARLfieldcurrentlylacksa rigorousdefinitionanddeeperunderstandingofheterogeneity. This paper systematically discusses heterogeneity in MARL from the perspectives of definition, quantification, and utilization. First, based on an agent-level modeling of MARL, we categorize heterogeneity into five types and provide mathematical definitions. Second, we define the concept of heterogeneity distance and propose a practical quantification method. Third, we design a heterogeneity-based multi-agent dynamic parameter sharing algorithm as an example of the application of our methodology. Case studies demonstrate that our method can effectively identify and quantify various types of agent heterogeneity. Experimental results show that the proposed algorithm, compared to other parameter sharing baselines, has better interpretability and stronger adaptability. The proposed methodology will help the MARL community gain a more comprehensive and profound understanding of heterogeneity, and further promote the development of practical algorithms. 1

EAAI Journal 2026 Journal Article

Integrated scheduling of pallets and vehicles for automated warehouses with multi-tier racks

  • Hanzhao Wu
  • Zhenyong Wu
  • Yuan Wang
  • Mark Goh

This paper studies the problem of integrated scheduling of pallets and transport vehicles in an automated warehouse with multi-layer racks and two storage locations. The considered automated transport vehicles include automated guided vehicles (AGVs), unit loaders, pallet lifts, and shuttles. A mixed integer programming (MIP) model is used to handle the inbound operations of pallet allocation, transportation and storage, and vehicle scheduling. A three-step heuristic algorithm is proposed to verify and supplement the proposed model. Computational experiments on large instances with more than three million constraints and variables show that the heuristic algorithm can obtain solutions in a reasonable time, significantly outperforming the exact MIP solver. The comparison results indicate that the proposed algorithm achieves excellent solution quality and convergence performance while significantly reducing the risk of falling into local optimum. In addition, sensitivity analysis is conducted to provide managerial insights to improve the overall operational efficiency of automated warehouses.

AAAI Conference 2026 Conference Paper

IPFormer: Instance Prompt-guided Transformer for Multi-modal Multi-shot Video Understanding

  • Yujia Liang
  • Jile Jiao
  • Xuetao Feng
  • Xinchen Liu
  • Kun Liu
  • Yuan Wang
  • Zixuan Ye
  • Hao Lu

Video Large Language Models (VideoLLMs), which adopt large language models for video understanding, have been demonstrated for single-shot videos. However, they usually struggle in multi-shot videos with frequent shot changes, varying camera angles, etc., which makes VideoLLMs hardly answer questions about multiple instances or shots over the whole video. We attribute this challenge to two issues: 1) the lack of multi-shot multi-instance annotations of existing datasets, and 2) the negligence of instance-aware modeling of current VideoLLMs. Therefore, we first introduce a new dataset termed MultiClip-Bench, featuring dense descriptions and question-answering pairs tailored for multi-shot and multi-instance scenarios. Moreover, since the existing VideoLLMs neglect the explicit modeling of instance-related features, we propose a novel Instance Prompt-guided Transformer, named IPFormer, to achieve instance-aware videounderstanding. In the IPFormer, we design a simple but effective instance-aware feature injection module, which encodes instance features as instance prompts via an attention-based connector. By this means, IPFormer can aggregate instance-specific information across multiple shots. Extensive experiments not only show that our dataset and model significantly improve multi-shot video understanding. but also show that our MultiClip-Bench can provide valuable training data and benchmarks for various video understanding tasks.

EAAI Journal 2026 Journal Article

Multi-objective optimization of rail welded joint grinding in railroad tracks via reinforcement learning

  • Tianci Gao
  • Yuan Wang
  • Xin Wang

Rail welded joints are prone to geometric defects that compromise operational safety and ride quality. Traditional grinding strategies often rely on fixed rules, which fail to adapt to the diverse and irregular profiles of weld defects, leading to either excessive material removal or inefficient maintenance. To address this issue, this paper develops a reinforcement learning-based multi-objective optimization framework using a dual-critic Deep Deterministic Policy Gradient (DDPG) to generate efficient grinding strategies. The proposed model simultaneously minimizes grinding volume and the number of grinding passes, addressing both material conservation and operational efficiency. A continuous-action actor network is used to predict the optimal grinding depths and locations, while two separate critic networks evaluate the trade-offs between the competing objectives. The model is trained on high-resolution field data collected from 55 real-world rail welded joints across high-speed, conventional, and subway lines. After 6000 training steps, both critic networks converged with stable critic losses and smooth policy gradients, ensuring accurate value estimations under different grinding conditions. Across 300 sampled trade-off weights, the model generates over 220 feasible grinding strategies per case, with defined Pareto-optimal fronts. Case studies demonstrate that the optimized strategies reduce grinding volume and passes, improving profile trough restoration from −0. 41 to −0. 18 mm. Therefore, the proposed method can serve as an interpretable and adaptive tool to support rail maintenance decision-making by offering context-specific grinding strategies—from aggressive single-pass interventions to multi-pass, low-depth approaches that prioritize rail longevity.

YNIMG Journal 2026 Journal Article

Standardized quantification of [18F]Florbetazine amyloid PET with the Centiloid scale

  • Meiqi Wu
  • Menglin Liang
  • Chenhui Mao
  • Liling Dong
  • Qi Ge
  • Yuying Li
  • Jingnan Wang
  • Chao Ren

C]PiB across different image-processing pipelines and effective image resolutions (EIRs). METHODS: C]PiB SUVR were evaluated under different EIRs. RESULTS: F]FBZ SUVR were observed across EIRs with the SPM pipeline, whereas regression parameters varied across EIRs with the FreeSurfer pipeline. CONCLUSION: F]FBZ demonstrated equal or improved quantification precision, supporting its broader use in clinical and research Aβ imaging.

AAAI Conference 2026 Conference Paper

TCoT: Trajectory Chain-of-Thoughts for Robotic Manipulation with Failure Recovery in Vision-Language-Action Model

  • Xiang Li
  • Ya-Li Li
  • Yuan Wang
  • Huaqiang Wang
  • Shengjin Wang

Recent advances in vision-language-action (VLA) models have demonstrated impressive generalization for robotic manipulation. However, these models often operate by directly mapping visual and linguistic inputs to subsequent actions, lacking intermediate task planning, along with failure detection and recovery ability. These limitations prevent them from effectively decomposing complex tasks, recognizing problems, and correcting erroneous actions, ultimately resulting in complete task failure. This significantly hinders their ability to perform long-horizon tasks and generalization ability. To this end, we introduce TCoT: Trajectory Chain-of-Thought, a unified VLA framework that enhances this direct mapping with trajectory planning as well as failure detection and recovery. TCoT leverages hierarchy trajectories as a precise and compact representation of CoT reasoning for manipulation: global planning provides a high-level, goal-oriented trajectory to guide the robot toward its task objective, while local planning focuses on real-time adjustments to address dynamic changes. Moreover, we designed the Global-Local Switching Recovery algorithm that detects and effectively recovers from failures. Experimental results reveal that TCoT surpasses the state-of-the-art methods across both real and simulated scenarios and exhibits superior generalization capabilities.

JBHI Journal 2026 Journal Article

TriCSART: Semi-Supervised Medical Image Segmentation with Triple-Level Contrastive Learning and Selective Active Re-Training

  • Shujian Gao
  • Yuan Wang
  • Jiacheng Yang
  • Mengwen Ye
  • Weifan Liu
  • Zekuan Yu

Semi-supervised medical image segmentation has garnered significant attention due to challenges of limited medical data accessibility and expensive annotation costs. However, existing studies face two critical challenges: 1) while contrastive learning has demonstrated potential in semi-supervised frameworks, prior implementations lack hierarchical modeling, failing to comprehensively integrate contrastive mechanisms across intra-, inter-, and memorybankdimensions; 2)conventional pseudo-labeling strategies inadequately address quality assessment, potentially propagating annotation biases through continual error accumulation. To address these issues, this paper introduces a Triple-Level Contrastive (TLC) Learning and Selective Active Re-Training (SART) strategy for medical image analysis. The proposed method adopts a teacher–student architecture with two main components: the TLC module and the SART module. The TLC module establishes multilevel semantic consistency across different views through three distinct, complementary loss functions, simultaneously enhancing inter-class discriminability and intra-class compactness. To further mitigate sample quality imbalance, the SART module introduces a metric-driven evaluation mechanism to automatically identify salient samples. Finally, these selected unlabeled samples are integrated with labeled data for re-training guided by calculated curriculum scores. Extensive experiments are conducted on five diverse bench marks, including four public datasets and one private CBCT dataset. The results demonstrate that our approach achieves state-of-the-art performance, consistently outperforming other semi-supervised segmentation strategies. Ablation studies further confirm the efficacy of each proposed component.

AAAI Conference 2026 Conference Paper

Unreal-MAP: Unreal-Engine-Based General Platform for Multi-agent Reinforcement Learning

  • Tianyi Hu
  • Qingxu Fu
  • Zhiqiang Pu
  • Yuan Wang
  • Tenghai Qiu

In this paper, we propose Unreal Multi-Agent Playground (Unreal-MAP), an MARL general platform based on the Unreal-Engine (UE). Unreal-MAP allows users to freely create multi-agent tasks using the vast visual and physical resources available in the UE community, and deploy state-of-the-art (SOTA) MARL algorithms within them. Unreal-MAP is user-friendly in terms of deployment, modification, and visualization, and all its components are open-source. We also develop an experimental framework compatible with algorithms ranging from rule-based to learning-based provided by third-party frameworks. Lastly, we deploy several SOTA algorithms in example tasks developed via Unreal-MAP, and conduct corresponding experimental analyses including a sim2real demo. We believe Unreal-MAP can play an important role in the MARL field by closely integrating existing algorithms with user-customized tasks, thus advancing the field of MARL.

AAAI Conference 2026 Conference Paper

VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

  • Jiazheng Xu
  • Yu Huang
  • Jiale Cheng
  • Yuanming Yang
  • Jiajun Xu
  • Yuan Wang
  • Wenbo Duan
  • Shen Yang

Visual generative models have achieved remarkable progress in synthesizing photorealistic images and videos, yet aligning their outputs with human preferences across critical dimensions remains a persistent challenge. Though reinforcement learning from human feedback offers promise for preference alignment, existing reward models for visual generation face limitations, including black-box scoring without interpretability and potentially resultant unexpected biases. We present VisionReward, a general framework for learning human visual preferences in both image and video generation. Specifically, we employ a hierarchical visual assessment framework to capture fine-grained human preferences, and leverages linear weighting to enable interpretable preference learning. Furthermore, we propose a multi-dimensional consistent strategy when using VisionReward as a reward model during preference optimization for visual generation. Experiments show that VisionReward can significantly outperform existing image and video reward models on both machine metrics and human evaluation. Notably, VisionReward surpasses VideoScore by 17.2% in preference prediction accuracy, and text-to-video models with VisionReward achieve a 31.6% higher pairwise win rate compared to the same models using VideoScore.

IJCAI Conference 2025 Conference Paper

A Survey of Optimization Modeling Meets LLMs: Progress and Future Directions

  • Ziyang Xiao
  • Jingrong Xie
  • Lilin Xu
  • Shisi Guan
  • Jingyan Zhu
  • Xiongwei Han
  • Xiaojin Fu
  • WingYin Yu

By virtue of its great utility in solving real-world problems, optimization modeling has been widely employed for optimal decision-making across various sectors, but it requires substantial expertise from operations research professionals. With the advent of large language models (LLMs), new opportunities have emerged to automate the procedure of mathematical modeling. This survey presents a comprehensive and timely review of recent advancements that cover the entire technical stack, including data synthesis and fine-tuning for the base model, inference frameworks, benchmark datasets, and performance evaluation. In addition, we conducted an in-depth analysis on the quality of benchmark datasets, which was found to have a surprisingly high error rate. We cleaned the datasets and constructed a new leaderboard with fair performance evaluation in terms of base LLM model and datasets. We also build an online portal that integrates resources of cleaned datasets, code and paper repository to benefit the community. Finally, we identify limitations in current methodologies and outline future research opportunities.

EAAI Journal 2025 Journal Article

An improved reinforcement learning-based differential evolution algorithm for combined economic and emission dispatch problems

  • Yuan Wang
  • Xiaobing Yu
  • Wen Zhang

To overcome challenges posed by escalating environmental pollution and climate change, the combined economic and emission dispatch problem is proposed to balance economic efficiency with emission cost. The primary objective of the problem is to ensure that emissions are minimized while optimal economic costs are achieved simultaneously. However, due to the nonlinear and nonconvex characteristics of the model, the optimization is confronted with many difficulties. Hence, an innovative improved reinforcement learning-based differential evolution algorithm is proposed in this article, with reinforcement learning seamlessly integrated into the differential evolution algorithm. Q-learning from reinforcement learning technique is utilized to dynamically adjust parameter settings and select appropriate mutation strategies, thereby boosting the algorithm's adaptability and overall performance. The effectiveness of the proposed algorithm is tested on thirty testing functions and combined economic and emission dispatch problems in comparison with the other five algorithms. According to the experimental results of testing functions, superior performance is consistently achieved by the proposed algorithm, with the highest adaptability exhibited and an average ranking of 1. 4167. Its superiority is further demonstrated through Wilcoxon tests on results of testing functions and combined economic and emission dispatch problems with the proportion of 100%, and the proposed algorithm is significantly better than other algorithms at a 0. 05 significance level. The superiority of the proposed algorithm in optimizing combined economic and emission dispatch problems demonstrates that the proposed algorithm is shown to be adaptable to complex optimization environments, which proves useful for industrial applications and artificial intelligence.

IROS Conference 2025 Conference Paper

BookBot: A Robotic Manipulation Benchmark for Voice-Driven Book Recognition and Grasping in Cluttered Environments

  • Huaqiang Wang
  • Yuan Wang
  • Xiang Li
  • Yali Li
  • Shengjin Wang

Books, as enduring repositories of cultural heritage as well as knowledge, play a fundamental role in human development. Although advances in embodied AI and robotics revolutionize automation in domains, e. g. , manufacturing and logistics, robotic book manipulation remains an underexplored frontier. Two primary bottlenecks impede progress: (1) scarcity of fine-grained annotated datasets for benchmarking robotic book manipulation, and (2) lack of unified perception-action frameworks capable of dynamically coupling multi-modal sensing and manipulation in real-world scenarios. To these issues, we present THU-Book, the first open-access benchmark featuring 643 3D scene captures, encompassing 11, 298 high-fidelity book instances with rich annotations to support tasks from book recognition and localization to grasping and repositioning. Building upon this foundation, we develop BookBot, a novel voice-interactive book manipulation pipeline to support cross-environmental, multilingual, and multi-categorical book manipulation. First, we utilize Large Language Models (LLMs) to parse and comprehend ambiguity in user instructions. We further propose an instance segmentation module combined with OCR tool to link language to visual instances. Finally, we introduce a PCA-based manipulation policy to refine the robotic grasp pose, utilizing the principal components of the books’ geometry, improving the precision and efficiency of grasping. Experiments conducted on the THU-Book benchmark validate the effectiveness of our BookBot. The dataset is available at https://github.com/wanghq-public/BookBot.

AAAI Conference 2025 Conference Paper

Exploring the Better Multimodal Synergy Strategy for Vision-Language Models

  • Xiaotian Yin
  • Xin Liu
  • Si Chen
  • Yuan Wang
  • Yuwen Pan
  • Tianzhu Zhang

Vision-Language models (VLMs) have shown great potential in enhancing open-world visual concept comprehension. Recent researches focus on an optimum multimodal collaboration strategy that significantly advances CLIP-based few-shot tasks. However, existing prompt-based solutions suffer from unidirectional information flow and increased parameters since they explicitly condition the vision prompts on textual prompts across different transformer layers using non-shareable coupling functions. To address this issue, we propose a Dual-shared mechanism based on LoRA (DsRA) that addresses VLM adaptation in low-data regimes. The proposed DsRA enjoys several merits. First, we design an inter-modal shared coefficient that focuses on capturing visual and textual shared patterns, ensuring effective mutual synergy between image and text features. Second, an intra-modal shared matrix is proposed to achieve efficient parameter fine-tuning by combining the different coefficients to generate layer-wise adapters placed in encoder layers. Our extensive experiments demonstrate that DsRA improves the generalizability under few-shot classification, base-to-new generalization, and domain generalization settings. Our code will be released soon.

IJCAI Conference 2025 Conference Paper

Human-Centric Foundation Models: Perception, Generation and Agentic Modeling

  • Shixiang Tang
  • Yizhou Wang
  • Lu Chen
  • Yuan Wang
  • Sida Peng
  • Dan Xu
  • Wanli Ouyang

Human understanding and generation are critical for modeling digital humans and humanoid embodiments. Recently, Human-centric Foundation Models (HcFMs)—inspired by the success of generalist models such as large language and vision models—have emerged to unify diverse human-centric tasks into a single framework, surpassing traditional task-specific approaches. In this survey, we present a comprehensive overview of HcFMs by proposing a taxonomy that categorizes current approaches into four groups: (1) Human-centric Perception Foundation Models that capture fine-grained features for multi-modal 2D and 3D understanding; (2) Human-centric AIGC Foundation Models that generate high-fidelity, diverse human-related content; (3) Unified Perception and Generation Models that integrate these capabilities to enhance both human understanding and synthesis; and (4) Human-centric Agentic Foundation Models that extend beyond perception and generation to learn human-like intelligence and interactive behaviors for humanoid embodied tasks. We review state-of-the-art techniques, discuss emerging challenges and future research directions. This survey aims to serve as a roadmap for researchers and practitioners working towards more robust, versatile, and intelligent digital human and embodiments modeling. Website is https: //github. com/HumanCentricModels/Awesome-Human-Centric-Foundation-Models/

AAAI Conference 2025 Conference Paper

LIBA: Language Instructed Multi-granularity Bridge Assistant for 3D Visual Grounding

  • Yuan Wang
  • Ya-Li Li
  • W U Eastman Z Y
  • Shengjin Wang

3D Vision Grounding (3D-VG) seeks to unravel referential language and identify targets in 3D physical world. Prevailing methods align with the 2D-VG's pipeline to pinpoint the referred object in a categorical multi-modal reasoning manner. However, the geometric complexities of 3D scenes and the nuanced syntactic structures of language, exacerbates the \textbf{granularity inconsistency} of point cloud and text features, hindering the development of 3D-VG systems in complex scenarios. Towards this issue, we propose LIBA, a Language-Instructed multi-granularity Bridge Assistant tailored for 3D-VG task. LIBA tackles this issue as follows. (1) \textit{How to establish a multi-granularity 3D vision-text feature alignment in a unified model}? We advance a bilateral Dynamic Bridge Adapter (DBA) build multi-granularity interaction of 3D vision and language backnones during feature extraction. We further develop the Language-aware Cross-scale Object Modulation (LCOM) module to integrate multi-scale point cloud features modulated by language information. (2) After aligning multi-modal features, \textit{how to fully harness language model's knowledge to bolster vision concepts understanding}? A LLM-guided Hierarchical Query Selection (LLM-HQS) module incorporates world knowledge of Large Language Model~(LLM) to ground the target referral via an Attribute-then-Relation reasoning process. In this manner, our LIBA inherits reasoning prowess and world knowledge of LLM to bridge point clouds and texts at multiple granularities. Experiments on ScanRefer and Nr3D/Sr3D benchmarks substantiate the superiority of our LIBA, trumping state-of-the-arts by a considerable margin.

AAAI Conference 2025 Conference Paper

Look Back for More: Harnessing Historical Sequential Updates for Personalized Federated Adapter Tuning

  • Danni Peng
  • Yuan Wang
  • Huazhu Fu
  • Jinpeng Jiang
  • Yong Liu
  • Rick Siow Mong Goh
  • Qingsong Wei

Personalized federated learning (PFL) studies effective model personalization to address the data heterogeneity issue among clients in traditional federated learning (FL). Existing PFL approaches mainly generate personalized models by relying solely on the clients' latest updated models while ignoring their previous updates, which may result in suboptimal personalized model learning. To bridge this gap, we propose a novel framework termed pFedSeq, designed for personalizing adapters to fine-tune a foundation model in FL. In pFedSeq, the server maintains and trains a sequential learner, which processes a sequence of past adapter updates from clients and generates calibrations for personalized adapters. To effectively capture the cross-client and cross-step relations hidden in previous updates and generate high-performing personalized adapters, pFedSeq adopts the powerful selective state space model (SSM) as the architecture of sequential learner. Through extensive experiments on four public benchmark datasets, we demonstrate the superiority of pFedSeq over state-of-the-art PFL methods.

IJCAI Conference 2024 Conference Paper

Aggregation and Purification: Dual Enhancement Network for Point Cloud Few-shot Segmentation

  • Guoxin Xiong
  • Yuan Wang
  • Zhaoyang Li
  • Wenfei Yang
  • Tianzhu Zhang
  • Xu Zhou
  • Shifeng Zhang
  • Yongdong Zhang

Point cloud few-shot semantic segmentation (PC-FSS) aims to segment objects within query samples of new categories given only a handful of annotated support samples. Although PC-FSS demonstrates enhanced category generalization capabilities compared to the fully supervised paradigm, the prevalent significant scene discrepancies, which can be systematically summarized into intra-semantic diversity and semantic inconsistency, have posed substantial challenges to the area. In this work, we design a novel Dual Enhancement Network (DENet) to comprehensively tackle different kinds of scene discrepancies in a coherent and synergistic framework. The proposed DENet enjoys several merits. First, we design a mutual aggregation module to reconcile the intrinsic tension between the support prototypes and query point features, and the intra-semantic diversity is diminished in a bidirectional manner. Second, the consistent purification strategy is introduced to eliminate ambiguous prototypes, thereby reducing the mismatches brought by semantic inconsistency. Extensive experiments on S3DIS and ScanNet under different settings demonstrate that DENet significantly outperforms previous SOTAs.

AIIM Journal 2024 Journal Article

Clinical knowledge-guided deep reinforcement learning for sepsis antibiotic dosing recommendations

  • Yuan Wang
  • Anqi Liu
  • Jucheng Yang
  • Lin Wang
  • Ning Xiong
  • Yisong Cheng
  • Qin Wu

Sepsis is the third leading cause of death worldwide. Antibiotics are an important component in the treatment of sepsis. The use of antibiotics is currently facing the challenge of increasing antibiotic resistance (Evans et al. , 2021). Sepsis medication prediction can be modeled as a Markov decision process, but existing methods fail to integrate with medical knowledge, making the decision process potentially deviate from medical common sense and leading to underperformance. (Wang et al. , 2021). In this paper, we use Deep Q-Network (DQN) to construct a Sepsis Anti-infection DQN (SAI-DQN) model to address the challenge of determining the optimal combination and duration of antibiotics in sepsis treatment. By setting sepsis clinical knowledge as reward functions to guide DQN complying with medical guidelines, we formed personalized treatment recommendations for antibiotic combinations. The results showed that our model had a higher average value for decision-making than clinical decisions. For the test set of patients, our model predicts that 79. 07% of patients will achieve a favorable prognosis with the recommended combination of antibiotics. By statistically analyzing decision trajectories and drug action selection, our model was able to provide reasonable medication recommendations that comply with clinical practices. Our model was able to improve patient outcomes by recommending appropriate antibiotic combinations in line with certain clinical knowledge.

NeurIPS Conference 2024 Conference Paper

Enhancing LLM Reasoning via Vision-Augmented Prompting

  • Ziyang Xiao
  • Dongxiang Zhang
  • Xiongwei Han
  • Xiaojin Fu
  • Yin Yu
  • Tao Zhong
  • Sai Wu
  • Yuan Wang

Verbal and visual-spatial information processing are two critical subsystems that activate different brain regions and often collaborate together for cognitive reasoning. Despite the rapid advancement of LLM-based reasoning, the mainstream frameworks, such as Chain-of-Thought (CoT) and its variants, primarily focus on the verbal dimension, resulting in limitations in tackling reasoning problems with visual and spatial clues. To bridge the gap, we propose a novel dual-modality reasoning framework called Vision-Augmented Prompting (VAP). Upon receiving a textual problem description, VAP automatically synthesizes an image from the visual and spatial clues by utilizing external drawing tools. Subsequently, VAP formulates a chain of thought in both modalities and iteratively refines the synthesized image. Finally, a conclusive reasoning scheme based on self-alignment is proposed for final result generation. Extensive experiments are conducted across four versatile tasks, including solving geometry problems, Sudoku, time series prediction, and travelling salesman problem. The results validated the superiority of VAP over existing LLMs-based reasoning frameworks.

AAAI Conference 2024 Conference Paper

Frequency Shuffling and Enhancement for Open Set Recognition

  • Lijun Liu
  • Rui Wang
  • Yuan Wang
  • Lihua Jing
  • Chuan Wang

Open-Set Recognition (OSR) aims to accurately identify known classes while effectively rejecting unknown classes to guarantee reliability. Most existing OSR methods focus on learning in the spatial domain, where subtle texture and global structure are potentially intertwined. Empirical studies have shown that DNNs trained in the original spatial domain are inclined to over-perceive subtle texture. The biased semantic perception could lead to catastrophic over-confidence when predicting both known and unknown classes. To this end, we propose an innovative approach by decomposing the spatial domain to the frequency domain to separately consider global (low-frequency) and subtle (high-frequency) information, named Frequency Shuffling and Enhancement (FreSH). To alleviate the overfitting of subtle texture, we introduce the High-Frequency Shuffling (HFS) strategy that generates diverse high-frequency information and promotes the capture of low-frequency invariance. Moreover, to enhance the perception of global structure, we propose the Low-Frequency Residual (LFR) learning procedure that constructs a composite feature space, integrating low-frequency and original spatial features. Experiments on various benchmarks demonstrate that the proposed FreSH consistently trumps the state-of-the-arts by a considerable margin.

AAAI Conference 2024 Conference Paper

Near-Optimal Resilient Aggregation Rules for Distributed Learning Using 1-Center and 1-Mean Clustering with Outliers

  • Yuhao Yi
  • Ronghui You
  • Hong Liu
  • Changxin Liu
  • Yuan Wang
  • Jiancheng Lv

Byzantine machine learning has garnered considerable attention in light of the unpredictable faults that can occur in large-scale distributed learning systems. The key to secure resilience against Byzantine machines in distributed learning is resilient aggregation mechanisms. Although abundant resilient aggregation rules have been proposed, they are designed in ad-hoc manners, imposing extra barriers on comparing, analyzing, and improving the rules across performance criteria. This paper studies near-optimal aggregation rules using clustering in the presence of outliers. Our outlier-robust clustering approach utilizes geometric properties of the update vectors provided by workers. Our analysis show that constant approximations to the 1-center and 1-mean clustering problems with outliers provide near-optimal resilient aggregators for metric-based criteria, which have been proven to be crucial in the homogeneous and heterogeneous cases respectively. In addition, we discuss two contradicting types of attacks under which no single aggregation rule is guaranteed to improve upon the naive average. Based on the discussion, we propose a two-phase resilient aggregation framework. We run experiments for image classification using a non-convex loss function. The proposed algorithms outperform previously known aggregation rules by a large margin with both homogeneous and heterogeneous data distributions among non-faulty workers. Code and appendix are available at https://github.com/jerry907/AAAI24-RASHB.

AAAI Conference 2024 Conference Paper

Pay Attention to Target: Relation-Aware Temporal Consistency for Domain Adaptive Video Semantic Segmentation

  • Huayu Mai
  • Rui Sun
  • Yuan Wang
  • Tianzhu Zhang
  • Feng Wu

Video semantic segmentation has achieved conspicuous achievements attributed to the development of deep learning, but suffers from labor-intensive annotated training data gathering. To alleviate the data-hunger issue, domain adaptation approaches are developed in the hope of adapting the model trained on the labeled synthetic videos to the real videos in the absence of annotations. By analyzing the dominant paradigm consistency regularization in the domain adaptation task, we find that the bottlenecks exist in previous methods from the perspective of pseudo-labels. To take full advantage of the information contained in the pseudo-labels and empower more effective supervision signals, we propose a coherent PAT network including a target domain focalizer and relation-aware temporal consistency. The proposed PAT network enjoys several merits. First, the target domain focalizer is responsible for paying attention to the target domain, and increasing the accessibility of pseudo-labels in consistency training. Second, the relation-aware temporal consistency aims at modeling the inter-class consistent relationship across frames to equip the model with effective supervision signals. Extensive experimental results on two challenging benchmarks demonstrate that our method performs favorably against state-of-the-art domain adaptive video semantic segmentation methods.

EAAI Journal 2024 Journal Article

Topological persistence guided knowledge distillation for wearable sensor data

  • Eun Som Jeon
  • Hongjun Choi
  • Ankita Shukla
  • Yuan Wang
  • Hyunglae Lee
  • Matthew P. Buman
  • Pavan Turaga

Deep learning methods have achieved a lot of success in various applications involving converting wearable sensor data to actionable health insights. A common application areas is activity recognition, where deep-learning methods still suffer from limitations such as sensitivity to signal quality, sensor characteristic variations, and variability between subjects. To mitigate these issues, robust features obtained by topological data analysis (TDA) have been suggested as a potential solution. However, there are two significant obstacles to using topological features in deep learning: (1) large computational load to extract topological features using TDA, and (2) different signal representations obtained from deep learning and TDA which makes fusion difficult. In this paper, to enable integration of the strengths of topological methods in deep-learning for time-series data, we propose to use two teacher networks — one trained on the raw time-series data, and another trained on persistence images generated by TDA methods. These two teachers are jointly used to distill a single student model, which utilizes only the raw time-series data at test-time. This approach addresses both issues. The use of KD with multiple teachers utilizes complementary information, and results in a compact model with strong supervisory features and an integrated richer representation. To assimilate desirable information from different modalities, we design new constraints, including orthogonality imposed on feature correlation maps for improving feature expressiveness and allowing the student to easily learn from the teacher. Also, we apply an annealing strategy in KD for fast saturation and better accommodation from different features, while the knowledge gap between the teachers and student is reduced. Finally, a robust student model is distilled, which can at test-time uses only the time-series data as an input, while implicitly preserving topological features. The experimental results demonstrate the effectiveness of the proposed method on wearable sensor data. The proposed method shows 71. 74% in classification accuracy on GENEActiv with WRN16-1 (1D CNNs) student, which outperforms baselines and takes much less processing time (less than 17 sec) than teachers on 6k testing samples.

IJCAI Conference 2023 Conference Paper

A New ANN-SNN Conversion Method with High Accuracy, Low Latency and Good Robustness

  • Bingsen Wang
  • Jian Cao
  • Jue Chen
  • Shuo Feng
  • Yuan Wang

Due to the advantages of low energy consumption, high robustness and fast inference speed, Spiking Neural Networks (SNNs), with good biological interpretability and the potential to be applied on neuromorphic hardware, are regarded as the third generation of Artificial Neural Networks (ANNs). Despite having so many advantages, the biggest challenge encountered by spiking neural networks is training difficulty caused by the non-differentiability of spike signals. ANN-SNN conversion is an effective method that solves the training difficulty by converting parameters in ANNs to those in SNNs through a specific algorithm. However, the ANN-SNN conversion method also suffers from accuracy degradation and long inference time. In this paper, we reanalyzed the relationship between Integrate-and-Fire (IF) neuron model and ReLU activation function, proposed a StepReLU activation function more suitable for SNNs under membrane potential encoding, and used it to train ANNs. Then we converted the ANNs to SNNs with extremely small conversion error and introduced leakage mechanism to the SNNs and get the final models, which have high accuracy, low latency and good robustness, and have achieved the state-of-the-art performance on various datasets such as CIFAR and ImageNet.

NeurIPS Conference 2023 Conference Paper

Focus on Query: Adversarial Mining Transformer for Few-Shot Segmentation

  • Yuan Wang
  • Naisong Luo
  • Tianzhu Zhang

Few-shot segmentation (FSS) aims to segment objects of new categories given only a handful of annotated samples. Previous works focus their efforts on exploring the support information while paying less attention to the mining of the critical query branch. In this paper, we rethink the importance of support information and propose a new query-centric FSS model Adversarial Mining Transformer (AMFormer), which achieves accurate query image segmentation with only rough support guidance or even weak support labels. The proposed AMFormer enjoys several merits. First, we design an object mining transformer (G) that can achieve the expansion of incomplete region activated by support clue, and a detail mining transformer (D) to discriminate the detailed local difference between the expanded mask and the ground truth. Second, we propose to train G and D via an adversarial process, where G is optimized to generate more accurate masks approaching ground truth to fool D. We conduct extensive experiments on commonly used Pascal-5i and COCO-20i benchmarks and achieve state-of-the-art results across all settings. In addition, the decent performance with weak support labels in our query-centric paradigm may inspire the development of more general FSS models.

AAAI Conference 2023 Conference Paper

Neural TSP Solver with Progressive Distillation

  • Dongxiang Zhang
  • Ziyang Xiao
  • Yuan Wang
  • Mingli Song
  • Gang Chen

Travelling salesman problem (TSP) is NP-Hard with exponential search space. Recently, the adoption of encoder-decoder models as neural TSP solvers has emerged as an attractive topic because they can instantly obtain near-optimal results for small-scale instances. Nevertheless, their training efficiency and solution quality degrade dramatically when dealing with large-scale problems. To address the issue, we propose a novel progressive distillation framework, by adopting curriculum learning to train TSP samples in increasing order of their problem size and progressively distilling high-level knowledge from small models to large models via a distillation loss. In other words, the trained small models are used as the teacher network to guide action selection when training large models. To accelerate training speed, we also propose a Delaunary-graph based action mask and a new attention-based decoder to reduce decoding cost. Experimental results show that our approach establishes clear advantages over existing encoder-decoder models in terms of training effectiveness and solution quality. In addition, we validate its usefulness as an initial solution generator for the state-of-the-art TSP solvers, whose probability of obtaining the optimal solution can be further improved in such a hybrid manner.

EAAI Journal 2022 Journal Article

Application of multi-objective particle swarm optimization based on short-term memory and K-means clustering in multi-modal multi-objective optimization

  • Yang Yang
  • Qianfeng Liao
  • Jiang Wang
  • Yuan Wang

To solve the multi-modal multi-objective optimization problems in which the same Pareto Front (PF) may correspond to multiple different Pareto Optimal Sets (PSs), an improved multi-objective particle swarm optimizer with short-term memory and K-means clustering (MOPSO-SMK) is proposed in this paper. According to the framework of multi-objective particle swarm optimization (MOPSO) algorithm, the designs of updating mechanism and population maintenance mechanism are the keys to obtain the optimal solutions. As a significant influence factor of the updating mechanism, the inertia weight has been discussed in this paper. In the improved algorithm, a new update model for the value of pbest based on short-term memory is proposed. The update strategies based on K-means clustering are adopted to obtain the better gbest and elite archive. 16 multi-modal multi-objective optimization functions are used to verify the feasibility and effectiveness of the proposed MOPSO-SMK. As the results show, MOPSO-SMK has more advantages in four indexes (1/PSP, 1/HV, IGDX, and IGDF) compared with other three multi-objective optimization algorithms.

IJCAI Conference 2022 Conference Paper

Estimation and Comparison of Linear Regions for ReLU Networks

  • Yuan Wang

We study the relationship between the arrangement of neurons and the complexity of the ReLU-activated neural networks measured by the number of linear regions. More specifically, we provide both theoretical and empirical evidence for the point of view that shallow networks tend to have higher complexity than deep ones when the total number of neurons is fixed. In the theoretical part, we prove that this is the case for networks whose neurons in the hidden layers are arranged in the forms of 1x2n, 2xn and nx2; in the empirical part, we implement an algorithm that precisely tracks (hence counts) all the linear regions, and run it on networks with various structures. Although the time complexity of the algorithm is quite high, we verify that the problem of calculating the number of linear regions of a ReLU network is itself NP-hard. So currently there is no surprisingly efficient way to solve it. Roughly speaking, in the algorithm we divide the linear regions into subregions called the "activation regions", which are convex and easy to propagate through the network. The relationship between the number of the linear regions and that of the activation regions is also discussed.

AIIM Journal 2022 Journal Article

Global and local attentional feature alignment for domain adaptive nuclei detection in histopathology images

  • Zhi Wang
  • Xiaoya Zhu
  • Ao Li
  • Yuan Wang
  • Gang Meng
  • Minghui Wang

Automated nuclei detection is crucial prerequisites for a number of histopathology related image analysis such as cancer diagnosis. Although existing deep learning based nuclei detection methods have achieved promising results, they cannot effectively deal with domain shift problem caused by different staining procedures and organ specific nuclear morphology. To handle this problem, in this paper a novel adversarial feature alignment method is proposed for domain adaptive nuclei detection, which includes both global alignment and local attentional alignment components to transfer the knowledge from source domain to target domain. Specifically, in local attentional alignment component, by using nuclei locations as guidance we extract local features and perform adversarial alignment. Furthermore, to address the issue that these local features from nuclei regions often contain insufficient information because of the small size of nuclei, we introduce an efficient location-aware self-attention (LocSA) module to refine local features by utilizing cues from all nuclei for obtaining discriminative features to perform successful feature alignment. Extensive experimental results are provided on two adaptation scenarios and our method demonstrates favorable performance against existing domain adaptation methods, which highlights the effectiveness of the proposed method for domain adaptive nuclei detection.

AAAI Conference 2022 Conference Paper

Learning to Detect 3D Facial Landmarks via Heatmap Regression with Graph Convolutional Network

  • Yuan Wang
  • Min Cao
  • Zhenfeng Fan
  • Silong Peng

3D facial landmark detection is extensively used in many research fields such as face registration, facial shape analysis, and face recognition. Most existing methods involve traditional features and 3D face models for the detection of landmarks, and their performances are limited by the hand-crafted intermediate process. In this paper, we propose a novel 3D facial landmark detection method, which directly locates the coordinates of landmarks from 3D point cloud with a wellcustomized graph convolutional network. The graph convolutional network learns geometric features adaptively for 3D facial landmark detection with the assistance of constructed 3D heatmaps, which are Gaussian functions of distances to each landmark on a 3D face. On this basis, we further develop a local surface unfolding and registration module to predict 3D landmarks from the heatmaps. The proposed method forms the first baseline of deep point cloud learning method for 3D facial landmark detection. We demonstrate experimentally that the proposed method exceeds the existing approaches by a clear margin on BU-3DFE and FRGC datasets for landmark localization accuracy and stability, and also achieves highprecision results on a recent large-scale dataset.

ICLR Conference 2021 Conference Paper

Anchor & Transform: Learning Sparse Embeddings for Large Vocabularies

  • Paul Pu Liang
  • Manzil Zaheer
  • Yuan Wang
  • Amr Ahmed 0001

Learning continuous representations of discrete objects such as text, users, movies, and URLs lies at the heart of many applications including language and user modeling. When using discrete objects as input to neural networks, we often ignore the underlying structures (e.g., natural groupings and similarities) and embed the objects independently into individual vectors. As a result, existing methods do not scale to large vocabulary sizes. In this paper, we design a simple and efficient embedding algorithm that learns a small set of anchor embeddings and a sparse transformation matrix. We call our method Anchor & Transform (ANT) as the embeddings of discrete objects are a sparse linear combination of the anchors, weighted according to the transformation matrix. ANT is scalable, flexible, and end-to-end trainable. We further provide a statistical interpretation of our algorithm as a Bayesian nonparametric prior for embeddings that encourages sparsity and leverages natural groupings among objects. By deriving an approximate inference algorithm based on Small Variance Asymptotics, we obtain a natural extension that automatically learns the optimal number of anchors instead of having to tune it as a hyperparameter. On text classification, language modeling, and movie recommendation benchmarks, we show that ANT is particularly suitable for large vocabulary sizes and demonstrates stronger performance with fewer parameters (up to 40x compression) as compared to existing compression baselines.

EAAI Journal 2020 Journal Article

Improved sine cosine algorithm combined with optimal neighborhood and quadratic interpolation strategy

  • Wen-yan Guo
  • Yuan Wang
  • Fang Dai
  • Peng Xu

The sine cosine algorithm (SCA) is a new population-based stochastic optimization algorithm, utilizing the oscillating property of the sine cosine function to balance the exploration and exploitation performance of SCA. A hybrid sine cosine algorithm based on the optimal neighborhood and quadratic interpolation strategy (QISCA) was proposed to overcome the shortcoming of updating the population guided by the global optimal individual in the sine cosine algorithm. The new algorithm uses a Stochastic Optimal Neighborhood for neighborhood updates, and it adopts a Quadratic Interpolation curve for individual updates. In addition, QISCA incorporates Quasi-Opposition Learning strategies to enhance the population’s global exploration capabilities, and improves the convergence speed and accuracy. The two simulation experiments of 23 benchmark functions and 30 latest CEC2017 test functions show that the new algorithm can better coordinate the exploration and exploitation capabilities and improve the global optimization ability, compared with the other improved sine cosine algorithm and the representative stochastic optimization algorithm. The three representative engineering problems validate the effectiveness of the new algorithm to solve practical problems.

AAAI Conference 2020 Short Paper

Neural Dynamics and Gamma Oscillation on a Hybrid Excitatory-Inhibitory Complex Network (Student Abstract)

  • Yuan Wang
  • Xia Shi
  • Bo Cheng
  • Junliang Chen

This paper investigates the neural dynamics and gamma oscillation on a complex network with excitatory and inhibitory neurons (E-I network), as such network is ubiquitous in the brain. The system consists of a small-world network of neurons, which are emulated by Izhikevich model. Moreover, mixed Regular Spiking (RS) and Chattering (CH) neurons are considered to imitate excitatory neurons, and Fast Spiking (FS) neurons are used to mimic inhibitory neurons. Besides, the relationship between synchronization and gamma rhythm is explored by adjusting the critical parameters of our model. Experiments visually demonstrate that the gamma oscillations are generated by synchronous behaviors of our neural network. We also discover that the Chattering(CH) excitatory neurons can make the system easier to synchronize.

YNIMG Journal 2018 Journal Article

Generalized Recurrent Neural Network accommodating Dynamic Causal Modeling for functional MRI analysis

  • Yuan Wang
  • Yao Wang
  • Yvonne W. Lui

Dynamic Causal Modeling (DCM) is an advanced biophysical model which explicitly describes the entire process from experimental stimuli to functional magnetic resonance imaging (fMRI) signals via neural activity and cerebral hemodynamics. To conduct a DCM study, one needs to represent the experimental stimuli as a compact vector-valued function of time, which is hard in complex tasks such as book reading and natural movie watching. Deep learning provides the state-of-the-art signal representation solution, encoding complex signals into compact dense vectors while preserving the essence of the original signals. There is growing interest in using Recurrent Neural Networks (RNNs), a major family of deep learning techniques, in fMRI modeling. However, the generic RNNs used in existing studies work as black boxes, making the interpretation of results in a neuroscience context difficult and obscure. In this paper, we propose a new biophysically interpretable RNN built on DCM, DCM-RNN. We generalize the vanilla RNN and show that DCM can be cast faithfully as a special form of the generalized RNN. DCM-RNN uses back propagation for parameter estimation. We believe DCM-RNN is a promising tool for neuroscience. It can fit seamlessly into classical DCM studies. We demonstrate face validity of DCM-RNN in two principal applications of DCM: causal brain architecture hypotheses testing and effective connectivity estimation. We also demonstrate construct validity of DCM-RNN in an attention-visual experiment. Moreover, DCM-RNN enables end-to-end training of DCM and representation learning deep neural networks, extending DCM studies to complex tasks.

AAAI Conference 2016 Conference Paper

Hashtag-Based Sub-Event Discovery Using Mutually Generative LDA in Twitter

  • Chen Xing
  • Yuan Wang
  • Jie Liu
  • Yalou Huang
  • Wei-Ying Ma

Sub-event discovery is an effective method for social event analysis in Twitter. It can discover sub-events from large amount of noisy event-related information in Twitter and semantically represent them. The task is challenging because tweets are short, informal and noisy. To solve this problem, we consider leveraging event-related hashtags that contain many locations, dates and concise sub-event related descriptions to enhance sub-event discovery. To this end, we propose a hashtag-based mutually generative Latent Dirichlet Allocation model(MGe-LDA). In MGe-LDA, hashtags and topics of a tweet are mutually generated by each other. The mutually generative process models the relationship between hashtags and topics of tweets, and highlights the role of hashtags as a semantic representation of the corresponding tweets. Experimental results show that MGe-LDA can significantly outperform state-of-the-art methods for sub-event discovery.

v2026.09.13