Arrow Research search

Author name cluster

Zhiyuan Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

17 papers
2 author rows

Possible papers

17

AAAI Conference 2026 Conference Paper

COIN: Uncertainty-Guarding Selective Question Answering for Foundation Models with Provable Risk Guarantees

  • Zhiyuan Wang
  • Jinhao Duan
  • Qingni Wang
  • Xiaofeng Zhu
  • Tianlong Chen
  • Xiaoshuang Shi
  • Kaidi Xu

Uncertainty quantification (UQ) in foundation models is crucial for identifying and mitigating hallucinations in automatically generated text. However, heuristic UQ approaches lack statistical guarantees for key metrics such as the false discovery rate (FDR) in selective prediction tasks. Previous research adopts the split conformal prediction (SCP) framework to ensure desired coverage of admissible answers by constructing data-driven prediction sets, yet these sets typically contain incorrect candidates, undermining their practical effectiveness. To address this, we introduce COIN, an uncertainty-guarding selection framework that calibrates statistically valid uncertainty thresholds to filter a single generated answer per question under user-specified FDR constraints. COIN estimates the empirical error rate on the calibration set and applies confidence interval methods such as Clopper–Pearson to establish a high-probability upper bound on the true error rate (i.e., FDR). This enables the selection of the largest threshold that ensures FDR control on test data while significantly increasing sample retention. We demonstrate COIN's robustness in risk control, strong test-time power in retaining admissible answers, and predictive efficiency under limited calibration data across both general and multimodal text generation tasks. Furthermore, we show that employing alternative UQ and upper bound construction strategies can further boost COIN's power performance, which underscores its extensibility and adaptability to diverse application scenarios.

AAAI Conference 2026 Conference Paper

MAPS: Multi-Agent Personality Shaping for Collaborative Reasoning

  • Jian Zhang
  • Zhiyuan Wang
  • Zhangqi Wang
  • Fangzhi Xu
  • Qika Lin
  • Lingling Zhang
  • Rui Mao
  • Erik Cambria

Collaborative reasoning with multiple agents offers the potential for more robust and diverse problem-solving. However, existing approaches often suffer from homogeneous agent behaviors and lack of reflective and rethinking capabilities. We propose Multi-Agent Personality Shaping ((MAPS), a novel framework that enhances reasoning through agent diversity and internal critique. Inspired by the Big Five personality theory, MAPS assigns distinct personality traits to individual agents, shaping their reasoning styles and promoting heterogeneous collaboration. To enable deeper and more adaptive reasoning, MAPS introduces a Critic agent that reflects on intermediate outputs, revisits flawed steps, and guides iterative refinement. This integration of personality-driven agent design and structured collaboration improves both reasoning depth and flexibility. Empirical evaluations across three benchmarks demonstrate the strong performance of MAPS, with further analysis confirming its generalizability across different large language models and validating the benefits of multi-agent collaboration.

NeurIPS Conference 2025 Conference Paper

A Plug-and-Play Query Synthesis Active Learning Framework for Neural PDE Solvers

  • Zhiyuan Wang
  • Jinwoo Go
  • Byung-Jun Yoon
  • Nathan Urban
  • Xiaoning Qian

In recent developments in scientific machine learning (SciML), neural surrogate solvers for partial differential equations (PDEs) have become powerful tools for accelerating scientific computation for various science and engineering applications. However, training neural PDE solvers often demands a large amount of high-fidelity PDE simulation data, which are expensive to generate. Active learning (AL) offers a promising solution by adaptively selecting training data from the PDE settings--including parameters, initial and boundary conditions--that are expected to be most informative to help reduce this data burden. In this work, we introduce PaPQS, a Plug-and-Play Query Synthesis AL framework that synthesizes informative PDE settings directly in the continuous design space. PaPQS optimizes the Expected Information Gain (EIG) while encouraging batch diversity, enabling model-aware exploration of the design space via backpropagation through the neural PDE solution trajectories. The framework is applicable to general PDE systems and surrogate architectures, and can be seamlessly integrated with existing AL strategies. Extensive experiments across different PDE systems demonstrate that our AL framework, PaPQS, consistently improves sample efficiency over existing AL baselines.

ECAI Conference 2025 Conference Paper

Cascaded Large-Scale TSP Solving with Unified Neural Guidance: Bridging Local and Population-Based Search

  • Haoze Lv
  • Wenjie Chen
  • Zhiyuan Wang
  • Shengcai Liu

The traveling salesman problem (TSP) is a fundamental NP-hard optimization problem. Over the past decades, traditional heuristic methods have achieved substantial success in solving TSP, yet their performance, particularly for large-scale instances, remains to be further improved. The advancement of deep learning technologies over the past decade has driven a growing number of attempts to solve TSP by leveraging neural guidance. However, these efforts predominantly focus on small-scale TSP instances, with limited improvements in solving performance for large-scale instances, revealing persistent scalability challenges. This work presents UNiCS, a novel unified neural-guided cascaded solver for solving large-scale TSP instances. UNiCS comprises a local search (LS) phase and a population-based search (PBS) phase, both guided by a learning component called unified neural guidance (UNG). Specifically, UNG guides solution generation across both phases and determines appropriate phase transition timing to effectively combine the complementary strengths of LS and PBS. While trained only on simple distributions with relatively small-scale TSP instances, UNiCS generalizes effectively to challenging TSP benchmarks containing much larger instances (10, 000-71, 009 nodes) with diverse node distributions entirely unseen during training. Experimental results on the large-scale TSP instances demonstrate that UNiCS consistently outperforms state-of-the-art methods, with its advantage remaining consistent across various runtime budgets.

NeurIPS Conference 2025 Conference Paper

Harnessing the Computation Redundancy in ViTs to Boost Adversarial Transferability

  • Jiani Liu
  • Zhiyuan Wang
  • Zeliang Zhang
  • Chao Huang
  • Susan Liang
  • Yunlong Tang
  • Chenliang Xu

Vision Transformers (ViTs) have demonstrated impressive performance across a range of applications, including many safety-critical tasks. Many previous studies have observed that adversarial examples crafted on ViTs exhibit higher transferability than those crafted on CNNs, indicating that ViTs contain structural characteristics favorable for transferable attacks. In this work, we take a further step to deeply investigate the role of computational redundancy brought by its unique characteristics in ViTs and its impact on adversarial transferability. Specifically, we identify two forms of redundancy, including the data-level and model-level, that can be harnessed to amplify attack effectiveness. Building on this insight, we design a suite of techniques, including attention sparsity manipulation, attention head permutation, clean token regularization, ghost MoE diversification, and learn to robustify before the attack. A dynamic online learning strategy is also proposed to fully leverage these operations to enhance the adversarial transferability. Extensive experiments on the ImageNet-1k dataset validate the effectiveness of our approach, showing that our methods significantly outperform existing baselines in both transferability and generality across diverse model architectures, including different variants of ViTs and mainstream Vision Large Language Models (VLLMs).

EAAI Journal 2025 Journal Article

Multi-source perception data fusion of vessels in visual occlusion scenarios: Leveraging prior knowledge of vessel motion

  • Wei He
  • Wenbo He
  • Jinyu Lei
  • Sitong Wan
  • Zhiyuan Wang

Automatic Identification System (AIS) and cameras are widely used in harbor and coastal supervision to detect ship movement. Integrating both can enhance the perception and monitoring of the navigation status of surrounding vessels. However, their integration faces several challenges. First, the two data sources have different coordinate systems and sampling frequencies. Additionally, in video data, visual occlusion during vessel encounters may lead to the loss or displacement of ship detection targets, significantly affecting the fusion of ship data. To address these issues, we incorporate prior knowledge of vessel motion to strengthen the model’s ability to identify and track occluded targets. Firstly, when detecting and tracking ships in video, we use occlusion prior knowledge and tracking results to evaluate and manage the occluded detection box areas. Secondly, when the occluded detection box disappears or undergoes severe deformation, we use the prior knowledge of ship motion characteristics from AIS and images to predict the detection box. Finally, we validated our improved method on the FVessel_v1. 0 dataset, confirming its accuracy in data fusion under occlusion conditions. Compared with the state-of-the-art ship data fusion algorithm, our method improved the Multiple Object Fusion Accuracy (MOFA), Identification Precision (IDP), Identification Recall (IDR), and Identification F1 score (IDPF1) metrics by 3. 01%, 1. 33%, 1. 57%, and 1. 33%, and reduced MOFA by 1. 85%. Additionally, by only changing the method of utilizing prior knowledge without altering the detection and tracking algorithms of the state-of-the-art ship data fusion algorithm, we improved the MOFA, IDP, IDR, and IDPF1 metrics by 2. 65%, 1. 94%, 0. 62%, and 0. 87%, respectively, and reduced MOFA by 1. 16%.

IJCAI Conference 2025 Conference Paper

PanComplex: Leveraging Complex-Valued Neural Networks for Enhanced Pansharpening

  • Chunhui Luo
  • Dong Li
  • Xiaoliang Ma
  • Xin Lu
  • Zhiyuan Wang
  • Jiangtong Tan
  • Xueyang Fu

Pansharpening combines panchromatic and low-resolution multispectral images to generate high-resolution multispectral images. Previous studies have explored the connection between pansharpening and the frequency domain, but mostly in the real-valued domain, leaving the complex domain relatively unexplored. To redefine the pansharpening task, we propose a complex-valued spatial-frequency dual-domain framework, PanComplex. To achieve this, we first establish complex representations and introduce basic complex operators tailored to pansharpening, enabling the transformation of multispectral real-valued signals into the complex domain for learning. We then model both spatial and frequency branches to capture global frequency features and local spatial features comprehensively. Finally, we employ a complex-based interaction module to fuse the spatial and frequency features, achieving complementary information across both domains. By using the representation power of the complex domain, PanComplex effectively extracts complementary features from PAN and MS images, thereby enhancing pansharpening performance. Experiments on multiple datasets demonstrate that our method achieves optimal performance with the fewest parameters and exhibits strong generalization ability to other tasks. The source code for this work is publicly available at https: //github. com/lch-ustc/PanComplex.

ICLR Conference 2025 Conference Paper

Sample then Identify: A General Framework for Risk Control and Assessment in Multimodal Large Language Models

  • Qingni Wang
  • Tiantian Geng
  • Zhiyuan Wang
  • Teng Wang 0007
  • Bo Fu
  • Feng Zheng 0001

Multimodal Large Language Models (MLLMs) exhibit promising advancements across various tasks, yet they still encounter significant trustworthiness issues. Prior studies apply Split Conformal Prediction (SCP) in language modeling to construct prediction sets with statistical guarantees. However, these methods typically rely on internal model logits or are restricted to multiple-choice settings, which hampers their generalizability and adaptability in dynamic, open-ended environments. In this paper, we introduce *TRON*, a **t**wo-step framework for **r**isk c**o**ntrol and assessme**n**t, applicable to any MLLM that supports sampling in both open-ended and closed-ended scenarios. *TRON* comprises two main components: (1) a novel conformal score to **sample** response sets of minimum size, and (2) a nonconformity score to **identify** high-quality responses based on self-consistency theory, controlling the error rates by two specific risk levels. Furthermore, we investigate semantic redundancy in prediction sets within open-ended contexts for the first time, leading to a promising evaluation metric for MLLMs based on average set size. Our comprehensive experiments across four Video Question-Answering (VideoQA) datasets utilizing eight MLLMs show that *TRON* achieves desired error rates bounded by two user-specified risk levels. Additionally, deduplicated prediction sets maintain adaptiveness while being more efficient and stable for risk assessment under different risk levels.

ICRA Conference 2025 Conference Paper

Winding Number-Guided Edge-Preserving Implicit Neural Representation of CAD Surfaces

  • Yuhang Cheng
  • Zhiyuan Wang
  • Jialan He
  • Xiaogang Wang

Implicit surface representations have emerged as a powerful tool for the task of 3D reconstruction due to their excellent performance. Yet, when the normal information cannot be available, the previous methods often lead to unsatisfactory reconstruction results, even failure. To this end, we propose a winding number—guided implicit surface reconstruction method, which mainly consists of a winding number—guided regularizer and a dynamic edge sampling strategy. Among them, the winding number-guided regularizer can effectively constrain the global normal consistency of the input raw data, as well as improve the unsatisfactory implicit surface reconstruction result caused by the unavailability of normal information. Meanwhile, in order to reduce the excessive smoothing at sharp edges of implicit surface, we proposed a dynamic edge sampling strategy for sampling near the sharp edge regions of 3D shape, which can effectively avoid the regularizer from smoothing all regions. Finally, we combine them with a simple data term for robust implicit surface reconstruction. Compared with the state-of-the-art methods, experimental results show that our method significantly improves the quality of 3D reconstruction results. In addition, since the winding number-guided regularizer effectively constraints the globally consistent normal of the input 3D raw data, our method can also receive an additional gift, namely the globally consistent normal estimation results of 3D raw data.

EAAI Journal 2025 Journal Article

Word-Sequence Entropy: Towards uncertainty estimation in free-form medical question answering applications and beyond

  • Zhiyuan Wang
  • Jinhao Duan
  • Chenxi Yuan
  • Qingyu Chen
  • Tianlong Chen
  • Yue Zhang
  • Ren Wang
  • Xiaoshuang Shi

Uncertainty estimation is crucial for the reliability of safety-critical human and artificial intelligence (AI) interaction systems, particularly in the domain of healthcare engineering. However, a robust and general uncertainty measure for free-form answers has not been well-established in open-ended medical question-answering (QA) tasks, where generative inequality introduces a large number of irrelevant words and sequences within the generated set for uncertainty quantification (UQ), which can lead to biases. This paper proposes Word-Sequence Entropy (WSE), which calibrates uncertainty at both the word and sequence levels based on semantic relevance, highlighting keywords and enlarging the generative probability of trustworthy responses when performing UQ. We compare WSE with six baseline methods on five free-form medical QA datasets, utilizing seven popular large language models (LLMs), and demonstrate that WSE exhibits superior performance in accurate UQ under two standard criteria for correctness evaluation. Additionally, in terms of the potential for real-world medical QA applications, we achieve a significant enhancement (e. g. , a 6. 36% improvement in model accuracy on the COVID-QA dataset) in the performance of LLMs when employing responses with lower uncertainty that are identified by WSE as final answers, without requiring additional task-specific fine-tuning or architectural modifications.

AAAI Conference 2023 Short Paper

DyCVAE: Learning Dynamic Causal Factors for Non-stationary Series Domain Generalization (Student Abstract)

  • Weifeng Zhang
  • Zhiyuan Wang
  • Kunpeng Zhang
  • Ting Zhong
  • Fan Zhou

Learning domain-invariant representations is a major task of out-of-distribution generalization. To address this issue, recent efforts have taken into accounting causality, aiming at learning the causal factors with regard to tasks. However, extending existing generalization methods for adapting non-stationary time series may be ineffective, because they fail to model the underlying causal factors due to temporal-domain shifts except for source-domain shifts, as pointed out by recent studies. To this end, we propose a novel model DyCVAE to learn dynamic causal factors. The results on synthetic and real datasets demonstrate the effectiveness of our proposed model for the task of generalization in time series domain.

AAAI Conference 2023 Short Paper

Learning Dynamic Temporal Relations with Continuous Graph for Multivariate Time Series Forecasting (Student Abstract)

  • Zhiyuan Wang
  • Fan Zhou
  • Goce Trajcevski
  • Kunpeng Zhang
  • Ting Zhong

The recent advance in graph neural networks (GNNs) has inspired a few studies to leverage the dependencies of variables for time series prediction. Despite the promising results, existing GNN-based models cannot capture the global dynamic relations between variables owing to the inherent limitation of their graph learning module. Besides, multi-scale temporal information is usually ignored or simply concatenated in prior methods, resulting in inaccurate predictions. To overcome these limitations, we present CGMF, a Continuous Graph learning method for Multivariate time series Forecasting (CGMF). Our CGMF consists of a continuous graph module incorporating differential equations to capture the long-range intra- and inter-relations of the temporal embedding sequence. We also introduce a controlled differential equation-based fusion mechanism that efficiently exploits multi-scale representations to form continuous evolutional dynamics and learn rich relations and patterns shared across different scales. Comprehensive experiments demonstrate the effectiveness of our method for a variety of datasets.

IJCAI Conference 2023 Conference Paper

VecoCare: Visit Sequences-Clinical Notes Joint Learning for Diagnosis Prediction in Healthcare Data

  • Yongxin Xu
  • Kai Yang
  • Chaohe Zhang
  • Peinie Zou
  • Zhiyuan Wang
  • Hongxin Ding
  • Junfeng Zhao
  • Yasha Wang

Due to the insufficiency of electronic health records (EHR) data utilized in practical diagnosis prediction scenarios, most works are devoted to learning powerful patient representations either from structured EHR data (e. g. , temporal medical events, lab test results, etc. ) or unstructured data (e. g. , clinical notes, etc. ). However, synthesizing rich information from both of them still needs to be explored. Firstly, the heterogeneous semantic biases across them heavily hinder the synthesis of representation spaces, which is critical for diagnosis prediction. Secondly, the intermingled quality of partial clinical notes leads to inadequate representations of to-be-predicted patients. Thirdly, typical attention mechanisms mainly focus on aggregating information from similar patients, ignoring important auxiliary information from others. To tackle these challenges, we propose a novel visit sequences-clinical notes joint learning approach, dubbed VecoCare. It performs a Gromov-Wasserstein Distance (GWD)-based contrastive learning task and an adaptive masked language model task in a sequential pre-training manner to reduce heterogeneous semantic biases. After pre-training, VecoCare further aggregates information from both similar and dissimilar patients through a dual-channel retrieval mechanism. We conduct diagnosis prediction experiments on two real-world datasets, which indicates that VecoCare outperforms state-of-the-art approaches. Moreover, the findings discovered by VecoCare are consistent with the medical researches.

AAAI Conference 2022 Short Paper

Large-Scale IP Usage Identification via Deep Ensemble Learning (Student Abstract)

  • Zhiyuan Wang
  • Fan Zhou
  • Kunpeng Zhang
  • Yong Wang

Understanding users’ behavior via IP addresses is essential towards numerous practical IP-based applications such as online content delivery, fraud prevention, and many others. Among which profiling IP address has been extensively studied, such as IP geolocation and anomaly detection. However, less is known about the scenario of an IP address, e. g. , dedicated enterprise network or home broadband. In this work, we initiate the first attempt to address a large-scale IP scenario prediction problem. Specifically, we collect IP scenario data from four regions and propose a novel deep ensemble learning-based model to learn IP assignment rules and complex feature interactions. Extensive experiments support that our method can make accurate IP scenario identification and generalize from data in one region to another.

NeurIPS Conference 2022 Conference Paper

Learning Latent Seasonal-Trend Representations for Time Series Forecasting

  • Zhiyuan Wang
  • Xovee Xu
  • Weifeng Zhang
  • Goce Trajcevski
  • Ting Zhong
  • Fan Zhou

Forecasting complex time series is ubiquitous and vital in a range of applications but challenging. Recent advances endeavor to achieve progress by incorporating various deep learning techniques (e. g. , RNN and Transformer) into sequential models. However, clear patterns are still hard to extract since time series are often composed of several intricately entangled components. Motivated by the success of disentangled variational autoencoder in computer vision and classical time series decomposition, we plan to infer a couple of representations that depict seasonal and trend components of time series. To achieve this goal, we propose LaST, which, based on variational inference, aims to disentangle the seasonal-trend representations in the latent space. Furthermore, LaST supervises and disassociates representations from the perspectives of themselves and input reconstruction, and introduces a series of auxiliary objectives. Extensive experiments prove that LaST achieves state-of-the-art performance on time series forecasting task against the most advanced representation learning and end-to-end forecasting models. For reproducibility, our implementation is publicly available on Github.

AAAI Conference 2022 Conference Paper

PrEF: Probabilistic Electricity Forecasting via Copula-Augmented State Space Model

  • Zhiyuan Wang
  • Xovee Xu
  • Goce Trajcevski
  • Kunpeng Zhang
  • Ting Zhong
  • Fan Zhou

Electricity forecasting has important implications for the key decisions in modern electricity systems, ranging from power generation, transmission, distribution and so on. In the literature, traditional statistic approaches, machine-learning methods and deep learning (e. g. , recurrent neural network) based models are utilized to model the trends and patterns in electricity time-series data. However, they are restricted either by their deterministic forms or by independence in probabilistic assumptions – thereby neglecting the uncertainty or significant correlations between distributions of electricity data. Ignoring these, in turn, may yield error accumulation, especially when relying on historical data and aiming at multi-step prediction. To overcome these, we propose a novel method named Probabilistic Electricity Forecasting (PrEF) by proposing a non-linear neural state space model (SSM) and incorporating copula-augmented mechanism into that, which can learn uncertainty-dependencies knowledge and understand interactive relationships between various factors from large-scale electricity time-series data. Our method distinguishes itself from existing models by its traceable inference procedure and its capability of providing high-quality probabilistic distribution predictions. Extensive experiments on two real-world electricity datasets demonstrate that our method consistently outperforms the alternatives.

LORI Conference 2009 Conference Paper

Existence of Satisfied Alternative and the Occurring of Morph-Dictator

  • Zhiyuan Wang

Abstract In social life of human being, individuals or collective group will be always confronted with choices. The occurring of choice implies that rational action agent (individual, collective group or social group in wide sense) must make a satisfying decision based on the alternatives set whose cardinal number is at least 2. Choice depends on preferences (Fishburn (1979)), so the nature of choice can be deemed to preference whether for individual or collective group. Preference is the ordering of alternatives given by rational agent according to his own will based on the sensibility and proneness. Preference can be crisp and fuzzy also.

v2026.09.13