Arrow Research search

Author name cluster

Ye Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

45 papers
2 author rows

Possible papers

45

EAAI Journal 2026 Journal Article

Artificial neural network-based pump operation policy models for energy cost minimisation

  • Lang Zheng
  • Wenyan Wu
  • Angus R. Simpson
  • Ye Wang

Efficient pump operation in water distribution systems (WDSs) is crucial for meeting end-users’ demands while minimising energy costs. Artificial neural network (ANN)-based operation policy models have been developed to address the limitations of existing methods in adjusting pump operations under dynamic conditions, where water demand and electricity tariffs vary. However, developing a suitable ANN-based policy model for a specific WDS is challenging due to the distinctly different learning patterns of ANN architectures and different operational characteristics of fixed and variable speed pumps, which collectively influence model performance. In this study, a modelling framework is proposed that couples a physics-based hydraulic model with an ANN-based operation policy model to inform decision-making under dynamic conditions. Two different ANNs, namely the radial basis function (RBF) and the multilayer perceptron (MLP), have been developed as policy models, and their strengths and limitations were evaluated through two WDSs - one with fixed speed pumps and another with a variable speed pump. The ANN-based policy models have been evaluated for their ability to minimise pumping energy costs and detailed operational behaviour. For fixed speed pumping system, the RBF-based policy model outperforms the MLP-based policy model in minimising energy costs and exhibits more advantageous operational behaviour under conditions driven by peak daily demand. For variable speed pumping system, both ANN-based policy models perform reasonably well in minimising energy costs. The RBF-based policy model remains more advantageous under conditions with peak daily demand, while the MLP-based policy model can be more advantageous under conditions with high electricity tariffs.

AAAI Conference 2026 Conference Paper

Chain-of-Thought Driven Adversarial Scenario Extrapolation for Robust Language Models

  • Md Rafi Ur Rashid
  • Vishnu Asutosh Dasu
  • Ye Wang
  • Gang Tan
  • Shagufta Mehnaz

Large Language Models (LLMs) exhibit impressive capabilities, but remain susceptible to a growing spectrum of safety risks, including jailbreaks, toxic content, hallucinations, and bias. Existing defenses often address only a single threat type or resort to rigid outright rejection, sacrificing user experience and failing to generalize across diverse and novel attacks. This paper introduces Adversarial Scenario Extrapolation (ASE), a novel inference-time computation framework that leverages Chain-of-Thought (CoT) reasoning to simultaneously enhance LLM robustness and seamlessness. ASE guides the LLM through a self-generative process of contemplating potential adversarial scenarios and formulating defensive strategies before generating a response to the user query. Comprehensive evaluation on four adversarial benchmarks with four latest LLMs shows that ASE achieves near-zero jailbreak attack success rates and minimal toxicity, while slashing outright rejections to <4%. ASE outperforms six state-of-the-art defenses in robustness-seamlessness trade-offs, with 92–99% accuracy on adversarial Q&A and 4–10× lower bias scores. By transforming adversarial perception into an intrinsic cognitive process, ASE sets a new paradigm for secure and natural human-AI interaction.

AAAI Conference 2026 Conference Paper

DeepPhy: Benchmarking Agentic VLMs on Physical Reasoning

  • Xinrun Xu
  • Pi Bu
  • Ye Wang
  • Börje F. Karlsson
  • Ziming Wang
  • Tengtao Song
  • Qi Zhu
  • Jun Song

Although Vision Language Models (VLMs) exhibit strong perceptual abilities and impressive visual reasoning, they struggle with attention to detail and precise action planning in complex, dynamic environments, leading to subpar performance. Real-world tasks typically require complex interactions, advanced spatial reasoning, long-term planning, and continuous strategy refinement, usually necessitating understanding the physics rules of the target scenario. However, evaluating these capabilities in real-world scenarios is often prohibitively expensive. To bridge this gap, we introduce DeepPHY, a novel benchmark framework designed to systematically evaluate VLMs' understanding and reasoning about fundamental physical principles through a series of challenging simulated environments. DeepPHY integrates multiple physical reasoning environments of varying difficulty levels and incorporates fine-grained evaluation metrics. Our evaluation finds that even state-of-the-art VLMs struggle to translate descriptive physical knowledge into precise, predictive control.

AAAI Conference 2026 Conference Paper

LatentLLM: Activation-Aware Transform to Multi-Head Latent Attention

  • Toshiaki Koike-Akino
  • Xiangyu Chen
  • Jing Liu
  • Ye Wang
  • Pu (Perry) Wang
  • Matthew Brand

Modern foundation models such as large language models (LLMs) require a massive amount of computational and memory resources. We propose a new framework to convert such LLMs into a reduced-dimension latent structure. Our method extends a local activation-aware tensor decomposition to a global attention-aware joint tensor decomposition. Our framework can significantly improve the model accuracy over the existing model compression methods when reducing the latent dimension to realize computationally/memory-efficient LLMs. We show the benefit on several benchmark including multi-modal reasoning tasks.

JBHI Journal 2026 Journal Article

Reliable Multimodal Cancer Survival Prediction With Confidence-Aware Risk Modeling

  • Xuping Xie
  • Ye Wang
  • Ziqi Zhao
  • Qixing Yang
  • Lan Huang
  • Fengfeng Zhou
  • Yan Wang

Multimodal survival methods that integrate histology whole-slide images and transcriptomic profiles hold significant promise for understanding patient prognostication and guiding personalized treatment strategies. However, existing approaches primarily focus on improving predictive performance through multimodal information fusion, often neglecting the reliability estimation of the prediction results and the inherent alignment noise across modalities. Thus, we propose ReCaSP, a novel and reliable cancer survival prediction framework that effectively integrates histology and transcriptomics data via multimodal alignment and fusion, providing the auxiliary confidence levels for survival predictions through a confidence-aware risk modeling mechanism. Specifically, our approach incorporates a fine-grained risk classifier that models risk labels jointly over both multiple time intervals and censorship status, utilizing evidential deep learning to yield fine-grained risk predictions accompanied by confidence scores. Additionally, to mitigate the inherent noise in multimodal data alignment, we introduce a cross-attention alignment module that effectively aligns histology data with transcriptomics data prior to multimodal fusion, thereby facilitating cross-modal interaction learning. Extensive experiments on five datasets demonstrate that ReCaSP significantly outperforms state-of-the-art methods, achieving a 4. 58% improvement in the overall C-Index.

AAAI Conference 2026 Conference Paper

Schema-Guided Event Reasoning: A Plug-and-Play Event Reasoning Framework Based on Large Language Models

  • Yuying Liu
  • Xuechen Zhao
  • Yanyi Huang
  • Ye Wang
  • Xin Song
  • Yue Zhang
  • Haiyan Liu
  • Bin Zhou

Recent advancements in Large Language Models have increasingly demonstrated their potential for event reasoning. However, LLMs still struggle with this task due to inadequate modeling of event structures. Although introducing schema knowledge has been shown to improve event reasoning performance, existing methods rely on predefined schema library, compromising their scalability and lightweight deployment. To address these challenges, we propose SGER, a plug-and-play Schema-Guided Event Reasoning framework. In the schema extraction stage, the model maps event descriptions with diverse surface forms to potential semantic structure representations, achieving an abstract transformation from instances to schemas. The schema prediction stage captures the potential associations between historical event schemas to make forward-looking inferences about possible future event schemas. In the event reasoning stage, we integrate historical events and predicted schemas into prompts to guide LLMs in generating specific, contextually consistent predicted events. Experimental evaluations demonstrate that our framework significantly improves event reasoning performance of LLMs.

JBHI Journal 2025 Journal Article

Cross-Interaction of Chinese Characters Structures and Boundary Features for Improving Clinical Named Entity Recognition

  • Ye Wang
  • Qi Wei
  • Hong Yu
  • Guoyin Wang
  • Chunmeng Shi
  • Dajiang Lei

In the natural language processing task of clinical named entity recognition (CNER), accurately identifying the boundaries and categories of medical entities is crucial. However, traditional methods struggle to recognize a large number of clinical terms and symbols that have never been encountered before, ultimately limiting the performance of CNER. Besides, there exist some easy-to-confuse Chinese clinical entities that are semantically similar but belong to quite different categories, such as “ 肺结节 ” (pulmonary nodules, a symptom entity) and “ 肺结核 ” (pulmonary tuberculosis, a disease entity), which can lead to entity misidentification. To address these problems, we propose a novel NER model called Cross-Interaction of Chinese characters structures and Boundary Features (CCS). The proposed model leverages Chinese character structural features and boundary information to comprehensively and accurately identify confusing entities. We further design a Cross-Attention mechanism to capture dependency relationships between different entities and radicals of characters, enhancing the model's semantic understanding of specialized terms and symbols, as well as improving its ability to recognize boundaries. Our experimental results show that our proposed model outperforms other state-of-the-art models on various public medical datasets, achieving significant improvements on the CCKS2020, CMeEE, CMI, and IMCS datasets, respectively.

AAAI Conference 2025 Conference Paper

Forget to Flourish: Leveraging Machine-Unlearning on Pretrained Language Models for Privacy Leakage

  • Md Rafi Ur Rashid
  • Jing Liu
  • Toshiaki Koike-Akino
  • Ye Wang
  • Shagufta Mehnaz

Fine-tuning large language models on private data for downstream applications poses significant privacy risks in potentially exposing sensitive information. Several popular community platforms now offer convenient distribution of a large variety of pre-trained models, allowing anyone to publish without rigorous verification. This scenario creates a privacy threat, as pre-trained models can be intentionally crafted to compromise the privacy of fine-tuning datasets. In this study, we introduce a novel poisoning technique that uses model-unlearning as an attack tool. This approach manipulates a pre-trained language model to increase the leakage of private data during the fine-tuning process. Our method enhances both membership inference and data extraction attacks while preserving model utility. Experimental results across different models, datasets, and fine-tuning setups demonstrate that our attacks significantly surpass baseline performance. This work serves as a cautionary note for users who download pretrained models from unverified sources, highlighting the potential risks involved.

NeurIPS Conference 2025 Conference Paper

FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction

  • Jiang Lin
  • Xinyu Chen
  • Song Wu
  • Zhiqiu Zhang
  • Jizhi Zhang
  • Ye Wang
  • Qiang Tang
  • Qian Wang

Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference cost due to dual-path denoising. We present \textbf{FreeControl}, a training-free framework for semantic structural control in diffusion models. Unlike prior methods that extract attention across multiple timesteps, FreeControl performs \textit{one-step attention extraction} from a single, optimally chosen timestep and reuses it throughout denoising. This enables efficient structural guidance without inversion or retraining. To further improve quality and stability, we introduce \textit{Latent-Condition Decoupling (LCD)}: a principled separation of the timestep condition and the noised latent used in attention extraction. LCD provides finer control over attention quality and eliminates structural artifacts. FreeControl also supports compositional control via reference images assembled from multiple sources, enabling intuitive scene layout design and stronger prompt alignment. FreeControl introduces a new paradigm for test-time control—enabling structurally and semantically aligned, visually coherent generation directly from raw images, with the flexibility for intuitive compositional design and compatibility with modern diffusion models at ~5\% additional cost.

NeurIPS Conference 2025 Conference Paper

Iterative Foundation Model Fine-Tuning on Multiple Rewards

  • Pouya M. Ghari
  • simone sciabola
  • Ye Wang

Fine-tuning foundation models has emerged as a powerful approach for generating objects with specific desired properties. Reinforcement learning (RL) provides an effective framework for this purpose, enabling models to generate outputs that maximize a given reward function. However, in many applications such as text generation and drug discovery, it can be suboptimal to optimize using a single reward signal, as multiple evaluation criteria are often necessary. This paper proposes a novel reinforcement learning-based method for fine-tuning foundation models using multiple reward signals. By employing an iterative fine-tuning strategy across these rewards, our approach generalizes state-of-the-art RL-based methods. We further provide a theoretical analysis that offers insights into the performance of multi-reward RL fine-tuning. Experimental results across diverse domains including text, biological sequence, and small molecule generation, demonstrate the effectiveness of the proposed algorithm compared to state-of-the-art baselines.

EAAI Journal 2025 Journal Article

Koopman-based predictive tracking control

  • Ye Wang
  • Yujia Yang
  • Ye Pu
  • Chris Manzie

Constraint handling during tracking operations is at the core of many real-world control implementations and is well understood when dynamic models of the underlying system exist, yet becomes more challenging when data-driven models are used to describe the nonlinear system at hand. We seek to combine the nonlinear modeling capabilities of a wide class of neural networks with the constraint-handling guarantees of model predictive control in a rigorous and online computationally tractable framework. The class of networks considered can be captured using Koopman operators, and are integrated into a Koopman-based predictive tracking control (KPTC) for nonlinear systems to track piecewise constant references. The effect of model mismatch between original nonlinear dynamics and its trained Koopman linear model is handled by using a constraint-tightening approach in the proposed KPTC controller. By choosing two Lyapunov functions, we prove that the solution is recursively feasible and input-to-state stable to a neighborhood of both online and offline optimal reachable steady outputs in the presence of bounded modeling errors under certain assumptions. The proposed approach has the advantage relative to existing model-based tracking approaches of enabling data-driven models to be utilized with explicit guarantees, while using efficient quadratic program solvers in online implementations. We demonstrate the proposed approach initially in simulations, and then experimentally to the problem of reference tracking by an autonomous ground vehicle.

IJCAI Conference 2025 Conference Paper

LivePoem: Improving the Learning Experience of Classical Chinese Poetry with AI-Generated Musical Storyboards

  • Qihao Liang
  • Xichu Ma
  • Torin Hopkins
  • Ye Wang

Textbook reading has long dominated classical poetry education in Chinese-speaking communities. However, research has shown that extensive text-based learning can lead to learner disengagement and a less pleasant experience. This paper aims to improve the experience of classical Chinese poetry learning by introducing LivePoem—a system that generates musical storyboards (storyboards with background music) as audiovisual aids to support poetry comprehension. We employ a pre-trained diffusion model for storyboard generation and train a prosody-based poem-to-melody generator using a Transformer model, both validated by standard objective metrics to ensure generation quality. Through a within-subjects study involving 25 non-native Chinese learners, we compared learning outcomes from textbook reading and musical storyboard viewing through standardised reading comprehension tests. Additionally, the learning experience was assessed by the Self-Assessment Manikin (SAM) and an inductive thematic analysis of learners' open-ended feedback. Experimental results show that musical storyboards retained the learning outcomes of textbooks, while more effectively engaging learners and providing a more pleasant learning experience.

AAAI Conference 2025 Conference Paper

LLM-DR: A Novel LLM-Aided Diffusion Model for Rule Generation on Temporal Knowledge Graphs

  • Kai Chen
  • Xin Song
  • Ye Wang
  • Liqun Gao
  • Aiping Li
  • Xiaojuan Zhao
  • Bin Zhou
  • Yalong Xie

Among various temporal knowledge graph (TKG) extrapolation methods, rule-based approaches stand out for their explicit rules and transparent reasoning paths. However, the vast search space for rule extraction poses a challenge in identifying high-quality logic rules. To navigate this challenge, we explore the use of generation models to generate new rules, thereby enriching our rule base and enhancing our reasoning capabilities. In this paper, we introduce LLM-DR, an innovative rule-based method for TKG extrapolation, which harnesses diffusion models to generate rules that are consistent with the distribution of the source data, while also amalgamating the rich semantic insights of Large Language Models (LLMs). Specifically, our LLM-DR generates semantically relevant and high-quality rules, employing conditional diffusion models in a classifier-free guidance fashion and refining them with LLM-based constraints. To assess rule efficacy, we meticulously design a coarse-to-fine evaluation strategy that initiates with coarse-grained filtering to eliminate less plausible rules and proceeds with fine-grained scoring to quantify the reliability of the retained. Extensive experiments demonstrate the promising capacity of our LLM-DR.

TMLR Journal 2025 Journal Article

On Memorization in Diffusion Models

  • Xiangming Gu
  • Chao Du
  • Tianyu Pang
  • Chongxuan Li
  • Min Lin
  • Ye Wang

Due to their capacity to generate novel and high-quality samples, diffusion models have attracted significant research interest in recent years. Notably, the typical training objective of diffusion models, i.e., denoising score matching, has a closed-form optimal solution that can only generate training-data replicating samples. This indicates that a memorization behavior is theoretically expected, which contradicts the common generalization ability of state-of-the-art diffusion models, and thus calls for a deeper understanding. Looking into this, we first observe that memorization behaviors tend to occur on smaller-sized datasets, which motivates our definition of effective model memorization (EMM), a metric measuring the maximum size of training data at which a model approximates its theoretical optimum. Then, we quantify the impact of the influential factors on these memorization behaviors in terms of EMM, focusing primarily on data distribution, model configuration, and training procedure. Besides comprehensive empirical results identifying the influential factors, we surprisingly find that conditioning training data on uninformative random labels can significantly trigger the memorization in diffusion models. Our study holds practical significance for diffusion model users and offers clues to theoretical research in deep generative models.

EAAI Journal 2025 Journal Article

Physics-informed neural networks in heat transfer-dominated multiphysics systems: A comprehensive review

  • Zhuang Zhao
  • Ye Wang
  • Weijian Zhang
  • Zhenggang Ba
  • Lin Sun

This article presents a comprehensive review of the role of Physics-Informed Neural Networks (PINNs) in engineering applications dominated by heat transfer, synthesizing bibliometric analysis with practical engineering perspectives. PINNs, which integrate separable neural networks and parameterized multiphysics models, offer a robust framework for solving both forward and inverse heat transfer problems. These models demonstrate particular efficacy in complex multiphysics environments, such as combustion modeling and geothermal forecasting, circumventing the need for extensive mesh generation required by traditional methods. Key applications explored include thermal management in electronics, monitoring of cogeneration systems, geothermal heat production, lifespan management of batteries and transformers, food drying processes, material thermal property identification and thermal oversight in additive manufacturing. Recent advancements, including the incorporation of numerical methods, adaptive activation functions, and transfer learning, have significantly improved the practicality of PINNs. However, their training cost remains a notable drawback when compared to the efficiency of conventional numerical simulation methods for regular, large-scale simulations. Emerging trends in deep learning and probabilistic approaches suggest promising future directions, such as potential integration with Graph Neural Networks and meta-learning to enhance modeling capabilities. While traditional numerical simulation methods remain preferable for short-term, well-defined scenarios, PINNs or hybrid approaches are recommended for long-term, innovative solutions requiring adaptability. This review provides a structured framework for engineers and researchers to apply PINNs in heat transfer-dominated applications and delineates future pathways for advancing data-driven, physics-constrained modeling strategies.

ICRA Conference 2025 Conference Paper

PlaneHEC: Efficient Hand-Eye Calibration for Multi-View Robotic Arm via Any Point Cloud Plane Detection

  • Ye Wang
  • Haodong Jing
  • Yang Liao
  • Yongqiang Ma
  • Nanning Zheng 0001

Hand-eye calibration is an important task in vision-guided robotic systems and is crucial for determining the transformation matrix between the camera coordinate system and the robot end-effector. Existing methods, for multi-view robotic systems, usually rely on accurate geometric models or manual assistance, generalize poorly, and can be very complicated and inefficient. Therefore, in this study, we propose PlaneHEC, a generalized hand-eye calibration method that does not require complex models and can be accomplished using only depth cameras, which achieves the optimal and fastest calibration results using arbitrary planar surfaces like walls and tables. PlaneHEC introduces hand-eye calibration equations based on planar constraints, which makes it strongly interpretable and generalizable. PlaneHEC also uses a comprehensive solution that starts with a closed-form solution and improves it with iterative optimization, which greatly improves accuracy. We comprehensively evaluated the performance of PlaneHEC in both simulated and real-world environments and compared the results with other point-cloud-based calibration methods, proving its superiority. Our approach achieves universal and fast calibration with an innovative design of computational models, providing a strong contribution to the development of multi-agent systems and embodied intelligence.

ICML Conference 2025 Conference Paper

Scaling Large Motion Models with Million-Level Human Motions

  • Ye Wang
  • Sipeng Zheng
  • Bin Cao
  • Qianshan Wei
  • Weishuai Zeng
  • Qin Jin
  • Zongqing Lu 0002

Inspired by the recent success of LLMs, the field of human motion understanding has increasingly shifted toward developing large motion models. Despite some progress, current efforts remain far from achieving truly generalist models, primarily due to the lack of massive high-quality data. To address this gap, we present MotionLib, the first million-level dataset for motion generation, which is at least 15$\times$ larger than existing counterparts and enriched with hierarchical text descriptions. Using MotionLib, we train a large motion model named Being-M0, demonstrating robust performance across a wide range of human activities, including unseen ones. Through systematic investigation, for the first time, we highlight the importance of scaling both data and model size for advancing motion generation, along with key insights to achieve this goal. To better integrate the motion modality, we propose Motionbook, an innovative motion encoding approach including (1) a compact yet lossless feature to represent motions; (2) a novel 2D lookup-free motion tokenizer that preserves fine-grained motion details while expanding codebook capacity, significantly enhancing the representational power of motion tokens. We believe this work lays the groundwork for developing more versatile and powerful motion generation models in the future. For further details, visit https: //beingbeyond. github. io/Being-M0/.

AAAI Conference 2025 Conference Paper

SigStyle: Signature Style Transfer via Personalized Text-to-Image Models

  • Ye Wang
  • Tongyuan Bai
  • Xuping Xie
  • Zili Yi
  • Yilin Wang
  • Rui Ma

Style transfer enables the seamless integration of artistic styles from a style image into a content image, resulting in visually striking and aesthetically enriched outputs. Despite numerous advances in this field, existing methods did not explicitly focus on the signature style, which represents the distinct and recognizable visual traits of the image such as geometric and structural patterns, color palettes and brush strokes etc. In this paper, we introduce SigStyle, a framework that leverages the semantic priors that embedded in a personalized text-to-image diffusion model to capture the signature style representation. This style capture process is powered by a hypernetwork that efficiently fine-tunes the diffusion model for any given single style image. Style transfer then is conceptualized as the reconstruction process of content image through learned style tokens from the personalized diffusion model. Additionally, to ensure the content consistency throughout the style transfer process, we introduce a time-aware attention swapping technique that incorporates content information from the original image into the early denoising steps of target image generation. Beyond enabling high-quality signature style transfer across a wide range of styles, SigStyle supports multiple interesting applications, such as local style transfer, texture transfer, style fusion and style-guided text-to-image generation. Quantitative and qualitative evaluations demonstrate our approach outperforms existing style transfer methods for recognizing and transferring the signature styles.

EAAI Journal 2025 Journal Article

Soft Prompt-tuning with Self-Resource Verbalizer for short text streams

  • Yi Zhu
  • Ye Wang
  • Yun Li
  • Jipeng Qiang
  • Yunhao Yuan

Short text streams such as real-time news and search snippets have attained vast amounts of attention and research in recent decades, the characteristics of high generation velocity, feature sparsity, and high ambiguity accentuate both the importance and challenges to language models. However, most of the existing short text stream classification methods can neither automatically select relevant knowledge components for arbitrary samples, nor expand knowledge internally instead of rely on external open knowledge base to address the inherent limitations of short text stream. In this paper, we propose a Soft Prompt-tuning with Self-Resource Verbalizer (SPSV for short) for short text stream classification, the soft prompt with self-resource knowledgeable expansion is conducted for updating label words space to address evolved semantic topics in the data streams. Specifically, the automatic constructed prompt is first generated to instruct the model prediction, which is optimized to address the problem of high velocity and topic drift in short text streams. Then, in each chunk, the projection between category names and label words space, i. e. verbalizer, is updated, which is constructed by internal knowledge expansion from the short text itself. Through comprehensive experiments on four well-known benchmark datasets, we validate the superb performance of our method compared to other short text stream classification and fine-tuning PLMs methods, which achieves up to more than 90% classification accuracy with the counts of data chunk increased.

NeurIPS Conference 2025 Conference Paper

Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding

  • Ye Wang
  • Ziheng Wang
  • Boshen Xu
  • Yang Du
  • Kejun Lin
  • Zihan Xiao
  • Zihao Yue
  • Jianzhong Ju

Temporal Video Grounding (TVG), the task of locating specific video segments based on language queries, is a core challenge in long-form video understanding. While recent Large Vision-Language Models (LVLMs) have shown early promise in tackling TVG through supervised fine-tuning (SFT), their ability to generalize remains limited. To address this, we propose a novel post-training framework that enhances the generalization capabilities of LVLMs via reinforcement learning (RL). Specifically, our contributions span three key directions: (1) Time-R1: we introduce a reasoning-guided post-training framework via RL with verifiable reward to enhance capabilities of LVLMs on the TVG task. (2) TimeRFT: we explore post-training strategies on our curated RL-friendly dataset, which trains the model to progressively comprehend more difficult samples, leading to better generalization. (3) TVGBench: we carefully construct a small but comprehensive and balanced benchmark suitable for LVLM evaluation, which is sourced from available public benchmarks. Extensive experiments demonstrate that Time-R1 achieves state-of-the-art performance across multiple downstream datasets using significantly less training data than prior LVLM approaches, while improving its general video understanding capabilities. Project Page: https: //xuboshen. github. io/Time-R1/.

NeurIPS Conference 2025 Conference Paper

Unifying Symbolic Music Arrangement: Track-Aware Reconstruction and Structured Tokenization

  • Longshen Ou
  • Jingwei Zhao
  • Ziyu Wang
  • Gus Xia
  • Qihao Liang
  • Torin Hopkins
  • Ye Wang

We present a unified framework for automatic multitrack music arrangement that enables a single pre-trained symbolic music model to handle diverse arrangement scenarios, including reinterpretation, simplification, and additive generation. At its core is a segment-level reconstruction objective operating on token-level disentangled content and style, allowing for flexible any-to-any instrumentation transformations at inference time. To support track-wise modeling, we introduce REMI-z, a structured tokenization scheme for multitrack symbolic music that enhances modeling efficiency and effectiveness for both arrangement tasks and unconditional generation. Our method outperforms task-specific state-of-the-art models on representative tasks in different arrangement scenarios---band arrangement, piano reduction, and drum arrangement, in both objective metrics and perceptual evaluations. Taken together, our framework demonstrates strong generality and suggests broader applicability in symbolic music-to-music transformation.

AAAI Conference 2024 Conference Paper

A Positive-Unlabeled Metric Learning Framework for Document-Level Relation Extraction with Incomplete Labeling

  • Ye Wang
  • Huazheng Pan
  • Tao Zhang
  • Wen Wu
  • Wenxin Hu

The goal of document-level relation extraction (RE) is to identify relations between entities that span multiple sentences. Recently, incomplete labeling in document-level RE has received increasing attention, and some studies have used methods such as positive-unlabeled learning to tackle this issue, but there is still a lot of room for improvement. Motivated by this, we propose a positive-augmentation and positive-mixup positive-unlabeled metric learning framework (P3M). Specifically, we formulate document-level RE as a metric learning problem. We aim to pull the distance closer between entity pair embedding and their corresponding relation embedding, while pushing it farther away from the none-class relation embedding. Additionally, we adapt the positive-unlabeled learning to this loss objective. In order to improve the generalizability of the model, we use dropout to augment positive samples and propose a positive-none-class mixup method. Extensive experiments show that P3M improves the F1 score by approximately 4-10 points in document-level RE with incomplete labeling, and achieves state-of-the-art results in fully labeled scenarios. Furthermore, P3M has also demonstrated robustness to prior estimation bias in incomplete labeled scenarios.

JMLR Journal 2024 Journal Article

Efficient Active Manifold Identification via Accelerated Iteratively Reweighted Nuclear Norm Minimization

  • Hao Wang
  • Ye Wang
  • Xiangyu Yang

This paper considers the problem of minimizing the sum of a smooth function and the Schatten-$p$ norm of the matrix. Our contribution involves proposing accelerated iteratively reweighted nuclear norm methods designed to solve the nonconvex low-rank minimization problem. Two major novelties characterize our approach. First, the proposed method possesses an active manifold identification property, enabling the provable identification of the correct rank of the stationary point within a finite number of iterations. Second, we introduce an adaptive updating strategy for smoothing parameters. This strategy automatically fixes parameters associated with zero singular values as constants upon detecting the correct rank while quickly driving the remaining parameters to zero. This adaptive behavior transforms the algorithm into one that effectively solves smooth problems after a few iterations, setting our work apart from existing iteratively reweighted methods for low-rank optimization. We prove the global convergence of the proposed algorithm, guaranteeing that every limit point of the iterates is a critical point. Furthermore, a local convergence rate analysis is provided under the Kurdyka-Łojasiewicz property. We conduct numerical experiments using both synthetic and real data to showcase our algorithm's efficiency and superiority over existing methods. [abs] [ pdf ][ bib ] &copy JMLR 2024. ( edit, beta )

IJCAI Conference 2024 Conference Paper

End-to-End Real-World Polyphonic Piano Audio-to-Score Transcription with Hierarchical Decoding

  • Wei Zeng
  • Xian He
  • Ye Wang

Piano audio-to-score transcription (A2S) is an important yet underexplored task with extensive applications for music composition, practice, and analysis. However, existing end-to-end piano A2S systems faced difficulties in retrieving bar-level information such as key and time signatures, and have been trained and evaluated with only synthetic data. To address these limitations, we propose a sequence-to-sequence (Seq2Seq) model with a hierarchical decoder that aligns with the hierarchical structure of musical scores, enabling the transcription of score information at both the bar and note levels by multi-task learning. To bridge the gap between synthetic data and recordings of human performance, we propose a two-stage training scheme, which involves pre-training the model using an expressive performance rendering (EPR) system on synthetic audio, followed by fine-tuning the model using recordings of human performance. To preserve the voicing structure for score reconstruction, we propose a pre-processing method for **Kern scores in scenarios with an unconstrained number of voices. Experimental results support the effectiveness of our proposed approaches, in terms of both transcription performance on synthetic audio data in comparison to the current state-of-the-art, and the first experiment on human recordings.

IJCAI Conference 2024 Conference Paper

Joint Domain Adaptive Graph Convolutional Network

  • Niya Yang
  • Ye Wang
  • Zhizhi Yu
  • Dongxiao He
  • Xin Huang
  • Di Jin

In the realm of cross-network tasks, graph domain adaptation is an effective tool due to its ability to transfer abundant labels from nodes in the source domain to those in the target domain. Existing adversarial domain adaptation methods mainly focus on domain-wise alignment. These approaches, while effective in mitigating the marginal distribution shift between the two domains, often ignore the integral aspect of structural alignment, potentially leading to negative transfer. To address this issue, we propose a joint adversarial domain adaptive graph convolutional network (JDA-GCN) that is uniquely augmented with structural graph alignment, so as to enhance the efficacy of knowledge transfer. Specifically, we construct a structural graph to delineate the interconnections among nodes within identical categories across the source and target domains. To further refine node representation, we integrate the local consistency matrix with the global consistency matrix, thereby leveraging the learning of the sub-structure similarity of nodes to enable more robust and effective representation of nodes. Empirical evaluation on diverse real-world datasets substantiates the superiority of our proposed method, marking a significant advancement over existing state-of-the-art graph domain adaptation algorithms.

AAAI Conference 2024 Conference Paper

MedBench: A Large-Scale Chinese Benchmark for Evaluating Medical Large Language Models

  • Yan Cai
  • Linlin Wang
  • Ye Wang
  • Gerard de Melo
  • Ya Zhang
  • Yanfeng Wang
  • Liang He

The emergence of various medical large language models (LLMs) in the medical domain has highlighted the need for unified evaluation standards, as manual evaluation of LLMs proves to be time-consuming and labor-intensive. To address this issue, we introduce MedBench, a comprehensive benchmark for the Chinese medical domain, comprising 40,041 questions sourced from authentic examination exercises and medical reports of diverse branches of medicine. In particular, this benchmark is composed of four key components: the Chinese Medical Licensing Examination, the Resident Standardization Training Examination, the Doctor In-Charge Qualification Examination, and real-world clinic cases encompassing examinations, diagnoses, and treatments. MedBench replicates the educational progression and clinical practice experiences of doctors in Mainland China, thereby establish- ing itself as a credible benchmark for assessing the mastery of knowledge and reasoning abilities in medical language learning models. We perform extensive experiments and conduct an in-depth analysis from diverse perspectives, which culminate in the following findings: (1) Chinese medical LLMs underperform on this benchmark, highlighting the need for significant advances in clinical knowledge and diagnostic precision. (2) Several general-domain LLMs surprisingly possess considerable medical knowledge. These findings elucidate both the capabilities and limitations of LLMs within the context of MedBench, with the ultimate goal of aiding the medical research community.

EAAI Journal 2024 Journal Article

Query-induced multi-task decomposition and enhanced learning for aspect-based sentiment quadruple prediction

  • Hua Zhang
  • Xiawen Song
  • Xiaohui Jia
  • Cheng Yang
  • Zeqi Chen
  • Bi Chen
  • Bo Jiang
  • Ye Wang

A complete sentiment analysis of product and service reviews has attracted growing concerns from merchants to enhance personalized marketing activities. Aspect sentiment quadruple prediction (ASQP) is a demanding and challenging task with the objective to predict four sentiment elements from given reviews. Existing methods for ASQP face certain issues, with pipeline-based non-generative approaches prone to error propagation and generative models at the potential risk of producing unexpected outputs or longer inference times. To avoid these shortcomings, we develop a novel end-to-end non-generative model for ASQP involving multi-task decomposition within machine reading comprehension (MRC) framework. Specifically, the ASQP task is decomposed into six query-induced subtasks by introducing task-specific question templates. The proposed model, named MRC-CLRI, is trained with multi-task joint learning. It also incorporates contrastive learning for category identification and sentiment classification to enhance the correlation of the six subtasks. To further promote the quadruple prediction, we present a refined inference algorithm in a bidirectional multi-turn inference procedure to effectively match aspect and opinion terms and optimize two inference hyperparameters: distance threshold and probability threshold. Experimental results exhibit superior performance compared to existing two non-generative and seven generative baselines. Our proposed MRC-CLRI, as a novel non-generative model, outperforms the best existing generative method by an average F1 score improvement of 1. 69% and the best previous non-generative method by an average F1 score improvement of 15. 77% across four review datasets. Ablation experiments further validate the efficacy of the designed contrastive learning and the refined inference algorithm.

EAAI Journal 2024 Journal Article

SOFT: Self-supervised sparse Optical Flow Transformer for video stabilization via quaternion

  • Naiyao Wang
  • Changdong Zhou
  • Rongfeng Zhu
  • Bo Zhang
  • Ye Wang
  • Hongbo Liu

Video stabilization is crucial for video representation learning, which suffers from the challenges such as the perception of unstable vision, the stripping and cognition of target motion features in complex scenes, the correction of the jittery camera systems trails. In this paper, we propose a Self-supervised sparse Optical Flow Transformer (SOFT) model, consisting of a self-supervised contrastive learning transformer network, a sparse optical flow perception network and a multimodal cognitive fusion network. The SOFT model takes advantage of optical flow to estimate motion. The sparse optical flow perception network perceiving partially sparse optical flow containing motion features. This serves as the input to the self-supervised contrastive learning transformer network for generating sparse optical flow features, which are fed into the multimodal cognitive fusion network together with the real and virtual camera pose for video frame warping. Experimental comparisons with state-of-the-art models on 4 metrics demonstrate the effectiveness of the SOFT model. It achieves the best performance with an average Stability of 0. 869 and average Distortion of 0. 993 across 6 categories videos, which shows that the SOFT model can effectively perceive the motion in the video and smooth the jitter track of videos.

NeurIPS Conference 2024 Conference Paper

Structured Multi-Track Accompaniment Arrangement via Style Prior Modelling

  • Jingwei Zhao
  • Gus Xia
  • Ziyu Wang
  • Ye Wang

In the realm of music AI, arranging rich and structured multi-track accompaniments from a simple lead sheet presents significant challenges. Such challenges include maintaining track cohesion, ensuring long-term coherence, and optimizing computational efficiency. In this paper, we introduce a novel system that leverages prior modelling over disentangled style factors to address these challenges. Our method presents a two-stage process: initially, a piano arrangement is derived from the lead sheet by retrieving piano texture styles; subsequently, a multi-track orchestration is generated by infusing orchestral function styles into the piano arrangement. Our key design is the use of vector quantization and a unique multi-stream Transformer to model the long-term flow of the orchestration style, which enables flexible, controllable, and structured music generation. Experiments show that by factorizing the arrangement task into interpretable sub-stages, our approach enhances generative capacity while improving efficiency. Additionally, our system supports a variety of music genres and provides style control at different composition hierarchies. We further show that our system achieves superior coherence, structure, and overall arrangement quality compared to existing baselines.

IJCAI Conference 2024 Conference Paper

Temporal Knowledge Graph Extrapolation via Causal Subhistory Identification

  • Kai Chen
  • Ye Wang
  • Xin Song
  • Siwei Chen
  • Han Yu
  • Aiping Li

Temporal knowledge graph extrapolation has become a prominent area of study interest in recent years. Numerous methods for extrapolation have been put forth, mining query-relevant information from history to generate forecasts. However, existing approaches normally do not discriminate between causal and non-causal effects in reasoning; instead, they focus on analyzing the statistical correlation between the future events to be predicted and the historical data given, which may be deceptive and hinder the model's capacity to learn real causal information that actually affects the reasoning conclusions. To tackle it, we propose a novel approach called Causal Subhistory Identification (CSI), which focuses on extracting the causal subhistory for reasoning purposes from a large amount of historical data. CSI can improve the clarity and transparency of the reasoning process and more effectively convey the logic behind conclusions by giving priority to the causal subhistory and eliminating non-causal correlations. Extensive experiments demonstrate the remarkable potential of our CSI in the following aspects: superiority, improvement, explainability, and robustness.

IJCAI Conference 2024 Conference Paper

XAI-Lyricist: Improving the Singability of AI-Generated Lyrics with Prosody Explanations

  • Qihao Liang
  • Xichu Ma
  • Finale Doshi-Velez
  • Brian Lim
  • Ye Wang

Explaining the singability of lyrics is an important but missing ability of language models (LMs) in song lyrics generation. This ability allows songwriters to quickly assess if LM-generated lyrics can be sung harmoniously with melodies and helps singers align lyrics with melodies during practice. This paper presents XAI-Lyricist, leveraging musical prosody to guide LMs in generating singable lyrics and providing human-understandable singability explanations. We employ a Transformer model to generate lyrics under musical prosody constraints and provide demonstrations of the lyrics' prosody patterns as singability explanations. XAI-Lyricist is evaluated by computational metrics (perplexity, prosody-BLEU) and a human-grounded study (human ratings, average time and number of attempts for singing). Experimental results show that musical prosody can significantly improve the singability of LM-generated lyrics. A controlled study with 14 singers also confirms the usefulness of the provided explanations in helping them to interpret lyrical singability faster than reading plain text lyrics.

AAMAS Conference 2023 Conference Paper

Deep Learning-Powered Iterative Combinatorial Auctions with Active Learning

  • Benjamin Estermann
  • Stefan Kramer
  • Roger Wattenhofer
  • Ye Wang

Deep learning-powered iterative combinatorial auctions (DL-ICA) are auctions that utilize machine learning techniques. Unlike traditional auctions, bidders in DL-ICA do not need to report the valuations for all bundles upfront. Instead, they report their value for certain bundles iteratively, and the allocation of the items is determined by solving a winner determination problem. During this process, the bidder profiles are modeled with neural networks. However, DL-ICA may not always achieve the optimal winner allocation due to the relatively low number of reported bundles, resulting in reduced economic efficiency. This paper proposes an algorithm that uses active learning for initial sampling strategies to improve the resulting economic efficiency (social welfare). The proposed algorithm outperforms previous studies in real-world combinatorial auction models across various domains while using fewer samples on average.

AAAI Conference 2023 Conference Paper

FedNP: Towards Non-IID Federated Learning via Federated Neural Propagation

  • Xueyang Wu
  • Hengguan Huang
  • Youlong Ding
  • Hao Wang
  • Ye Wang
  • Qian Xu

Traditional federated learning (FL) algorithms, such as FedAvg, fail to handle non-i.i.d data because they learn a global model by simply averaging biased local models that are trained on non-i.i.d local data, therefore failing to model the global data distribution. In this paper, we present a novel Bayesian FL algorithm that successfully handles such a non-i.i.d FL setting by enhancing the local training task with an auxiliary task that explicitly estimates the global data distribution. One key challenge in estimating the global data distribution is that the data are partitioned in FL, and therefore the ground-truth global data distribution is inaccessible. To address this challenge, we propose an expectation-propagation-inspired probabilistic neural network, dubbed federated neural propagation (FedNP), which efficiently estimates the global data distribution given non-i.i.d data partitions. Our algorithm is sampling-free and end-to-end differentiable, can be applied with any conventional FL frameworks and learns richer global data representation. Experiments on both image classification tasks with synthetic non-i.i.d image data partitions and real-world non-i.i.d speech recognition tasks demonstrate that our framework effectively alleviates the performance deterioration caused by non-i.i.d data.

IJCAI Conference 2023 Conference Paper

Q&amp; A: Query-Based Representation Learning for Multi-Track Symbolic Music re-Arrangement

  • Jingwei Zhao
  • Gus Xia
  • Ye Wang

Music rearrangement is a common music practice of reconstructing and reconceptualizing a piece using new composition or instrumentation styles, which is also an important task of automatic music generation. Existing studies typically model the mapping from a source piece to a target piece via supervised learning. In this paper, we tackle rearrangement problems via self-supervised learning, in which the mapping styles can be regarded as conditions and controlled in a flexible way. Specifically, we are inspired by the representation disentanglement idea and propose Q&A, a query-based algorithm for multi-track music rearrangement under an encoder-decoder framework. Q&A learns both a content representation from the mixture and function (style) representations from each individual track, while the latter queries the former in order to rearrange a new piece. Our current model focuses on popular music and provides a controllable pathway to four scenarios: 1) re-instrumentation, 2) piano cover generation, 3) orchestration, and 4) voice separation. Experiments show that our query system achieves high-quality rearrangement results with delicate multi-track structures, significantly outperforming the baselines.

IJCAI Conference 2022 Conference Paper

Adversarial Bi-Regressor Network for Domain Adaptive Regression

  • Haifeng Xia
  • Pu Wang
  • Toshiaki Koike-Akino
  • Ye Wang
  • Philip Orlik
  • Zhengming Ding

Domain adaptation (DA) aims to transfer the knowledge of a well-labeled source domain to facilitate unlabeled target learning. When turning to specific tasks such as indoor (Wi-Fi) localization, it is essential to learn a cross-domain regressor to mitigate the domain shift. This paper proposes a novel method Adversarial Bi-Regressor Network (ABRNet) to seek more effective cross- domain regression model. Specifically, a discrepant bi-regressor architecture is developed to maximize the difference of bi-regressor to discover uncertain target instances far from the source distribution, and then an adversarial training mechanism is adopted between feature extractor and dual regressors to produce domain-invariant representations. To further bridge the large domain gap, a domain- specific augmentation module is designed to synthesize two source-similar and target-similar inter- mediate domains to gradually eliminate the original domain mismatch. The empirical studies on two cross-domain regressive benchmarks illustrate the power of our method on solving the domain adaptive regression (DAR) problem.

NeurIPS Conference 2022 Conference Paper

Extrapolative Continuous-time Bayesian Neural Network for Fast Training-free Test-time Adaptation

  • Hengguan Huang
  • Xiangming Gu
  • Hao Wang
  • Chang Xiao
  • Hongfu Liu
  • Ye Wang

Human intelligence has shown remarkably lower latency and higher precision than most AI systems when processing non-stationary streaming data in real-time. Numerous neuroscience studies suggest that such abilities may be driven by internal predictive modeling. In this paper, we explore the possibility of introducing such a mechanism in unsupervised domain adaptation (UDA) for handling non-stationary streaming data for real-time streaming applications. We propose to formulate internal predictive modeling as a continuous-time Bayesian filtering problem within a stochastic dynamical system context. Such a dynamical system describes the dynamics of model parameters of a UDA model evolving with non-stationary streaming data. Building on such a dynamical system, we then develop extrapolative continuous-time Bayesian neural networks (ECBNN), which generalize existing Bayesian neural networks to represent temporal dynamics and allow us to extrapolate the distribution of model parameters before observing the incoming data, therefore effectively reducing the latency. Remarkably, our empirical results show that ECBNN is capable of continuously generating better distributions of model parameters along the time axis given historical data only, thereby achieving (1) training-free test-time adaptation with low latency, (2) gradually improved alignment between the source and target features and (3) gradually improved model performance over time during the real-time testing stage.

AAAI Conference 2022 Conference Paper

MOST-GAN: 3D Morphable StyleGAN for Disentangled Face Image Manipulation

  • Safa C. Medin
  • Bernhard Egger
  • Anoop Cherian
  • Ye Wang
  • Joshua B. Tenenbaum
  • Xiaoming Liu
  • Tim K. Marks

Recent advances in generative adversarial networks (GANs) have led to remarkable achievements in face image synthesis. While methods that use style-based GANs can generate strikingly photorealistic face images, it is often difficult to control the characteristics of the generated faces in a meaningful and disentangled way. Prior approaches aim to achieve such semantic control and disentanglement within the latent space of a previously trained GAN. In contrast, we propose a framework that a priori models physical attributes of the face such as 3D shape, albedo, pose, and lighting explicitly, thus providing disentanglement by design. Our method, MOST-GAN, integrates the expressive power and photorealism of stylebased GANs with the physical disentanglement and flexibility of nonlinear 3D morphable models, which we couple with a state-of-the-art 2D hair manipulation network. MOST-GAN achieves photorealistic manipulation of portrait images with fully disentangled 3D control over their physical attributes, enabling extreme manipulation of lighting, facial expression, and pose variations up to full profile view.

TMLR Journal 2022 Journal Article

Unsupervised Mismatch Localization in Cross-Modal Sequential Data with Application to Mispronunciations Localization

  • Wei Wei
  • Hengguan Huang
  • Xiangming Gu
  • Hao Wang
  • Ye Wang

Content mismatch usually occurs when data from one modality is translated to another, e.g. language learners producing mispronunciations (errors in speech) when reading a sentence (target text) aloud. However, most existing alignment algorithms assume that the content involved in the two modalities is perfectly matched, thus leading to difficulty in locating such mismatch between speech and text. In this work, we develop an unsupervised learning algorithm that can infer the relationship between content-mismatched cross-modal sequential data, especially for speech-text sequences. More specifically, we propose a hierarchical Bayesian deep learning model, dubbed mismatch localization variational autoencoder (ML-VAE), which decomposes the generative process of the speech into hierarchically structured latent variables, indicating the relationship between the two modalities. Training such a model is very challenging due to the discrete latent variables with complex dependencies involved. To address this challenge, we propose a novel and effective training procedure that alternates between estimating the hard assignments of the discrete latent variables over a specifically designed mismatch localization finite-state acceptor (ML-FSA) and updating the parameters of neural networks. In this work, we focus on the mismatch localization problem for speech and text, and our experimental results show that ML-VAE successfully locates the mismatch between text and speech, without the need for human annotations for model training.

TCS Journal 2021 Journal Article

Star-critical Ramsey number of large cycle and book of different orders

  • Yan Li
  • Yusheng Li
  • Ye Wang

For graphs F, G and H, let F → ( G, H ) signify that any red/blue edge coloring of F contains either a red G or a blue H. The Ramsey number R ( G, H ) is defined to be the minimum r such that K r → ( G, H ), and the star-critical Ramsey number R S ( G, H ) is defined to be the maximum t such that K r ∖ K 1, t → ( G, H ), where r = R ( G, H ). In this note, we shall determine R S ( B n, C m ) for almost same orders.

JBHI Journal 2021 Journal Article

Universal Physiological Representation Learning With Soft-Disentangled Rateless Autoencoders

  • Mo Han
  • Ozan Ozdenizci
  • Toshiaki Koike-Akino
  • Ye Wang
  • Deniz Erdogmus

Human computer interaction (HCI) involves a multidisciplinary fusion of technologies, through which the control of external devices could be achieved by monitoring physiological status of users. However, physiological biosignals often vary across users and recording sessions due to unstable physical/mental conditions and task-irrelevant activities. To deal with this challenge, we propose a method of adversarial feature encoding with the concept of a Rateless Autoencoder (RAE), in order to exploit disentangled, nuisance-robust, and universal representations. We achieve a good trade-off between user-specific and task-relevant features by making use of the stochastic disentanglement of the latent representations by adopting additional adversarial networks. The proposed model is applicable to a wider range of unknown users and tasks as well as different classifiers. Results on cross-subject transfer evaluations show the advantages of the proposed framework, with up to an 11. 6% improvement in the average subject-transfer classification accuracy.

TCS Journal 2020 Journal Article

Complete bipartite graphs deleted in Ramsey graphs

  • Yan Li
  • Yusheng Li
  • Ye Wang

For graphs F, G and H, let F → ( G, H ) signify that any red/blue edge coloring of F contains either a red G or a blue H. The Ramsey number R ( G, H ) is defined as min ⁡ { r | K r → ( G, H ) }. In this note, we consider an optimization problem to bound the complete bipartite-critical Ramsey number R Λ ( G, H ) defined as max ⁡ { t | K r ∖ K t, t → ( G, H ) } where r = R ( G, H ) and Λ is a set of K t, t, and determine R Λ ( G, H ) for some pairs ( G, H ).

TCS Journal 2020 Journal Article

Maximum star deleted from Ramsey graphs of book and tree

  • Ye Wang
  • Yusheng Li
  • Yan Li

For graphs F, G and H, let F → ( G, H ) signify that any red/blue edge coloring of F contains a red G or a blue H. Define the Ramsey number R ( G, H ) to be the smallest r such that K r → ( G, H ). In this note, we consider an optimization problem to find the star-critical Ramsey number R S ( G, H ) defined as max ⁡ { n | K r ∖ K 1, n → ( G, H ) } by showing that for n ≥ 3 m, R S ( B m, T n ) = n − 2, where r = R ( G, H ).

IJCAI Conference 2019 Conference Paper

Discovering Regularities from Traditional Chinese Medicine Prescriptions via Bipartite Embedding Model

  • Chunyang Ruan
  • Jiangang Ma
  • Ye Wang
  • Yanchun Zhang
  • Yun Yang

Regularities analysis for prescriptions is a significant task for traditional Chinese medicine (TCM), both in inheritance of clinical experience and in improvement of clinical quality. Recently, many methods have been proposed for regularities discovery, but this task is challenging due to the quantity, sparsity and free-style of prescriptions. In this paper, we address the specific problem of regularities discovery and propose a graph embedding based framework for regularities discovery for massive prescriptions. We model this task as a relation prediction in which the correlation of two herbs or of herb and symptom are incorporated to characterize the different relationships. Specifically, we first establish a heterogeneous network with herbs and symptoms as its nodes. We develop a bipartite embedding model termed HS2Vec to detect regularities, which explores multiple relations of herbherb, and herb-symptom based on the heterogeneous network. Experiments on four real-world datasets demonstrate that the proposed framework is very effective for regularities discovery.

IJCAI Conference 2018 Conference Paper

Non-decreasing Payment Rules for Combinatorial Auctions

  • Vitor Bosshard
  • Ye Wang
  • Sven Seuken

Combinatorial auctions are used to allocate resources in domains where bidders have complex preferences over bundles of goods. However, the behavior of bidders under different payment rules is not well understood, and there has been limited success in finding Bayes-Nash equilibria of such auctions due to the computational difficulties involved. In this paper, we introduce non-decreasing payment rules. Under such a rule, the payment of a bidder cannot decrease when he increases his bid, which is a natural and desirable property. VCG-nearest, the payment rule most commonly used in practice, violates this property and can thus be manipulated in surprising ways. In contrast, we show that many other payment rules are non-decreasing. We also show that a non-decreasing payment rule imposes a structure on the auction game that enables us to search for an approximate Bayes-Nash equilibrium much more efficiently than in the general case. Finally, we introduce the utility planes BNE algorithm, which exploits this structure and outperforms a state-of-the-art algorithm by multiple orders of magnitude.

NeurIPS Conference 2015 Conference Paper

Probabilistic Curve Learning: Coulomb Repulsion and the Electrostatic Gaussian Process

  • Ye Wang
  • David Dunson

Learning of low dimensional structure in multidimensional data is a canonical problem in machine learning. One common approach is to suppose that the observed data are close to a lower-dimensional smooth manifold. There are a rich variety of manifold learning methods available, which allow mapping of data points to the manifold. However, there is a clear lack of probabilistic methods that allow learning of the manifold along with the generative distribution of the observed data. The best attempt is the Gaussian process latent variable model (GP-LVM), but identifiability issues lead to poor performance. We solve these issues by proposing a novel Coulomb repulsive process (Corp) for locations of points on the manifold, inspired by physical models of electrostatic interactions among particles. Combining this process with a GP prior for the mapping function yields a novel electrostatic GP (electroGP) process. Focusing on the simple case of a one-dimensional manifold, we develop efficient inference algorithms, and illustrate substantially improved performance in a variety of experiments including filling in missing frames in video.

v2026.09.13