Arrow Research search

Author name cluster

Xinyu Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
2 author rows

Possible papers

10

AAAI Conference 2026 Conference Paper

MetaGPT: A Large Vision-Language Model for Meme Metaphor Understanding

  • Bo Xu
  • Chenyuan Wang
  • Xinyu Chen
  • Hongfei Lin
  • Feng Xia

Meme is an expressive medium that often conveys rich emotions and intentions. Recent studies have confirmed the critical role of metaphors in meme understanding. However, existing metaphor research heavily relies on manual annotations, and mainstream vision-language models (VLMs) still struggle with the recognition and comprehension of metaphors. To address these challenges, we introduce MetaGPT, the first vision-language model specifically designed for meme metaphor understanding. MetaGPT is capable of identifying and extracting metaphors in memes, and generating accurate meme interpretations. Furthermore, we construct a dedicated dataset for meme understanding, MUnd, which comprises approximately 32,000 high-quality question-answer (QA) pairs across three core tasks: metaphor detection, metaphor domain extraction, and meme interpretation. Based on MUnd, we further propose an evaluation benchmark for meme understanding and conduct a comprehensive assessment of existing VLMs. Experimental results reveal that current models still face challenges in metaphor comprehension, while MetaGPT consistently outperforms them across all tasks, highlighting its potential in advancing meme understanding.

AAAI Conference 2026 Conference Paper

Minimum-Length Conformal Prediction Sets for Ordinal Classification

  • Zijian Zhang
  • Xinyu Chen
  • Yuanjie Shi
  • Liyuan Lillian Ma
  • Zifan Xu
  • Yan Yan

Ordinal classification has been widely applied in many high-stakes applications, e.g., medical imaging and diagnosis, where reliable uncertainty quantification (UQ) is essential for decision making. Conformal prediction (CP) is a general UQ framework that provides statistically valid guarantees, which is especially useful in practice. However, prior ordinal CP methods mainly focus on heuristic algorithms or restrictively require the underlying model to predict a unimodal distribution over ordinal labels. Consequently, they provide limited insight into coverage–efficiency trade-offs, or a model-agnostic and distribution-free nature favored by CP methods. To this end, we fill this gap by propose an ordinal-CP method that is model-agnostic and provides instance-level optimal prediction intervals. Specifically, we formulate conformal ordinal classification as a minimum-length covering problem at the instance level. To solve this problem, we develop a sliding-window algorithm that is optimal on each calibration data, with only a linear time complexity in K, the # of label candidates. The local optimality per instance further also improves predictive efficiency in expectation. Moreover, we propose a length-regularized variant that shrinks prediction set size while preserving coverage. Experiments on four benchmark datasets from diverse domains are conducted to demonstrate the significantly improved predictive efficiency of the proposed methods over baselines (by 15%↓ on average over four datasets).

EAAI Journal 2025 Journal Article

A novel lightweight semantic segmentation network for line-structured light measurement under strong reflection noise

  • Xinyu Chen
  • Chen Fang
  • Yan Ren
  • Ailing Hu

Line-structured light measurement is a critical technology for realizing industrial automation. However, in practical applications, the noise from multiple reflections between object surfaces can cause non-ideal distributions of light stripes, which poses challenges for three-dimensional (3-D) measurement accuracy and correctness. To alleviate this problem, a topology-based lightweight segmentation network based on the geometric features is proposed in this paper, and is integrated into the line-structured light 3-D measurement system under the premise of ensuring accuracy and low latency. Experimental results demonstrate that the proposed method can accurately extract laser stripes under strong reflection noise, with a processing speed of about 22. 65 frames per second (FPS) on central processing unit (CPU) and 94. 70 FPS on graphics processing unit (GPU).

NeurIPS Conference 2025 Conference Paper

FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction

  • Jiang Lin
  • Xinyu Chen
  • Song Wu
  • Zhiqiu Zhang
  • Jizhi Zhang
  • Ye Wang
  • Qiang Tang
  • Qian Wang

Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference cost due to dual-path denoising. We present \textbf{FreeControl}, a training-free framework for semantic structural control in diffusion models. Unlike prior methods that extract attention across multiple timesteps, FreeControl performs \textit{one-step attention extraction} from a single, optimally chosen timestep and reuses it throughout denoising. This enables efficient structural guidance without inversion or retraining. To further improve quality and stability, we introduce \textit{Latent-Condition Decoupling (LCD)}: a principled separation of the timestep condition and the noised latent used in attention extraction. LCD provides finer control over attention quality and eliminates structural artifacts. FreeControl also supports compositional control via reference images assembled from multiple sources, enabling intuitive scene layout design and stronger prompt alignment. FreeControl introduces a new paradigm for test-time control—enabling structurally and semantically aligned, visually coherent generation directly from raw images, with the flexibility for intuitive compositional design and compatibility with modern diffusion models at ~5\% additional cost.

NeurIPS Conference 2025 Conference Paper

Safe RLHF-V: Safe Reinforcement Learning from Multi-modal Human Feedback

  • Jiaming Ji
  • Xinyu Chen
  • Rui Pan
  • Han Zhu
  • Jiahao Li
  • Donghai Hong
  • Boyuan Chen
  • Jiayi Zhou

Multimodal large language models (MLLMs) are essential for building general-purpose AI assistants; however, they pose increasing safety risks. How can we ensure safety alignment of MLLMs to prevent undesired behaviors? Going further, it is critical to explore how to fine-tune MLLMs to preserve capabilities while meeting safety constraints. Fundamentally, this challenge can be formulated as a min-max optimization problem. However, existing datasets have not yet disentangled single preference signals into explicit safety constraints, hindering systematic investigation in this direction. Moreover, it remains an open question whether such constraints can be effectively incorporated into the optimization process for multi-modal models. In this work, we present the first exploration of the Safe RLHF-V -- the first multimodal safety alignment framework. The framework consists of: (I) BeaverTails-V, the first open-source dataset featuring dual preference annotations for helpfulness and safety, supplemented with multi-level safety labels (minor, moderate, severe); (II) Beaver-Guard-V, a multi-level guardrail system to proactively defend against unsafe queries and adversarial attacks. Applying the guard model over five rounds of filtering and regeneration significantly enhances the precursor model’s overall safety by an average of 40. 9%. (II) Based on dual preference, we initiate the first exploration of multi-modal safety alignment within a constrained optimization. Experimental results demonstrate that Safe RLHF effectively improves both model helpfulness and safety. Specifically, Safe RLHF-V enhances model safety by 34. 2% and helpfulness by 34. 3%.

ECAI Conference 2024 Conference Paper

Enhancing Discourse Coherence to Improve Cross-Document Event Coreference Resolution

  • Xinyu Chen
  • Sheng Xu 0006
  • Peifeng Li
  • Qiaoming Zhu

Cross-Document Event Coreference Resolution (CD-ECR) is a task of grouping event mentions across multiple documents that refer to the same real-world events. In contrast to within-document event mentions, which are linked by rich, coherent contexts, cross-document event mentions lack such contexts, making it challenging for the model to establish a connection between two event mentions in different documents. To address this issue, we propose a novel mechanism of enhancing discourse coherence to boost CD-ECR. Specifically, we introduce a new task, ECD-CoE (Event-oriented Cross-Document Coherence Enhancement), which selects coherent sentences that form a coherent text for two cross-document event mentions. We then use this coherent text to represent the event mentions and resolve coreferent events. Experimental results on both the ECB+ and GVC datasets indicate that our proposed method outperforms several state-of-the-art baselines.

JMLR Journal 2024 Journal Article

Optimal Weighted Random Forests

  • Xinyu Chen
  • Dalei Yu
  • Xinyu Zhang

The random forest (RF) algorithm has become a very popular prediction method for its great flexibility and promising accuracy. In RF, it is conventional to put equal weights on all the base learners (trees) to aggregate their predictions. However, the predictive performance of different trees within the forest can vary significantly due to the randomization of the embedded bootstrap sampling and feature selection. In this paper, we focus on RF for regression and propose two optimal weighting algorithms, namely the 1 Step Optimal Weighted RF (1step-WRF$_\mathrm{opt}$) and 2 Steps Optimal Weighted RF (2steps-WRF$_\mathrm{opt}$), that combine the base learners through the weights determined by weight choice criteria. Under some regularity conditions, we show that these algorithms are asymptotically optimal in the sense that the resulting squared loss and risk are asymptotically identical to those of the infeasible but best possible weighted RF. Numerical studies conducted on real-world data sets and semi-synthetic data sets indicate that these algorithms outperform the equal-weight forest and two other weighted RFs proposed in the existing literature in most cases. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2024. ( edit, beta )

AAMAS Conference 2023 Conference Paper

Improving Quantal Cognitive Hierarchy Model Through Iterative Population Learning

  • Yuhong Xu
  • Shih-Fen Cheng
  • Xinyu Chen

In this paper, we propose to enhance the state-of-the-art quantal cognitive hierarchy (QCH) model with iterative population learning (IPL) to estimate the empirical distribution of agents’ reasoning levels and fit human agents’ behavioral data. We apply our approach to a real-world dataset from the Swedish lowest unique positive integer (LUPI) game and show that our proposed approach outperforms the theoretical Poisson Nash equilibrium predictions and the QCH approach by 49. 8% and 46. 6% in Wasserstein distance respectively. Our approach also allows us to explicitly measure an agent’s reasoning level distribution, which is not previously possible.

YNIMG Journal 2022 Journal Article

Human visual processing during walking: Dissociable pre- and post-stimulus influences

  • Xinyu Chen
  • Liyu Cao
  • Barbara F Haendel

Walking influences visual processing but the underlying mechanism remains poorly understood. In this study, we investigated the influence of walking on pre-stimulus and stimulus-induced visual neural activity and behavioural performance in a discrimination task while participants were standing or freely walking. The results showed dissociable pre- and post-stimulus influences by the movement state. Walking was associated with a reduced pre-stimulus alpha power, which predicted enhanced N1 and decreased P3 components during walking. This pre-stimulus alpha activity was additionally modulated by time on the task, which was paralleled by a similar behavioural modulation. In contrast, the post-stimulus alpha power was reduced in its modulation due to stimulus onset during walking but showed no evidence of modulation by time on the task. Additionally, stimulus parameters (eccentricity, laterality, distractor presence significantly influenced post-stimulus alpha power, whereas the visually evoked components showed no evidence of such an influence. There was further no evidence of a correlation between pre-stimulus and post stimulus alpha power. We conclude that walking has two dissociable influences on visual processing: while the walking induced reduction in alpha power suggests an attentional state change that relates to visual awareness, the post-stimulus influence on alpha power modulation indicates changed spatial visual processing during walking.

IS Journal 2004 Journal Article

Design and Evaluation of a Fault-Tolerant Mobile-Agent System

  • M.R. Lyu
  • Xinyu Chen
  • Tsz Yeung Wong

The mobile agents create a new paradigm for data exchange and resource sharing in rapidly growing and continually changing computer networks. In a distributed system, failures can occur in any software or hardware component. A mobile agent can get lost when its hosting server crashes during execution, or it can get dropped in a congested network. Therefore, survivability and fault tolerance are vital issues for deploying mobile-agent systems. This fault tolerance approach deploys three kinds of cooperating agents to detect server and agent failures and recover services in mobile-agent systems. An actual agent is a common mobile agent that performs specific computations for its owner. Witness agents monitor the actual agent and detect whether it's lost. A probe recovers the failed actual agent and the witness agents. A peer-to-peer message-passing mechanism stands between each actual agent and its witness agents to perform failure detection and recovery through time-bounded information exchange; a log records the actual agent's actions. When failures occur, the system performs rollback recovery to abort uncommitted actions. Moreover, our method uses checkpointed data to recover the lost actual agent.

v2026.09.13