Arrow Research search

Author name cluster

Tiejun Zhao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

18 papers
2 author rows

Possible papers

18

AAAI Conference 2026 Conference Paper

Long-form RewardBench: Evaluating Reward Models for Long-form Generation

  • Hui Huang
  • Yancheng He
  • Wei Liu
  • Muyun Yang
  • Jiaheng Liu
  • Kehai Chen
  • Bing Xu
  • Conghui Zhu

The widespread adoption of reinforcement learning-based alignment highlights the growing importance of reward models. Various benchmarks have been built to evaluate reward models in various domains and scenarios. However, a significant gap remains in assessing reward models for long-form generation, despite its critical role in real-world applications. To bridge this, we introduce Long-form RewardBench, the first reward modeling testbed specifically designed for long-form generation. Our benchmark encompasses five key subtasks: QA, RAG, Chat, Writing, and Reasoning. We collected instruction and preference data through a meticulously designed multi-stage data collection process, and conducted extensive experiments on 20+ mainstream reward models, including both classifiers and generative models. Our findings reveal that current models still lack long-form reward modeling capabilities. Furthermore, we designed a novel Long-form Needle-in-a-Haystack Test, which revealed a correlation between reward modeling performance and the error's position within a response, as well as the overall response length, with distinct characteristics observed between classification and generative models. Finally, we demonstrate that classifier exhibit better generalizability compared to generative models trained on the same data. As the first benchmark for long-form reward modeling, this work aims to offer a robust platform for visualizing progress in this crucial area.

AAAI Conference 2026 Conference Paper

Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response Theory

  • Hongli Zhou
  • Hui Huang
  • Ziqing Zhao
  • Lvyuan Han
  • Huicheng Wang
  • Kehai Chen
  • Muyun Yang
  • Wei Bao

The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model capabilities. This paper provides a critical analysis of benchmark effectiveness, examining mainstream prominent LLM benchmarks using results from diverse models. We first propose Pseudo-Siamese Network for Item Response Theory (PSN-IRT), an enhanced Item Response Theory framework that incorporates a rich set of item parameters within an IRT-grounded architecture. PSN-IRT can be utilized for accurate and reliable estimations of item characteristics and model abilities. Based on PSN-IRT, we conduct extensive analysis on 11 LLM benchmarks comprising 41,871 items, revealing significant and varied shortcomings in their measurement quality. Furthermore, we demonstrate that leveraging PSN-IRT is able to construct smaller benchmarks while maintaining stronger alignment with human preference.

AAAI Conference 2025 Conference Paper

Look Before You Leap: Enhance Attention and Vigilance Regarding Harmful Content with GuidelineLLM

  • Shaoqing Zhang
  • Zhuosheng Zhang
  • Kehai Chen
  • Rongxiang Weng
  • Muyun Yang
  • Tiejun Zhao
  • Min Zhang

Despite being empowered with alignment mechanisms, large language models (LLMs) are increasingly vulnerable to emerging jailbreak attacks that can compromise their alignment mechanisms. This vulnerability poses significant risks to real-world applications. Existing work faces challenges in both training efficiency and generalization capabilities (i.e., Reinforcement Learning from Human Feedback and Red-Teaming). Developing effective strategies to enable LLMs to resist continuously evolving jailbreak attempts represents a significant challenge. To address this challenge, we propose a novel defensive paradigm called GuidelineLLM, which assists LLMs in recognizing queries that may have harmful content. Before LLMs respond to a query, GuidelineLLM first identifies potential risks associated with the query, summarizes these risks into guideline suggestions, and then feeds these guidelines to the responding LLMs. Importantly, our approach eliminates the necessity for additional safety fine-tuning of the LLMs themselves; only the GuidelineLLM requires fine-tuning. This characteristic enhances the general applicability of GuidelineLLM across various LLMs. Experimental results demonstrate that GuidelineLLM can significantly reduce the attack success rate (ASR) against LLM (an average reduction of 34.17% ASR) while maintaining the usefulness of LLM in handling benign queries.

IJCAI Conference 2025 Conference Paper

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

  • Shaojun E
  • Yuchen Yang
  • Jiaheng Wu
  • Yan Zhang
  • Tiejun Zhao
  • Ziyan Chen

In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively correlated with the coupling between visual encoders and large language models. Existing approaches often face issues such as vector gaps or semantic disparities, resulting in information loss during the propagation process. To address these issues, we propose MAGE (Multimodal Alignment and Generation Enhancement), a novel framework that bridges the semantic spaces of vision and text through an innovative alignment mechanism. By introducing the Intelligent Alignment Network (IAN), MAGE achieves dimensional and semantic alignment. To reduce the gap between synonymous heterogeneous data, we employ a training strategy that combines cross-entropy and mean squared error, significantly enhancing the alignment effect. Moreover, to enhance MAGE’s “Any-to-Any” capability, we developed a fine-tuning dataset for multimodal tool-calling instructions to expand the model’s output capability boundaries. Finally, our proposed multimodal large model architecture, MAGE, achieved significantly better performance compared to similar works across various evaluation benchmarks, including MME, MMBench, and SEED. Complete code and appendix are available at: https: //github. com/GTCOM-NLP/MAGE

NeurIPS Conference 2025 Conference Paper

Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning

  • Yihong Tang
  • Kehai Chen
  • Muyun Yang
  • Zheng-Yu Niu
  • Jing Li
  • Tiejun Zhao
  • Min Zhang

The advancement of Large Language Models (LLMs) has spurred significant interest in Role-Playing Agents (RPAs) for applications such as emotional companionship and virtual interaction. However, recent RPAs are often built on explicit dialogue data, lacking deep, human-like internal thought processes, resulting in superficial knowledge and style expression. While Large Reasoning Models (LRMs) can be employed to simulate character thought, their direct application is hindered by attention diversion (i. e. , RPAs forget their role) and style drift (i. e. , overly formal and rigid reasoning rather than character-consistent reasoning). To address these challenges, this paper introduces a novel Role-Aware Reasoning (RAR) method, which consists of two important stages: Role Identity Activation (RIA) and Reasoning Style Optimization (RSO). RIA explicitly guides the model with character profiles during reasoning to counteract attention diversion, and then RSO aligns reasoning style with the character and scene via LRM distillation to mitigate style drift. Extensive experiments demonstrate that the proposed RAR significantly enhances the performance of RPAs by effectively addressing attention diversion and style drift.

AAAI Conference 2024 Conference Paper

Enhancing Bilingual Lexicon Induction via Bi-directional Translation Pair Retrieving

  • Qiuyu Ding
  • Hailong Cao
  • Tiejun Zhao

Most Bilingual Lexicon Induction (BLI) methods retrieve word translation pairs by finding the closest target word for a given source word based on cross-lingual word embeddings (WEs). However, we find that solely retrieving translation from the source-to-target perspective leads to some false positive translation pairs, which significantly harm the precision of BLI. To address this problem, we propose a novel and effective method to improve translation pair retrieval in cross-lingual WEs. Specifically, we consider both source-side and target-side perspectives throughout the retrieval process to alleviate false positive word pairings that emanate from a single perspective. On a benchmark dataset of BLI, our proposed method achieves competitive performance compared to existing state-of-the-art (SOTA) methods. It demonstrates effectiveness and robustness across six experimental languages, including similar language pairs and distant language pairs, under both supervised and unsupervised settings.

AAAI Conference 2024 Conference Paper

Spot the Error: Non-autoregressive Graphic Layout Generation with Wireframe Locator

  • Jieru Lin
  • Danqing Huang
  • Tiejun Zhao
  • Dechen Zhan
  • Chin-Yew Lin

Layout generation is a critical step in graphic design to achieve meaningful compositions of elements. Most previous works view it as a sequence generation problem by concatenating element attribute tokens (i.e., category, size, position). So far the autoregressive approach (AR) has achieved promising results, but is still limited in global context modeling and suffers from error propagation since it can only attend to the previously generated tokens. Recent non-autoregressive attempts (NAR) have shown competitive results, which provides a wider context range and the flexibility to refine with iterative decoding. However, current works only use simple heuristics to recognize erroneous tokens for refinement which is inaccurate. This paper first conducts an in-depth analysis to better understand the difference between the AR and NAR framework. Furthermore, based on our observation that pixel space is more sensitive in capturing spatial patterns of graphic layouts (e.g., overlap, alignment), we propose a learning-based locator to detect erroneous tokens which takes the wireframe image rendered from the generated layout sequence as input. We show that it serves as a complementary modality to the element sequence in object space and contributes greatly to the overall performance. Experiments on two public datasets show that our approach outperforms both AR and NAR baselines. Extensive studies further prove the effectiveness of different modules with interesting findings. Our code will be available at https://github.com/ffffatgoose/SpotError.

AAAI Conference 2021 Conference Paper

Document-Level Relation Extraction with Reconstruction

  • Wang Xu
  • Kehai Chen
  • Tiejun Zhao

In document-level relation extraction (DocRE), graph structure is generally used to encode relation information in the input document to classify the relation category between each entity pair, and has greatly advanced the DocRE task over the past several years. However, the learned graph representation universally models relation information between all entity pairs regardless of whether there are relationships between these entity pairs. Thus, those entity pairs without relationships disperse the attention of the encoder-classifier DocRE for ones with relationships, which may further hind the improvement of DocRE. To alleviate this issue, we propose a novel encoder-classifierreconstructor model for DocRE. The reconstructor manages to reconstruct the ground-truth path dependencies from the graph representation, to ensure that the proposed DocRE model pays more attention to encode entity pairs with relationships in the training. Furthermore, the reconstructor is regarded as a relationship indicator to assist relation classification in the inference, which can further improve the performance of DocRE model. Experimental results on a large-scale DocRE dataset show that the proposed model can significantly improve the accuracy of relation extraction on a strong heterogeneous graph-based baseline. The code is publicly available at https: //github. com/xwjim/DocRE-Rec.

ECAI Conference 2020 Conference Paper

Multimodal Matching Transformer for Live Commenting

  • Chaoqun Duan
  • Lei Cui 0001
  • Shuming Ma
  • Furu Wei
  • Conghui Zhu
  • Tiejun Zhao

Automatic live commenting aims to provide real-time comments on videos for viewers. It encourages users engagement on online video sites, and is also a good benchmark for video-to-text generation. Recent work on this task adopts encoder-decoder models to generate comments. However, these methods do not model the interaction between videos and comments explicitly, so they tend to generate popular comments that are often irrelevant to the videos. In this work, we aim to improve the relevance between live comments and videos by modeling the cross-modal interactions among different modalities. To this end, we propose a multimodal matching transformer to capture the relationships among comments, vision, and audio. The proposed model is based on the transformer framework and can iteratively learn the attention-aware representations for each modality. We evaluate the model on a publicly available live commenting dataset. Experiments show that the multimodal matching transformer model outperforms the state-of-the-art methods.

IJCAI Conference 2018 Conference Paper

Attention-Fused Deep Matching Network for Natural Language Inference

  • Chaoqun Duan
  • Lei Cui
  • Xinchi Chen
  • Furu Wei
  • Conghui Zhu
  • Tiejun Zhao

Natural language inference aims to predict whether a premise sentence can infer another hypothesis sentence. Recent progress on this task only relies on a shallow interaction between sentence pairs, which is insufficient for modeling complex relations. In this paper, we present an attention-fused deep matching network (AF-DMN) for natural language inference. Unlike existing models, AF-DMN takes two sentences as input and iteratively learns the attention-aware representations for each side by multi-level interactions. Moreover, we add a self-attention mechanism to fully exploit local context information within each sentence. Experiment results show that AF-DMN achieves state-of-the-art performance and outperforms strong baselines on Stanford natural language inference (SNLI), multi-genre natural language inference (MultiNLI), and Quora duplicate questions datasets.

IJCAI Conference 2018 Conference Paper

Point Set Registration for Unsupervised Bilingual Lexicon Induction

  • Hailong Cao
  • Tiejun Zhao

Inspired by the observation that word embeddings exhibit isomorphic structure across languages, we propose a novel method to induce a bilingual lexicon from only two sets of word embeddings, which are trained on monolingual source and target data respectively. This is achieved by formulating the task as point set registration which is a more general problem. We show that a transformation from the source to the target embedding space can be learned automatically without any form of cross-lingual supervision. By properly adapting a traditional point set registration model to make it be suitable for processing word embeddings, we achieved state-of-the-art performance on the unsupervised bilingual lexicon induction task. The point set registration problem has been well-studied and can be solved by many elegant models, we thus opened up a new opportunity to capture the universal lexical semantic structure across languages.

AAAI Conference 2018 Conference Paper

Syntax-Directed Attention for Neural Machine Translation

  • Kehai Chen
  • Rui Wang
  • Masao Utiyama
  • Eiichiro Sumita
  • Tiejun Zhao

Attention mechanism, including global attention and local attention, plays a key role in neural machine translation (NMT). Global attention attends to all source words for word prediction. In comparison, local attention selectively looks at fixed-window source words. However, alignment weights for the current target word often decrease to the left and right by linear distance centering on the aligned source position and neglect syntax distance constraints. In this paper, we extend the local attention with syntax-distance constraint, which focuses on syntactically related source words with the predicted target word to learning a more effective context vector for predicting translation. Moreover, we further propose a double context NMT architecture, which consists of a global context vector and a syntax-directed context vector from the global attention, to provide more translation performance for NMT from source representation. The experiments on the largescale Chinese-to-English and English-to-German translation tasks show that the proposed approach achieves a substantial and significant improvement over the baseline system.

AAAI Conference 2018 Conference Paper

Table-to-Text: Describing Table Region With Natural Language

  • Junwei Bao
  • Duyu Tang
  • Nan Duan
  • Zhao Yan
  • Yuanhua Lv
  • Ming Zhou
  • Tiejun Zhao

In this paper, we present a generative model to generate a natural language sentence describing a table region, e. g. , a row. The model maps a row from a table to a continuous vector and then generates a natural language sentence by leveraging the semantics of a table. To deal with rare words appearing in a table, we develop a flexible copying mechanism that selectively replicates contents from the table in the output sequence. Extensive experiments demonstrate the accuracy of the model and the power of the copying mechanism. On two synthetic datasets, WIKIBIO and SIMPLEQUESTIONS, our model improves the current state-of-the-art BLEU-4 score from 34. 70 to 40. 26 and from 33. 32 to 39. 12, respectively. Furthermore, we introduce an open-domain dataset WIK- ITABLETEXT including 13, 318 explanatory sentences for 4, 962 tables. Our model achieves a BLEU-4 score of 38. 23, which outperforms template based and language model based approaches.

AAAI Conference 2017 Conference Paper

Deterministic Attention for Sequence-to-Sequence Constituent Parsing

  • Chunpeng Ma
  • Lemao Liu
  • Akihiro Tamura
  • Tiejun Zhao
  • Eiichiro Sumita

The sequence-to-sequence model is proven to be extremely successful in constituent parsing. It relies on one key technique, the probabilistic attention mechanism, to automatically select the context for prediction. Despite its successes, the probabilistic attention model does not always select the most important context. For example, the headword and boundary words of a subtree have been shown to be critical when predicting the constituent label of the subtree, but this contextual information becomes increasingly difficult to learn as the length of the sequence increases. In this study, we proposed a deterministic attention mechanism that deterministically selects the important context and is not affected by the sequence length. We implemented two different instances of this framework. When combined with a novel bottom-up linearization method, our parser demonstrated better performance than that achieved by the sequence-to-sequence parser with probabilistic attention mechanism.

AAAI Conference 2017 Conference Paper

Translation Prediction with Source Dependency-Based Context Representation

  • Kehai Chen
  • Tiejun Zhao
  • Muyun Yang
  • Lemao Liu

Learning context representations is very promising to improve translation results, particularly through neural networks. Previous efforts process the context words sequentially and neglect their internal syntactic structure. In this paper, we propose a novel neural network based on bi-convolutional architecture to represent the source dependency-based context for translation prediction. The proposed model is able to not only encode the long-distance dependencies but also capture the functional similarities for better translation prediction (i. e. , ambiguous words translation and word forms translation). Examined by a largescale Chinese-English translation task, the proposed approach achieves a significant improvement (of up to +1. 9 BLEU points) over the baseline system, and meanwhile outperforms a number of context-enhanced comparison system.

IJCAI Conference 2013 Conference Paper

Fusion of Word and Letter Based Metrics for Automatic MT Evaluation

  • Muyun Yang
  • Junguo Zhu
  • Sheng Li
  • Tiejun Zhao

With the progress in machine translation, it becomes more subtle to develop the evaluation metric capturing the systems’ differences in comparison to the human translations. In contrast to the current efforts in leveraging more linguistic information to depict translation quality, this paper takes the thread of combining language independent features for a robust solution to MT evaluation metric. To compete with finer granularity of modeling brought by linguistic features, the proposed method augments the word level metrics by a letter based calculation. An empirical study is then conducted over WMT data to train the metrics by ranking SVM. The results reveal that the integration of current language independent metrics can generate well enough performance for a variety of languages. Time-split data validation is promising as a better training setting, though the greedy strategy also works well.

IJCAI Conference 2013 Conference Paper

Learning Domain Differences Automatically for Dependency Parsing Adaptation

  • Mo Yu
  • Tiejun Zhao
  • Yalong Bai

In this paper, we address the relation between domain differences and domain adaptation for dependency parsing. Our quantitative analyses showed that it is the inconsistent behavior of same features cross-domain, rather than word or feature coverage, that is the major cause of performances decrease of out-domain model. We further studied those ambiguous features in depth and found that the set of ambiguous features is small and has concentric distributions. Based on the analyses, we proposed a DA method. The DA method can automatically learn which features are ambiguous cross domain according to errors made by out-domain model on in-domain training data. Our method is also extended to utilize multiple out-domain models. The results of dependency parser adaptation from WSJ to Genia and Question bank showed that our method achieved significant improvements on small in-domain datasets where DA is mostly in need. Additionally, we achieved improvement on the published best results of CoNLL07 shared task on domain adaptation, which confirms the significance of our analyses and our method.

YNIMG Journal 2008 Journal Article

BOLD study of stimulation-induced neural activity and resting-state connectivity in medetomidine-sedated rat

  • Fuqiang Zhao
  • Tiejun Zhao
  • Lei Zhou
  • Qiulin Wu
  • Xiaoping Hu

Functional magnetic resonance imaging (fMRI) in anesthetized-animals is critical in studying the mechanisms of fMRI and investigating animal models of various diseases. Medetomidine was recently introduced for independent anesthesia for longitudinal (survival) fMRI studies in rats. Since stimulation-induced fMRI signal is anesthesia-dependent and its characteristics in rats under medetomidine are not fully elucidated, the blood oxygenation level dependent (BOLD) fMRI response to electrical forepaw stimulation under medetomidine was systematically investigated at 9. 4 T. Robust activations in contralateral primary somatosensory cortex (SI) and thalamus were observed and peaked at the stimulus frequency of 9 Hz. The response in SI saturates at the stimulus strength of 4 mA while that in thalamus monotonically increases. In addition to fMRI data acquired with the forepaw stimulation, data were also acquired during the resting-state to investigate the synchronization of low frequency fluctuations (LFF) in the BOLD signal (<0. 08 Hz) in different brain regions. LFF during resting-state have been observed to be synchronized between functionally related brain regions in human subjects while its origin is not fully understood. LFF have not been extensively studied or widely reported in anesthetized-animals. In our data, synchronized LFF of BOLD signals are found in clustered, bilaterally symmetric regions, including SI and caudate–putamen and the magnitude of the LFF is ∼1. 5%, comparable to the stimulation-induced BOLD signals. Similar to resting-state data reported in human subjects, LFF in rats under medetomidine likely reflect functional connectivity of these brain regions.

v2026.09.13