Arrow Research search

Author name cluster

Jing Yu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

18 papers
2 author rows

Possible papers

18

AAAI Conference 2026 Conference Paper

Manipulation Intention Understanding for Zero-Shot Composed Image Retrieval

  • Yuanmin Tang
  • Jing Yu
  • Keke Gai
  • Gang Xiong
  • Gaopeng Gou
  • Meikang Qiu
  • Qi Wu

Zero-shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with varied visual manipulation intents across domains, scenes, objects, and attributes. A key challenge is that existing datasets contain limited intent-relevant annotations, making it hard for models to infer human intent from textual modifications. We introduce an intent-centric image–text dataset generated via reasoning by a Multimodal Large Language Model (MLLM) to better train ZS-CIR models for human manipulation intent understanding. Building on this dataset, we propose De-MINDS, a framework that distills the MLLM’s reasoning ability to capture manipulation intent and enhance models’ comprehension of modified text. A simple mapping network translates image information into language space and combines it with the manipulation text to form a query. De-MINDS then extracts intention-relevant information from this query and encodes it as pseudo-word tokens for accurate ZS-CIR. Across four ZS-CIR tasks, De-MINDS shows strong generalization and improves over existing methods by 2.15% to 4.05%, establishing new state-of-the-art results with comparable inference time.

YNIMG Journal 2025 Journal Article

ERP-based interbrain causal model reveals closed-loop information interaction in interpersonal negotiations

  • Yuqin Li
  • Genon Sarah
  • Chunli Chen
  • Lin Jiang
  • Baodan Chen
  • Rihui Li
  • Zhen Liang
  • Jing Yu

Uncovering the interbrain neural mechanisms underlying interpersonal negotiation offers insight into social decision-making dynamics in resource allocation. In this study, we used EEG hyperscanning alongside an iterated ultimatum game to investigate interbrain coupling and dyadic exchange behavior during negotiation. Frontal cortex event-related potentials (ERPs) revealed the distinct neural responses driven by partners' behavioral cues: the proposer's N200 differed significantly for fair versus unfair offers, and the responder's feedback-related negativity (FRN) showed a trend toward significance for the same contrast, while the proposer's N500 varied between acceptance and rejection feedback. Our analysis introduced a novel causal model based on directional phase transfer entropy (dPTE) and time-varying ERP amplitudes, illustrating directed neural processes driven by social exchange, where the proposer's brain activity initially exerts a causal impact on the responder's, whose feedback in turn influences the proposer, creating a closed-loop interaction that drives adaptive negotiation strategies. Additionally, our prediction model with autoregression with exogenous input, which incorporated these causal links between brains, demonstrated higher accuracy than single-brain or reverse causal models, underscoring the significance of dynamic interbrain coupling in interpersonal coordination. This causal model provides a mechanistic explanation of how proposer-responder pairs perceive and adapt to each other's decisions, facilitating shared attention and behavioral coordination in reciprocal, asymmetric negotiations. These findings offer a novel theoretical framework for studying complex social behaviors through interbrain dynamics and may inspire future applications in enhancing cooperative decision-making processes.

AAAI Conference 2025 Conference Paper

MFL-Owner: Ownership Protection for Multi-modal Federated Learning via Orthogonal Transform Watermark

  • Keke Gai
  • Dongjue Wang
  • Jing Yu
  • Mohan Wang
  • Liehuang Zhu
  • Qi Wu

Multi-modal Federated Learning (MFL) is a distributed machine learning paradigm that enables multiple participants with multi-modal data to collaboratively train a global model for multi-modal tasks without sharing their local data. MFL typically deploys the trained global model as an Embedding-as-a-Service (EaaS), allowing participants to obtain embeddings for downstream tasks. However, it increases the risk of unauthorized copying and leakage of the model. Protecting the ownership of the MFL model while maintaining model performance is challenging. In this paper, we propose the first general model ownership protection framework for MFL, named MFL-Owner. MFL-Owner decouples the watermarking process from the model training process and addresses both ownership verification and traceability, effectively safeguarding the interests of the MFL collective. MFL-Owner leverages the concept of orthogonal transformations by incorporating a linear transformation matrix with orthogonal constraints into the model, achieving high-quality ownership verification and traceability with minimal impact on model performance. To enhance the practicality of the watermark and prevent conflicts among multiple clients during tracing, we propose a trigger dataset selection method based on out-of-distribution data combined with Gaussian noise perturbation. Our experiments on multiple datasets demonstrate that MFL-Owner is effective for model ownership verification and traceability for MFL.

ICLR Conference 2025 Conference Paper

SaMer: A Scenario-aware Multi-dimensional Evaluator for Large Language Models

  • Kehua Feng
  • Keyan Ding
  • Jing Yu
  • Yiwen Qu
  • Zhiwen Chen 0002
  • Chengfei Lv
  • Gang Yu
  • Qiang Zhang 0026

Evaluating the response quality of large language models (LLMs) for open-ended questions poses a significant challenge, especially given the subjectivity and multi-dimensionality of "quality" in natural language generation. Existing LLM evaluators often neglect that different scenarios require distinct evaluation criteria. In this work, we propose **SaMer**, a scenario-aware multi-dimensional evaluator designed to provide both overall and fine-grained assessments of LLM-generated responses. Unlike fixed-dimension evaluation approaches, SaMer adapts to different scenarios by automatically identifying and prioritizing relevant evaluation dimensions tailored to the given query. To achieve this, we construct a large-scale fine-grained preference dataset spanning multiple real-world scenarios, each with distinct evaluation dimensions. We then leverage a text embedding model combined with three specialized heads to predict the appropriate evaluation dimensions and corresponding scores, as well as the respective weights that contribute to the overall score. The resulting model offers fine-grained and interpretable evaluations and shows robust adaptability across diverse scenarios. Extensive experiments on eight single rating and pairwise comparison datasets demonstrate that SaMer outperforms existing baselines in a variety of evaluation tasks, showcasing its robustness, versatility, and generalizability.

IJCAI Conference 2025 Conference Paper

Secure and Efficient Watermarking for Latent Diffusion Models in Model Distribution Scenarios

  • Liangqi Lei
  • Keke Gai
  • Jing Yu
  • Liehuang Zhu
  • Qi Wu

Latent diffusion models have exhibited considerable potential in generative tasks. Watermarking is considered to be an alternative to safeguard the copyright of generative models and prevent their misuse. However, in the context of model distribution scenarios, the accessibility of models to large scale of model users brings new challenges to the security, efficiency and robustness of existing watermark solutions. To address these issues, we propose a secure and efficient watermarking solution. A new security mechanism is designed to prevent watermark leakage and watermark escape, which considers watermark randomness and watermark-model association as two constraints for mandatory watermark injection. To reduce the time cost of training the security module, watermark injection and the security mechanism are decoupled, ensuring that fine-tuning VAE only accomplishes the security mechanism without the burden of learning watermark patterns. A watermark distribution-based verification strategy is proposed to enhance the robustness against diverse attacks in the model distribution scenarios. Experimental results prove that our watermarking consistently outperforms existing six baselines on effectiveness and robustness against ten image processing attacks and adversarial attacks, while enhancing security in the distribution scenarios. The code is available at https: //anonymous. 4open. science/r/DistriMark-F11F/.

YNIMG Journal 2024 Journal Article

Are older adults less generous? Age differences in emotion-related social decision making

  • Hong-Zhou Xu
  • Xue-Rui Peng
  • Shen-Yin Huan
  • Jia-Jie Xu
  • Jing Yu
  • Qing-Guo Ma

In social interaction, age-related differences in emotional processing may lead to varied social decision making between young and older adults. However, previous studies of social decision making have paid less attention to the interactants' emotions, leaving age differences and underlying neural mechanisms unexplored. To address this gap, the present study combined functional and structural magnetic resonance imaging, employing a modified dictator game task with recipients displaying either neutral or sad facial expressions. Behavioral results indicated that although older adults' overall allocations did not differ significantly from those of young adults, older adults' allocations showing a decrease in emotion-related generosity compared to young adults. Using representational similarity analysis, we found that older adults showed reduced neural representations of recipients' emotions and gray matter volume in the right anterior cingulate gyrus (ACC), right insula, and left dorsomedial prefrontal cortex (DMPFC) compared to young adults. More importantly, mediation analyses indicated that age influenced allocations not only through serial mediation of neural representations of the right insula and left DMPFC, but also through serial mediation of the mean gray matter volume of the right ACC and left DMPFC. This study identifies the potential neural pathways through which age affects emotion-related social decision making, advancing our understanding of older adults' social interaction behavior that they may not be less generous unless confronted with individuals with specific emotions.

AAAI Conference 2024 Conference Paper

Context-I2W: Mapping Images to Context-Dependent Words for Accurate Zero-Shot Composed Image Retrieval

  • Yuanmin Tang
  • Jing Yu
  • Keke Gai
  • Jiamin Zhuang
  • Gang Xiong
  • Yue Hu
  • Qi Wu

Different from the Composed Image Retrieval task that requires expensive labels for training task-specific models, Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manipulation intent that could be related to domain, scene, object, and attribute. The key challenge for ZS-CIR tasks is to learn a more accurate image representation that has adaptive attention to the reference image for various manipulation descriptions. In this paper, we propose a novel context-dependent mapping network, named Context-I2W, for adaptively converting description-relevant Image information into a pseudo-word token composed of the description for accurate ZS-CIR. Specifically, an Intent View Selector first dynamically learns a rotation rule to map the identical image to a task-specific manipulation view. Then a Visual Target Extractor further captures local information covering the main targets in ZS-CIR tasks under the guidance of multiple learnable queries. The two complementary modules work together to map an image to a context-dependent pseudo-word token without extra supervision. Our model shows strong generalization ability on four ZS-CIR tasks, including domain conversion, object composition, object manipulation, and attribute manipulation. It obtains consistent and significant performance boosts ranging from 1.88% to 3.60% over the best methods and achieves new state-of-the-art results on ZS-CIR. Our code is available at https://anonymous.4open.science/r/Context-I2W-4224/.

JBHI Journal 2024 Journal Article

Memory-Based Cross-Modal Semantic Alignment Network for Radiology Report Generation

  • Yitian Tao
  • Liyan Ma
  • Jing Yu
  • Han Zhang

Generating radiology reports automatically reduces the workload of radiologists and helps the diagnoses of specific diseases. Many existing methods take this task as modality transfer process. However, since the key information related to disease accounts for a small proportion in both image and report, it is hard for the model to learn the latent relation between the radiology image and its report, thus failing to generate fluent and accurate radiology reports. To tackle this problem, we propose a memory-based cross-modal semantic alignment model (MCSAM) following an encoder-decoder paradigm. MCSAM includes a well initialized long-term clinical memory bank to learn disease-related representations as well as prior knowledge for different modalities to retrieve and use the retrieved memory to perform feature consolidation. To ensure the semantic consistency of the retrieved cross modal prior knowledge, a cross-modal semantic alignment module (SAM) is proposed. SAM is also able to generate semantic visual feature embeddings which can be added to the decoder and benefits report generation. More importantly, to memorize the state and additional information while generating reports with the decoder, we use learnable memory tokens which can be seen as prompts. Extensive experiments demonstrate the promising performance of our proposed method which generates state-of-the-art performance on the MIMIC-CXR dataset.

NeurIPS Conference 2024 Conference Paper

PGN: The RNN's New Successor is Effective for Long-Range Time Series Forecasting

  • Yuxin Jia
  • Youfang Lin
  • Jing Yu
  • Shuo Wang
  • Tianhao Liu
  • Huaiyu Wan

Due to the recurrent structure of RNN, the long information propagation path poses limitations in capturing long-term dependencies, gradient explosion/vanishing issues, and inefficient sequential execution. Based on this, we propose a novel paradigm called Parallel Gated Network (PGN) as the new successor to RNN. PGN directly captures information from previous time steps through the designed Historical Information Extraction (HIE) layer and leverages gated mechanisms to select and fuse it with the current time step information. This reduces the information propagation path to $\mathcal{O}(1)$, effectively addressing the limitations of RNN. To enhance PGN's performance in long-range time series forecasting tasks, we propose a novel temporal modeling framework called Temporal PGN (TPGN). TPGN incorporates two branches to comprehensively capture the semantic information of time series. One branch utilizes PGN to capture long-term periodic patterns while preserving their local characteristics. The other branch employs patches to capture short-term information and aggregate the global representation of the series. TPGN achieves a theoretical complexity of $\mathcal{O}(\sqrt{L})$, ensuring efficiency in its operations. Experimental results on five benchmark datasets demonstrate the state-of-the-art (SOTA) performance and high efficiency of TPGN, further confirming the effectiveness of PGN as the new successor to RNN in long-range time series forecasting. The code is available in this repository: https: //github. com/Water2sea/TPGN.

ICRA Conference 2024 Conference Paper

Risk-Inspired Aerial Active Exploration for Enhancing Autonomous Driving of UGV in Unknown Off-Road Environments

  • Rongchuan Wang
  • Mengyin Fu
  • Jing Yu
  • Yi Yang 0009
  • Wenjie Song 0001

Unknown area exploration is a crucial but challenging task for autonomous driving of unmanned ground vehicles (UGV) in unknown off-road environments. However, the exploration efficiency of a single UGV is low due to its limited sensing range. To solve this problem, this paper proposes a risk-inspired aerial active exploration system, which utilizes the flexibility and field of view advantages of Unmanned Aerial Vehicles (UAV) to guide the UGV in unknown off-road environments. Firstly, a fast terrain risk mapping method that can be used for both UAV and UGV is developed. This method efficiently combines quadtree and hash table data structure to enable UAV to analyze large scale terrain point cloud in real time. Based on the risk mapping result, a risk-inspired active exploration method is proposed to actively search a safe reference path for the UGV, which introduces terrain risk information into the process of travel point selection. Finally, the reference path is gradually generated and optimized, so that the UGV can safely and smoothly follow the path to the target location. Compared with single UGV exploration system, our approach reduces the overall path risk by 26. 8% in simulated experiments, showing that the proposed system can enhance autonomous driving of the UGV and help it effectively avoid high-risk areas in unknown off-road environments.

TCS Journal 2022 Journal Article

Web of things based social media fake news classification with feature extraction using pre-trained convoluted recurrent network with deep fuzzy learning

  • Tingyu Li
  • Jing Yu
  • Haiping Zhang

Fake news has been surfacing often and in great quantity online recently due to burgeoning development of online social networks for various economic and political goals. Online social network users can quickly become infected by this fake news by using deceptive language, and it has already had a significant impact on offline culture. Finding bogus news quickly is a crucial step in raising the credibility of information in online social networks. This study proposes a fuzzy logic-based web of things categorization system for detection of fake news. Deep learning algorithms were used to perform the feature extraction. The first feature extraction step has been completed. Convoluted recurrent network that has been pre-trained (Pre-Tr_Conv_ReNet). High dimensionality indexing was applied following the extraction of the characteristics. The feature, whether it be an image or text, is then identified using a ranking method based on indexing metrics. The learned file has been annotated using deep fuzzy learning (De_Fuz_Lear) to distinguish fraudulent web content after this training. Then decision was taken and the fake news was detected. The simulation results give the classified outputs. For this obtained output the parametric analysis has been done.

IJCAI Conference 2021 Conference Paper

CogTree: Cognition Tree Loss for Unbiased Scene Graph Generation

  • Jing Yu
  • Yuan Chai
  • Yujing Wang
  • Yue Hu
  • Qi Wu

Scene graphs are semantic abstraction of images that encourage visual understanding and reasoning. However, the performance of Scene Graph Generation (SGG) is unsatisfactory when faced with biased data in real-world scenarios. Conventional debiasing research mainly studies from the view of balancing data distribution or learning unbiased models and representations, ignoring the correlations among the biased classes. In this work, we analyze this problem from a novel cognition perspective: automatically building a hierarchical cognitive structure from the biased predictions and navigating that hierarchy to locate the relationships, making the tail relationships receive more attention in a coarse-to-fine mode. To this end, we propose a novel debiasing Cognition Tree (CogTree) loss for unbiased SGG. We first build a cognitive structure CogTree to organize the relationships based on the prediction of a biased SGG model. The CogTree distinguishes remarkably different relationships at first and then focuses on a small portion of easily confused ones. Then, we propose a debiasing loss specially for this cognitive structure, which supports coarse-to-fine distinction for the correct relationships. The loss is model-agnostic and consistently boosting the performance of several state-of-the-art models. The code is available at: https: //github. com/CYVincent/Scene-Graph-Transformer-CogTree.

AAAI Conference 2021 Conference Paper

Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-Encoder

  • Yajing Sun
  • Yong Shan
  • Chengguang Tang
  • Yue Hu
  • Yinpei Dai
  • Jing Yu
  • Jian Sun
  • Fei Huang

It is important for task-oriented dialogue systems to discover the dialogue structure (i. e. the general dialogue flow) from dialogue corpora automatically. Previous work models dialogue structure by extracting latent states for each utterance first and then calculating the transition probabilities among states. These two-stage methods ignore the contextual information when calculating the probabilities, which makes the transitions between the states ambiguous. This paper proposes a conversational graph (CG) to represent deterministic dialogue structure where nodes and edges represent the utterance and context information respectively. An unsupervised Edge- Enhanced Graph Auto-Encoder (EGAE) architecture is designed to model local-contextual and global-structural information for conversational graph learning. Furthermore, a selfsupervised objective is introduced with the response selection task to guide the unsupervised learning of the dialogue structure. Experimental results on several public datasets demonstrate that the novel model outperforms several alternatives in aggregating utterances with similar semantics. The effectiveness of the learned dialogue structured is also verified by more than 5% joint accuracy improvement in the downstream task of low resource dialogue state tracking.

IJCAI Conference 2020 Conference Paper

DAM: Deliberation, Abandon and Memory Networks for Generating Detailed and Non-repetitive Responses in Visual Dialogue

  • Xiaoze Jiang
  • Jing Yu
  • Yajing Sun
  • Zengchang Qin
  • Zihao Zhu
  • Yue Hu
  • Qi Wu

Visual Dialogue task requires an agent to be engaged in a conversation with human about an image. The ability of generating detailed and non-repetitive responses is crucial for the agent to achieve human-like conversation. In this paper, we propose a novel generative decoding architecture to generate high-quality responses, which moves away from decoding the whole encoded semantics towards the design that advocates both transparency and flexibility. In this architecture, word generation is decomposed into a series of attention-based information selection steps, performed by the novel recurrent Deliberation, Abandon and Memory (DAM) module. Each DAM module performs an adaptive combination of the response-level semantics captured from the encoder and the word-level semantics specifically selected for generating each word. Therefore, the responses contain more detailed and non-repetitive descriptions while maintaining the semantic accuracy. Furthermore, DAM is flexible to cooperate with existing visual dialogue encoders and adaptive to the encoder structures by constraining the information selection mode in DAM. We apply DAM to three typical encoders and verify the performance on the VisDial v1. 0 dataset. Experimental results show that the proposed models achieve new state-of-the-art performance with high-quality responses. The code is available at https: //github. com/JXZe/DAM.

AAAI Conference 2020 Conference Paper

DualVD: An Adaptive Dual Encoding Model for Deep Visual Understanding in Visual Dialogue

  • Xiaoze Jiang
  • Jing Yu
  • Zengchang Qin
  • Yingying Zhuang
  • Xingxing Zhang
  • Yue Hu
  • Qi Wu

Different from Visual Question Answering task that requires to answer only one question about an image, Visual Dialogue involves multiple questions which cover a broad range of visual content that could be related to any objects, relationships or semantics. The key challenge in Visual Dialogue task is thus to learn a more comprehensive and semantic-rich image representation which may have adaptive attentions on the image for variant questions. In this research, we propose a novel model to depict an image from both visual and semantic perspectives. Specifically, the visual view helps capture the appearance-level information, including objects and their relationships, while the semantic view enables the agent to understand high-level visual semantics from the whole image to the local regions. Futhermore, on top of such multiview image features, we propose a feature selection framework which is able to adaptively capture question-relevant information hierarchically in fine-grained level. The proposed method achieved state-of-the-art results on benchmark Visual Dialogue datasets. More importantly, we can tell which modality (visual or semantic) has more contribution in answering the current question by visualizing the gate values. It gives us insights in understanding of human cognition in Visual Dialogue.

NeurIPS Conference 2020 Conference Paper

End-to-End Learning and Intervention in Games

  • Jiayang Li
  • Jing Yu
  • Yu Nie
  • Zhaoran Wang

In a social system, the self-interest of agents can be detrimental to the collective good, sometimes leading to social dilemmas. To resolve such a conflict, a central designer may intervene by either redesigning the system or incentivizing the agents to change their behaviors. To be effective, the designer must anticipate how the agents react to the intervention, which is dictated by their often unknown payoff functions. Therefore, learning about the agents is a prerequisite for intervention. In this paper, we provide a unified framework for learning and intervention in games. We cast the equilibria of games as individual layers and integrate them into an end-to-end optimization framework. To enable the backward propagation through the equilibria of games, we propose two approaches, respectively based on explicit and implicit differentiation. Specifically, we cast the equilibria as the solutions to variational inequalities (VIs). The explicit approach unrolls the projection method for solving VIs, while the implicit approach exploits the sensitivity of the solutions to VIs. At the core of both approaches is the differentiation through a projection operator. Moreover, we establish the correctness of both approaches and identify the conditions under which one approach is more desirable than the other. The analytical results are validated using several real-world problems.

AAAI Conference 2020 Conference Paper

History-Adaption Knowledge Incorporation Mechanism for Multi-Turn Dialogue System

  • Yajing Sun
  • Yue Hu
  • Luxi Xing
  • Jing Yu
  • Yuqiang Xie

Keeping the conversation consistent and avoiding its repetition are two key factors to construct an intelligent multiturn knowledge-grounded dialogue system. Although some works tend to combine history with external knowledge such as personal background information to boost dialogue quality, they are prone to ignore the fact that incorporating the same knowledge multiple times into the conversation leads to repetition. The main reason is the lack of effective control over the use of knowledge on the conversation level. So we design a history-adaption knowledge incorporation mechanism to build an effective multi-turn dialogue model. Our proposed model addresses repetition by recurrently updating the knowledge from the conversation level and progressively incorporating it into the history step-by-step. And the knowledge-grounded history representation also enhances the conversation consistency. Experimental results show that our proposed model significantly outperforms several retrievalbased models on some benchmark datasets. The human evaluation demonstrates that our model can maintain conversation consistent and reduce conversation repetition.

IJCAI Conference 2020 Conference Paper

Mucko: Multi-Layer Cross-Modal Knowledge Reasoning for Fact-based Visual Question Answering

  • Zihao Zhu
  • Jing Yu
  • Yujing Wang
  • Yajing Sun
  • Yue Hu
  • Qi Wu

Fact-based Visual Question Answering (FVQA) requires external knowledge beyond the visible content to answer questions about an image. This ability is challenging but indispensable to achieve general VQA. One limitation of existing FVQA solutions is that they jointly embed all kinds of information without fine-grained selection, which introduces unexpected noises for reasoning the final answer. How to capture the question-oriented and information-complementary evidence remains a key challenge to solve the problem. In this paper, we depict an image by a multi-modal heterogeneous graph, which contains multiple layers of information corresponding to the visual, semantic and factual features. On top of the multi-layer graph representations, we propose a modality-aware heterogeneous graph convolutional network to capture evidence from different layers that is most relevant to the given question. Specifically, the intra-modal graph convolution selects evidence from each modality and cross-modal graph convolution aggregates relevant information across different graph layers. By stacking this process multiple times, our model performs iterative reasoning across three modalities and predicts the optimal answer by analyzing all question-oriented evidence. We achieve a new state-of-the-art performance on the FVQA task and demonstrate the effectiveness and interpretability of our model with extensive experiments.

v2026.09.13