Arrow Research search

Author name cluster

Ziqi Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

15 papers
2 author rows

Possible papers

15

AAAI Conference 2026 Conference Paper

MMhops-R1: Multimodal Multi-hop Reasoning

  • Tao Zhang
  • Ziqi Zhang
  • Zongyang Ma
  • Yuxin Chen
  • Bing Li
  • Chunfeng Yuan
  • Guangting Wang
  • Fengyun Rao

The ability to perform multi-modal multi-hop reasoning by iteratively integrating information across various modalities and external knowledge is critical for addressing complex real-world challenges. However, existing Multi-modal Large Language Models (MLLMs) are predominantly limited to single-step reasoning, as existing benchmarks lack the complexity needed to evaluate and drive multi-hop abilities. To bridge this gap, we introduce MMhops, a novel, large-scale benchmark designed to systematically evaluate and foster multi-modal multi-hop reasoning. MMhops dataset comprises two challenging task formats, Bridging and Comparison, which necessitate that models dynamically construct complex reasoning chains by integrating external knowledge. To tackle the challenges posed by MMhops, we propose MMhops-R1, a novel multi-modal Retrieval-Augmented Generation (mRAG) framework for dynamic reasoning. Our framework utilizes reinforcement learning to optimize the model for autonomously planning reasoning paths, formulating targeted queries, and synthesizing multi-level information. Comprehensive experiments demonstrate that MMhops-R1 significantly outperforms strong baselines on MMhops, highlighting that dynamic planning and multi-modal knowledge integration are crucial for complex reasoning. Moreover, MMhops-R1 demonstrates strong generalization to tasks requiring fixed-hop reasoning, underscoring the robustness of our dynamic planning approach.

AAAI Conference 2026 Conference Paper

PatchET: Learning Enzyme Temperature Properties Through Patch-Based Neural Architectures

  • Ziqi Zhang
  • Runze Yang
  • Longbing Cao
  • Zhaohong Deng

Understanding enzyme thermal properties is essential for biotechnology and protein engineering, yet experimental measurements of attributes such as temperature optimum, stability, and range remain labor-intensive and costly. Prior studies have shown that specific regions within enzyme sequences disproportionately influence thermal behavior—an aspect often overlooked by existing deep learning models. In this work, we introduce PatchET, a biologically inspired deep learning model that predicts enzyme thermal properties directly from amino acid sequences. PatchET employs a dual-stage, patch-based architecture that captures both intra-patch local features and inter-patch global dependencies, reflecting the hierarchical nature of protein thermal adaptation. Alongside the model, we curate a comprehensive benchmark, including a refined dataset for temperature optimum and the first publicly available dataset for temperature range prediction. PatchET achieves state-of-the-art performance across three key tasks—temperature optimum, stability, and range—and serves as the first dedicated model for temperature range prediction. Extensive ablation studies further validate the effectiveness of our architectural design. Together, PatchET and the accompanying benchmark provide a unified and generalizable framework for modeling enzyme thermal properties, offering new tools for the rational design of thermostable enzymes.

NeurIPS Conference 2025 Conference Paper

AegisGuard: RL-Guided Adapter Tuning for TEE-Based Efficient & Secure On-Device Inference

  • CHE WANG
  • Ziqi Zhang
  • Yinggui Wang
  • Tiantong Wang
  • Yurong Hao
  • Jianbo Gao
  • Tao Wei
  • Yang Cao

On-device large models (LMs) reduce cloud dependency but expose proprietary model weights to the end-user, making them vulnerable to white-box model stealing (MS) attacks. A common defense is TEE-Shielded DNN Partition (TSDP), which places all trainable LoRA adapters (fine tuned on private data) inside a trusted execution environment (TEE). However, this design suffers from excessive host-to-TEE communication latency. We propose AegisGuard, a fine tuning and deployment framework that selectively shields the MS sensitive adapters while offloading the rest to the GPU, balancing security and efficiency. AegisGuard integrates two key components: i) RL-based Sensitivity Measurement (RSM), which injects Gaussian noise during training and applies a lightweight reinforcement learning to rank adapters based on their impact on model stealing; and (ii) Shielded-Adapter Compression (SAC), which structurally prunes the selected adapters to reduce both parameter size and intermediate feature maps, further lowering TEE computation and data transfer costs. Extensive experiments demonstrate that AegisGuard achieves black-box level MS resilience (surrogate accuracy around 39%, matching fully shielded baselines), while reducing end-to-end inference latency by 2–3× and cutting TEE memory usage by 4× compared to state-of-the-art TSDP methods.

PRL Workshop 2025 Workshop Paper

ContextFormer: Stitching via Expert Calibration

  • Ziqi Zhang
  • Jingzehua Xu
  • Jinxin Liu
  • Zifeng Zhuang
  • Donglin Wang
  • Miao Liu
  • Shuai Zhang

Offline reinforcement learning (RL) algorithms can learn better decision-making compared to behavior policies by stitching the suboptimal trajectories to derive more optimal ones. Meanwhile, Decision Transformer (DT) abstracts the RL as sequence modeling, showcasing competitive performance on offline RL benchmarks. However, recent studies demonstrate that DT lacks of stitching capacity, thus exploiting stitching capability for DT is vital to further improve its performance. In order to endow stitching capability to DT, we abstract trajectory stitching as divergent sequential expert matching and introduce our approach, ContextFormer, which integrates contextual information-based imitation learning (IL) and sequence modeling to stitch sub-optimal trajectory fragments by emulating the representations of a limited number of expert trajectories. To validate our approach, we conduct experiments from two perspectives: 1) We conduct extensive experiments on D4RL benchmarks under the settings of IL, and experimental results demonstrate ContextFormer can achieve competitive performance in multiple IL settings. 2) More importantly, we conduct a comparison of ContextFormer with various competitive DT variants using identical training datasets. The experimental results unveiled ContextFormer’s superiority, as it outperformed all other variants, showcasing its remarkable performance.

IROS Conference 2025 Conference Paper

Design and Performance Analysis of a Pipeline Crawling Robot Based on Spring-Roll Dielectric Elastomer Actuators

  • Qinghai Zhang
  • Wei Yu
  • Ziqi Zhang
  • Jianghua Zhao
  • Shijie Guo

With the increasing complexity of pipeline systems in various industrial and environmental applications, there is a critical need for flexible and efficient robotic solutions that can navigate and inspect confined spaces. This paper introduces a lightweight pipeline crawling robot based on spring-roll dielectric elastomer actuators (DEAs). Inspired by the adaptability of caterpillars, the robot combines anisotropic friction feet with a spring-roll DEA structure to achieve high-speed movement. It operates effectively in pipes with diameters ranging from 16 mm to 20 mm, reaching a maximum speed of 357 mm/s (5. 95 BL/s) under a 3. 5 kV driving voltage. The optimized design enhances actuator performance and friction distribution, significantly outperforming existing soft crawling robots. This innovation demonstrates great potential for high-speed, lightweight pipeline inspection applications and advances the field of soft robotics for diverse industrial tasks.

IJCAI Conference 2025 Conference Paper

Outstanding Orthodontist: No More Artifactual Teeth in Talking Face

  • Zibo Su
  • Ziqi Zhang
  • Kun Wei
  • Xu Yang
  • Cheng Deng

Audio-driven talking face synthesis (TFS) enables the creation of realistic speaking videos by combining a single facial image with a speech audio clip. Unlike other facial features that naturally deform during speech, teeth represent unique rigid structures whose shape and size should remain constant throughout the video sequence. However, current methods often produce temporal inconsistencies and artifacts in the teeth region, resulting in a less realistic appearance of the generated videos. To address this, we propose OrthoNet, a plug-and-play framework designed to eliminate unrealistic teeth effects in audio-driven TFS. Our method introduces a Detail-oriented Teeth Aligner module, designed to preserve teeth details and adapt to their shape. It works with a Memory-guided Teeth Stabilizer that integrates a long-term memory bank for global teeth structure and a short-term memory module for local temporal dynamics. Through this framework, OrthoNet acts like an orthodontist for existing Audio2Video methods, ensuring that teeth maintain natural rigidity and temporal consistency even under varying degrees of teeth occlusion. Extensive experiments demonstrate that our method makes the teeth in generated videos appear more natural during speech, significantly enhancing the temporal consistency and structural stability of audio-driven video generation.

NeurIPS Conference 2025 Conference Paper

SEC-bench: Automated Benchmarking of LLM Agents on Real-World Software Security Tasks

  • Hwiwon Lee
  • Ziqi Zhang
  • Hanxiao Lu
  • LINGMING ZHANG

Rigorous security-focused evaluation of large language model (LLM) agents is imperative for establishing trust in their safe deployment throughout the software development lifecycle. However, existing benchmarks largely rely on synthetic challenges or simplified vulnerability datasets that fail to capture the complexity and ambiguity encountered by security engineers in practice. We introduce SEC-bench, the first fully automated benchmarking framework for evaluating LLM agents on authentic security engineering tasks. SEC-bench employs a novel multi-agent scaffold that automatically constructs code repositories with harnesses, reproduces vulnerabilities in isolated environments, and generates gold patches for reliable evaluation. Our framework automatically creates high-quality software vulnerability datasets with reproducible artifacts at a cost of only $0. 87 per instance. Using SEC-bench, we implement two critical software security tasks to rigorously evaluate LLM agents' capabilities: proof-of-concept (PoC) generation and vulnerability patching. A comprehensive evaluation of state-of-the-art LLM code agents reveals significant performance gaps, achieving at most 18. 0% success in PoC generation and 34. 0% in vulnerability patching on our complete dataset. These results highlight the crucial steps needed toward developing LLM agents that are more practical, intelligent, and autonomous for security engineering.

NeurIPS Conference 2024 Conference Paper

ACFun: Abstract-Concrete Fusion Facial Stylization

  • Jiapeng Ji
  • Kun Wei
  • Ziqi Zhang
  • Cheng Deng

Owing to advancements in image synthesis techniques, stylization methodologies for large models have garnered remarkable outcomes. However, when it comes to processing facial images, the outcomes frequently fall short of expectations. Facial stylization is predominantly challenged by two significant hurdles. Firstly, obtaining a large dataset of high-quality stylized images is difficult. The scarcity and diversity of artistic styles make it impractical to compile comprehensive datasets for each style. Secondly, while many methods can transfer colors and strokes from style images, these elements alone cannot fully capture a specific style, which encompasses both concrete and abstract visual elements. Additionally, facial stylization often alters the visual features of the face, making it challenging to balance these changes with the need to retain facial information. To address these issues, we propose a novel method called ACFun, which uses only one style image and one facial image for facial stylization. ACFun comprises an Abstract Fusion Module (AFun) and a Concrete Fusion Module (CFun), which separately learn the abstract and concrete features of the style and face. We also design a Face and Style Imagery Alignment Loss to align the style image with the face image in the latent space. Finally, we generate styled facial images from noise directly to complete the facial stylization task. Experiments show that our method outperforms others in facial stylization, producing highly artistic and visually pleasing results.

AAAI Conference 2024 Conference Paper

Beyond OOD State Actions: Supported Cross-Domain Offline Reinforcement Learning

  • Jinxin Liu
  • Ziqi Zhang
  • Zhenyu Wei
  • Zifeng Zhuang
  • Yachen Kang
  • Sibo Gai
  • Donglin Wang

Offline reinforcement learning (RL) aims to learn a policy using only pre-collected and fixed data. Although avoiding the time-consuming online interactions in RL, it poses challenges for out-of-distribution (OOD) state actions and often suffers from data inefficiency for training. Despite many efforts being devoted to addressing OOD state actions, the latter (data inefficiency) receives little attention in offline RL. To address this, this paper proposes the cross-domain offline RL, which assumes offline data incorporate additional source-domain data from varying transition dynamics (environments), and expects it to contribute to the offline data efficiency. To do so, we identify a new challenge of OOD transition dynamics, beyond the common OOD state actions issue, when utilizing cross-domain offline data. Then, we propose our method BOSA, which employs two support-constrained objectives to address the above OOD issues. Through extensive experiments in the cross-domain offline RL setting, we demonstrate BOSA can greatly improve offline data efficiency: using only 10% of the target data, BOSA could achieve 74.4% of the SOTA offline RL performance that uses 100% of the target data. Additionally, we also show BOSA can be effortlessly plugged into model-based offline RL and noising data augmentation techniques (used for generating source-domain data), which naturally avoids the potential dynamics mismatch between target-domain data and newly generated source-domain data.

ICML Conference 2024 Conference Paper

GroupCover: A Secure, Efficient and Scalable Inference Framework for On-device Model Protection based on TEEs

  • Zheng Zhang
  • Na Wang 0003
  • Ziqi Zhang
  • Yao Zhang
  • Tianyi Zhang
  • Jianwei Liu 0001
  • Ye Wu

Due to the high cost of training DNN models, how to protect the intellectual property of DNN models, especially when the models are deployed to users’ devices, is becoming an important topic. One practical solution is to use Trusted Execution Environments (TEEs) and researchers have proposed various model obfuscation solutions to make full use of the high-security guarantee of TEEs and the high performance of collocated GPUs. In this paper, we first identify a common vulnerability, namely the fragility of randomness, that is shared by existing TEE-based model obfuscation solutions. This vulnerability benefits model-stealing attacks and allows the adversary to recover about 97% of the secret model. To improve the security of TEE-shielded DNN models, we further propose a new model obfuscation approach GroupCover, which uses sufficient randomization and mutual covering obfuscation to protect model weights. Experimental results demonstrate that GroupCover can achieve a comparable security level as the upper-bound (black-box protection), which is remarkably over 3x compared with existing solutions. Besides, GroupCover introduces 19% overhead and negligible accuracy loss compared to model unprotected scheme.

ICML Conference 2024 Conference Paper

Reinformer: Max-Return Sequence Modeling for Offline RL

  • Zifeng Zhuang
  • Dengyun Peng
  • Jinxin Liu
  • Ziqi Zhang
  • Donglin Wang

As a data-driven paradigm, offline reinforcement learning (RL) has been formulated as sequence modeling that conditions on the hindsight information including returns, goal or future trajectory. Although promising, this supervised paradigm overlooks the core objective of RL that maximizes the return. This overlook directly leads to the lack of trajectory stitching capability that affects the sequence model learning from sub-optimal data. In this work, we introduce the concept of max-return sequence modeling which integrates the goal of maximizing returns into existing sequence models. We propose Rein for ced Trans for mer ( Rein for mer ), indicating the sequence model is reinforced by the RL objective. Rein for mer additionally incorporates the objective of maximizing returns in the training phase, aiming to predict the maximum future return within the distribution. During inference, this in-distribution maximum return will guide the selection of optimal actions. Empirically, Rein for mer is competitive with classical RL methods on the D4RL benchmark and outperforms state-of-the-art sequence model particularly in trajectory stitching ability. Code is public at https: //github. com/Dragon-Zhuang/Reinformer.

EAAI Journal 2024 Journal Article

Research on data-driven model for power grid fault diagnosis fusing topological quantification information

  • Xu Zhang
  • Zirui Wang
  • Mingxuan Du
  • Xuekui Mao
  • Ruiting Ding
  • Haoran Yu
  • Ziqi Zhang

In increasingly complex power grid operation scenarios and fault modes, rapid and accurate fault identification is highly important for improving the reliability of power systems. Faced with the massive amount of available power grid data, the rapid development of artificial intelligence technology provides a powerful tool for power grid fault diagnosis. However, existing data-driven diagnosis methods lack quantitative power grid topology change representations, cannot integrate power grid fault topology with alarm information, and have limited effectiveness in terms of diagnosing complex faults. To address these issues, a data-driven power grid fault diagnosis model that integrates topological quantitative features is proposed. By studying the changes in the topological connection relationships of equipment before and after a power grid fault and the topological connectivity in a power failure zone, based on the basic topological network indicators in graph theory, a representation of the quantitative features of the power grid fault topology is achieved, and these quantitative features and the alarm information are integrated to construct a data-driven power grid fault diagnosis model. The model uses the light gradient boosting machine algorithm to determine the types of complex faults and accurately identify faulty equipment, addressing the lack of model diagnosis effects considered in previous studies. Finally, the accuracy and effectiveness of the model are verified by using simulated fault cases.

AAAI Conference 2024 Conference Paper

Set Prediction Guided by Semantic Concepts for Diverse Video Captioning

  • Yifan Lu
  • Ziqi Zhang
  • Chunfeng Yuan
  • Peng Li
  • Yan Wang
  • Bing Li
  • Weiming Hu

Diverse video captioning aims to generate a set of sentences to describe the given video in various aspects. Mainstream methods are trained with independent pairs of a video and a caption from its ground-truth set without exploiting the intra-set relationship, resulting in low diversity of generated captions. Different from them, we formulate diverse captioning into a semantic-concept-guided set prediction (SCG-SP) problem by fitting the predicted caption set to the ground-truth set, where the set-level relationship is fully captured. Specifically, our set prediction consists of two synergistic tasks, i.e., caption generation and an auxiliary task of concept combination prediction providing extra semantic supervision. Each caption in the set is attached to a concept combination indicating the primary semantic content of the caption and facilitating element alignment in set prediction. Furthermore, we apply a diversity regularization term on concepts to encourage the model to generate semantically diverse captions with various concept combinations. These two tasks share multiple semantics-specific encodings as input, which are obtained by iterative interaction between visual features and conceptual queries. The correspondence between the generated captions and specific concept combinations further guarantees the interpretability of our model. Extensive experiments on benchmark datasets show that the proposed SCG-SP achieves state-of-the-art (SOTA) performance under both relevance and diversity metrics.

NeurIPS Conference 2023 Conference Paper

Exploiting Contextual Objects and Relations for 3D Visual Grounding

  • Li Yang
  • Chunfeng Yuan
  • Ziqi Zhang
  • Zhongang Qi
  • Yan Xu
  • Wei Liu
  • Ying Shan
  • Bing Li

3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to understand and engage with the real-world environment. However, this task is challenging due to the necessity to capture 3D contextual information to distinguish target objects from complex 3D scenes. The absence of annotations for contextual objects and relations further exacerbates the difficulties. In this paper, we propose a novel model, CORE-3DVG, to address these challenges by explicitly learning about contextual objects and relations. Our method accomplishes 3D visual grounding via three sequential modular networks, including a text-guided object detection network, a relation matching network, and a target identification network. During training, we introduce a pseudo-label self-generation strategy and a weakly-supervised method to facilitate the learning of contextual objects and relations, respectively. The proposed techniques allow the networks to focus more effectively on referred objects within 3D scenes by understanding their context better. We validate our model on the challenging Nr3D, Sr3D, and ScanRefer datasets and demonstrate state-of-the-art performance. Our code will be public at https: //github. com/yangli18/CORE-3DVG.

PRL Workshop 2021 Workshop Paper

Extending Graph Neural Networks for Generalized Stochastic Planning

  • Ziqi Zhang
  • Florian Geißer

Probabilistic planning problems are often formulated in terms of a domain class that describes the general problem structure, and in terms of an instantiation of the domain, which yields a particular MDP. Most algorithms in probabilistic planning focus on computing policies for a single MDP. A relational policy represents a solution to any MDP induced from the same domain class. We present a graph neural network architecture that is based on a representation of the relation between types of a given domain and which allows us to generalize from small instances to large instances of the same domain class. Unlike other work, we do not impose restrictions on the structure of the domain. We conduct a preliminary study which compares the relational policies obtained from a network trained on small instances against a policy computed by a state-of-the-art domain-independent planner. The evaluation shows that the network generalizes well across instances of a domain, and is even able to outperform the instance-dependent policy in some of the benchmarks.

v2026.09.13