Arrow Research search

Author name cluster

Jiawei Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

32 papers
2 author rows

Possible papers

32

AAAI Conference 2026 Conference Paper

Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

  • Liyang Chen
  • Tianxiang Ma
  • Jiawei Liu
  • Bingchuan Li
  • Zhuowei Chen
  • Lijie Liu
  • Xu He
  • Gen Li

Human-Centric Video Generation (HCVG) methods seek to synthesize human videos from multimodal inputs, including text, images, and audio. Existing methods struggle to effectively coordinate these heterogeneous modalities due to two challenges: the scarcity of modality-complete data and the difficulty of jointly modeling triplet conditions without performance degradation. In this work, we present HuMo, a unified HCVG framework for collaborative multimodal control. For the first challenge, we construct an incomplete-yet-complementary dataset for improved data utilization efficiency and training scalability. For the second challenge, we propose a two-stage progressive multimodal training paradigm with task-specific strategies at each stage. In the first stage, to balance the text-following and subject-preservation abilities, we adopt the minimal-invasive image injection strategy. In the second stage, to enhance audio-visual sync, we propose a focus-by-predicting strategy that implicitly guides the model to associate audio with facial regions. For joint learning of controllabilities across multi-modal inputs, we progressively incorporate the audio-visual sync task, building on previously acquired capabilities. During inference, for flexible and fine-grained multimodal control, we design a stage-adaptive Classifier-Free Guidance strategy that dynamically adjusts guidance weights across denoising steps. Extensive experimental results demonstrate that HuMo surpasses specialized state-of-the-art methods in sub-tasks, establishing a unified framework for collaborative multimodal-conditioned HCVG.

AAAI Conference 2026 Conference Paper

Unleashing the Potential of Large Language Models for Text-to-Image Generation Through Autoregressive Representation Alignment

  • Xing Xie
  • Jiawei Liu
  • Ziyue Lin
  • Huijie Fan
  • Zhi Han
  • Yandong Tang
  • Liangqiong Qu

We present Autoregressive Representation Alignment (ARRA), a new training framework that unlocks global-coherent text-to-image generation in autoregressive LLMs without architectural modifications. Different from prior works that require complex architectural redesigns, ARRA aligns LLM's hidden states with visual representations from external visual foundational models via a global visual alignment loss and a hybrid token,. This token enforces dual constraints: local next-token prediction and global semantic distillation, enabling LLMs to implicitly learn spatial and contextual coherence while retaining their original autoregressive paradigm. Extensive experiments validate ARRA's plug-and-play versatility. When training T2I LLMs from scratch, ARRA reduces FID by 16.6% (ImageNet), 12.0% (LAION-COCO) for autoregressive LLMs like LlamaGen, without modifying original architecture and inference mechanism. For training from text-generation-only LLMs, ARRA reduces FID by 25.5% (MIMIC-CXR), 8.8% (DeepEyeNet) for advanced LLMs like Chameleon. For domain adaptation, ARRA aligns general-purpose LLMs with specialized models (e.g., BioMedCLIP), achieving an 18.6% FID reduction over direct fine-tuning on medical imaging (MIMIC-CXR). These results demonstrate that training objective redesign, rather than architectural modifications, can resolve cross-modal global coherence challenges. ARRA offers a complementary paradigm for advancing autoregressive models.

EAAI Journal 2025 Journal Article

A unified rotating machinery health management framework leveraging large language models for diverse components, conditions, and tasks

  • Haotian Peng
  • Jie Gao
  • Jiawei Liu
  • Jinsong Du
  • Wei Wang

This study introduces the Rotating Machinery Large Language Model (RotLLM), a unified framework for rotating machinery health management that integrates deep learning with large language models (LLMs) to address diverse operational conditions, components, and health management tasks. RotLLM employs a novel Spectral Folding Network (SFN) to transform vibration spectrum into a unified feature space that preserves essential health state information. A dedicated projection layer then maps these features into the semantic domain of an LLM. The framework is trained using a three-stage strategy: first, pre-training the encoder on the Large-scale Multimodal Rotating Machinery (LMR) dataset, which comprises 237, 298 vibration samples collected under hundreds of operating conditions; second, initializing the projection layer with textual health state labels; and finally, fine-tuning using parameter-efficient Low-Rank Adaptation (LoRA) with high-quality corpus for various health management tasks. Experimental evaluations demonstrate that RotLLM achieves state-of-the-art performance in fault classification, maintains strong robustness under noisy conditions, and delivers rapid multi-task inference with minimal computational overhead. The framework consistently outperforms conventional methods, enabling efficient, accurate, and context-aware health management for rotating machinery across diverse conditions and tasks. The dataset and source code are open-sourced (https: //github. com/SIA-IDE/RotLLM), fostering collaboration, reproducibility, and broader adoption in industrial prognostics research.

AAAI Conference 2025 Conference Paper

BearLLM: A Prior Knowledge-Enhanced Bearing Health Management Framework with Unified Vibration Signal Representation

  • Haotian Peng
  • Jiawei Liu
  • Jinsong Du
  • Jie Gao
  • Wei Wang

We propose a bearing health management framework leveraging large language models (BearLLM), a novel multimodal model that unifies multiple bearing-related tasks by processing user prompts and vibration signals. Specifically, we introduce a prior knowledge-enhanced unified vibration signal representation to handle various working conditions across multiple datasets. This involves adaptively sampling the vibration signals based on the sampling rate of the sensor, incorporating the frequency domain to unify input dimensions, and using a fault-free reference signal as an auxiliary input. To extract features from vibration signals, we first train a fault classification network, then convert and align the extracted features into word embedding, and finally concatenate these with text embedding as input to an LLM. To evaluate the performance of the proposed method, we constructed the first large-scale multimodal bearing health management (MBHM) dataset, including paired vibration signals and textual descriptions. With our unified vibration signal representation, BearLLM using one set of pre-trained weights achieves state-of-the-art performance on nine publicly available fault diagnosis benchmarks, outperforming specific methods designed for individual datasets. We provide a dataset, our model, and code to inspire future research on building more capable industrial multimodal models.

IROS Conference 2025 Conference Paper

Dual-Arm Teleoperated Robotic Microsurgery System with Live Volumetric OCT Image Feedback

  • Jiawei Liu
  • Guangshen Ma
  • Genggeng Zhou
  • Haochi Pan
  • Colin Lam
  • Catherine Jin
  • Nita Valikodath
  • Mark Draelos

In microsurgery, surgeons frequently encounter challenges due to the need for exceptional precision and dexterity, the lack of depth perception for micro-scale surgical maneuvers, and the inevitable effects of fatigue and hand tremor. In surgical robotics, conventional intraoperative perception systems normally provide real-time image feedback, but depth and volumetric information is typically lacking. To overcome these challenges, we propose a teleoperated robotic system with two arms to provide high-fidelity intraoperative volumetric imaging during micro-scale tissue manipulation. This system incorporates an optical coherence tomography sensor for real-time 3D visualization and a dual-arm teleoperated robot system controlled by haptic input devices for accurate and precise manipulation. We characterize the system’s performance through a precision positioning task and a vessel following task in a retinal model, which shows average positioning errors of approximately 232 μm and 83 μm, respectively. We demonstrate the fully integrated system through the completion of an eggshell membrane peeling task that simulates retinal membrane peeling.

NeurIPS Conference 2025 Conference Paper

Fact-R1: Towards Explainable Video Misinformation Detection with Deep Reasoning

  • Fanrui Zhang
  • Dian Li
  • Qiang Zhang
  • Junxiong Lin
  • Jiahong Yan
  • Jiawei Liu
  • Zheng-Jun Zha

The rapid spread of multimodal misinformation on social media has raised growing concerns, while research on video misinformation detection remains limited due to the lack of large-scale, diverse datasets. Existing methods often overfit to rigid templates and lack deep reasoning over deceptive content. To address these challenges, we introduce FakeVV, a large-scale benchmark comprising over 100, 000 video-text pairs with fine-grained, interpretable annotations. In addition, we further propose Fact-R1, a novel framework that integrates deep reasoning with collaborative rule-based reinforcement learning. Fact-R1 is trained through a three-stage process: (1) misinformation long-Chain-of-Thought (CoT) instruction tuning, (2) preference alignment via Direct Preference Optimization (DPO), and (3) Group Relative Policy Optimization (GRPO) using a novel verifiable reward function. This enables Fact-R1 to exhibit emergent reasoning behaviors comparable to those observed in advanced text-based reinforcement learning systems, but in the more complex multimodal misinformation setting. Our work establishes a new paradigm for misinformation detection, bridging large-scale video understanding, reasoning-guided alignment, and interpretable verification.

AAAI Conference 2025 Conference Paper

HOIMamba: Efficient Mamba-based Disentangled Progressive Learning for HOI Detection

  • Yongchao Xu
  • Jiawei Liu
  • Sen Tao
  • Qiang Zhang
  • Zheng-Jun Zha

Human-object interaction (HOI) detection aims to detect the spatial positions of human-object pairs and recognize their interactions. Existing single-branch, two-branch, and three-branch methods are challenging to make an appropriate trade-off on efficiency, multi-task decoupling, and collaborative learning, while they fail to identify rare and complex interaction categories effectively as well. In this work, we propose a novel Efficient Mamba-based Disentangled Progressive Learning (HOIMamba) for HOI Detection to absorb the advantages of the existing three approaches and adaptively aggregate multi-level interaction semantics guided by cross-task bidirectional information contexts. Specifically, HOIMamba builds an efficient and effective decoder through cascaded Low-Rank Adaptations (LoRAs), with high efficiency, thorough decoupling of tasks, and good multi-task collaborative learning. Furthermore, to alleviate the recognition problem of interactions in difficult HOI samples, a novel Mamba-based comprehensive progressive learning strategy with Cross-enhance Mamba (CEM) blocks and Detection Context Propagation (DCP) blocks is designed to gradually excavate interaction-related discriminative cues from four levels. CEM blocks automatically aggregate context to generate diverse task-shared semantics and simultaneously realize the cross-task interaction between human and object branches, guiding the interaction branch to extract more expressive HOI representation. DCP blocks further transfer the comprehensive interaction context to human and object branches to achieve rich and effective information exchange, facilitating the model to discover more HOI instances. Extensive experimental results on two standard benchmarks demonstrate the effectiveness of our HOIMamba.

AAAI Conference 2025 Conference Paper

Interweaving Memories of a Siamese Large Language Model

  • Xin Song
  • Zhikai Xue
  • Guoxiu He
  • Jiawei Liu
  • Wei Lu

Parameter-efficient fine-tuning (PEFT) methods optimize large language models (LLMs) by modifying or introducing a small number of parameters to enhance alignment with downstream tasks. However, they can result in catastrophic forgetting, where LLMs prioritize new knowledge at the expense of comprehensive world knowledge. A promising approach to mitigate this issue is to recall prior memories based on the original knowledge. To this end, we propose a model-agnostic PEFT framework, IMSM, which Interweaves Memories of a Siamese Large Language Model. Specifically, our siamese LLM is equipped with an existing PEFT method. Given an incoming query, it generates two distinct memories based on the pre-trained and fine-tuned parameters. IMSM then incorporates an interweaving mechanism that regulates the contributions of both original and enhanced memories when generating the next token. This framework is theoretically applicable to all open-source LLMs and existing PEFT methods. We conduct extensive experiments across various benchmark datasets, evaluating the performance of popular open-source LLMs using the proposed IMSM, in comparison to both classical and leading PEFT methods. Our findings indicate that IMSM maintains comparable time and space efficiency to backbone PEFT methods while significantly improving performance and effectively mitigating catastrophic forgetting.

IJCAI Conference 2025 Conference Paper

Learnable Frequency Decomposition for Image Forgery Detection and Localization

  • Dong Li
  • Jiayíng Zhu
  • Yidi Liu
  • Xin Lu
  • Xueyang Fu
  • Jiawei Liu
  • Aiping Liu
  • Zheng-Jun Zha

Concern for image authenticity spurs research in image forgery detection and localization (IFDL). Most deep learning-based methods focus primarily on spatial domain modeling and have not fully explored frequency domain strategies. In this paper, we observe and analyze the frequency characteristic changes caused by image tampering. Observations indicate that manipulation traces are especially prominent in phase components and span both low and high-frequency bands. Based on these findings, we propose a forensic frequency decomposition network (F2D-Net), which incorporates deep Fourier transforms and leverages both phase information and high and low-frequency components to enhance IFDL. Specifically, F2D-Net consists of the Spectral Decomposition Subnetwork (SDSN) and the Frequency Separation Subnetwork (FSSN). The former decomposes the image into amplitude and phase, focusing on learning the semantic content in the phase spectrum to identify forged objects, thus improving forgery detection accuracy. The latter further adaptively decomposes the output of the SDSN to obtain corresponding high and low frequencies, and applies a divide-and-conquer strategy to refine each frequency band, mitigating the optimization difficulties caused by coupled forgery traces across different frequencies, thereby better capturing the pixels belonging to the forged object to improve localization accuracy. Experiments on multiple datasets demonstrate that our method outperforms state-of-the-art image forgery detection and localization techniques both qualitatively and quantitatively.

NeurIPS Conference 2025 Conference Paper

PurpCode: Reasoning for Safer Code Generation

  • Jiawei Liu
  • Nirav Diwan
  • Zhe Wang
  • Haoyu Zhai
  • Xiaona Zhou
  • Kiet Nguyen
  • Tianjiao Yu
  • Muntasir Wahed

We introduce PurpCode, the first post-training recipe for training safe code reasoning models towards generating secure code and defending against malicious cyberactivities. PurpCode trains a reasoning model in two stages: (i) Rule Learning, which explicitly teaches the model to reference cybersafety rules to generate vulnerability-free code and to avoid facilitating malicious cyberactivities; and (ii) Reinforcement Learning, which optimizes model safety and preserves model utility through diverse, multi-objective reward mechanisms. To empower the training pipelines with comprehensive cybersafety data, we conduct internal red-teaming to synthesize comprehensive and high-coverage prompts based on real-world tasks for inducing unsafe cyberactivities in the model. Based on PurpCode, we develop a reasoning-based coding model, namely PurpCode-32B, which demonstrates state-of-the-art cybersafety, outperforming various frontier models. Moreover, our alignment method decreases the model overrefusal rates in both general and cybersafety-specific scenarios, while preserving model utility in both code generation and common security knowledge.

AAAI Conference 2025 Conference Paper

Reducing AUV Energy Consumption Through Dynamic Sensor Directions Switching via Deep Reinforcement Learning

  • Jiawei Liu
  • Yuanbo Xu
  • Shanshan Song
  • Lu Jiang

Autonomous underwater vehicle (AUV) is crucial for marine applications such as ocean data collection, pollution monitoring, and navigation. However, their limited energy resources constrain their operational duration, posing a significant challenge for long-term operations. Due to the complex and unpredictable nature of the underwater environment, AUVs allocate energy to their sensing systems to sense the surrounding environment and avoid obstacles. Existing methods focus on reducing energy consumption on AUV computing and movement, neglecting sensing energy consumption and few attempts have been made to balance the AUV energy and sensing ability with a flexible sensing system. Along these lines, we consider both AUV energy consumption and flexible sensing abilities, and propose a deep reinforcement learning-based method to Reduce Energy Consumption by AUV Sensing system (RECS). Specifically, we build an AUV sensing system in a 2-dimension space, with controllable 8-direction sensing abilities to collect the environment information dynamically. Then we divide the underwater environment into several areas and assign weights on the edges of areas based on the AUV planned path. Additionally, we dynamically switch the sensors in different directions and radii to sense the edges of the area where the AUV is located. The Artificial Potential Field (APF) method is employed to re-plan the AUV path to avoid obstacles and reach the target point effectively. Experimental results demonstrate that compared to full sensors on, our method reduces energy consumption by 53.48% and is capable of generalizing to varying environments and varying sensing system radii.

NeurIPS Conference 2025 Conference Paper

Reliable Lifelong Multimodal Editing: Conflict-Aware Retrieval Meets Multi-Level Guidance

  • Qiang Zhang
  • Fanrui Zhang
  • Jiawei Liu
  • Ming Hu
  • Junjun He
  • Zheng-Jun Zha

The dynamic nature of real-world information demands efficient knowledge editing in multimodal large language models (MLLMs) to ensure continuous knowledge updates. However, existing methods often struggle with precise matching in large-scale knowledge retrieval and lack multi-level guidance for coordinated editing, leading to less reliable outcomes. To tackle these challenges, we propose CARML, a novel retrieval-augmented editing framework that integrates conflict-aware dynamic retrieval with multi-level implicit and explicit guidance for reliable lifelong multimodal editing. Specifically, CARML introduces intra-modal uncertainty and inter-modal conflict quantification to dynamically integrate multi-channel retrieval results, so as to pinpoint the most relevant knowledge to the incoming edit samples. Afterwards, an edit scope classifier discerns whether the edit sample semantically aligns with the edit scope of the retrieved knowledge. If deemed in-scope, CARML refines the retrieved knowledge into information-rich continuous prompt prefixes, serving as the implicit knowledge guide. These prefixes not only include static knowledge prompt that capture key textual semantics but also incorporate token-level, context-aware dynamic prompt to explore fine-grained cross-modal associations between the edit sample and retrieved knowledge. To further enhance reliability, CARML incorporates a "hard correction" mechanism, leveraging explicit label knowledge to adjust the model’s output logits. Extensive experiments across multiple MLLMs and datasets indicate the superior performance of CARML in lifelong multimodal editing scenarios.

NeurIPS Conference 2025 Conference Paper

Squared families are useful conjugate priors

  • Russell Tsuchida
  • Jiawei Liu
  • Cheng Soon Ong
  • Dino Sejdinovic

Squared families of probability distributions have been studied and applied in numerous machine learning contexts. Typically, they appear as likelihoods, where their advantageous computational, geometric and statistical properties are exploited for fast estimation algorithms, representational properties and statistical guarantees. Here, we investigate the use of squared families as prior beliefs in Bayesian inference. We find that they can form helpful conjugate families, often allowing for closed-form and tractable Bayesian inference and marginal likelihoods. We apply such conjugate families to Bayesian regression in feature space using end-to-end learnable neural network features. Such a setting allows for a rich multi-modal alternative to Gaussian processes with neural network features, often called deep kernel learning. We demonstrate our method on few shot learning, outperforming existing neural methods based on Gaussian processes and normalising flows.

ICRA Conference 2024 Conference Paper

Bevel-Tip Needle Deflection Modeling, Simulation, and Validation in Multi-Layer Tissues

  • Yanzhou Wang
  • Lidia Al-Zogbi
  • Guanyun Liu
  • Jiawei Liu
  • Junichi Tokuda
  • Axel Krieger
  • Iulian I. Iordachita

Percutaneous needle insertions are commonly performed for diagnostic and therapeutic purposes as an effective alternative to more invasive surgical procedures. However, the outcome of needle-based approaches relies heavily on the accuracy of needle placement, which remains a challenge even with robot assistance and medical imaging guidance due to needle deflection caused by contact with soft tissues. In this paper, we present a novel mechanics-based 2D bevel-tip needle model that can account for the effect of nonlinear strain-dependent behavior of biological soft tissues under compression. Real-time finite element simulation allows multiple control inputs along the length of the needle with full three-degree-of-freedom (DOF) planar needle motions. Cross-validation studies using custom-designed multi-layer tissue phantoms as well as heterogeneous chicken breast tissues result in less than 1mm in-plane errors for insertions reaching depths of up to 61 mm, demonstrating the validity and generalizability of the proposed method.

JBHI Journal 2024 Journal Article

Convolutional Transformer-Based Cross Subject Model for SSVEP-Based BCI Classification

  • Jiawei Liu
  • Ruimin Wang
  • Yuankui Yang
  • Yuan Zong
  • Yue Leng
  • Wenming Zheng
  • Sheng Ge

Steady-state visual evoked potential (SSVEP) is a commonly used brain-computer interface (BCI) paradigm. The performance of cross-subject SSVEP classification has a strong impact on SSVEP-BCI. This study designed a cross subject generalization SSVEP classification model based on an improved transformer structure that uses domain generalization (DG). The global receptive field of multi-head self-attention is used to learn the global generalized SSVEP temporal information across subjects. This is combined with a parallel local convolution module, designed to avoid oversmoothing the oscillation characteristics of temporal SSVEP data and better fit the feature. Moreover, to improve the cross-subject calibration-free SSVEP classification performance, an DG method named StableNet is combined with the proposed convolutional transformer structure to form the DG-Conformer method, which can eliminate spurious correlations between SSVEP discriminative information and background noise to improve cross-subject generalization. Experiments on two public datasets, Benchmark and BETA, demonstrated the outstanding performance of the proposed DG-Conformer compared with other calibration-free methods, FBCCA, tt-CCA, Compact-CNN, FB-tCNN, and SSVEPNet. Additionally, DG-Conformer outperforms the classic calibration-required algorithms eCCA, eTRCA and eSSCOR when calibration is used. An incomplete partial stimulus calibration scheme was also explored on the Benchmark dataset, and it was demonstrated to be a potential solution for further high-performance personalized SSVEP-BCI with quick calibration.

YNICL Journal 2024 Journal Article

Genetic and vascular risk factors for ischemic stroke and cortical morphometry in individuals without a history of stroke: A UK Biobank observational cohort study

  • Jiawei Liu
  • Yingying Xie
  • Feng Liu
  • Wen Qin
  • Chunshui Yu

BACKGROUND: Stroke risk factors may contribute to cognitive decline and dementia by altering brain tissue integrity. If their effects on brain are nonnegligible, the target regions for stroke rehabilitation with brain stimulation identified by cross-sectional case-control studies may be biased due to the pre-existing brain differences caused by these risk factors. Here, we investigated the effects of stroke risk factors on cortical thickness (CT) and surface area (SA) in individuals without a history of stroke. METHODS: ), systolic blood pressure (SBP), diastolic blood pressure (DBP), glycated hemoglobin (HbA1c), triglycerides (TG), and low-density lipoprotein (LDL) on CT and SA of 62 cerebral regions. We excluded non-Caucasian participants and participants with missing data, unqualified brain images, or a history of stroke or any other brain diseases. We constructed a multivariate linear regression model for each phenotype to simultaneously test the effect of each factor and interaction between factors. The results were verified by sensitivity analyses of SDP or DBP input and adjusting for body-mass index, high-density lipoprotein cholesterol, or smoking and alcohol intake. By excluding participants with abnormal blood pressure, glucose, or lipid, we tested whether vascular risk factor within normal range also affected cortical phenotypes. To determine clinical relevance of our findings, we also investigated the effects of stroke risk factors and cortical phenotypes on cognitive decline assessed by fluid intelligence score (FIQ) and the mediation of cortical phenotype for the association between stroke risk factor and FIQ. RESULTS: and SBP with cognitive decline were mediated by CT phenotypes. CONCLUSIONS: Stroke risk factors have substantial effects on cortical morphometry and cognitive decline in middle-aged and older people, which should be considered in the prevention of dementia and in the identification of target regions for stroke rehabilitation with brain stimulation.

NeurIPS Conference 2024 Conference Paper

Lips Are Lying: Spotting the Temporal Inconsistency between Audio and Visual in Lip-Syncing DeepFakes

  • Weifeng Liu
  • Tianyi She
  • Jiawei Liu
  • Boheng Li
  • Dongyu Yao
  • Ziyou Liang
  • Run Wang

In recent years, DeepFake technology has achieved unprecedented success in high-quality video synthesis, but these methods also pose potential and severe security threats to humanity. DeepFake can be bifurcated into entertainment applications like face swapping and illicit uses such as lip-syncing fraud. However, lip-forgery videos, which neither change identity nor have discernible visual artifacts, present a formidable challenge to existing DeepFake detection methods. Our preliminary experiments have shown that the effectiveness of the existing methods often drastically decrease or even fail when tackling lip-syncing videos. In this paper, for the first time, we propose a novel approach dedicated to lip-forgery identification that exploits the inconsistency between lip movements and audio signals. We also mimic human natural cognition by capturing subtle biological links between lips and head regions to boost accuracy. To better illustrate the effectiveness and advances of our proposed method, we create a high-quality LipSync dataset, AVLips, by employing the state-of-the-art lip generators. We hope this high-quality and diverse dataset could be well served the further research on this challenging and interesting field. Experimental results show that our approach gives an average accuracy of more than 95. 3% in spotting lip-syncing videos, significantly outperforming the baselines. Extensive experiments demonstrate the capability to tackle deepfakes and the robustness in surviving diverse input transformations. Our method achieves an accuracy of up to 90. 2% in real-world scenarios (e. g. , WeChat video call) and shows its powerful capabilities in real scenario deployment. To facilitate the progress of this research community, we release all resources at https: //github. com/AaronComo/LipFD.

IJCAI Conference 2024 Conference Paper

Natural Language-centered Inference Network for Multi-modal Fake News Detection

  • Qiang Zhang
  • Jiawei Liu
  • Fanrui Zhang
  • Jingyi Xie
  • Zheng-Jun Zha

The proliferation of fake news with image and text in the internet has triggered widespread concern. Existing research has made important contributions in cross-modal information interaction and fusion, but fails to fundamentally address the modality gap among news image, text, and news-related external knowledge representations. In this paper, we propose a novel Natural Language-centered Inference Network (NLIN) for multi-modal fake news detection by aligning multi-modal news content with the natural language space and introducing an encoder-decoder architecture to fully comprehend the news in-context. Specifically, we first unify multi-modal news content into textual modality by converting news images and news-related external knowledge into plain textual content. Then, we design a multi-modal feature reasoning module, which consists of a multi-modal encoder, a unified-modal context encoder and an inference decoder with prompt phrase. This framework not only fully extracts the latent representation of cross-modal news content, but also utilizes the prompt phrase to stimulate the powerful in-context learning ability of the pre-trained large language model to reason about the truthfulness of the news content. In addition, to support the research in the field of multi-modal fake news detection, we produce a challenging large scale, multi-platform, multi-domain multi-modal Chinese Fake News Detection (CFND) dataset. Extensive experiments show that our CFND dataset is challenging and the proposed NLIN outperforms state-of-the-art methods.

EAAI Journal 2024 Journal Article

Recognizing wearable upper-limb rehabilitation gestures by a hybrid multi-feature neural network

  • Shu Wang
  • Jiawei Liu
  • Shen Chen
  • Shanshan Wang
  • Yuxin Peng
  • Changbo Liao
  • Li Liu

Stroke remains a leading cause of disability, presenting significant challenges to individuals and society. The post-stroke rehabilitation process demands prolonged professional training and evaluation. To tackle the issue of limited resources hindering patients’ access to frequent assessments and to facilitate personalized rehabilitation treatment, wearable technology has emerged as a promising solution. However, current wearable-based approaches often rely solely on raw sensor data fed directly into deep neural networks, which may not effectively capture intricate temporal relationships unless they incorporate the knowledge typically employed in clinical analysis. In this study, we introduce a hybrid multi-feature neural network that combines manually designed features, commonly utilized in clinical analysis, with latent features generated by deep networks. By explicitly considering the motion context and spatio-temporal relations among multiple body parts in the upper limb, our model can accurately detect their real-time motions. Empirical evaluations on our proprietary dataset reveal that the accuracy for subject-dependent and subject-independent experiments on 8 coarse-grained actions is 0. 9849 and 0. 9871, respectively, while for 24 fine-grained actions, they are 0. 9724 and 0. 9829, respectively. These results indicate that our model exhibits superior performance compared to other methods, contributing to the advancement of stroke rehabilitation and personalized therapy utilizing wearable systems.

NeurIPS Conference 2024 Conference Paper

SelfCodeAlign: Self-Alignment for Code Generation

  • Yuxiang Wei
  • Federico Cassano
  • Jiawei Liu
  • Yifeng Ding
  • Naman Jain
  • Zachary Mueller
  • Harm de Vries
  • Leandro Von Werra

Instruction tuning is a supervised fine-tuning approach that significantly improves the ability of large language models (LLMs) to follow human instructions. For programming tasks, most models are finetuned with costly human-annotated instruction-response pairs or those generated by large, proprietary LLMs, which may not be permitted. We propose SelfCodeAlign, the first fully transparent and permissive pipeline for self-aligning code LLMs without extensive human annotations or distillation. SelfCodeAlign employs the same base model for inference throughout the data generation process. It first extracts diverse coding concepts from high-quality seed snippets to generate new tasks. It then samples multiple responses per task, pairs each with test cases, and validates them in a sandbox environment. Finally, passing examples are selected for instruction tuning. In our primary experiments, we use SelfCodeAlign with CodeQwen1. 5-7B to generate a dataset of 74k instruction-response pairs. Finetuning on this dataset leads to a model that achieves a 67. 1 pass@1 on HumanEval+, surpassing CodeLlama-70B-Instruct despite being ten times smaller. Across all benchmarks, this finetuned model consistently outperforms the original version trained with OctoPack, the previous state-of-the-art method for instruction tuning without human annotations or distillation. Additionally, we show that SelfCodeAlign is effective across LLMs of various sizes, from 3B to 33B, and that the base models can benefit more from alignment with their own data distribution. We further validate each component’s effectiveness in our pipeline, showing that SelfCodeAlign outperforms both direct distillation from GPT-4o and leading GPT-3. 5-based distillation methods, such as OSS-Instruct and Evol-Instruct. SelfCodeAlign has also led to the creation of StarCoder2-Instruct, the first fully transparent, permissively licensed, and self-aligned code LLM that achieves state-of-the-art coding performance. Overall, SelfCodeAlign shows for the first time that a strong instruction-tuned code LLM can result from self-alignment rather than distillation.

IROS Conference 2024 Conference Paper

Tracking Tumors under Deformation from Partial Point Clouds using Occupancy Networks

  • Pit Henrich
  • Jiawei Liu
  • Jiawei Ge 0001
  • Samuel Schmidgall
  • Lauren M. Shepard
  • Ahmed Ezzat Ghazi
  • Franziska Mathis-Ullrich
  • Axel Krieger

To track tumors during surgery, information from preoperative CT scans is used to determine their position. However, as the surgeon operates, the tumor may be deformed which presents a major hurdle for accurately resecting the tumor, and can lead to surgical inaccuracy, increased operation time, and excessive margins. This issue is particularly pronounced in robot-assisted partial nephrectomy (RAPN), where the kidney undergoes significant deformations during operation. Toward addressing this, we introduce a occupancy network-based method for the localization of tumors within kidney phantoms undergoing deformations at interactive speeds. We validate our method by introducing a 3D hydrogel kidney phantom embedded with exophytic and endophytic renal tumors. It closely mimics real tissue mechanics to simulate kidney deformation during in vivo surgery, providing excellent contrast and clear delineation of tumor margins to enable automatic threshold-based segmentation. Our findings indicate that the proposed method can localize tumors in moderately deforming kidneys with a margin of 6mm to 10mm, while providing essential volumetric 3D information at over 60Hz. This capability directly enables downstream tasks such as robotic resection.

EAAI Journal 2023 Journal Article

A review of wearable sensors based fall-related recognition systems

  • Jiawei Liu
  • Xiaohu Li
  • Shanshan Huang
  • Rui Chao
  • Zhidong Cao
  • Shu Wang
  • Aiguo Wang
  • Li Liu

Falls are an important factor in significantly deteriorating quality of life of older adults, consequently leading to both physical and psychological harm. A wearable-based fall-related recognition system (WFRS) indeed facilitates the prediction, detection, and classification of fall events in helping fallers. Previous studies have provided a relatively comprehensive introduction to WFRSs from the perspective of sensor types and recognition algorithms. However, while these studies provide a clear technical direction, how to choose the appropriate technology for each phase of the experiment is a stumbling block for newly interested researchers. Accordingly, a comprehensive review article covering the mainstream technologies of WFRSs is imperative and meaningful. This review analyzes 48 state-of-the-art researches in WFRSs from three databases (i. e. , IEEE Explorer, ScienceDirect, and MDPI) and introduces the pipeline techniques that consist of data acquisition, preprocessing, feature extraction, model training, and evaluation. Specifically, we first analyze the pros and cons of the use of different number of sensors for data collection. We then introduce the widely used preprocessing techniques including filtering and data augmentation. Afterwards, we detail the extraction of various features and illustrate methods for the selection, training, and evaluation of fall recognition models. We finally discuss factors affecting the overall performance of a model and offer suggestions for future research.

AAAI Conference 2023 Conference Paper

A Speaker Turn-Aware Multi-Task Adversarial Network for Joint User Satisfaction Estimation and Sentiment Analysis

  • Kaisong Song
  • Yangyang Kang
  • Jiawei Liu
  • Xurui Li
  • Changlong Sun
  • Xiaozhong Liu

User Satisfaction Estimation is an important task and increasingly being applied in goal-oriented dialogue systems to estimate whether the user is satisfied with the service. It is observed that whether the user’s needs are met often triggers various sentiments, which can be pertinent to the successful estimation of user satisfaction, and vice versa. Thus, User Satisfaction Estimation (USE) and Sentiment Analysis (SA) should be treated as a joint, collaborative effort, considering the strong connections between the sentiment states of speakers and the user satisfaction. Existing joint learning frameworks mainly unify the two highly pertinent tasks over cascade or shared-bottom implementations, however they fail to distinguish task-specific and common features, which will produce sub-optimal utterance representations for downstream tasks. In this paper, we propose a novel Speaker Turn-Aware Multi-Task Adversarial Network (STMAN) for dialogue-level USE and utterance-level SA. Specifically, we first introduce a multi-task adversarial strategy which trains a task discriminator to make utterance representation more task-specific, and then utilize a speaker-turn aware multi-task interaction strategy to extract the common features which are complementary to each task. Extensive experiments conducted on two real-world service dialogue datasets show that our model outperforms several state-of-the-art methods.

IROS Conference 2023 Conference Paper

Development and Evaluation of a Single-arm Robotic System for Autonomous Suturing

  • Jiawei Liu
  • Michael Kam
  • Justin D. Opfermann
  • Zheyuan Zhang
  • Michael H. Hsieh
  • Jin U. Kang
  • Axel Krieger

This article introduces a novel suture managing device (SMD) and new suture management controller to enable single-arm suture management during autonomous suturing with the Smart Tissue Autonomous Robot (STAR). The primary function of the SMD is to tension and manage the suture thread, a task that was previously carried out by a second manipulator or a human assistant. The SMD and its controller are integrated into STAR's autonomous suturing workflow. Experiments were conducted to quantify the tensioning force of SMD and to evaluate the suture quality of the new single-arm system. The prototype of SMD achieves 1. 67N tensioning force with suturing time of 29. 1±0. 42 seconds per stitch. Our study results demonstrate that the single-arm STAR system with SMD achieves equivalent performance to our previous works in suturing efficiency where suture management was performed with either a dual-armed robotic system or by a human surgical assistant. The study's findings contribute to the field of medical robotics and to our knowledge represent the first known instance of single-arm suturing with suture management during autonomous anastomosis.

NeurIPS Conference 2023 Conference Paper

Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation

  • Jiawei Liu
  • Chunqiu Steven Xia
  • Yuyao Wang
  • LINGMING ZHANG

Program synthesis has been long studied with recent approaches focused on directly using the power of Large Language Models (LLMs) to generate code. Programming benchmarks, with curated synthesis problems and test-cases, are used to measure the performance of various LLMs on code synthesis. However, these test-cases can be limited in both quantity and quality for fully assessing the functional correctness of the generated code. Such limitation in the existing benchmarks begs the following question: In the era of LLMs, is the code generated really correct? To answer this, we propose EvalPlus – a code synthesis evaluation framework to rigorously benchmark the functional correctness of LLM-synthesized code. EvalPlus augments a given evaluation dataset with large amounts of test-cases newly produced by an automatic test input generator, powered by both LLM- and mutation-based strategies. While EvalPlus is general, we extend the test-cases of the popular HumanEval benchmark by 80x to build HumanEval+. Our extensive evaluation across 26 popular LLMs (e. g. , GPT-4 and ChatGPT) demonstrates that HumanEval+ is able to catch significant amounts of previously undetected wrong code synthesized by LLMs, reducing the pass@k by up-to 19. 3-28. 9%. We also surprisingly found that test insufficiency can lead to mis-ranking. For example, both WizardCoder-CodeLlama and Phind-CodeLlama now outperform ChatGPT on HumanEval+, while none of them could on HumanEval. Our work not only indicates that prior popular code synthesis evaluation results do not accurately reflect the true performance of LLMs for code synthesis, but also opens up a new direction to improve such programming benchmarks through automated testing. We have open-sourced our tools, enhanced datasets as well as all LLM-generated code at https: //github. com/evalplus/evalplus to facilitate and accelerate future LLM-for-code research.

AAAI Conference 2022 Conference Paper

Debiased Batch Normalization via Gaussian Process for Generalizable Person Re-identification

  • Jiawei Liu
  • Zhipeng Huang
  • Liang Li
  • Kecheng Zheng
  • Zheng-Jun Zha

Generalizable person re-identification aims to learn a model with only several labeled source domains that can perform well on unseen domains. Without access to the unseen domain, the feature statistics of the batch normalization (BN) layer learned from a limited number of source domains is doubtlessly biased for unseen domain. This would mislead the feature representation learning for unseen domain and deteriorate the generalizaiton ability of the model. In this paper, we propose a novel Debiased Batch Normalization via Gaussian Process approach (GDNorm) for generalizable person reidentification, which models the feature statistic estimation from BN layers as a dynamically self-refining Gaussian process to alleviate the bias to unseen domain for improving the generalization. Specifically, we establish a lightweight model with multiple set of domain-specific BN layers to capture the discriminability of individual source domain, and learn the corresponding parameters of the domain-specific BN layers. These parameters of different source domains are employed to deduce a Gaussian process. We randomly sample several paths from this Gaussian process served as the BN estimations of potential new domains outside of existing source domains, which can further optimize these learned parameters from source domains, and estimate more accurate Gaussian process by them in return, tending to real data distribution. Even without a large number of source domains, GDNorm can still provide debiased BN estimation by using the mean path of the Gaussian process, while maintaining low computational cost during testing. Extensive experiments demonstrate that our GDNorm effectively improves the generalization ability of the model on unseen domain.

AAAI Conference 2022 Conference Paper

Modality-Adaptive Mixup and Invariant Decomposition for RGB-Infrared Person Re-identification

  • Zhipeng Huang
  • Jiawei Liu
  • Liang Li
  • Kecheng Zheng
  • Zheng-Jun Zha

RGB-infrared person re-identification is an emerging crossmodality re-identification task, which is very challenging due to significant modality discrepancy between RGB and infrared images. In this work, we propose a novel modalityadaptive mixup and invariant decomposition (MID) approach for RGB-infrared person re-identification towards learning modality-invariant and discriminative representations. MID designs a modality-adaptive mixup scheme to generate suitable mixed modality images between RGB and infrared images for mitigating the inherent modality discrepancy at the pixel-level. It formulates modality mixup procedure as Markov decision process, where an actor-critic agent learns dynamical and local linear interpolation policy between different regions of cross-modality images under a deep reinforcement learning framework. Such policy guarantees modality-invariance in a more continuous latent space and avoids manifold intrusion by the corrupted mixed modality samples. Moreover, to further counter modality discrepancy and enforce invariant visual semantics at the feature-level, MID employs modality-adaptive convolution decomposition to disassemble a regular convolution layer into modalityspecific basis layers and a modality-shared coefficient layer. Extensive experimental results on two challenging benchmarks demonstrate superior performance of MID over stateof-the-art methods.

AAAI Conference 2021 Conference Paper

Time to Transfer: Predicting and Evaluating Machine-Human Chatting Handoff

  • Jiawei Liu
  • Zhe Gao
  • Yangyang Kang
  • Zhuoren Jiang
  • Guoxiu He
  • Changlong Sun
  • Xiaozhong Liu
  • Wei Lu

Is chatbot able to completely replace the human agent? The short answer could be – “it depends. .. ”. For some challenging cases, e. g. , dialogue’s topical spectrum spreads beyond the training corpus coverage, the chatbot may malfunction and return unsatisfied utterances. This problem can be addressed by introducing the Machine-Human Chatting Handoff (MHCH) which enables human-algorithm collaboration. To detect the normal/transferable utterances, we propose a Difficulty-Assisted Matching Inference (DAMI) network, utilizing difficulty-assisted encoding to enhance the representations of utterances. Moreover, a matching inference mechanism is introduced to capture the contextual matching features. A new evaluation metric, Golden Transfer within Tolerance (GT-T), is proposed to assess the performance by considering the tolerance property of the MHCH. To provide insights into the task and validate the proposed model, we collect two new datasets. Extensive experimental results are presented and contrasted against a series of baseline models to demonstrate the efficacy of our model on MHCH.

IJCAI Conference 2020 Conference Paper

Co-Saliency Spatio-Temporal Interaction Network for Person Re-Identification in Videos

  • Jiawei Liu
  • Zheng-Jun Zha
  • Xierong Zhu
  • Na Jiang

Person re-identification aims at identifying a certain pedestrian across non-overlapping camera networks. Video-based person re-identification approaches have gained significant attention recently, expanding image-based approaches by learning features from multiple frames. In this work, we propose a novel Co-Saliency Spatio-Temporal Interaction Network (CSTNet) for person re-identification in videos. It captures the common salient foreground regions among video frames and explores the spatial-temporal long-range context interdependency from such regions, towards learning discriminative pedestrian representation. Specifically, multiple co-saliency learning modules within CSTNet are designed to utilize the correlated information across video frames to extract the salient features from the task-relevant regions and suppress background interference. Moreover, multiple spatial-temporal interaction modules within CSTNet are proposed, which exploit the spatial and temporal long-range context interdependencies on such features and spatial-temporal information correlation, to enhance feature representation. Extensive experiments on two benchmarks have demonstrated the effectiveness of the proposed method.

IJCAI Conference 2020 Conference Paper

Decorrelated Clustering with Data Selection Bias

  • Xiao Wang
  • Shaohua Fan
  • Kun Kuang
  • Chuan Shi
  • Jiawei Liu
  • Bai Wang

Most of existing clustering algorithms are proposed without considering the selection bias in data. In many real applications, however, one cannot guarantee the data is unbiased. Selection bias might bring the unexpected correlation between features and ignoring those unexpected correlations will hurt the performance of clustering algorithms. Therefore, how to remove those unexpected correlations induced by selection bias is extremely important yet largely unexplored for clustering. In this paper, we propose a novel Decorrelation regularized K-Means algorithm (DCKM) for clustering with data selection bias. Specifically, the decorrelation regularizer aims to learn the global sample weights which are capable of balancing the sample distribution, so as to remove unexpected correlations among features. Meanwhile, the learned weights are combined with k-means, which makes the reweighted k-means cluster on the inherent data distribution without unexpected correlation influence. Moreover, we derive the updating rules to effectively infer the parameters in DCKM. Extensive experiments results on real world datasets well demonstrate that our DCKM algorithm achieves significant performance gains, indicating the necessity of removing unexpected feature correlations induced by selection bias when clustering.

IJCAI Conference 2020 Conference Paper

Multi-Scale Spatial-Temporal Integration Convolutional Tube for Human Action Recognition

  • Haoze Wu
  • Jiawei Liu
  • Xierong Zhu
  • Meng Wang
  • Zheng-Jun Zha

Applying multi-scale representations leads to consistent performance improvements on a wide range of image recognition tasks. However, with the addition of the temporal dimension in video domain, directly obtaining layer-wise multi-scale spatial-temporal features will add a lot extra computational cost. In this work, we propose a novel and efficient Multi-Scale Spatial-Temporal Integration Convolutional Tube (MSTI) aiming at achieving accurate recognition of actions with lower computational cost. It firstly extracts multi-scale spatial and temporal features through the multi-scale convolution block. Considering the interaction of different-scales representations and the interaction of spatial appearance and temporal motion, we employ the cross-scale attention weighted blocks to perform feature recalibration by integrating multi-scale spatial and temporal features. An end-to-end deep network, MSTI-Net, is also presented based on the proposed MSTI tube for human action recognition. Extensive experimental results show that our MSTI-Net significantly boosts the performance of existing convolution networks and achieves state-of-the-art accuracy on three challenging benchmarks, i. e. , UCF-101, HMDB-51 and Kinetics-400, with much fewer parameters and FLOPs.

IJCAI Conference 2019 Conference Paper

Mutually Reinforced Spatio-Temporal Convolutional Tube for Human Action Recognition

  • Haoze Wu
  • Jiawei Liu
  • Zheng-Jun Zha
  • Zhenzhong Chen
  • Xiaoyan Sun

Recent works use 3D convolutional neural networks to explore spatio-temporal information for human action recognition. However, they either ignore the correlation between spatial and temporal features or suffer from high computational cost by spatio-temporal features extraction. In this work, we propose a novel and efficient Mutually Reinforced Spatio-Temporal Convolutional Tube (MRST) for human action recognition. It decomposes 3D inputs into spatial and temporal representations, mutually enhances both of them by exploiting the interaction of spatial and temporal information and selectively emphasizes informative spatial appearance and temporal motion, meanwhile reducing the complexity of structure. Moreover, we design three types of MRSTs according to the different order of spatial and temporal information enhancement, each of which contains a spatio-temporal decomposition unit, a mutually reinforced unit and a spatio-temporal fusion unit. An end-to-end deep network, MRST-Net, is also proposed based on the MRSTs to better explore spatio-temporal information in human actions. Extensive experiments show MRST-Net yields the best performance, compared to state-of-the-art approaches.

v2026.09.13