Arrow Research search

Author name cluster

Jian Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

61 papers
2 author rows

Possible papers

61

EAAI Journal 2026 Journal Article

An ultra-efficient edge-based wearable system for real-time and remote blood pressure monitoring

  • Wei Xiang
  • Jian Liu
  • Shuaicong Hu
  • HaiHui Zhang
  • Chao Huang
  • Cuiwei Yang

Continuous blood pressure (BP) monitoring is crucial for health management, yet existing methods struggle with efficiency and adaptability in home and clinical environments. To address this, we propose the edge-based blood pressure estimation system (Edge-BP), an ultra-efficient wearable system for real-time, cuffless BP monitoring. First, we present the Cascaded Depthwise Separable Residual Network (CDS-Net), which employs a cascaded depthwise separable residual architecture and attention mechanisms to estimate BP from photoplethysmography (PPG) and electrocardiography (ECG) signals. Furthermore, we propose a progressive distillation-pruning framework, a model compression method that integrates dependency graph–guided structured pruning with dual-teacher knowledge distillation, substantially improving model compression efficiency. We also develop a wearable device for synchronized PPG and ECG acquisition, leveraging cross-database transfer learning to improve adaptability. The optimized model is deployed on a neural processing unit (NPU) and integrated with fourth-generation (4G) communication, enabling remote monitoring and automatic alerts for abnormal BP detection. CDS-Net exhibits superior performance in estimating systolic blood pressure (SBP) and diastolic blood pressure (DBP), with mean absolute error (MAE) 3. 52 mm of mercury (mmHg) and 2. 00 mmHg, respectively. Even after the model is compressed—reducing parameters by 94. 4 % and computational complexity by 93. 7 %—it still achieves MAEs 4. 12 mmHg and 2. 86 mmHg for SBP and DBP estimation. The testing results meet standards set by the British Hypertension Society and the Institute of Electrical and Electronics Engineers. This study provides a comprehensive solution for continuous BP monitoring in both home and clinical settings, paving the way for future advancements in wearable, physiological signal-based health management.

EAAI Journal 2026 Journal Article

Automatic and accurate characterization of rock fractures based on Deep Learning

  • Jian Liu
  • Omar AlDajani
  • Bing Q. Li
  • Herbert Einstein
  • Zili Li

Fractures play a pivotal role in the hydromechanical behavior of rocks. Researchers have tried to understand the fundamental mechanisms and behavior of fractures, both theoretically and experimentally. Fractures produced in laboratory tests can be characterized using photographic images at a range of scales. Traditionally, identifying and extracting the temporal-spatial characteristics of fractures rely on meticulous manual labeling, entailing significant time and labor. Recently, Deep Learning (DL) methods have been deployed for automatic fracture extraction and characterization of fracture evolution from rock images. However, they generally necessitate a large dataset to train DL models for extracting fractures. More importantly, they cannot distinguish existing fractures or background noise and erroneously identify a continuous fracture as multiple disconnected segments, leading to significant errors in quantifying fracture evolution. This study introduces a new Deep Learning-based approach for mapping temporal sequence of images, the ‘Hybrid Fracture Mapping Method’ (HFMM). It incorporates a small training dataset using systematic sampling and employs an advanced denoising method. It is validated on a hydraulic fracturing image dataset from Massachusetts Institute of Technology (MIT) rock mechanics laboratory. The results demonstrate that the HFMM shows a notable improvement compared to the conventional DL methods in mapping hydraulic fracture evolution. It achieved a reduction in Mean Absolute Error (MAE) for quantifying fracture length and number of fracture branches, decreasing the MAE by 81 %–86 % and 69 %–93 %, respectively.

YNIMG Journal 2026 Journal Article

Dynamic reorganization of functional connectome gradients reveals time-specific recovery patterns after stroke

  • Qingwen Chen
  • Tao Zhong
  • Jian Liu
  • Dajun Yan
  • Han Gao

Stroke, a leading cause of death and disability worldwide, severely disrupts brain functional organization and cognitive abilities. Previous research has mainly focused on discrete functional network changes post-stroke, but how stroke affects whole-brain functional hierarchy and its relationship to cognitive recovery remains poorly understood. In this exploratory longitudinal study, we used connectome gradient mapping in 33 patients with first-ever stroke and 21 healthy controls to examine how stroke affects large-scale functional network organization across an early post-stroke(∼2 weeks), a subacute stage (3 months), and a chronic stage (1 year). By projecting functional connectivity patterns onto a low-dimensional gradient space, we found that although overall gradient structure remained relatively stable at the group level, individual patients exhibited significant deviations (EDfunc) from the healthy topology, most prominently at the early post-stroke across visual, somatomotor, ventral attention, and control networks. Furthermore, EDfunc showed time-specific associations with cognitive functions: broad negative correlations with visuospatial attention in the early post-stroke, transitioning to more selective associations with motor and attention measures in the chronic stage. In addition, dynamic interhemispheric functional imbalances emerged in the subacute and chronic stages. Taken together, these findings provide preliminary, hypothesis-generating evidence for dynamic reorganization of whole-brain functional hierarchy following stroke, and suggest that connectome gradient analysis and EDfunc may offer a sensitive framework for monitoring recovery and informing individualized rehabilitation strategies, pending confirmation in larger multi-center cohorts.

AIIM Journal 2026 Journal Article

Hierarchical classification for differential diagnosis of fever of unknown origin: A multi-task learning approach with self-adaptive representation sharing

  • Zhixiao Wang
  • Yu Tian
  • Jian Liu
  • Tianshu Zhou
  • Yunqing Qiu
  • Jingsong Li

Leveraging label dependencies as prior knowledge during both training and testing has proven valuable across diverse domains such as image annotation and text categorization. In our previous research, we successfully reframed the clinical challenge of aiding decision-making for patients with fever of unknown origin (FUO) as a hierarchical classification problem, validating its feasibility through local methods. However, these approaches still encounter challenges, including high training costs and potential error propagation during predictions. Moreover, existing global approaches for exploiting label dependencies impose strict prerequisites—such as fixed data modalities, manual specification of information-sharing directions, and equal-length label sequences—that limit their applicability to FUO etiologies. In this paper, we introduce a novel global hierarchical classification method based on a multi-task learning architecture for the early diagnosis of FUO patients. Our method leverages multimodal clinical data and a predefined label hierarchy and comprises three key components: a task decomposition strategy employing End-of-Sequence (EOS) markers (with each parent node in the label hierarchy corresponding to an individual classification task), a multimodal data feature extraction and fusion module, and a self-adaptive representation sharing module (Sa-RSM). We evaluated our approach on an experimental dataset extracted from electronic health records (EHRs) of a large-scale tertiary hospital in China, spanning January 2011 to October 2020 and comprising 34, 051 hospital admissions of 30, 794 FUO patients. Our results clearly demonstrate that the proposed method not only achieves superior predictive performance but also proactively halts predictions at coarser-grained classification tasks. Moreover, even in cases of misclassification, our method exhibits lower mistake severity, underscoring its potential clinical utility.

EAAI Journal 2026 Journal Article

International roughness index prediction based on research institute of highway ministry of transport track data using multiple relation graphs and temporal graph attention network

  • Jian Liu
  • Chunru Cheng
  • Zhen Wang
  • Shuhan Yang
  • Ebenezer O. Fanijo
  • Linbing Wang

International Roughness Index (IRI) is a comprehensive indicator used to evaluate overall pavement condition. Although the implemented artificial intelligence approaches for modeling IRI progression over time have achieved significant accuracy by capturing the temporal dependencies of IRI, existing models fail to fully capture the implicit interdependencies between IRI and other pavement performance indicators, such as rutting and cracking. This study proposes a novel IRI predictive model that can simultaneously capture the temporal dependencies of pavement performance and the interdependent relationships between IRI and other pavement indicators. Specifically, three relation graphs are first constructed by conducting Granger causality, transfer entropy causality, and similarity analysis of historical pavement performance data. External variables, including traffic loading, climate information, and pavement structural and material properties, are included as node features within these relation graphs. Node embeddings output from three Graph Attention Networks are then processed by three parallel Long Short-Term Memory and subsequently passed through multi-head attention layers. Finally, the performance of the proposed model was evaluated using data from the Research Institute of Highway Ministry of Transport Track and compared with that of state-of-the-art models. The result shows the causal relationship between rutting and IRI is the strongest, followed by Falling Weight Deflectometer measurements, with little causal influence from Surface Macrotexture Depth on IRI. Moreover, compared to existing models, the proposed model demonstrates superior accuracy and stability, achieving R2 values of 0. 9518, 0. 9496, 0. 9405, 0. 9305, 0. 9135, and 0. 8909 across prediction horizons of 1, 2, 4, 6, 8, and 12, respectively.

AIIM Journal 2026 Journal Article

PDAFormer 3+: A full-scale connected modified transformer with parallel dual attention for 3D medical image segmentation

  • Jinhui Zhang
  • Yueyang Gao
  • Jian Liu
  • Duanduan Chen

Medical image segmentation is essential for enhancing diagnostic and therapeutic accuracy, improving healthcare efficiency, and advancing medical research. In recent years, transformers have gained increasing attention in medical image segmentation owing to their ability to capture long-range dependencies, effectively compensating for the limitations of convolutional neural networks (CNNs) in global context modeling. This paper proposes PDAFormer 3+, a full-scale connected 3D medical image segmentation framework that integrates a parallel dual-attention modified transformer. Specifically, we introduce a parallel dual attention (PDA) mechanism to replace the conventional self-attention mechanism in transformers, enabling parallel modeling of global dependencies in both spatial and channel dimensions. In addition, the multi-layer perceptron (MLP) in transformers is replaced with a residual convolution block (RCB) as the feed-forward network, reducing computational complexity while enhancing the local representations. Inspired by U-Net 3+, we further incorporate full-scale features at each network layer and design a convolution excitation module (CEM) to enhance fused features. A deep supervision strategy is employed to perform multi-level representation learning from the aggregated feature maps. Notably, the introduction of linear mappings and convolution modules enables the modified transformer to be applied to large-scale 3D medical images, significantly reducing the computational burden of the network. Extensive experiments on Synapse, Automated Cardiac Diagnosis Challenge (ACDC), and type-B aortic dissection (type-B AD) show that PDAFormer 3+ achieves 86. 90%, 92. 54%, and 91. 93% mean Dice Similarity Coefficient (DSC), respectively, while maintaining strong efficiency. Overall, PDAFormer 3+ couples the complementary strengths of CNNs (local detail) and transformers (global context) to deliver accurate and efficient 3D medical image segmentation. Our code is publicly available at https: //github. com/BitGyy/PDAFormer.

AAAI Conference 2026 Conference Paper

TextShield-R1: Reinforced Reasoning for Tampered Text Detection

  • Chenfan Qu
  • Yiwu Zhong
  • Jian Liu
  • Xuekang Zhu
  • Bohan Yu
  • Lianwen Jin

The growing prevalence of tampered images poses serious security threats, highlighting the urgent need for reliable detection methods. Multimodal large language models (MLLMs) demonstrate strong potential in analyzing tampered images and generating interpretations. However, they still struggle with identifying micro-level artifacts, exhibit low accuracy in localizing tampered text regions, and heavily rely on expensive annotations for forgery interpretation. To this end, we introduce TextShield-R1, the first reinforcement learning based MLLM solution for tampered text detection and reasoning. Specifically, our approach introduces Forensic Continual pre-training, an easy-to-hard curriculum that well prepares the MLLM for tampered text detection by harnessing the large-scale cheap data from natural image forensic and OCR tasks. During fine-tuning, we perform Group Relative Policy Optimization with novel reward functions to reduce annotation dependency and improve reasoning capabilities. At inference time, we enhance localization accuracy via OCR Rectification, a method that leverages the MLLM’s strong text recognition abilities to refine its predictions. Furthermore, to support rigorous evaluation, we introduce Text Forensics Reasoning (TFR) benchmark, comprising over 45k real and tampered images across 16 languages, 10 tampering techniques, and diverse domains. Rich reasoning-style annotations are included, allowing for comprehensive assessment. Our TFR benchmark simultaneously addresses seven major limitations of existing benchmarks and enables robust evaluation under cross-style, cross-method, and cross-language conditions. Extensive experiments demonstrate that TextShield-R1 significantly advances the state of the art in interpretable tampered text detection.

JBHI Journal 2026 Journal Article

TinnitusLLM: A Multimodal Large Language Model Framework for Tinnitus Diagnosis Through EEG-fMRI Fusion Learning

  • Yipeng Du
  • Xiaohui Chen
  • Zewei Liu
  • Zhengwu Liu
  • Ngai Wong
  • Chi Zhang
  • Jian Chen
  • Zhiwei Ding

Accurate tinnitus diagnosis is crucial for enabling timely therapeutic intervention and longitudinal treatment monitoring. While non-invasive neuroimaging modalities-particularly electroencephalography (EEG) with millisecond temporal resolution and functional magnetic resonance imaging (fMRI) with millimeter spatial resolution- provide complementary neural features, existing diagnostic approaches remain constrained to unimodal analysis of EEG or fMRI data, inherently limiting diagnostic precision and clinical generalizability. This paper introduces TinnitusLLM, the first multimodal large language model (LLM) framework that synergistically integrates EEG and fMRI features for tinnitus diagnosis. To enable LLM-based interpretation of neural signals, this framework integrates three key components: (1) a neuroinspired positional encoding mechanism that injects neurophysiological priors into the embedding space, enabling neurologically grounded, dynamic positional mapping of EEG and fMRI tokens; (2) multimodal autoregressive pretraining on more than 500 hours of EEG and 250 hours of fMRI data to learn causally informed predictive representations; and (3) fine-tuning with a cross-modal, subject-invariant adversarial learning strategy that enforces subject-independent constraints in the shared cross-modal feature space, thereby substantially improving diagnostic robustness across subjects. We validate TinnitusLLM through comprehensive experiments on a rigorously collected multimodal dataset containing 20 participants. Quantitative evaluations demonstrate that TinnitusLLM achieves superior cross-subject diagnostic accuracy compared to the state-of-the-art baseline methods. These results underscore TinnitusLLM's potential as a clinically viable framework for objective tinnitus assessment through multimodal neural decoding.

JBHI Journal 2026 Journal Article

Unleashing the Power of Pretrained Transformer for Dense Prediction in Physiological Signals

  • Qihan Hu
  • Daomiao Wang
  • Hong Wu
  • Jian Liu
  • Cuiwei Yang

The physiological signals obtained from advanced sensors, combined with deep learning techniques for classification and regression tasks, have become a core driving force in enhancing smart healthcare. Recently, dense prediction tasks for physiological signals—aimed at generating predictions that are closely aligned with the input signal to enable fine-grained analysis—have garnered increasing attention. The UNet family, often combined with sophisticated task-specific customizations, has become a popular choice to improve prediction performance. However, pretrained Transformers have recently revolutionized deep learning due to their powerful transferability and effectiveness. In this work, we aim to harness the power of pretrained Transformers for dense prediction, eliminating the need for extensive task-specific architecture design. We propose a simple yet universal encoder-decoder architecture that utilizes a pretrained Transformer encoder and a lightweight convolutional Restormer decoder for dense prediction on physiological signals. To optimize the trade-off between model performance and computational efficiency, we incorporate knowledge distillation (KD). Our experiments focus on four representative dense prediction tasks: blood pressure waveform (BPW) estimation, PPG-to-ECG (P2E) reconstruction, denoising, and fiducial point localization. The results show that our proposed architecture outperforms state-of-the-art models, validating the potential of pretrained Transformers in enhancing physiological signal processing and medical diagnostics. This approach marks a significant step forward in optimizing both the performance and efficiency of dense prediction tasks.

NeurIPS Conference 2025 Conference Paper

Auto-Connect: Connectivity-Preserving RigFormer with Direct Preference Optimization

  • jingfeng Guo
  • Jian Liu
  • Jinnan Chen
  • Shiwei Mao
  • Changrong Hu
  • Puhua Jiang
  • Junlin Yu
  • Jing Xu

We introduce Auto-Connect, a novel approach for automatic rigging that explicitly preserves skeletal connectivity through a connectivity-preserving tokenization scheme. Unlike previous methods that predict bone positions represented as two joints or first predict points before determining connectivity, our method employs special tokens to define endpoints for each joint's children and for each hierarchical layer, effectively automating connectivity relationships. This approach significantly enhances topological accuracy by integrating connectivity information directly into the prediction framework. To further guarantee high-quality topology, we implement a topology-aware reward function that quantifies topological correctness, which is then utilized in a post-training phase through reward-guided Direct Preference Optimization. Additionally, we incorporate implicit geodesic features for latent top-$k$ bone selection, which substantially improves skinning quality. By leveraging geodesic distance information within the model's latent space, our approach intelligently determines the most influential bones for each vertex, effectively mitigating common skinning artifacts. This combination of connectivity-preserving tokenization, reward-guided fine-tuning, and geodesic-aware bone selection enables our model to consistently generate more anatomically plausible skeletal structures with superior deformation properties.

IROS Conference 2025 Conference Paper

EmoRLTalk: Speech-Driven Emotional Facial Animation With Offline Reinforcement Learning

  • Gaofeng Liu
  • Xuetong Li
  • Ruoyu Gao
  • Ye Yuan
  • Jian Liu
  • Hengsen Li
  • Hong Huo
  • Tao Fang

In recent years, significant breakthroughs have been made in audio-guided 3D facial animation. However, existing methods mainly focus on lip shape and audio consistency and still face key challenges to achieve alignment between facial emotions and speech emotions. To overcome this limitation, we introduce EmoRLTalk, a novel framework that integrates offline reinforcement learning to implicitly capture the intricate relationship between 3D facial landmarks and blendshape parameters, thereby enhancing the granularity of emotional expression. Furthermore, we harness the strength of conditional diffusion models to synthesize facial motions that are emotionally coherent with the input speech. Additionally, based on the multi-task learning paradigm, we construct a collaborative training framework of a regression main task and a classification sub-task. Specifically, we use emotion classification of blendshape as a sub-task to further improve the model’s ability to express facial emotions. To further enhance system controllability, we integrate the ControlNet module, allowing users to achieve precise facial expression control. Extensive experiments demonstrate that EmoRLTalk achieves superior emotional expressiveness and lip-sync performance compared to previous approaches.

EAAI Journal 2025 Journal Article

Fast and intelligent measurement of the ventilation resistance coefficient for the whole mine based on sparse measurement points

  • Dong Wang
  • Jian Liu
  • Lijun Deng
  • Peng Cao
  • Li Liu

Artificial intelligence is playing an important role in mine ventilation engineering, especially in ensuring safe mine production. The mine ventilation resistance coefficient (MVRC) is the core and basic parameter of a mine ventilation system. It is crucial to quickly and accurately obtain the ventilation resistance coefficient (VRC) of the whole mine for the scientific, safe, and intelligent management of the mine ventilation system. To solve the time-consuming and laborious problem of the traditional mine ventilation resistance measurement method, we propose a fast and intelligent measurement method to obtain the whole mine's VRC based on an artificial intelligence differential evolution algorithm and sparse measurement points. The VRC was experimentally measured to verify the validity of the intelligent measurement method and the reliability of the model. The relative error of the air volume at the observation points of the solved results is less than 6 %. The fast intelligent measurement of the MVRC of the Longshou mine was carried out. The results were applied to develop an emergency plan for addressing the insufficient air supply in the ventilation system caused by the collapse of the mine's main blind return shaft and validated through engineering practice. After field practice, the relative error between the predicted and tested air return volume of the 10-row inclined shaft was 3. 23 %. It is verified that the results obtained using this method can solve mine ventilation system problems with relatively high accuracy, significantly reducing both the testing workload and time required for the mine ventilation resistance measurements.

NeurIPS Conference 2025 Conference Paper

FerretNet: Efficient Synthetic Image Detection via Local Pixel Dependencies

  • Shuqiao Liang
  • Jian Liu
  • Chen Renzhang
  • Quanlong Guan

The increasing realism of synthetic images generated by advanced models such as VAEs, GANs, and LDMs poses significant challenges for synthetic image detection. To address this issue, we explore two artifact types introduced during the generation process: (1) latent distribution deviations and (2) decoding-induced smoothing effects, which manifest as inconsistencies in local textures, edges, and color transitions. Leveraging local pixel dependencies (LPD) properties rooted in Markov Random Fields, we reconstruct synthetic images using neighboring pixel information to expose disruptions in texture continuity and edge coherence. Building upon LPD, we propose FerretNet, a lightweight neural network with only 1. 1M parameters that delivers efficient and robust synthetic image detection. Extensive experiments demonstrate that FerretNet—trained exclusively on the 4-class ProGAN dataset—achieves an average accuracy of 97. 1% on an open-world benchmark comprising 22 generative models. Our code and datasets are publicly available at https: //github. com/xigua7105/FerretNet.

NeurIPS Conference 2025 Conference Paper

ForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization

  • Bo Du
  • Xuekang Zhu
  • Xiaochen Ma
  • Chenfan Qu
  • Kaiwen Feng
  • Zhe Yang
  • Chi-Man Pun
  • Jian Liu

The field of Fake Image Detection and Localization (FIDL) is highly fragmented, encompassing four domains: deepfake detection (Deepfake), image manipulation detection and localization (IMDL), artificial intelligence-generated image detection (AIGC), and document image manipulation localization (Doc). Although individual benchmarks exist in some domains, a unified benchmark for all domains in FIDL remains blank. The absence of a unified benchmark results in significant domain silos, where each domain independently constructs its datasets, models, and evaluation protocols without interoperability, preventing cross-domain comparisons and hindering the development of the entire FIDL field. To close the domain silo barrier, we propose ForensicHub, the first unified benchmark & codebase for all-domain fake image detection and localization. Considering drastic variations on dataset, model, and evaluation configurations across all domains, as well as the scarcity of open-sourced baseline models and the lack of individual benchmarks in some domains, ForensicHub: i) proposes a modular and configuration-driven architecture that decomposes forensic pipelines into interchangeable components across datasets, transforms, models, and evaluators, allowing flexible composition across all domains; ii) fully implements 10 baseline models (3 of which are reproduced from scratch), 6 backbones, 2 new benchmarks for AIGC and Doc, and integrates 2 existing benchmarks of DeepfakeBench and IMDLBenCo through an adapter-based design; iii) establishes an image forensic fusion protocol evaluation mechanism that supports unified training and testing of diverse forensic models across tasks; iv) conducts indepth analysis based on the ForensicHub, offering 8 key actionable insights into FIDL model architecture, dataset characteristics, and evaluation standards. Specifically, ForensicHub includes 4 forensic tasks, 23 datasets, 42 baseline models, 6 backbones, 11 GPU-accelerated pixel- and image-level evaluation metrics, and realizes 16 kinds of cross-domain evaluations. ForensicHub represents a significant leap forward in breaking the domain silos in the FIDL field and inspiring future breakthroughs. Code is available at: https: //github. com/scu-zjz/ForensicHub.

AAAI Conference 2025 Conference Paper

KGCRR: An Effective Metric-Driven Knowledge Graph Completion Framework by Designing a Novel Upper Bound Function with Adaptive Approximation to Reciprocal Rank

  • Kuan Xu
  • Kuo Yang
  • Jian Liu
  • Xiangkui Lu
  • Jun Wu
  • Xuezhong Zhou

Knowledge Graph Embedding (KGE) methods have achieved great success in predicting missing links in knowledge graphs, a task also known as Knowledge Graph Completion (KGC). Under this task, the Reciprocal Rank (RR) of ground-truth items serve as a key indicator for evaluating the method’s performance. However, most existing studies have overlooked the inconsistency between the ranking metric, RR, and the optimization objective functions, resulting in sub-optimal KGC performance. To address this issue, we propose a KGC framework called KGCRR by designing a novel upper bound function named CRR. By introducing the parameter-pressure ρ to shift the sigmoid function, CRR achieves a better approximation to RR compared with existing objective functions. We theoretically proved that by adjusting ρ, CRR can achieve a more effective approximation to RR. By narrowing the discrepancy with RR and alleviating the gradient vanishing issue associated with the direct optimization of RR loss, CRR demonstrates an advantage in optimizing RR. CRR serves as a plug-and-play objective, capable of seamless integration into various KGE methods. Through extensive experiments conducted on FB15k-237 and WN18RR datasets, we have obtained promising results, with an average improvement of 19.06% in MRR, indicating that CRR significantly enhances the performance of existing methods.

EAAI Journal 2025 Journal Article

Memory guided representation learning for cross-domain face anti-spoofing

  • Pengchao Deng
  • Yanhui Zhou
  • Zhiheng Fu
  • Jian Liu
  • Shengjun Xu
  • Chenyang Ge
  • Farid Boussaid
  • Mohammed Bennamoun

Addressing generalized Face Anti-Spoofing is challenging due to the wide variety of spoofing techniques, variations in environmental conditions, and the diversity of devices used to capture images. Most approaches to enhancing the generalization ability of systems often manipulate image statistics to normalize images into a uniform representation space. However, this approach can restrict the capacity of the system to represent images accurately because image statistics vary significantly across different domains, each with unique characteristics. Recognizing that each source domain possesses distinct characteristics, we introduce an innovative approach based on memory guided representation learningto represent these characteristics in separate latent spaces. Specifically, we present a dual component framework comprising a Memory Guided Clustering Representation (MGCR)generator and a Memory Guided Mapping Representation (MGMR)classifier. Additionally, we create a memory bank filled with template and meta features, which are refined over time using a momentum update mechanism. The MGCR component employs clustering to allow new, unseen deep features close to the most similar template feature, thereby creating generalized clues for the generator. Meanwhile, the MGMR process leverages meta features to dynamically represent new domain spaces through linear combinations, bridging the divide between known and unknown domains. Our experimental results show that the proposed method significantly enhances the effectiveness under cross-domain scenarios, outperforming existing techniques. Especially in the “Leave One Out”setting, its average Half Total Error Rate exceeds that of other methods, reaching less than 10%.

NeurIPS Conference 2025 Conference Paper

Mesh-RFT: Enhancing Mesh Generation via Fine-grained Reinforcement Fine-Tuning

  • Jian Liu
  • Jing Xu
  • Song Guo
  • Jing Li
  • jingfeng Guo
  • Jiaao Yu
  • Haohan Weng
  • Biwen Lei

Existing pretrained models for 3D mesh generation often suffer from data biases and produce low-quality results, while global reinforcement learning (RL) methods rely on object-level rewards that struggle to capture local structure details. To address these challenges, we present $\textbf{Mesh-RFT}$, a novel fine-grained reinforcement fine-tuning framework that employs Masked Direct Preference Optimization (M-DPO) to enable localized refinement via quality-aware face masking. To facilitate efficient quality evaluation, we introduce an objective topology-aware scoring system to evaluate geometric integrity and topological regularity at both object and face levels through two metrics: Boundary Edge Ratio (BER) and Topology Score (TS). By integrating these metrics into a fine-grained RL strategy, Mesh-RFT becomes the first method to optimize mesh quality at the granularity of individual faces, resolving localized errors while preserving global coherence. Experiment results show that our M-DPO approach reduces Hausdorff Distance (HD) by 24. 6\% and improves Topology Score (TS) by 3. 8\% over pre-trained models, while outperforming global DPO methods with a 17. 4\% HD reduction and 4. 9\% TS gain. These results demonstrate Mesh-RFT’s ability to improve geometric integrity and topological regularity, achieving new state-of-the-art performance in production-ready mesh generation.

JBHI Journal 2025 Journal Article

Multimodal Drug Target Binding Affinity Prediction Using Graph Local Substructure

  • Xun Peng
  • Chunping Ouyang
  • Yongbin Liu
  • Ying Yu
  • Jian Liu
  • Min Chen

Predicting the binding affinity of drug target is essential to reduce drug development costs and cycles. Recently, several deep learning-based methods have been proposed to utilize the structural or sequential information of drugs and targets to predict the drug-target binding affinity (DTA). However, methods that rely solely on sequence features do not consider hydrogen atom data, which may result in information loss. Graph-based methods may contain information that is not directly related to the prediction process. Additionally, the lack of structured division can limit the representation of characteristics. To address these issues, we propose a multimodal DTA prediction model using graph local substructures, called MLSDTA. This model comprehensively integrates the graph and sequence modal information from drugs and targets, achieving multimodal fusion through a cross-attention approach for multimodal features. Additionally, adaptive structure aware pooling is applied to generate graphs containing local substructural information. The model also utilizes the DropNode strategy to enhance the distinctions between different molecules. Experiments on two benchmark datasets have shown that MLSDTA outperforms current state-of-the-art models, demonstrating the feasibility of MLSDTA.

JBHI Journal 2025 Journal Article

PCLT-PPI: Predicting Multi-Type Interactions Between Proteins Based on Point Cloud Structure and Local Topology Preservation

  • Minglei Li
  • Yurui Hou
  • Shuqin Wang
  • Jinmao Wei
  • Jian Liu

Protein-protein interactions (PPIs) play a crucial role in cellular biochemical reactions. Computationally mining PPI can help us better understand cellular regulatory mechanisms. Most existing methods focus on the linear structure of proteins, ignoring the influence of native spatial structure on their properties. Furthermore, when neural networks are used to learn protein embeddings, the nonlinear transformations may change the topological relationships between proteins. To address the above issues, we propose a PPI prediction method based on protein point cloud structure and local topology preservation, naming it PCLT-PPI. It extracts structural features from protein point cloud structures and relational features through graph neural networks. Throughout the process, PCLT-PPI maintains the local topology of proteins in their origin and embedding spaces. Experimental results show that, under three test set partition modes (Random, BFS, DFS) and four evaluation metrics (F1, AUC, AUPR, Hamming Loss), PCLT-PPI performs better than several state-of-the-art PPI prediction methods, especially when predicting protein PPIs that are not visible during training, exhibiting stronger robustness and higher generalization ability. The results also demonstrate that point cloud structure and local topology preservation can improve PPI prediction performance, which may provide a reference for subsequent related research.

AAAI Conference 2025 Conference Paper

PhysAug: A Physical-guided and Frequency-based Data Augmentation for Single-Domain Generalized Object Detection

  • Xiaoran Xu
  • Jiangang Yang
  • Wenhui Shi
  • Siyuan Ding
  • Luqing Luo
  • Jian Liu

Single-Domain Generalized Object Detection (S-DGOD) aims to train on a single source domain for robust performance across a variety of unseen target domains by taking advantage of an object detector. Existing S-DGOD approaches often rely on data augmentation strategies, including a composition of visual transformations, to enhance the detector's generalization ability. However, the absence of real-world prior knowledge hinders data augmentation from contributing to the diversity of training data distributions. To address this issue, we propose PhysAug, a novel physical model-based non-ideal imaging condition data augmentation method, to enhance the adaptability of the S-DGOD tasks. Drawing upon the principles of atmospheric optics, we develop a universal perturbation model that serves as the foundation for our proposed PhysAug. Given that visual perturbations typically arise from the interaction of light with atmospheric particles, the image frequency spectrum is harnessed to simulate real-world variations during training. This approach fosters the detector to learn domain-invariant representations, thereby enhancing its ability to generalize across various settings. Without altering the network architecture or loss function, our approach significantly outperforms the state-of-the-art across various S-DGOD datasets. In particular, it achieves a substantial improvement of 7.3% and 7.2% over the baseline on DWD and Cityscape-C, highlighting its enhanced generalizability in real-world settings.

EAAI Journal 2025 Journal Article

Pixel-Level Semantics Boosted Fine-Grained Bird Image Classification

  • Haoxiang Ma
  • Yongjian Deng
  • Bochen Xie
  • Jian Liu
  • Hai Liu
  • Youfu Li
  • Zhen Yang

Fine-grained bird image classification (FBIC) is crucial for endangered bird conservation and biodiversity research. However, existing methods often struggle to capture detailed features and manage the interference caused by complex backgrounds. To address these challenges, we propose a novel Pixel-Level Semantic Boosted Fine-Grained Bird Image Classification (PFIC) framework, which enhances fine-grained bird image classification by incorporating pixel-level semantic information. PFIC consists of two core components: the Grouped Detail Enhancement (GDE) module and the Background–Foreground Enhancement (BFE) strategy. GDE integrates multi-level pixel-level semantic information, derived from a segmentation feature extractor, into classification features via two submodules: grouped aggregation and detail enhancement. This approach enhances the model’s ability to capture fine-grained details. BFE augments training samples by restricting background ranges and applying random shifts to foreground objects, thereby improving the model’s capability to recognize foreground objects in complex environments. Experimental results demonstrate that our method achieves state-of-the-art performance on the CUB-200-2011 and NABirds datasets. Additionally, further experiments on the Stanford Cars dataset validate the framework’s potential for generalization to other fine-grained image classification tasks.

NeurIPS Conference 2025 Conference Paper

PSBench: a large-scale benchmark for estimating the accuracy of protein complex structural models

  • Pawan Neupane
  • Jian Liu
  • Jianlin Cheng

Predicting protein complex structures is essential for protein function analysis, protein design, and drug discovery. While AI methods like AlphaFold can predict accurate structural models for many protein complexes, reliably estimating the quality of these predicted models (estimation of model accuracy, or EMA) for model ranking and selection remains a major challenge. A key barrier to developing effective machine learning-based EMA methods is the lack of large, diverse, and well-annotated datasets for training and evaluation. To address this gap, we introduce PSBench, a benchmark suite comprising five large-scale, labeled datasets, four of which were generated during the 15th and 16th community-wide Critical Assessment of Protein Structure Prediction (CASP15 and CASP16), and one curated for new Protein Data Bank (PDB) entries deposited between July 2024 and August 2025. PSBench includes over 1. 4 million structural models covering a wide range of protein sequence lengths, complex stoichiometries, functional classes, and modeling difficulties. Each model is annotated with multiple complementary quality scores at the global, local, and interface levels. PSBench also provides multiple evaluation metrics and baseline EMA methods to facilitate rigorous comparisons. To demonstrate PSBench’s utility, we trained and evaluated GATE, a graph transformer-based EMA method, on the CASP15 data. GATE was blindly tested in CASP16 (2024), where it ranked among the top-performing EMA methods. These results highlight PSBench as a valuable resource for advancing EMA research in protein complex modeling. PSBench is publicly available at: https: //github. com/BioinfoMachineLearning/PSBench.

JBHI Journal 2025 Journal Article

Valence-Arousal Disentangled Representation Learning for Emotion Recognition in SSVEP-Based BCIs

  • Yipeng Du
  • Jie Chen
  • Zhengwu Liu
  • Ngai Wong
  • Chi Zhang
  • Zhiwei Ding
  • Jian Liu
  • Edith C.H. Ngai

Steady state visually evoked potential (SSVEP)-based brain-computer interfaces (BCIs), which are widely used in rehabilitation and disability assistance, can benefit from real-time emotion recognition to enhance human–machine interaction. However, the learned discri-minative latent representations in SSVEP-BCIs may generalize in an unintended direction, which can lead to reduced accuracy in detecting emotional states. In this paper, we introduce a Valence-Arousal Disentangled Representation Learning (VADL) method, drawing inspir-ation from the classical two-dimensional emotional model, to enhance the performance and generalization of emotion recognition within SSVEP-BCIs. VADL distinctly disentangles the latent variables of valence and arousal information to improve accuracy. It utilizes the structured state space duality model to thoroughly extract global emotional features. Additionally, we propose a Multisubject Gradient Blending training strategy that individually tailors the learning pace of reconstruction and discrimination tasks within VADL on-the-fly. To verify the feasibility of our method, we have developed a comprehensive database comprising 23 subjects, in which both the emotional states and SSVEPs were effectively elicited. Experimental results indicate that VADL surpasses existing state-of-the-art benchmark algorithms.

JBHI Journal 2024 Journal Article

A Lightweight Hybrid Model Using Multiscale Markov Transition Field for Real-Time Quality Assessment of Photoplethysmography Signals

  • Jian Liu
  • Shuaicong Hu
  • Ya'nan Wang
  • Qihan Hu
  • Daomiao Wang
  • Cuiwei Yang

Objective: The proliferation of wearable devices has escalated the standards for photoplethysmography (PPG) signal quality. This study introduces a lightweight model to address the imperative need for precise, real-time evaluation of PPG signal quality, followed by its deployment and validation utilizing our integrated upper computer and hardware system. Methods: Multiscale Markov Transition Fields (MMTF) are employed to enrich the morphological information of the signals, serving as the input for our proposed hybrid model (HM). HM undergoes initial pre-training utilizing the MIMIC-III and UCI databases, followed by fine-tuning the Queensland dataset. Knowledge distillation (KD) then transfers the large-parameter model's knowledge to the lightweight hybrid model (LHM). LHM is subsequently deployed on the upper computer for real-time signal quality assessment. Results: HM achieves impressive accuracies of 99. 1% and 96. 0% for binary and ternary classification, surpassing current state-of-the-art methods. LHM, with only 0. 2 M parameters (0. 44% of HM), maintains high accuracy despite a 2. 6% drop. It achieves an inference speed of 0. 023 s per image, meeting real-time display requirements. Furthermore, LHM attains a 97. 7% accuracy on a self-created database. HM outperforms current methods in PPG signal quality accuracy, demonstrating the effectiveness of our approach. Additionally, LHM substantially reduces parameter count while maintaining high accuracy, enhancing efficiency and practicality for real-time applications. Conclusion: The proposed methodology demonstrates the capability to achieve high-precision and real-time assessment of PPG signal quality, and its practical validation has been successfully conducted during deployment. Significance: This study contributes a convenient and accurate solution for the real-time evaluation of PPG signals, offering extensive application potential.

EAAI Journal 2024 Journal Article

A two-stage framework for pixel-level pavement surface crack detection

  • Feng Guo
  • Jian Liu
  • Quanyi Xie
  • Huayang Yu

Surface crack is one of the most common distresses of pavement structure, impacting its serviceability and sustainability. Over the past decade, many efforts have been devoted to developing computer vision-based models (e. g. , image processing- or deep learning-based) for the automatic detection of pavement surface crack. However, there is a great gap between the public image data taken by the phone or other portable devices and the real-world image data acquired by the linear array charge-coupled device (CCD) camera, which usually is high resolution and contains limited crack pixels per image. To improve the pavement surface crack detection efficiency and accuracy in engineering practice, we propose a novel two-stage framework for automatic pavement surface detection at the pixel level. In stage I, the images concluding pixel cracks are selected using a convolutional neural network (CNN)-based classification network. In stage II, the selected images are processed with our proposed separation-combination strategy and the second version of crack transformer (CTv2) for pavement surface crack detection at the pixel level. Comprehensive experimental investigation and comparison have been conducted on training performance and visualization results, validating the superiority of the developed framework. It paves the way for the large-scale application of automatic pavement crack detection in an efficient manner.

NeurIPS Conference 2024 Conference Paper

AFBench: A Large-scale Benchmark for Airfoil Design

  • Jian Liu
  • Jianyu Wu
  • Hairun Xie
  • Guoqing Zhang
  • Jing Wang
  • Wei Liu
  • Wanli Ouyang
  • Junjun Jiang

Data-driven generative models have emerged as promising approaches towards achieving efficient mechanical inverse design. However, due to prohibitively high cost in time and money, there is still lack of open-source and large-scale benchmarks in this field. It is mainly the case for airfoil inverse design, which requires to generate and edit diverse geometric-qualified and aerodynamic-qualified airfoils following the multimodal instructions, \emph{i. e. ,} dragging points and physical parameters. This paper presents the open-source endeavors in airfoil inverse design, \emph{AFBench}, including a large-scale dataset with 200 thousand airfoils and high-quality aerodynamic and geometric labels, two novel and practical airfoil inverse design tasks, \emph{i. e. ,} conditional generation on multimodal physical parameters, controllable editing, and comprehensive metrics to evaluate various existing airfoil inverse design methods. Our aim is to establish \emph{AFBench} as an ecosystem for training and evaluating airfoil inverse design methods, with a specific focus on data-driven controllable inverse design models by multimodal instructions capable of bridging the gap between ideas and execution, the academic research and industrial applications. We have provided baseline models, comprehensive experimental observations, and analysis to accelerate future research. Our baseline model is trained on an RTX 3090 GPU within 16 hours. The codebase, datasets and benchmarks will be available at \url{https: //hitcslj. github. io/afbench/}.

TCS Journal 2024 Journal Article

Constructions of rotation symmetric Boolean functions satisfying almost all cryptographic criteria

  • Lei Sun
  • Zexia Shi
  • Jian Liu
  • Fang-Wei Fu

Constructions of Boolean functions with various cryptographic properties have always been an important challenge in cryptography. This paper proposes systematic constructions of even-variable rotation symmetric Boolean functions satisfying almost all cryptographic criteria, that is, resiliency, optimal algebraic degree, strict avalanche criterion, high nonlinearity, nonexistence of nonzero linear structures, good global avalanche characteristics. Moreover, some of the constructions also have high algebraic immunity. This is the first time that Boolean functions having such cryptographic properties are obtained, which can be considered as good candidates for the design of real-life encryption schemes.

AAAI Conference 2024 Conference Paper

Divide-and-Aggregate Learning for Evaluating Performance on Unlabeled Data

  • Shuyu Miao
  • Jian Liu
  • Lin Zheng
  • Hong Jin

Artificial Intelligence (AI) models have become an integral part of modern society, significantly improving human lives. However, ensuring the reliability and safety of these models is of paramount importance. One critical aspect is the continuous monitoring and verification of model performance to prevent any potential risks. Real-time online evaluation of AI models is necessary to maintain their effectiveness and mitigate any harm caused by performance degradation. The traditional approach to model evaluation involves supervised methods that rely on manual labeling to compare results with model predictions. Unfortunately, this method is not suitable for online model monitoring due to its inherent lag and high cost. While there have been attempts to explore free-label model evaluation, these approaches often consider only the global features of the entire dataset. Additionally, they can only perform model evaluation based on a single dimension of model confidence or features. In this paper, we propose a novel approach called Divide-and-Aggregate Learning (DAL) for unsupervised model evaluation. Our method addresses the limitations of previous approaches by dividing the output of the model into buckets, capturing local information of the distribution. We then aggregate this local information to obtain global information and further represent the relationship between the distribution and model performance. Importantly, our method can simultaneously handle the confidence distribution and feature distribution of the model output. Extensive experiments have been conducted to demonstrate the effectiveness of our DAL model. The results show that our approach outperforms previous methods on four widely used datasets. We will make our source code publicly available.

IJCAI Conference 2024 Conference Paper

EAB-FL: Exacerbating Algorithmic Bias through Model Poisoning Attacks in Federated Learning

  • Syed Irfan Ali Meerza
  • Jian Liu

Federated Learning (FL) is a technique that allows multiple parties to train a shared model collaboratively without disclosing their private data. It has become increasingly popular due to its distinct privacy advantages. However, FL models can suffer from biases against certain demographic groups (e. g. , racial and gender groups) due to the heterogeneity of data and party selection. Researchers have proposed various strategies for characterizing the group fairness of FL algorithms to address this issue. However, the effectiveness of these strategies in the face of deliberate adversarial attacks has not been fully explored. Although existing studies have revealed various threats (e. g. , model poisoning attacks) against FL systems caused by malicious participants, their primary aim is to decrease model accuracy, while the potential of leveraging poisonous model updates to exacerbate model unfairness remains unexplored. In this paper, we propose a new type of model poisoning attack, EAB-FL, with a focus on exacerbating group unfairness while maintaining a good level of model utility. Extensive experiments on three datasets demonstrate the effectiveness and efficiency of our attack, even with state-of-the-art fairness optimization algorithms and secure aggregation rules employed. We hope this work will help the community fully understand the attack surfaces of current FL systems and facilitate corresponding mitigation to improve their resilience.

YNICL Journal 2024 Journal Article

Large-scale effective connectivity analysis reveals the existence of two mutual inhibitory systems in patients with major depression

  • Jia Wang
  • Baojuan Li
  • Jian Liu
  • Jiaming Li
  • Adeel Razi
  • Kaizhong Zheng
  • Baoyu Yan
  • Huaning Wang

It is posited that cognitive and affective dysfunction in patients with major depression disorder (MDD) may be caused by dysfunctional signal propagation in the brain. By leveraging dynamic causal modeling, we investigated large-scale directed signal propagation (effective connectivity) among distributed large-scale brain networks with 43 MDD patients and 56 healthy controls. The results revealed the existence of two mutual inhibitory systems: the anterior default mode network, auditory network, sensorimotor network, salience network and visual networks formed an "emotional" brain, while the posterior default mode network, central executive networks, cerebellum and dorsal attention network formed a "rational brain". These two networks exhibited excitatory intra-system connectivity and inhibitory inter-system connectivity. Patients were characterized by potentiated intra-system connections within the "emotional/sensory brain", as well as over-inhibition of the "rational brain" by the "emotional/sensory brain". The hierarchical architecture of the large-scale effective connectivity networks was then analyzed using a PageRank algorithm which revealed a shift of the controlling role of the "rational brain" to the "emotional/sensory brain" in the patients. These findings inform basic organization of distributed large-scale brain networks and furnish a better characterization of the neural mechanisms of depression, which may facilitate effective treatment.

AIIM Journal 2024 Journal Article

Non-invasive fractional flow reserve derived from reduced-order coronary model and machine learning prediction of stenosis flow resistance

  • Yili Feng
  • Ruisen Fu
  • Hao Sun
  • Xue Wang
  • Yang Yang
  • Chuanqi Wen
  • Yaodong Hao
  • Yutong Sun

Background and objective Recently, computational fluid dynamics enables the non-invasive calculation of fractional flow reserve (FFR) based on 3D coronary model, but it is time-consuming. Currently, machine learning technique has emerged as an efficient and reliable approach for prediction, which allows saving a lot of analysis time. This study aimed at developing a simplified FFR prediction model for rapid and accurate assessment of functional significance of stenosis. Methods A reduced-order lumped parameter model (LPM) of coronary system and cardiovascular system was constructed for rapidly simulating coronary flow, in which a machine learning model was embedded for accurately predicting stenosis flow resistance at a given flow from anatomical features of stenosis. Importantly, the LPM was personalized in both structures and parameters according to coronary geometries from computed tomography angiography and physiological measurements such as blood pressure and cardiac output for personalized simulations of coronary pressure and flow. Coronary lesions with invasive FFR ≤ 0. 80 were defined as hemodynamically significant. Results A total of 91 patients (93 lesions) who underwent invasive FFR were involved in FFR derived from machine learning (FFRML) calculation. Of the 93 lesions, 27 lesions (29. 0%) showed lesion-specific ischemia. The average time of FFRML simulation was about 10 min. On a per-vessel basis, the FFRML and FFR were significantly correlated (r = 0. 86, p < 0. 001). The diagnostic accuracy, sensitivity, specificity, positive predictive value and negative predictive value were 91. 4%, 92. 6%, 90. 9%, 80. 6% and 96. 8%, respectively. The area under the receiver-operating characteristic curve of FFRML was 0. 984. Conclusion In this selected cohort of patients, the FFRML improves the computational efficiency and ensures the accuracy. The favorable performance of FFRML approach greatly facilitates its potential application in detecting hemodynamically significant coronary stenosis in future routine clinical practice.

ICML Conference 2024 Conference Paper

Position: TrustLLM: Trustworthiness in Large Language Models

  • Yue Huang 0001
  • Lichao Sun 0001
  • Haoran Wang 0005
  • Siyuan Wu 0001
  • Qihui Zhang
  • Yuan Li
  • Chujie Gao
  • Yixin Huang

Large language models (LLMs) have gained considerable attention for their excellent natural language processing capabilities. Nonetheless, these LLMs present many challenges, particularly in the realm of trustworthiness. This paper introduces TrustLLM, a comprehensive study of trustworthiness in LLMs, including principles for different dimensions of trustworthiness, established benchmark, evaluation, and analysis of trustworthiness for mainstream LLMs, and discussion of open challenges and future directions. Specifically, we first propose a set of principles for trustworthy LLMs that span eight different dimensions. Based on these principles, we further establish a benchmark across six dimensions including truthfulness, safety, fairness, robustness, privacy, and machine ethics. We then present a study evaluating 16 mainstream LLMs in TrustLLM, consisting of over 30 datasets. Our findings firstly show that in general trustworthiness and capability (i. e. , functional effectiveness) are positively related. Secondly, our observations reveal that proprietary LLMs generally outperform most open-source counterparts in terms of trustworthiness, raising concerns about the potential risks of widely accessible open-source LLMs. However, a few open-source LLMs come very close to proprietary ones, suggesting that open-source models can achieve high levels of trustworthiness without additional mechanisms like moderator, offering valuable insights for developers in this field. Thirdly, it is important to note that some LLMs may be overly calibrated towards exhibiting trustworthiness, to the extent that they compromise their utility by mistakenly treating benign prompts as harmful and consequently not responding. Besides these observations, we’ve uncovered key insights into the multifaceted trustworthiness in LLMs. We emphasize the importance of ensuring transparency not only in the models themselves but also in the technologies that underpin trustworthiness. We advocate that the establishment of an AI alliance between industry, academia, the open-source community to foster collaboration is imperative to advance the trustworthiness of LLMs.

EAAI Journal 2024 Journal Article

Research on predictive modeling method of loader working resistance in a sensor-less environment

  • Shaojie Wang
  • Shuilin Huang
  • Liang Hou
  • Tianlin Hu
  • Jifang Li
  • Jian Liu

In view of the inconvenient installation and high cost of the current multi-sensor data prediction methods for predicting loader working resistance, this study proposes a method oriented towards predicting loader working resistance in environments with fewer sensors. First, building on previous research (Wu et al. , 2023), non-essential sensor features are removed by a maximum information coefficient (MIC)method that incorporates expert experience. Second, the Optuna automation framework is embedded to realize the training and testing of the proposed method and compare its prediction performance with other popular methods. Finally, in order to verify its generalization performance, it is validated using loader operation data under different working conditions. The results of this study demonstrate that the proposed method effectively and accurately characterizes the work resistance of loaders under operating conditions. With short testing times and excellent generalization performance, the method proves highly applicable and valuable.

IJCAI Conference 2024 Conference Paper

Strengthening Layer Interaction via Dynamic Layer Attention

  • Kaishen Wang
  • Xun Xia
  • Jian Liu
  • Zhang Yi
  • Tao He

In recent years, employing layer attention to enhance interaction among hierarchical layers has proven to be a significant advancement in building network structures. In this paper, we delve into the distinction between layer attention and the general attention mechanism, noting that existing layer attention methods achieve layer interaction on fixed feature maps in a static manner. These static layer attention methods limit the ability for context feature extraction among layers. To restore the dynamic context representation capability of the attention mechanism, we propose a Dynamic Layer Attention (DLA) architecture. The DLA comprises dual paths, where the forward path utilizes an improved recurrent neural network block, named Dynamic Sharing Unit (DSU), for context feature extraction. The backward path updates features using these shared context representations. Finally, the attention mechanism is applied to these dynamically refreshed feature maps among layers. Experimental results demonstrate the effectiveness of the proposed DLA architecture, outperforming other state-of-the-art methods in image recognition and object detection tasks. Additionally, the DSU block has been evaluated as an efficient plugin in the proposed DLA architecture. The code is available at https: //github. com/tunantu/Dynamic-Layer-attention.

ICML Conference 2024 Conference Paper

SyCoCa: Symmetrizing Contrastive Captioners with Attentive Masking for Multimodal Alignment

  • Ziping Ma 0002
  • Furong Xu
  • Jian Liu
  • Ming Yang 0007
  • Qingpei Guo

Multimodal alignment between language and vision is the fundamental topic in current vision-language model research. Contrastive Captioners (CoCa), as a representative method, integrates Contrastive Language-Image Pretraining (CLIP) and Image Caption (IC) into a unified framework, resulting in impressive results. CLIP imposes a bidirectional constraints on global representations of entire images and sentences. Although IC conducts an unidirectional image-to-text generation on local representation, it lacks any constraint on local text-to-image reconstruction, which limits the ability to understand images at a fine-grained level when aligned with texts. To achieve multimodal alignment from both global and local perspectives, this paper proposes Symmetrizing Contrastive Captioners (SyCoCa), which introduces bidirectional interactions on images and texts across the global and local representation levels. Specifically, we expand a Text-Guided Masked Image Modeling (TG-MIM) head based on ITC and IC heads. The improved SyCoCa further leverages textual cues to reconstruct contextual images and visual cues to predict textual contents. When implementing bidirectional local interactions, the local contents of images tend to be cluttered or unrelated to their textual descriptions. Thus, we employ an attentive masking strategy to select effective image patches for interaction. Extensive experiments on five vision-language tasks, including image-text retrieval, image-captioning, visual question answering, and zero-shot/finetuned image classification, validate the effectiveness of our proposed method.

AAAI Conference 2024 Conference Paper

Video Event Extraction with Multi-View Interaction Knowledge Distillation

  • Kaiwen Wei
  • Runyan Du
  • Li Jin
  • Jian Liu
  • Jianhua Yin
  • Linhao Zhang
  • Jintao Liu
  • Nayu Liu

Video event extraction (VEE) aims to extract key events and generate the event arguments for their semantic roles from the video. Despite promising results have been achieved by existing methods, they still lack an elaborate learning strategy to adequately consider: (1) inter-object interaction, which reflects the relation between objects; (2) inter-modality interaction, which aligns the features from text and video modality. In this paper, we propose a Multi-view Interaction with knowledge Distillation (MID) framework to solve the above problems with the Knowledge Distillation (KD) mechanism. Specifically, we propose the self-Relational KD (self-RKD) to enhance the inter-object interaction, where the relation between objects is measured by distance metric, and the high-level relational knowledge from the deeper layer is taken as the guidance for boosting the shallow layer in the video encoder. Meanwhile, to improve the inter-modality interaction, the Layer-to-layer KD (LKD) is proposed, which integrates additional cross-modal supervisions (i.e., the results of cross-attention) with the textual supervising signal for training each transformer decoder layer. Extensive experiments show that without any additional parameters, MID achieves the state-of-the-art performance compared to other strong methods in VEE.

JBHI Journal 2023 Journal Article

A Privacy-Preserving Medical Data Sharing Scheme Based on Blockchain

  • Guangquan Xu
  • Chen Qi
  • Wenyu Dong
  • Lixiao Gong
  • Shaoying Liu
  • Si Chen
  • Jian Liu
  • Xi Zheng

With the increasing penetration of the Internet of things (IoT) into people's lives, the limitations of traditional medical systems are emerging. First, the typical way of handling sensitive information can easily lead to privacy disclosure. Second, the medical system is relatively isolated. It is difficult for one medical system to share data with another, and the scope of users' activities is limited within the system boundary. To solve these two problems, we propose a new privacy-preserving medical data-sharing scheme by introducing the authorization mechanism and attribute-based encryption (ABE) based on blockchain, which breaks system boundaries and realizes data sharing among several medical institutions. ABE is used to realize scalable access control. In addition, doctors can share their knowledge to diagnose users by introducing many-to-many matching, which means that patients' health data can be represented by multiple keywords and doctors' expertise can be represented by multiple interests. We provide the correctness and security analysis of our scheme and implement a prototype tool on Ethereum. The experimental results show that our scheme solves the contradiction between the privacy preservation of medical data and the necessity of data sharing.

AAAI Conference 2023 Conference Paper

Generating Transferable 3D Adversarial Point Cloud via Random Perturbation Factorization

  • Bangyan He
  • Jian Liu
  • Yiming Li
  • Siyuan Liang
  • Jingzhi Li
  • Xiaojun Jia
  • Xiaochun Cao

Recent studies have demonstrated that existing deep neural networks (DNNs) on 3D point clouds are vulnerable to adversarial examples, especially under the white-box settings where the adversaries have access to model parameters. However, adversarial 3D point clouds generated by existing white-box methods have limited transferability across different DNN architectures. They have only minor threats in real-world scenarios under the black-box settings where the adversaries can only query the deployed victim model. In this paper, we revisit the transferability of adversarial 3D point clouds. We observe that an adversarial perturbation can be randomly factorized into two sub-perturbations, which are also likely to be adversarial perturbations. It motivates us to consider the effects of the perturbation and its sub-perturbations simultaneously to increase the transferability for sub-perturbations also contain helpful information. In this paper, we propose a simple yet effective attack method to generate more transferable adversarial 3D point clouds. Specifically, rather than simply optimizing the loss of perturbation alone, we combine it with its random factorization. We conduct experiments on benchmark dataset, verifying our method's effectiveness in increasing transferability while preserving high efficiency.

JBHI Journal 2023 Journal Article

Integrating Medical Domain Knowledge for Early Diagnosis of Fever of Unknown Origin: An Interpretable Hierarchical Multimodal Neural Network Approach

  • Zhixiao Wang
  • Jian Liu
  • Yu Tian
  • Tianshu Zhou
  • Qianghua Liu
  • Yunqing Qiu
  • Jingsong Li

Accurate and interpretable differential diagnostic technologies are crucial for supporting clinicians in decision-making and treatment-planning for patients with fever of unknown origin (FUO). Existing solutions commonly address the diagnosis of FUO by transforming it into a multi-classification task. However, after the emergence of COVID-19 pandemic, clinicians have recognized the heightened significance of early diagnosis in patients with FUO, particularly for practical needs such as early triage. This has resulted in increased demands for identifying a wider range of etiologies, shorter observation windows, and better model interpretability. In this article, we propose an interpretable hierarchical multimodal neural network framework (iHMNNF) to facilitate early diagnosis of FUO by incorporating medical domain knowledge and leveraging multimodal clinical data. The iHMNNF comprises a top-down hierarchical reasoning framework (Td-HRF) built on the class hierarchy of FUO etiologies, five local attention-based multimodal neural networks (La-MNNs) trained for each parent node of the class hierarchy, and an interpretable module based on layer-wise relevance propagation (LRP) and attention mechanism. Experimental datasets were collected from electronic health records (EHRs) at a large-scale tertiary grade-A hospital in China, comprising 34, 051 hospital admissions of 30, 794 FUO patients from January 2011 to October 2020. Our proposed La-MNNs achieved area under the receiver operating characteristic curve (AUROC) values ranging from 0. 7809 to 0. 9035 across all five decomposed tasks, surpassing competing machine learning (ML) and single-modality deep learning (DL) methods while also providing enhanced interpretability. Furthermore, we explored the feasibility of identifying FUO etiologies using only the first N -hour time series data obtained after admission.

NeurIPS Conference 2023 Conference Paper

Label-efficient Segmentation via Affinity Propagation

  • Wentong Li
  • Yuqian Yuan
  • Song Wang
  • Wenyu Liu
  • Dongqi Tang
  • Jian Liu
  • Jianke Zhu
  • Lei Zhang

Weakly-supervised segmentation with label-efficient sparse annotations has attracted increasing research attention to reduce the cost of laborious pixel-wise labeling process, while the pairwise affinity modeling techniques play an essential role in this task. Most of the existing approaches focus on using the local appearance kernel to model the neighboring pairwise potentials. However, such a local operation fails to capture the long-range dependencies and ignores the topology of objects. In this work, we formulate the affinity modeling as an affinity propagation process, and propose a local and a global pairwise affinity terms to generate accurate soft pseudo labels. An efficient algorithm is also developed to reduce significantly the computational cost. The proposed approach can be conveniently plugged into existing segmentation networks. Experiments on three typical label-efficient segmentation tasks, i. e. box-supervised instance segmentation, point/scribble-supervised semantic segmentation and CLIP-guided semantic segmentation, demonstrate the superior performance of the proposed approach.

EAAI Journal 2023 Journal Article

Multi-layer additive tensor decomposition of infrared video for titanium alloy tensile testing

  • Tao Zhang
  • Jian Liu
  • Yibo Ai
  • Weidong Zhang

Infrared video (in mathematics terms, tensor) has been widely used in the tensile testing of metallic materials, such as titanium alloy and steel. The infrared video of the tensile testing process can effectively and efficiently determine the properties of metallic materials, e. g. , Young’s modulus, Poisson’s ratio, yield strength, etc. The infrared video with structural properties, such as smoothness and sparsity, can be used to characterize the tensile testing process. To extract the features in the infrared video with structural properties, we propose a multi-layer additive tensor decomposition (MLATD) method based on regularization tensor regression for tensile testing. It decomposes a tensor into three classes of components: the multi-smooth layers (including background and foreground), the sparse layers (including between-tensor and in-tensor), and the noise layer. The scree plot is proposed to determine the number of multi-smooth layers, which is a downward curve of the difference between the smooth layers and the sparse layer. The alternating direction method of multipliers (ADMM) algorithm is proposed to solve the proposed method. The decomposition results of the simulation data and real-world case study revealed that the proposed method outperforms the existing state-of-the-art methods.

TMLR Journal 2023 Journal Article

On the Robustness of Dataset Inference

  • Sebastian Szyller
  • Rui Zhang
  • Jian Liu
  • N Asokan

Machine learning (ML) models are costly to train as they can require a significant amount of data, computational resources and technical expertise. Thus, they constitute valuable intellectual property that needs protection from adversaries wanting to steal them. Ownership verification techniques allow the victims of model stealing attacks to demonstrate that a suspect model was in fact stolen from theirs. Although a number of ownership verification techniques based on watermarking or fingerprinting have been proposed, most of them fall short either in terms of security guarantees (well-equipped adversaries can evade verification) or computational cost. A fingerprinting technique, Dataset Inference (DI) has been shown to offer better robustness and efficiency than prior methods. The authors of DI provided a correctness proof for linear (suspect) models. However, in a subspace of the same setting, we prove that DI suffers from high false positives (FPs) -- it can incorrectly identify an independent model trained with non-overlapping data from the same distribution as stolen. We further prove that DI also triggers FPs in realistic, non-linear suspect models. We then confirm empirically that DI in the black-box setting leads to FPs, with high confidence. Second, we show that DI also suffers from false negatives (FNs) -- an adversary can fool DI by regularising a stolen model's decision boundaries using adversarial training, thereby leading to an FN. To this end, we demonstrate that black-box DI fails to identify a model adversarially trained from a stolen dataset -- the setting where DI is the hardest to evade. Finally, we discuss the implications of our findings, the viability of fingerprinting-based ownership verification in general, and suggest directions for future work.

JBHI Journal 2023 Journal Article

Predicting Drug-Protein Interactions by Self-Adaptively Adjusting the Topological Structure of the Heterogeneous Network

  • Rong Tang
  • Chang Sun
  • Jipeng Huang
  • Minglei Li
  • Jinmao Wei
  • Jian Liu

Many powerful computational methods based on graph neural networks (GNNs) have been proposed to predict drug-protein interactions (DPIs). It can effectively reduce laboratory workload and the cost of drug discovery and drug repurposing. However, many clinical functions of drugs and proteins are unknown due to their unobserved indications. Therefore, it is difficult to establish a reliable drug-protein heterogeneous network that can describe the relationships between drugs and proteins based on the available information. To solve this problem, we propose a DPI prediction method that can self-adaptively adjust the topological structure of the heterogeneous networks, and name it SATS. SATS establishes a representation learning module based on graph attention network to carry out the drug-protein heterogeneous network. It can self-adaptively learn the relationships among the nodes based on their attributes and adjust the topological structure of the network according to the training loss of the model. Finally, SATS predicts the interaction propensity between drugs and proteins based on their embeddings. The experimental results show that SATS can effectively improve the topological structure of the network. The performance of SATS outperforms several state-of-the-art DPI prediction methods under various evaluation metrics. These prove that SATS is useful to deal with incomplete data and unreliable networks. The case studies on the top section of the prediction results further demonstrate that SATS is powerful for discovering novel DPIs.

JBHI Journal 2023 Journal Article

Semi-Supervised Learning for Low-Cost Personalized Obstructive Sleep Apnea Detection Using Unsupervised Deep Learning and Single-Lead Electrocardiogram

  • Shuaicong Hu
  • Ya'nan Wang
  • Jian Liu
  • Cuiwei Yang
  • Aiguo Wang
  • Kuanzheng Li
  • Wenxin Liu

Objective: Obstructive sleep apnea (OSA) is a common sleep-related breathing disorder that can lead to a wide range of health issues if left untreated. This study aims to address the lack of research on personalized models for single-lead electrocardiogram (ECG)-based OSA detection, by proposing an automatic semi-supervised algorithm for automated low-cost personalization fine-tuning. Methods: We utilize a convolutional neural network (CNN)-based auto-encoder (AE) with a modified training objective to detect anomalous region of OSA. An indicator based on model outputs is utilized as a benchmark measure to assign pseudo-labels with confidence to each sample. Finally, we perform validation of the semi-supervised algorithm on the same database and cross-database scenarios. Results: By introducing semi-supervised personalization, the accuracy, AUC, and mean absolute error (MAE) of the general model (GM) of 35 subjects from the same database are improved from 86. 3%, 0. 915, and 5. 178 to 90. 3%, 0. 948, and 2. 593. Simultaneously, in the validation of 25 subjects from a cross-database, the accuracy, AUC, and MAE of the GM are enhanced from 75. 6%, 0. 800, and 9. 149 to 84. 3%, 0. 881, and 3. 509. Conclusion: The improved version of AE demonstrates excellent adaptability in identifying abnormal features in OSA, employing a data-driven approach to assign pseudo-labels for unknown data automatically. Additionally, leveraging the pseudo-labels through a semi-supervised fine-tuning strategy provides a solution to overcome the limitation of clinical annotations, facilitating low-cost implementation of personalized models. Significance: The semi-supervised approach proposed in this article provides a high-performance and annotation-free solution for personalized adjustment of automatic OSA detection.

IJCAI Conference 2022 Conference Paper

Low-Resource NER by Data Augmentation With Prompting

  • Jian Liu
  • Yufeng Chen
  • Jinan Xu

Named entity recognition (NER) is a fundamental information extraction task that seeks to identify entity mentions of certain types in text. Despite numerous advances, the existing NER methods rely on extensive supervision for model training, which struggle in a low-resource scenario with limited training data. In this paper, we propose a new data augmentation method for low-resource NER, by eliciting knowledge from BERT with prompting strategies. Particularly, we devise a label-conditioned word replacement strategy that can produce more label-consistent examples by capturing the underlying word-label dependencies, and a prompting with question answering method to generate new training data from unlabeled texts. The experimental results have widely confirmed the effectiveness of our approach. Particularly, in a low-resource scenario with only 150 training sentences, our approach outperforms previous methods without data augmentation by over 40% in F1 and prior best data augmentation methods by over 2. 0% in F1. Furthermore, our approach also fits with a zero-shot scenario, yielding promising results without using any human-labeled data for the task.

AAAI Conference 2022 Conference Paper

Privacy-Preserving Face Recognition in the Frequency Domain

  • Yinggui Wang
  • Jian Liu
  • Man Luo
  • Le Yang
  • Li Wang

Some applications require performing face recognition (FR) on third-party servers, which could be accessed by attackers with malicious intents to compromise the privacy of users’ face information. This paper advocates a practical privacypreserving frequency-domain FR scheme without key management. The new scheme first collects the components with the same frequency from different blocks of a face image to form component channels. Only part of the channels are retained and fed into the analysis network that performs an interpretable privacy-accuracy trade-off analysis to identify channels important for face image visualization but not crucial for maintaining high FR accuracy. For this purpose, the loss function of the analysis network consists of the empirical FR error loss and a face visualization penalty term, and the network is trained in an end-to-end manner. We find that with the developed analysis network, more than 94% of the image energy can be dropped while the face recognition accuracy stays almost undegraded. In order to further protect the remaining frequency components, we propose a fast masking method. Effectiveness of the new scheme in removing the visual information of face images while maintaining their distinguishability is validated over several large face datasets. Results show that the proposed scheme achieves a recognition performance and inference time comparable to ArcFace operating on original face images directly.

NeurIPS Conference 2022 Conference Paper

SNN-RAT: Robustness-enhanced Spiking Neural Network through Regularized Adversarial Training

  • Jianhao Ding
  • Tong Bu
  • Zhaofei Yu
  • Tiejun Huang
  • Jian Liu

Spiking neural networks (SNNs) are promising to be widely deployed in real-time and safety-critical applications with the advance of neuromorphic computing. Recent work has demonstrated the insensitivity of SNNs to small random perturbations due to the discrete internal information representation. The variety of training algorithms and the involvement of the temporal dimension pose more threats to the robustness of SNNs than that of typical neural networks. We account for the vulnerability of SNNs by constructing adversaries based on different differentiable approximation techniques. By deriving a Lipschitz constant specifically for the spike representation, we first theoretically answer the question of how much adversarial invulnerability is retained in SNNs. Hence, to defend against the broad attack methods, we propose a regularized adversarial training scheme with low computational overheads. SNNs can benefit from the constraint of the perturbed spike distance's amplification and the generalization on multiple adversarial $\epsilon$-neighbourhoods. Our experiments on the image recognition benchmarks have proven that our training scheme can defend against powerful adversarial attacks crafted from strong differentiable approximations. To be specific, our approach makes the black-box attacks of the Projected Gradient Descent attack nearly ineffective. We believe that our work will facilitate the spread of SNNs for safety-critical applications and help understand the robustness of the human brain.

IJCAI Conference 2021 Conference Paper

Cross-Domain Slot Filling as Machine Reading Comprehension

  • Mengshi Yu
  • Jian Liu
  • Yufeng Chen
  • Jinan Xu
  • Yujie Zhang

With task-oriented dialogue systems being widely applied in everyday life, slot filling, the essential component of task-oriented dialogue systems, is required to be quickly adapted to new domains that contain domain-specific slots with few or no training data. Previous methods for slot filling usually adopt sequence labeling framework, which, however, often has limited ability when dealing with the domain-specific slots. In this paper, we take a new perspective on cross-domain slot filling by framing it as a machine reading comprehension (MRC) problem. Our approach firstly transforms slot names into well-designed queries, which contain rich informative prior knowledge and are very helpful for the detection of domain-specific slots. In addition, we utilize the large-scale MRC dataset for pre-training, which further alleviates the data scarcity problem. Experimental results on SNIPS and ATIS datasets show that our approach consistently outperforms the existing state-of-the-art methods by a large margin.

IJCAI Conference 2021 Conference Paper

Discourse-Level Event Temporal Ordering with Uncertainty-Guided Graph Completion

  • Jian Liu
  • Jinan Xu
  • Yufeng Chen
  • Yujie Zhang

Learning to order events at discourse-level is a crucial text understanding task. Despite many efforts for this task, the current state-of-the-art methods rely heavily on manually designed features, which are costly to produce and are often specific to tasks/domains/datasets. In this paper, we propose a new graph perspective on the task, which does not require complex feature engineering but can assimilate global features and learn inter-dependencies effectively. Specifically, in our approach, each document is considered as a temporal graph, in which the nodes and edges represent events and event-event relations respectively. In this sense, the temporal ordering task corresponds to constructing edges for an empty graph. To train our model, we design a graph mask pre-training mechanism, which can learn inter-dependencies of temporal relations by learning to recover a masked edge following graph topology. In the testing stage, we design an certain-first strategy based on model uncertainty, which can decide the prediction orders and reduce the risk of error propagation. The experimental results demonstrate that our approach outperforms previous methods consistently and can meanwhile maintain good global consistency.

AAAI Conference 2021 Conference Paper

Enabling Fast and Universal Audio Adversarial Attack Using Generative Model

  • Yi Xie
  • Zhuohang Li
  • Cong Shi
  • Jian Liu
  • Yingying Chen
  • Bo Yuan

Recently, the vulnerability of deep neural network (DNN)based audio systems to adversarial attacks has obtained increasing attention. However, the existing audio adversarial attacks allow the adversary to possess the entire user’s audio input as well as granting sufficient time budget to generate the adversarial perturbations. These idealized assumptions, however, make the existing audio adversarial attacks mostly impossible to be launched in a timely fashion in practice (e. g. , playing unnoticeable adversarial perturbations along with user’s streaming input). To overcome these limitations, in this paper we propose fast audio adversarial perturbation generator (FAPG), which uses generative model to generate adversarial perturbations for the audio input in a single forward pass, thereby drastically improving the perturbation generation speed. Built on the top of FAPG, we further propose universal audio adversarial perturbation generator (UAPG), a scheme to craft universal adversarial perturbation that can be imposed on arbitrary benign audio input to cause misclassification. Extensive experiments on DNN-based audio systems show that our proposed FAPG can achieve high success rate with up to 214× speedup over the existing audio adversarial attack methods. Also our proposed UAPG generates universal adversarial perturbations that can achieve much better attack performance than the state-of-the-art solutions.

IJCAI Conference 2021 Conference Paper

Improving Stylized Neural Machine Translation with Iterative Dual Knowledge Transfer

  • Xuanxuan Wu
  • Jian Liu
  • Xinjie Li
  • Jinan Xu
  • Yufeng Chen
  • Yujie Zhang
  • Hui Huang

Stylized neural machine translation (NMT) aims to translate sentences of one style into sentences of another style, which is essential for the application of machine translation in a real-world scenario. However, a major challenge in this task is the scarcity of high-quality parallel data which is stylized paired. To address this problem, we propose an iterative dual knowledge transfer framework that utilizes informal training data of machine translation and formality style transfer data to create large-scale stylized paired data, for the training of stylized machine translation model. Specifically, we perform bidirectional knowledge transfer between translation model and text style transfer model iteratively through knowledge distillation. Then, we further propose a data-refinement module to process the noisy synthetic parallel data generated during knowledge transfer. Experiment results demonstrate the effectiveness of our method, achieving an improvement over the existing best model by 5 BLEU points on MTFC dataset. Meanwhile, extensive analyses illustrate our method can also improve the accuracy of formality style transfer.

AAAI Conference 2021 Conference Paper

Natural Language Inference in Context – Investigating Contextual Reasoning over Long Texts

  • Hanmeng Liu
  • Leyang Cui
  • Jian Liu
  • Yue Zhang

Natural language inference (NLI) is a fundamental NLP task, investigating the entailment relationship between two texts. Popular NLI datasets present the task at sentence-level. While adequate for testing semantic representations, they fall short for testing contextual reasoning over long texts, which is a natural part of the human inference process. We introduce ConTRoL, a new dataset for ConTextual Reasoning over Long Texts. Consisting of 8, 325 expert-designed “contexthypothesis” pairs with gold labels, ConTRoL is a passagelevel NLI dataset with a focus on complex contextual reasoning types such as logical reasoning. It is derived from competitive selection and recruitment test (verbal reasoning test) for police recruitment, with expert level quality. Compared with previous NLI benchmarks, the materials in ConTRoL are much more challenging, involving a range of reasoning types. Empirical results show that state-of-the-art language models perform by far worse than educated humans. Our dataset can also serve as a testing-set for downstream tasks like checking the factual correctness of summaries.

AAAI Conference 2020 Conference Paper

Adversarial Training Based Multi-Source Unsupervised Domain Adaptation for Sentiment Analysis

  • Yong Dai
  • Jian Liu
  • Xiancong Ren
  • Zenglin Xu

Multi-source unsupervised domain adaptation (MS-UDA) for sentiment analysis (SA) aims to leverage useful information in multiple source domains to help do SA in an unlabeled target domain that has no supervised information. Existing algorithms of MS-UDA either only exploit the shared features, i. e. , the domain-invariant information, or based on some weak assumption in NLP, e. g. , smoothness assumption. To avoid these problems, we propose two transfer learning frameworks based on the multi-source domain adaptation methodology for SA by combining the source hypotheses to derive a good target hypothesis. The key feature of the first framework is a novel Weighting Scheme based Unsupervised Domain Adaptation framework (WS-UDA), which combine the source classifiers to acquire pseudo labels for target instances directly. While the second framework is a Two-Stage Training based Unsupervised Domain Adaptation framework (2ST-UDA), which further exploits these pseudo labels to train a target private extractor. Importantly, the weights assigned to each source classifier are based on the relations between target instances and source domains, which measured by a discriminator through the adversarial training. Furthermore, through the same discriminator, we also fulfill the separation of shared features and private features. Experimental results on two SA datasets demonstrate the promising performance of our frameworks, which outperforms unsupervised state-of-the-art competitors.

IROS Conference 2020 Conference Paper

Designing A Dummy Skin by Evaluating Contacts between A Human Hand and A Robot End Tip

  • Yumena Iki
  • Yoji Yamada
  • Yasuhiro Akiyama
  • Shogo Okamoto
  • Jian Liu

Many manufacturing industries have a high demand for the construction of collaborative operation systems using industrial robots. Although there is a preexisting set of safety verification data in ISO/TS 15066 for collaborative operations, there is no established testing method for safety validation. To establish a testing method, it is effective to use a dummy that has mechanical properties similar to those of a human. However, there is no parametric study that exists for designing a dummy that represents the static and dynamic mechanical properties and the nonlinearity of the static mechanical properties. In this study, static and dynamic experiments were conducted to obtain the mechanical stiffness of human subjects and the contact force transitions during the dynamic contact between a robot system and a human. Subsequently, the same experiment was conducted using the proposed dummy. The biofidelity of the dummy was examined by comparing the parameters of a viscoelastic model. This study contributes to increasing the safety of collaborative operations by developing a dummy that can be used for risk assessments in collaborative operations.

IJCAI Conference 2020 Conference Paper

Knowledge Enhanced Event Causality Identification with Mention Masking Generalizations

  • Jian Liu
  • Yubo Chen
  • Jun Zhao

Identifying causal relations of events is a crucial language understanding task. Despite many efforts for this task, existing methods lack the ability to adopt background knowledge, and they typically generalize poorly to new, previously unseen data. In this paper, we present a new method for event causality identification, aiming to address limitations of previous methods. On the one hand, our model can leverage external knowledge for reasoning, which can greatly enrich the representation of events; On the other hand, our model can mine event-agnostic, context-specific patterns, via a mechanism called event mention masking generalization, which can greatly enhance the ability of our model to handle new, previously unseen cases. In experiments, we evaluate our model on three benchmark datasets and show our model outperforms previous methods by a significant margin. Moreover, we perform 1) cross-topic adaptation, 2) exploiting unseen predicates, and 3) cross-task adaptation to evaluate the generalization ability of our model. Experimental results show that our model demonstrates a definite advantage over previous methods.

IJCAI Conference 2020 Conference Paper

LogiQA: A Challenge Dataset for Machine Reading Comprehension with Logical Reasoning

  • Jian Liu
  • Leyang Cui
  • Hanmeng Liu
  • Dandan Huang
  • Yile Wang
  • Yue Zhang

Machine reading is a fundamental task for testing the capability of natural language understand- ing, which is closely related to human cognition in many aspects. With the rising of deep learning techniques, algorithmic models rival human performances on simple QA, and thus increasingly challenging machine reading datasets have been proposed. Though various challenges such as evidence integration and commonsense knowledge have been integrated, one of the fundamental capabilities in human reading, namely logical reasoning, is not fully investigated. We build a comprehensive dataset, named LogiQA, which is sourced from expert-written questions for testing human Logical reasoning. It consists of 8, 678 QA instances, covering multiple types of deductive reasoning. Results show that state-of-the-art neural models perform by far worse than human ceiling. Our dataset can also serve as a benchmark for reinvestigating logical AI under the deep learning NLP setting. The dataset is freely available at https: //github. com/lgw863/LogiQA-dataset.

AAAI Conference 2019 Conference Paper

Exploiting the Ground-Truth: An Adversarial Imitation Based Knowledge Distillation Approach for Event Detection

  • Jian Liu
  • Yubo Chen
  • Kang Liu

The ambiguity in language expressions poses a great challenge for event detection. To disambiguate event types, current approaches rely on external NLP toolkits to build knowledge representations. Unfortunately, these approaches work in a pipeline paradigm and suffer from error propagation problem. In this paper, we propose an adversarial imitation based knowledge distillation approach, for the first time, to tackle the challenge of acquiring knowledge from rawsentences for event detection. In our approach, a teacher module is first devised to learn the knowledge representations from the ground-truth annotations. Then, we set up a student module that only takes the raw-sentences as the input. The student module is taught to imitate the behavior of the teacher under the guidance of an adversarial discriminator. By this way, the process of knowledge distillation from rawsentence has been implicitly integrated into the feature encoding stage of the student module. To the end, the enhanced student is used for event detection, which processes raw texts and requires no extra toolkits, naturally eliminating the error propagation problem faced by pipeline approaches. We conduct extensive experiments on the ACE 2005 datasets, and the experimental results justify the effectiveness of our approach.

ICRA Conference 2018 Conference Paper

Caging Loops in Shape Embedding Space: Theory and Computation

  • Jian Liu
  • Shiqing Xin
  • Zengfu Gao
  • Kai Xu 0004
  • Changhe Tu
  • Baoquan Chen

We propose to synthesize feasible caging grasps for a target object through computing Caging Loops, a closed curve defined in the shape embedding space of the object. Different from the traditional methods, our approach decouples caging loops from the surface geometry of target objects through working in the embedding space. This enables us to synthesize caging loops encompassing multiple topological holes, instead of always tied with one specific handle which could be too small to be graspable by the robot gripper. Our method extracts caging loops through a topological analysis of the distance field defined for the target surface in the embedding space, based on a rigorous theoretical study on the relation between caging loops and the field topology. Due to the decoupling, our method can tolerate incomplete and noisy surface geometry of an unknown target object captured on-the-fly. We implemented our method with a robotic gripper and demonstrate through extensive experiments that our method can synthesize reliable grasps for objects with complex surface geometry and topology and in various scales.

AAAI Conference 2018 Conference Paper

Event Detection via Gated Multilingual Attention Mechanism

  • Jian Liu
  • Yubo Chen
  • Kang Liu
  • Jun Zhao

Identifying event instance in text plays a critical role in building NLP applications such as Information Extraction (IE) system. However, most existing methods for this task focus only on monolingual clues of a specific language and ignore the massive information provided by other languages. Data scarcity and monolingual ambiguity hinder the performance of these monolingual approaches. In this paper, we propose a novel multilingual approach — dubbed as Gated MultiLingual Attention (GMLATT) framework — to address the two issues simultaneously. In specific, to alleviate data scarcity problem, we exploit the consistent information in multilingual data via context attention mechanism. Which takes advantage of the consistent evidence in multilingual data other than learning only from monolingual data. To deal with monolingual ambiguity problem, we propose gated cross-lingual attention to exploit the complement information conveyed by multilingual data, which is helpful for the disambiguation. The cross-lingual attention gate serves as a sentinel modelling the confidence of the clues provided by other languages and controls the information integration of various languages. We have conducted extensive experiments on the ACE 2005 benchmark. Experimental results show that our approach significantly outperforms state-of-the-art methods.

YNICL Journal 2016 Journal Article

Repeated acupuncture treatments modulate amygdala resting state functional connectivity of depressive patients

  • Xiaoyun Wang
  • Zengjian Wang
  • Jian Liu
  • Jun Chen
  • Xian Liu
  • Guangning Nie
  • Joon-Seok Byun
  • Yilin Liang

As a widely-applied alternative therapy, acupuncture is gaining popularity in Western society. One challenge that remains, however, is incorporating it into mainstream medicine. One solution is to combine acupuncture with other conventional, mainstream treatments. In this study, we investigated the combination effect of acupuncture and the antidepressant fluoxetine, as well as its underlying mechanism using resting state functional connectivity (rsFC) in patients with major depressive disorders. Forty-six female depressed patients were randomized into a verum acupuncture plus fluoxetine or a sham acupuncture plus fluoxetine group for eight weeks. Resting-state fMRI data was collected before the first and last treatments. Results showed that compared with those in the sham acupuncture treatment, verum acupuncture treatment patients showed 1) greater clinical improvement as indicated by Montgomery-Åsberg Depression Rating Scale (MADRS) and Self-Rating Depression Scale (SDS) scores; 2) increased rsFC between the left amygdala and subgenual anterior cingulate cortex (sgACC)/preguenual anterior cingulate cortex (pgACC); 3) increased rsFC between the right amygdala and left parahippocampus (Para)/putamen (Pu). The strength of the amygdala-sgACC/pgACC rsFC was positively associated with corresponding clinical improvement (as indicated by a negative correlation with MADRS and SDS scores). Our findings demonstrate the additive effect of acupuncture to antidepressant treatment and suggest that this effect may be achieved through the limbic system, especially the amygdala and the ACC.

TCS Journal 2013 Journal Article

On the relationships between perfect nonlinear functions and universal hash families

  • Jian Liu
  • Lusheng Chen

In this paper, the relationships between perfect nonlinear (in brief, PN) functions and optimal universal hash families are discussed. We point out the equivalence of constructions between them, i. e. , from PN functions, one can obtain optimal universal hash families and vice versa. As an application of our construction, a message authentication code is proposed, which provides better resistance to substitution attack than a known construction given by Carlet et al. in 2006. More generally, the connections between functions with given differential uniformity and some universal hash families are studied.

v2026.09.13