Arrow Research search

Author name cluster

Jia Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

57 papers
2 author rows

Possible papers

57

JBHI Journal 2026 Journal Article

ECG-AuxNet: A Dual-Branch Spatial-Temporal Feature Fusion Framework with Auxiliary Learning for Enhanced Cardiac Disease Diagnosis

  • Ruiqi Shen
  • Yanan Wang
  • Chunge Cao
  • Shuaicong Hu
  • Jia Liu
  • Hongyu Wang
  • Gaoyan Zhong
  • Cuiwei Yang

Objective: Multiple limitations exist in current automated ECG analysis, including insufficient feature integration across leads, limited interpretability, poor generalization, and inadequate handling of class imbalance. To address these challenges, we develop a novel dual-branch framework that comprehensively captures spatial-temporal features for cardiac disease diagnosis. Methods: ECG-AuxNet combines a Multi-scale Transformer Attention CNN for spatial feature extraction and a GRU network for temporal dependency modeling. A Dual-stage Cross-Attention Fusion module integrates features from both branches, while a Feature Space Reconstruction (FSR) auxiliary task is introduced as a manifold regularizer to enhance feature discrimination. The framework was evaluated on PTB-XL (15, 709 ECGs) and validated in real-world clinical scenarios (SXMU-2k, 1, 673 ECGs). Results: For class-imbalanced disease recognition (NORM, CD, MI, STTC), ECG-AuxNet attained 78. 34% F1-score on PTB-XL and 82. 63% F1-score on SXMU-2k, outperforming 9 baseline models. FSR significantly improved feature discrimination by 11. 7%, enhancing class boundary clarity and classification accuracy. Grad-CAM analysis revealed attention patterns that precisely match cardiologists' diagnostic focus areas. Conclusion: ECG-AuxNet effectively integrates spatial-temporal features through auxiliary learning, achieving robust generalizability in cardiac disease diagnosis with interpretability aligned with clinical expertise.

AAAI Conference 2026 Conference Paper

Multi-Objective Bilevel Learning

  • Zhiyao Zhang
  • Zhuqing Liu
  • Xin Zhang
  • Wen-Yen Chen
  • Jiyan Yang
  • Jia Liu

As machine learning (ML) applications grow increasingly complex in recent years, modern ML frameworks often need to address multiple potentially conflicting objectives with coupled decision variables across different layers. This creates a compelling need for multi-objective bilevel learning (MOBL). So far, however, the field of MOBL remains in its infancy and many important problems remain under-explored. This motivates us to fill this gap and systematically investigate the theoretical and algorithmic foundation of MOBL. Specifically, we consider MOBL problems with multiple conflicting objectives guided by preferences at the upper-level subproblem, where part of the inputs depend on the optimal solution of the lower-level subproblem. Our goal is to develop efficient MOBL optimization algorithms to (1) identify a preference-guided Pareto-stationary solution with low oracle complexity; and (2) enable systematic Pareto front exploration. To this end, we propose a unifying algorithmic framework called weighted-Chebyshev multi-hyper-gradient-descent (WC-MHGD) for both deterministic and stochastic settings with finite-time Pareto-stationarity convergence rate guarantees, which not only implies low oracle complexity but also induces systematic Pareto front exploration. We further conduct extensive experiments to confirm our theoretical results.

JBHI Journal 2026 Journal Article

Multi-Task Learning for OSA Detection and Sleep Staging via Multi-Scale Modeling

  • Zhiya Wang
  • Tian Yang
  • Yunfeng Zhu
  • Jia Liu
  • Peter A. Cistulli
  • Wei Chen

Obstructive sleep apnea (OSA) and sleep fragmentation are closely linked physiological phenomena that play crucial roles in the diagnosis and management of sleep disorders. While numerous deep learning models have been developed for either OSA detection or sleep stage classification, few attempts have been made to address both tasks simultaneously. To this end, we propose MT-TASPPNet (Multi-Task Triple Atrous Spatial Pyramid Pooling Network), a unified multi-modal multi-task network that jointly performs automatic OSA event detection and sleep staging. The model integrates modality-specific feature extractors for EEG, ECG, and airflow signals, and employs Atrous Spatial Pyramid Pooling modules in both the modality-specific and shared representation pathways to capture multi-scale temporal-frequency patterns. Additionally, an EOG-guided prior mechanism is incorporated to enhance the discrimination of subtle sleep stages. We use a 3-min input window (1-min target with ±1-min context) and evaluate our method on three large-scale datasets: SHHS1, SHHS2, and Sydney Sleep Biobank. The model achieves OSA detection accuracy between 0. 798 and 0. 884 (MF1: 0. 772 to 0. 821), and sleep staging accuracy between 0. 776 and 0. 834 (MF1: 0. 735 to 0. 749, $\mathcal {K}$: 0. 697 to 0. 77). Notably, the model maintains consistent performance despite data heterogeneity and individual variability. These results validate the stability and adaptability of MT-TASPPNet in clinical settings, paving the way for efficient and scalable multi-task sleep analysis systems.

AAAI Conference 2026 Conference Paper

Resource Efficient Sleep Staging via Multi-Level Masking and Prompt Learning

  • Lejun Ai
  • Yulong Li
  • Haodong Yi
  • Jixuan Xie
  • Yue Wang
  • Jia Liu
  • Min Chen
  • Rui Wang

Automatic sleep staging plays a vital role in assessing sleep quality and diagnosing sleep disorders. Most existing methods rely heavily on long and continuous EEG recordings, which poses significant challenges for data acquisition in resource-constrained systems, such as wearable or home-based monitoring systems. In this paper, we propose the task of resource-efficient sleep staging, which aims to reduce the amount of signal collected per sleep epoch while maintaining reliable classification performance. To solve this task, we adopt the masking and prompt learning strategy and propose a novel framework called Mask-Aware Sleep Staging (MASS). Specifically, we design a multi-level masking strategy to promote effective feature modeling under partial and irregular observations. To mitigate the loss of contextual information introduced by masking, we further propose a hierarchical prompt learning mechanism that aggregates unmasked data into a global prompt, serving as a semantic anchor for guiding both patch-level and epoch-level feature modeling. MASS is evalutaed on four datasets, demonstrating state-of-the-art performance, especially when the amount of data is very limited. This result highlights its potential for efficient and scalable deployment in real-world low-resource sleep monitoring environments.

AAAI Conference 2026 Conference Paper

SCo-Cloud: Satellite Constellation Collaboration for Cloud-Aware Onboard-Computed Imaging and Transmission

  • Jia Liu
  • Qian Li
  • Yongqi Li
  • Cheng Ji
  • Shangguang Wang

Satellite-acquired optical remote sensing imagery is extensively applied in time-critical applications like traffic surveillance and evaluation of natural disasters. However, clouds, as a common atmospheric phenomenon, frequently obscure observation. Current approaches aim to restore visibility in cloud-obscured regions, yet they typically fall short in the presence of dense cloud cover, which are exceedingly prevalent in remote sensing imagery. Alternative approaches rely on the satellite revisit cycle, frequently surpassing ten days, a duration impractical for genuine application scenarios due to target changes and bandwidth limitations. To address these issues, this paper proposes SCo-Cloud, a novel satellite constellation collaboration framework for cloud-aware onboard-computed imaging and transmission, which consists of Center-Sat and Edge-Sats. We propose onboard thin cloud removal and re-imaging region location models to locate the impact of clouds. We further design a novel multi-satellite scheduling strategy to eliminate clouds. The models above are integrated within the Center-Sat, with the nearby Edge-Sats collaborating in tandem to execute re-imaging assignments. Furthermore, to facilitate in-depth research, we have meticulously developed a cloud-covered target detection dataset. Comprehensive experiments have conclusively demonstrated that SCo-Cloud effectively surpasses the limitations inherent in current approaches, providing accurate and timely responses within the domain of Earth observation.

YNIMG Journal 2026 Journal Article

Trait motivation is associated with fusiform face area morphometry

  • Hangshek Lau
  • Yiying Song
  • Shan Xu
  • Jia Liu

Trait motivation is fundamental in shaping human behaviors. Previous studies have primarily focused on their impact on affective and motivational processing, with their role in perceptual processes less investigated. The present study takes face perception, a crucial and well-studied perceptual process, as a representative specimen to examine the perceptual effect of trait motivation. We investigated whether the behavioral activation system (BAS) and the behavioral inhibition system (BIS) were associated with structural characteristics of the inferior temporal face-selective regions as well as face recognition performance. With a sample of Chinese young adults (N = 264), voxel-based morphometry revealed that BIS scores correlated with greater gray matter volume in the fusiform face area. Further, a higher BIS score was associated with slightly better performance in face recognition. These findings provide novel evidence that trait motivation, particularly behavioral inhibition, is linked to both the structure and function of the face processing system. This underlines the intrinsic coupling between motivational and perceptual systems, blurring the presumed divide between affective and perceptual processes.

IROS Conference 2025 Conference Paper

A Soft Active Surface Gripper for Safe In Hand Manipulation of Fragile Objects

  • Sheng Xiang
  • Jiahao Li
  • Yinqi Zhang
  • Zhong Wei
  • Jia Liu
  • Yang Yang 0002

This paper introduces a soft active surface gripper designed to manipulate fragile objects safely. This gripper consists of two fingers, each equipped with two compliant pneumatic actuators and a soft active surface. The gripper utilizes the elastic belt as its soft active surface, which is driven by a motor, and the opening angle of the elastic band is controlled by pneumatic actuators. The novel design allows for the passive deformation of both the soft active surface and the compliant pneumatic actuator, enabling adaptation to various object shapes and demonstrating superior handling capabilities for delicate items. By synchronizing the opening and closing of the pneumatic fingers with the conveying motion of the active surface, the active surface gripper realizes three degrees of freedom (DOF) for in-plane manipulation, specifically two translational movements and one rotational movement. A prototype gripper has been designed and fabricated for in-plane manipulation experiments with fragile objects, including strawberries, miniature cupcakes, and pears. Experimental results demonstrate that the gripper can execute in-plane in-hand manipulation of fragile objects with varying geometries and dimensions while maintaining secure and robust handling, preventing object slippage and preserving surface integrity without causing damage.

EAAI Journal 2025 Journal Article

High-resolution multi-view stereo with multi-scale feature fusion

  • Dapeng Chen
  • Qi Jia
  • Hao Wu
  • Da Yu
  • Nanxuan Huang
  • Jia Liu

To enhance the handling of three-dimensional reconstruction for large-scale scenes and high-resolution images, we introduce a novel multi-view high-resolution three-dimensional reconstruction approach. Our proposed method integrates a Feature Pyramid Network with the Swin Transformer for improved performance. We integrate the Swin Transformer into the feature pyramid. This integration aims to establish long-range feature dependencies, facilitate information exchange between different input positions, and enhance the global consistency of feature representation. This improves the efficiency of the feature extraction stage. Following this, we apply cost volume regularization to mitigate noise and compute depth maps. A Depth Optimization Module is employed to refine the predicted depth maps, thereby enhancing their precision. Experimental results demonstrate the efficacy of our method in generating more accurate depth information, particularly in predicting high-resolution depth maps. Our approach utilizes these depth predictions to generate point clouds, enabling precise matching and reconstruction of multi-view images. Experiments conducted on public datasets validate the effectiveness and superiority of our proposed method.

AAAI Conference 2025 Conference Paper

In-Dataset Trajectory Return Regularization for Offline Preference-based Reinforcement Learning

  • Songjun Tu
  • Jingbo Sun
  • Qichao Zhang
  • Yaocheng Zhang
  • Jia Liu
  • Ke Chen
  • Dongbin Zhao

Offline preference-based reinforcement learning (PbRL) typically operates in two phases: first, use human preferences to learn a reward model and annotate rewards for a reward-free offline dataset; second, learn a policy by optimizing the learned reward via offline RL. However, accurately modeling step-wise rewards from trajectory-level preference feedback presents inherent challenges. The reward bias introduced, particularly the overestimation of predicted rewards, leads to optimistic trajectory stitching, which undermines the pessimism mechanism critical to the offline RL phase. To address this challenge, we propose In-Dataset Trajectory Return Regularization (DTR) for offline PbRL, which leverages conditional sequence modeling to mitigate the risk of learning inaccurate trajectory stitching under reward bias. Specifically, DTR employs Decision Transformer and TD-Learning to strike a balance between maintaining fidelity to the behavior policy with high in-dataset trajectory returns and selecting optimal actions based on high reward labels. Additionally, we introduce an ensemble normalization technique that effectively integrates multiple reward models, balancing the trade-off between reward differentiation and accuracy. Empirical evaluations on various benchmarks demonstrate the superiority of DTR over other state-of-the-art baselines.

AAAI Conference 2025 Conference Paper

Learning Complexity of Gradient Descent and Conjugate Gradient Algorithms

  • Xianqi Jiao
  • Jia Liu
  • Zhiping Chen

Gradient Descent (GD) and Conjugate Gradient (CG) methods are among the most effective iterative algorithms for solving unconstrained optimization problems, particularly in machine learning and statistical modeling, where they are employed to minimize cost functions. In these algorithms, tunable parameters, such as step sizes or conjugate parameters, play a crucial role in determining key performance metrics, like runtime and solution quality. In this work, we introduce a framework that models algorithm selection as a statistical learning problem, and thus learning complexity can be estimated by the pseudo-dimension of the algorithm group. We first propose a new cost measure for unconstrained optimization algorithms, inspired by the concept of primal-dual integral in mixed-integer linear programming. Based on the new cost measure, we derive an improved upper bound for the pseudo-dimension of gradient descent algorithm group by discretizing the set of step size configurations. Moreover, we generalize our findings from gradient descent algorithm to the conjugate gradient algorithm group for the first time, and prove the existence a learning algorithm capable of probabilistically identifying the optimal algorithm with a sufficiently large sample size.

ICML Conference 2025 Conference Paper

Look Twice Before You Answer: Memory-Space Visual Retracing for Hallucination Mitigation in Multimodal Large Language Models

  • Xin Zou 0001
  • Yizhou Wang
  • Yibo Yan
  • Yuanhuiyi Lyu
  • Kening Zheng
  • Sirui Huang
  • Junkai Chen
  • Peijie Jiang

Despite their impressive capabilities, Multimodal Large Language Models (MLLMs) are prone to hallucinations, i. e. , the generated content that is nonsensical or unfaithful to input sources. Unlike in LLMs, hallucinations in MLLMs often stem from the sensitivity of text decoder to visual tokens, leading to a phenomenon akin to "amnesia" about visual information. To address this issue, we propose MemVR, a novel decoding paradigm inspired by common cognition: when the memory of an image seen the moment before is forgotten, people will look at it again for factual answers. Following this principle, we treat visual tokens as supplementary evidence, re-injecting them into the MLLM through Feed Forward Network (FFN) as “key-value memory” at the middle trigger layer. This look-twice mechanism occurs when the model exhibits high uncertainty during inference, effectively enhancing factual alignment. Comprehensive experimental evaluations demonstrate that MemVR significantly mitigates hallucination across various MLLMs and excels in general benchmarks without incurring additional time overhead.

AAAI Conference 2025 Conference Paper

PSMGD: Periodic Stochastic Multi-Gradient Descent for Fast Multi-Objective Optimization

  • Mingjing Xu
  • Peizhong Ju
  • Jia Liu
  • Haibo Yang

Multi-objective optimization (MOO) lies at the core of many machine learning (ML) applications that involve multiple, potentially conflicting objectives (e.g., multi-task learning, multi-objective reinforcement learning, among many others). Despite the long history of MOO, recent years have witnessed a surge in interest within the ML community in the development of gradient manipulation algorithms for MOO, thanks to the availability of gradient information in many ML problems. However, existing gradient manipulation methods for MOO often suffer from long training times, primarily due to the need for computing dynamic weights by solving an additional optimization problem to determine a common descent direction that can decrease all objectives simultaneously. To address this challenge, we propose a new and efficient algorithm called Periodic Stochastic Multi-Gradient Descent (PSMGD) to accelerate MOO. PSMGD is motivated by the key observation that dynamic weights across objectives exhibit small changes under minor updates over short intervals during the optimization process. Consequently, our PSMGD algorithm is designed to periodically compute these dynamic weights and utilizes them repeatedly, thereby effectively reducing the computational overload. Theoretically, we prove that PSMGD can achieve state-of-the-art convergence rates for strongly-convex, general convex, and non-convex functions. Additionally, we introduce a new computational complexity measure, termed backpropagation complexity, and demonstrate that PSMGD could achieve an objective-independent backpropagation complexity. Through extensive experiments, we verify that PSMGD can provide comparable or superior performance to state-of-the-art MOO algorithms while significantly reducing training time.

YNIMG Journal 2025 Journal Article

The cognitive critical brain: Modulation of criticality in perception-related cortical regions

  • Xingyu Liu
  • Xiaotian Fei
  • Jia Liu

The constantly evolving world necessitates a brain that can swiftly adapt and respond to rapid changes. The brain, conceptualized as a system performing cognitive functions through collective neural activity, has been shown to maintain a resting state characterized by near-critical neural dynamics, positioning it to effectively respond to external stimuli. However, how near-criticality is dynamically modulated during task performance remains insufficiently understood. In this study, we utilized the prototypical Ising Hamiltonian model to investigate the modulation of near-criticality in neural activity at the cortical subsystem level during perceptual tasks. Specifically, we simulated 2D-Ising models in silico using structural MRI data and empirically estimated the system's state in vivo using functional MRI data. We first replicated previous findings that the resting state is typically near-critical as captured by the Ising model. Importantly, we observed heterogeneous changes in criticality across cortical subsystems during a naturalistic movie-watching task, with visual and auditory regions fine-tuned closer to criticality. A more fine-grained analysis of the ventral temporal cortex during an object recognition task further revealed that only regions selectively responsive to a specific object category were tuned closer to criticality when processing that object category. In conclusion, our study provides empirical evidence from the domain of perception supporting the cognitive critical brain hypothesis that modulating the criticality of subsystems within the brain's hierarchical and modular organization may be a fundamental mechanism for achieving diverse cognitive functions.

AAAI Conference 2025 Conference Paper

Towards a Multimodal Large Language Model with Pixel-Level Insight for Biomedicine

  • Xiaoshuang Huang
  • Lingdong Shen
  • Jia Liu
  • Fangxin Shang
  • Hongxiang Li
  • Haifeng Huang
  • Yehui Yang

In recent years, Multimodal Large Language Models (MLLM) have achieved notable advancements, demonstrating the feasibility of developing an intelligent biomedical assistant. However, current biomedical MLLMs predominantly focus on image-level understanding and restrict interactions to textual commands, thus limiting their capability boundaries and the flexibility of usage. In this paper, we introduce a novel end-to-end multimodal large language model for the biomedical domain, named MedPLIB, which possesses pixel-level understanding. Excitingly, it supports visual question answering (VQA), arbitrary pixel-level prompts (points, bounding boxes, and free-form shapes), and pixel-level grounding. We propose a novel Mixture-of-Experts (MoE) multi-stage training strategy, which divides MoE into separate training phases for a visual-language expert model and a pixel-grounding expert model, followed by fine-tuning using MoE. This strategy effectively coordinates multitask learning while maintaining the computational cost at inference equivalent to that of a single expert model. To advance the research of biomedical MLLMs, we introduce the Medical Complex Vision Question Answering Dataset (MeCoVQA), which comprises an array of 8 modalities for complex medical imaging question answering and image region understanding. Experimental results indicate that MedPLIB has achieved state-of-the-art outcomes across multiple medical visual language tasks. More importantly, in zero-shot evaluations for the pixel grounding task, MedPLIB leads the best small and large models by margins of 19.7 and 15.6 respectively on the mDice metric.

AAAI Conference 2024 Conference Paper

DrFuse: Learning Disentangled Representation for Clinical Multi-Modal Fusion with Missing Modality and Modal Inconsistency

  • Wenfang Yao
  • Kejing Yin
  • William K. Cheung
  • Jia Liu
  • Jing Qin

The combination of electronic health records (EHR) and medical images is crucial for clinicians in making diagnoses and forecasting prognoses. Strategically fusing these two data modalities has great potential to improve the accuracy of machine learning models in clinical prediction tasks. However, the asynchronous and complementary nature of EHR and medical images presents unique challenges. Missing modalities due to clinical and administrative factors are inevitable in practice, and the significance of each data modality varies depending on the patient and the prediction target, resulting in inconsistent predictions and suboptimal model performance. To address these challenges, we propose DrFuse to achieve effective clinical multi-modal fusion. It tackles the missing modality issue by disentangling the features shared across modalities and those unique within each modality. Furthermore, we address the modal inconsistency issue via a disease-wise attention layer that produces the patient- and disease-wise weighting for each modality to make the final prediction. We validate the proposed method using real-world large-scale datasets, MIMIC-IV and MIMIC-CXR. Experimental results show that the proposed method significantly outperforms the state-of-the-art models.

ICLR Conference 2024 Conference Paper

Multi-granularity Correspondence Learning from Long-term Noisy Videos

  • Yijie Lin 0001
  • Jie Zhang
  • Zhenyu Huang 0005
  • Jia Liu
  • Zujie Wen
  • Xi Peng 0001

Existing video-language studies mainly focus on learning short video clips, leaving long-term temporal dependencies rarely explored due to over-high computational cost of modeling long videos. To address this issue, one feasible solution is learning the correspondence between video clips and captions, which however inevitably encounters the multi-granularity noisy correspondence (MNC) problem. To be specific, MNC refers to the clip-caption misalignment (coarse-grained) and frame-word misalignment (fine-grained), hindering temporal learning and video understanding. In this paper, we propose NOise Robust Temporal Optimal traNsport (Norton) that addresses MNC in a unified optimal transport (OT) framework. In brief, Norton employs video-paragraph and clip-caption contrastive losses to capture long-term dependencies based on OT. To address coarse-grained misalignment in video-paragraph contrast, Norton filters out the irrelevant clips and captions through an alignable prompt bucket and realigns asynchronous clip-caption pairs based on transport distance. To address the fine-grained misalignment, Norton incorporates a soft-maximum operator to identify crucial words and key frames. Additionally, Norton exploits the potential faulty negative samples in clip-caption contrast by rectifying the alignment target with OT assignment to ensure precise temporal modeling. Extensive experiments on video retrieval, videoQA, and action segmentation verify the effectiveness of our method. Code is available at https://lin-yijie.github.io/projects/Norton.

AAMAS Conference 2024 Conference Paper

Sample and Communication Efficient Fully Decentralized MARL Policy Evaluation via a New Approach: Local TD Update

  • Hairi
  • Zifan Zhang
  • Jia Liu

In actor-critic framework for fully decentralized multi-agent reinforcement learning (MARL), one of the key components is the MARL policy evaluation (PE) problem, where a set of 𝑁 agents work cooperatively to evaluate the value function of the global states for a given policy through communicating with their neighbors. In MARL-PE, a critical challenge is how to lower the sample and communication complexities, which are defined as the number of training samples and communication rounds needed to converge to some 𝜖-stationary point. To lower communication complexity in MARL-PE, a “natural” idea is to perform multiple local TD-update steps between each consecutive rounds of communication to reduce the communication frequency. However, the validity of the local TD-update approach remains unclear due to the potential “agentdrift” phenomenon resulting from heterogeneous rewards across agents in general. This leads to an interesting open question: Can the local TD-update approach entail low sample and communication complexities? In this paper, we make the first attempt to answer this fundamental question. We focus on the setting of MARL-PE with average reward, which is motivated by many multi-agent network optimization problems. Our theoretical and experimental results confirm that allowing multiple local TD-update steps is indeed an effective approach in lowering the sample and communication complexities of MARL-PE compared to consensus-based MARL-PE algorithms. Specifically, the local TD-update steps between two consecutive communication rounds can be as large as O(1/𝜖1/2 log (1/𝜖)) in order to converge to an 𝜖-stationary point of MARL-PE. Moreover, we show theoretically that in order to reach the optimal sample complexity, the communication complexity of local TD-update approach is O(1/𝜖1/2 log (1/𝜖)).

AAAI Conference 2024 Conference Paper

SpectralNeRF: Physically Based Spectral Rendering with Neural Radiance Field

  • Ru Li
  • Jia Liu
  • Guanghui Liu
  • Shengping Zhang
  • Bing Zeng
  • Shuaicheng Liu

In this paper, we propose SpectralNeRF, an end-to-end Neural Radiance Field (NeRF)-based architecture for high-quality physically based rendering from a novel spectral perspective. We modify the classical spectral rendering into two main steps, 1) the generation of a series of spectrum maps spanning different wavelengths, 2) the combination of these spectrum maps for the RGB output. Our SpectralNeRF follows these two steps through the proposed multi-layer perceptron (MLP)-based architecture (SpectralMLP) and Spectrum Attention UNet (SAUNet). Given the ray origin and the ray direction, the SpectralMLP constructs the spectral radiance field to obtain spectrum maps of novel views, which are then sent to the SAUNet to produce RGB images of white-light illumination. Applying NeRF to build up the spectral rendering is a more physically-based way from the perspective of ray-tracing. Further, the spectral radiance fields decompose difficult scenes and improve the performance of NeRF-based methods. Comprehensive experimental results demonstrate the proposed SpectralNeRF is superior to recent NeRF-based methods when synthesizing new views on synthetic and real datasets. The codes and datasets are available at https://github.com/liru0126/SpectralNeRF.

EAAI Journal 2023 Journal Article

Area and power optimization for Fixed Polarity Reed–Muller logic circuits based on Multi-strategy Multi-objective Artificial Bee Colony algorithm

  • Dongge Qin
  • Zhenxue He
  • Xiaojun Zhao
  • Jia Liu
  • Fan Zhang
  • Limin Xiao

Area and power optimization of Fixed Polarity Reed–Muller (FPRM) circuits has received a lot of attention. Polarity optimization for FPRM circuits is essentially a binary multi-objective optimization problem. However, the existing area and power optimization approaches for FPRM logic circuits rarely produce a frontier and a greater number of Pareto optimal solutions. In this paper, a Multi-strategy Multi-objective Artificial Bee Colony (MMABC) algorithm is proposed to solve the binary multi-objective optimization problem. The main innovation of MMABC can be summarized as follows: a flexible foraging behavior strategy for employed bees is proposed to improve the searching ability of the algorithm; a genetic retention evolution for onlooker bees is proposed to improve the quality of the population; an efficient transform strategy is proposed to help the algorithm to jump out the local optimal and increase convergence speed. Moreover, we propose an area and power optimization approach for FPRM logic circuits, which uses the MMABC to search for the polarities (i. e. , Pareto optimal solutions) with smaller area and lower power. Experimental results demonstrated the effectiveness and superiority of our approach in optimizing area and power of FPRM logic circuits.

EAAI Journal 2023 Journal Article

Continual learning classification method with human-in-the-loop based on the artificial immune system

  • Jia Liu
  • Dong Li
  • Wangweiyi Shan
  • Shulin Liu

Currently, most classification algorithms may suffer from serious misclassification when faced with unbalanced or new data, for lacking feedback and update capability. Many existing classification methods learn from the continuous learning mechanism in the biological immune system to realize the learning of new data. However, these methods in classifying new data, because of the lack of feedback, lead to a search for a long time, even may deviate from the correct results. Humans can make the immune response proceed to a certain extent as they wish through a variety of intervention techniques. Motivated by this, this paper proposes a continual learning classification method with human-in-the-loop (H-CLCM) based on the artificial immune system. H-CLCM adjusts parameters online corresponding to misidentified data by integrating human experience during the testing phase. This enables it not only to converge to an accurate prediction model at minimal cost but also to have the capability to learn new classes of data without retraining the classifier. Experiments on 10 benchmark datasets demonstrate the performances and advantages of the proposed method. The experimental results show that H-CLCM has superior classification performance than the other approaches under the same data.

NeurIPS Conference 2023 Conference Paper

Federated Multi-Objective Learning

  • Haibo Yang
  • Zhuqing Liu
  • Jia Liu
  • Chaosheng Dong
  • Michinari Momma

In recent years, multi-objective optimization (MOO) emerges as a foundational problem underpinning many multi-agent multi-task learning applications. However, existing algorithms in MOO literature remain limited to centralized learning settings, which do not satisfy the distributed nature and data privacy needs of such multi-agent multi-task learning applications. This motivates us to propose a new federated multi-objective learning (FMOL) framework with multiple clients distributively and collaboratively solving an MOO problem while keeping their training data private. Notably, our FMOL framework allows a different set of objective functions across different clients to support a wide range of applications, which advances and generalizes the MOO formulation to the federated learning paradigm for the first time. For this FMOL framework, we propose two new federated multi-objective optimization (FMOO) algorithms called federated multi-gradient descent averaging (FMGDA) and federated stochastic multi-gradient descent averaging (FSMGDA). Both algorithms allow local updates to significantly reduce communication costs, while achieving the {\em same} convergence rates as those of their algorithmic counterparts in the single-objective federated learning. Our extensive experiments also corroborate the efficacy of our proposed FMOO algorithms.

AAAI Conference 2023 Conference Paper

Robust Domain Adaptation for Machine Reading Comprehension

  • Liang Jiang
  • Zhenyu Huang
  • Jia Liu
  • Zujie Wen
  • Xi Peng

Most domain adaptation methods for machine reading comprehension (MRC) use a pre-trained question-answer (QA) construction model to generate pseudo QA pairs for MRC transfer. Such a process will inevitably introduce mismatched pairs (i.e., Noisy Correspondence) due to i) the unavailable QA pairs in target documents, and ii) the domain shift during applying the QA construction model to the target domain. Undoubtedly, the noisy correspondence will degenerate the performance of MRC, which however is neglected by existing works. To solve such an untouched problem, we propose to construct QA pairs by additionally using the dialogue related to the documents, as well as a new domain adaptation method for MRC. Specifically, we propose Robust Domain Adaptation for Machine Reading Comprehension (RMRC) method which consists of an answer extractor (AE), a question selector (QS), and an MRC model. Specifically, RMRC filters out the irrelevant answers by estimating the correlation to the document via the AE, and extracts the questions by fusing the candidate questions in multiple rounds of dialogue chats via the QS. With the extracted QA pairs, MRC is fine-tuned and provides the feedback to optimize the QS through a novel reinforced self-training method. Thanks to the optimization of the QS, our method will greatly alleviate the noisy correspondence problem caused by the domain shift. To the best of our knowledge, this could be the first study to reveal the influence of noisy correspondence in domain adaptation MRC models and show a feasible solution to achieve the robustness against the mismatched pairs. Extensive experiments on three datasets demonstrate the effectiveness of our method.

NeurIPS Conference 2022 Conference Paper

A Stochastic Linearized Augmented Lagrangian Method for Decentralized Bilevel Optimization

  • Songtao Lu
  • Siliang Zeng
  • Xiaodong Cui
  • Mark Squillante
  • Lior Horesh
  • Brian Kingsbury
  • Jia Liu
  • Mingyi Hong

Bilevel optimization has been shown to be a powerful framework for formulating multi-task machine learning problems, e. g. , reinforcement learning (RL) and meta-learning, where the decision variables are coupled in both levels of the minimization problems. In practice, the learning tasks would be located at different computing resource environments, and thus there is a need for deploying a decentralized training framework to implement multi-agent and multi-task learning. We develop a stochastic linearized augmented Lagrangian method (SLAM) for solving general nonconvex bilevel optimization problems over a graph, where both upper and lower optimization variables are able to achieve a consensus. We also establish that the theoretical convergence rate of the proposed SLAM to the Karush-Kuhn-Tucker (KKT) points of this class of problems is on the same order as the one achieved by the classical distributed stochastic gradient descent for only single-level nonconvex minimization problems. Numerical results tested on multi-agent RL problems showcase the superiority of SLAM compared with the benchmarks.

JBHI Journal 2022 Journal Article

GNN-Based Depression Recognition Using Spatio-Temporal Information: A fNIRS Study

  • Qiao Yu
  • Rui Wang
  • Jia Liu
  • Long Hu
  • Min Chen
  • Zhongchun Liu

In recent years, depression has become an increasingly serious problem globally. Previous studies of automatic depression recognition based on functional near-Infrared spectroscopy (fNIRS) or other brain imaging techniques have shown potential to serve as auxiliary diagnosis methods that provide assistance to medical professionals. Recently, some studies have found that, besides directly using the data themselves (temporal data), the use of functional connectivity among channels (spatial data) also can be effective. In this paper, we propose a method based on Graph Neural Network (GNN) that combines both temporal and spatial features of fNIRS data for automatic depression recognition. Specifically, fNIRS data of 96 subjects were collected and pre-processed. Basic statistical metrics of each channel were extracted as temporal features, and channel connectivity (coherence and correlation) were calculated as spatial features. Point-biserial analysis was conducted on these features and depression labels as a data-driven motivation. For classification, we considered data of each subject as a graph, with temporal features as node features and spatial features as edge weights. The graphs were fed into GNNs for training and testing. Experimental results showed that our GNN-based methods realized the best depression recognition performance compared with classical machine-learning methods regarding accuracy, F1 score, and precision, especially in F1 score for over 10%.

JBHI Journal 2022 Journal Article

Pulse Taking by a Piezoelectric Film Sensor via Mode Energy Ratio Analysis Helps Identify Pregnancy Status

  • Jing Nie
  • Lulu Zhang
  • Jia Liu
  • Yaqin Wang

In order to solve the problem of non-invasive diagnosis and monitoring of women during pregnancy, a piezoelectric film pulse sensing system combined with the mode energy ratio (MER) analysis is utilized to detect human pulses to reveal pregnant conditions. Inspired by traditional Chinese medicine (TCM), pulse diagnosis has a history of more than 2, 500 years. The life energy of the human body helps the diagnosis of the disease through the circulation of blood vessels connected to the organs. A PVDF piezoelectric film sensor is used to emulate the pulse taking process in TCM to record the pulse signals. And the algorithm of MER is proposed based on empirical mode decomposition (EMD). Through the MER analysis of 83 female volunteers with different pregnancy statuses, the identification and warning of pregnancy status and physical health indicators are realized.

NeurIPS Conference 2022 Conference Paper

SAGDA: Achieving $\mathcal{O}(\epsilon^{-2})$ Communication Complexity in Federated Min-Max Learning

  • Haibo Yang
  • Zhuqing Liu
  • Xin Zhang
  • Jia Liu

Federated min-max learning has received increasing attention in recent years thanks to its wide range of applications in various learning paradigms. Similar to the conventional federated learning for empirical risk minimization problems, communication complexity also emerges as one of the most critical concerns that affects the future prospect of federated min-max learning. To lower the communication complexity of federated min-max learning, a natural approach is to utilize the idea of infrequent communications (through multiple local updates) same as in conventional federated learning. However, due to the more complicated inter-outer problem structure in federated min-max learning, theoretical understandings of communication complexity for federated min-max learning with infrequent communications remain very limited in the literature. This is particularly true for settings with non-i. i. d. datasets and partial client participation. To address this challenge, in this paper, we propose a new algorithmic framework called \ul{s}tochastic \ul{s}ampling \ul{a}veraging \ul{g}radient \ul{d}escent \ul{a}scent ($\mathsf{SAGDA}$), which i) assembles stochastic gradient estimators from randomly sampled clients as control variates and ii) leverages two learning rates on both server and client sides. We show that $\mathsf{SAGDA}$ achieves a linear speedup in terms of both the number of clients and local update steps, which yields an $\mathcal{O}(\epsilon^{-2})$ communication complexity that is orders of magnitude lower than the state of the art. Interestingly, by noting that the standard federated stochastic gradient descent ascent (FSGDA) is in fact a control-variate-free special version of $\mathsf{SAGDA}$, we immediately arrive at an $\mathcal{O}(\epsilon^{-2})$ communication complexity result for FSGDA. Therefore, through the lens of $\mathsf{SAGDA}$, we also advance the current understanding on communication complexity of the standard FSGDA method for federated min-max learning.

NeurIPS Conference 2022 Conference Paper

Taming Fat-Tailed (“Heavier-Tailed” with Potentially Infinite Variance) Noise in Federated Learning

  • Haibo Yang
  • Peiwen Qiu
  • Jia Liu

In recent years, federated learning (FL) has emerged as an important distributed machine learning paradigm to collaboratively learn a global model with multiple clients, while keeping data local and private. However, a key assumption in most existing works on FL algorithms' convergence analysis is that the noise in stochastic first-order information has a finite variance. Although this assumption covers all light-tailed (i. e. , sub-exponential) and some heavy-tailed noise distributions (e. g. , log-normal, Weibull, and some Pareto distributions), it fails for many fat-tailed noise distributions (i. e. , ``heavier-tailed'' with potentially infinite variance) that have been empirically observed in the FL literature. To date, it remains unclear whether one can design convergent algorithms for FL systems that experience fat-tailed noise. This motivates us to fill this gap in this paper by proposing an algorithmic framework called $\mathsf{FAT}$-$\mathsf{Clipping}~$ (\ul{f}ederated \ul{a}veraging with \ul{t}wo-sided learning rates and \ul{clipping}), which contains two variants: $\mathsf{FAT}$-$\mathsf{Clipping}~$ per-round ($\mathsf{FAT}$-$\mathsf{Clipping}$-$\mathsf{PR}$) and $\mathsf{FAT}$-$\mathsf{Clipping}~$ per-iteration ($\mathsf{FAT}$-$\mathsf{Clipping}$-$\mathsf{PI}$). Specifically, for the largest $\alpha \in (1, 2]$ such that the fat-tailed noise in FL still has a bounded $\alpha$-moment, we show that both variants achieve $\mathcal{O}((mT)^{\frac{2-\alpha}{\alpha}})$ and $\mathcal{O}((mT)^{\frac{1-\alpha}{3\alpha-2}})$ convergence rates in the strongly-convex and general non-convex settings, respectively, where $m$ and $T$ are the numbers of clients and communication rounds. Moreover, at the expense of more clipping operations compared to $\mathsf{FAT}$-$\mathsf{Clipping}$-$\mathsf{PR}$, $\mathsf{FAT}$-$\mathsf{Clipping}$-$\mathsf{PI}~$ further enjoys a linear speedup effect with respect to the number of local updates at each client and being lower-bound-matching (i. e. , order-optimal). Collectively, our results advance the understanding of designing efficient algorithms for FL systems that exhibit fat-tailed first-order oracle information.

JBHI Journal 2021 Journal Article

A Data-Driven Approach to Transfer Function Analysis for Superior Discriminative Power: Optimized Assessment of Dynamic Cerebral Autoregulation

  • Jia Liu
  • Zhen-Ni Guo
  • David Simpson
  • Pandeng Zhang
  • Chang Liu
  • Jia-Ning Song
  • Xinyi Leng
  • Yi Yang

Transfer function analysis (TFA) is extensively used to assess human physiological functions. However, extracting parameters from TFA is not usually optimized for detecting impaired function. In this study, we propose to use data-driven approaches to improve the performance of TFA in assessing blood flow control in the brain (dynamic cerebral autoregulation, dCA). Data were collected from two distinct groups of subjects deemed to have normal and impaired dCA. Continuous arterial blood pressure (ABP) and cerebral blood flow velocity (CBFV) were simultaneously recorded for approximately 10 mins in 82 subjects (including 41 healthy controls) to give 328 labeled samples of the TFA variables. The recordings were further divided into 4, 294 short data segments to generate 17, 176 unlabeled samples of the TFA variables. We optimized TFA post-processing with a generic semi-supervised learning strategy and a novel semi-supervised stacked ensemble learning (SSEL) strategy for classification into normal and impaired dCA. The generic strategy led to a performance with no significant difference to that of the conventional dCA analysis methods, whereas the proposed new strategy boosted the performance of TFA to an accuracy of 93. 3%. To our knowledge, this is the best dCA discrimination performance obtained to date and the first attempt at optimizing TFA through machine learning techniques. Equivalent methods can potentially also be applied to assessing a wide spectrum of other human physiological functions.

YNIMG Journal 2021 Journal Article

Development of navigation network revealed by resting-state and task-state functional connectivity

  • Xin Hao
  • Taicheng Huang
  • Yiying Song
  • Xiangzhen Kong
  • Jia Liu

Humans possess the essential capacity to navigate in environment, supported by multiple brain regions constituting the navigation network. Recent studies on development of the navigation network mainly examined activation changes in the medial temporal regions. It is unclear how the large-scale organization of the whole navigation network develops and whether the network organizations under resting-state and task-state develop differently. We addressed these questions by examining functional connectivity (FC) of the navigation network in 122 children (10-13 years) and 260 adults. First, we identified a modular structure in the navigation network during resting-state that included a ventral and a dorsal module. Then, we found that the intrinsic modular structure was strengthened from children to adults, that is, adults showed stronger FC within the ventral module and weaker FC between ventral and dorsal modules than children. Further, the intrinsic modular structure was loosened when performing scene-viewing task, that is, both adults and children showed decreased within-ventral FC and increased between-module FC during task- than resting-state. Finally, the task-modulated FC changes were greater in adults than in children. In sum, our study reveals age-related changes in the navigation network organization as increasing modularity under resting-state and increasing flexibility under task-state.

AAAI Conference 2021 System Paper

IFDDS: An Anti-fraud Outbound Robot

  • Zihao Wang
  • Minghui Yang
  • Chunxiang Jin
  • Jia Liu
  • Zujie Wen
  • Saishuai Liu
  • Zhe Zhang

With the rapid growth of internet finance and e-payment, payment fraud has attracted increasing attention. To prevent customers from being cheated, systems often block risky payments depending on a risk factor. However, this may also inadvertently block cases which are not actually risky. To solve this problem, we present IFDDS, a system that proactively chats with customers through intelligent speech interaction to precisely determine the actual payment risk. Our system adopts imitation learning to learn dialogue policies. In addition, it encompasses a dialogue risk detection module which identifies fraud probability every turn based on the dialogue state. We create a web-based user interface which simulates a practical voice-based dialogue system.

YNIMG Journal 2021 Journal Article

Quantifying the variability of neural activation in working memory: A functional probabilistic atlas

  • Chen Chen
  • Ying Zhang
  • Zonglei Zhen
  • Yiying Song
  • Siyuan Hu
  • Jia Liu

Working memory is a fundamental cognitive ability that allows the maintenance and manipulation of information for a brief period of time. Previous studies found a set of brain regions activated during working memory tasks, such as the prefrontal and parietal cortex. However, little is known about the variability of neural activation in working memory. Here, we used functional magnetic resonance imaging to quantify individual, hemispheric, and sex differences of working memory activation in a large cohort of healthy adults (N = 477). We delineated subject-specific activated regions in each individual, including the frontal pole, middle frontal gyrus, frontal eye field, superior parietal lobule, insular, precuneus, and anterior cingulate cortex. A functional probabilistic atlas was created to quantify individual variability in working memory regions. More than 90% of the participants activated all seven regions in both hemispheres, but the intersection of regions across participants was markedly less (50%), indicating significant individual differences in working memory activations. Moreover, we found hemispheric and sex differences in activation location, extent, and magnitude. Most activation regions were larger in the right than in the left hemisphere, but the magnitude of activation did not follow a similar pattern. Men showed more extensive and stronger activations than women. Taken together, our functional probabilistic atlas quantified variabilities of neural activation in working memory, providing a robust spatial reference for standardization of functional localization.

NeurIPS Conference 2021 Conference Paper

Sample Complexity Bounds for Active Ranking from Multi-wise Comparisons

  • Wenbo Ren
  • Jia Liu
  • Ness Shroff

We study the sample complexity (i. e. , the number of comparisons needed) bounds for actively ranking a set of $n$ items from multi-wise comparisons. Here, a multi-wise comparison takes $m$ items as input and returns a (noisy) result about the best item (the winner feedback) or the order of these items (the full-ranking feedback). We consider two basic ranking problems: top-$k$ items selection and full ranking. Unlike previous works that study ranking from multi-wise comparisons, in this paper, we do not require any parametric model or assumption and work on the fundamental setting where each comparison returns the correct result with probability $1$ or a certain probability larger than $\frac{1}{2}$. This paper helps understand whether and to what degree utilizing multi-wise comparisons can reduce the sample complexity for the ranking problems compared to ranking from pairwise comparisons. Specifically, under the winner feedback setting, one can reduce the sample complexity for top-$k$ selection up to an $m$ factor and that for full ranking up to a $\log{m}$ factor. Under the full-ranking feedback setting, one can reduce the sample complexity for top-$k$ selection up to an $m$ factor and that for full ranking up to an $m\log{m}$ factor. We also conduct numerical simulations to confirm our theoretical results.

NeurIPS Conference 2021 Conference Paper

STEM: A Stochastic Two-Sided Momentum Algorithm Achieving Near-Optimal Sample and Communication Complexities for Federated Learning

  • Prashant Khanduri
  • Pranay Sharma
  • Haibo Yang
  • Mingyi Hong
  • Jia Liu
  • Ketan Rajawat
  • Pramod Varshney

Federated Learning (FL) refers to the paradigm where multiple worker nodes (WNs) build a joint model by using local data. Despite extensive research, for a generic non-convex FL problem, it is not clear, how to choose the WNs' and the server's update directions, the minibatch sizes, and the local update frequency, so that the WNs use the minimum number of samples and communication rounds to achieve the desired solution. This work addresses the above question and considers a class of stochastic algorithms where the WNs perform a few local updates before communication. We show that when both the WN's and the server's directions are chosen based on certain stochastic momentum estimator, the algorithm requires $\tilde{\mathcal{O}}(\epsilon^{-3/2})$ samples and $\tilde{\mathcal{O}}(\epsilon^{-1})$ communication rounds to compute an $\epsilon$-stationary solution. To the best of our knowledge, this is the first FL algorithm that achieves such {\it near-optimal} sample and communication complexities simultaneously. Further, we show that there is a trade-off curve between local update frequencies and local minibatch sizes, on which the above sample and communication complexities can be maintained. {Finally, we show that for the classical FedAvg (a. k. a. Local SGD, which is a momentum-less special case of the STEM), a similar trade-off curve exists, albeit with worse sample and communication complexities. Our insights on this trade-off provides guidelines for choosing the four important design elements for FL algorithms, the update frequency, directions, and minibatch sizes to achieve the best performance. }

NeurIPS Conference 2021 Conference Paper

Taming Communication and Sample Complexities in Decentralized Policy Evaluation for Cooperative Multi-Agent Reinforcement Learning

  • Xin Zhang
  • Zhuqing Liu
  • Jia Liu
  • Zhengyuan Zhu
  • Songtao Lu

Cooperative multi-agent reinforcement learning (MARL) has received increasing attention in recent years and has found many scientific and engineering applications. However, a key challenge arising from many cooperative MARL algorithm designs (e. g. , the actor-critic framework) is the policy evaluation problem, which can only be conducted in a {\em decentralized} fashion. In this paper, we focus on decentralized MARL policy evaluation with nonlinear function approximation, which is often seen in deep MARL. We first show that the empirical decentralized MARL policy evaluation problem can be reformulated as a decentralized nonconvex-strongly-concave minimax saddle point problem. We then develop a decentralized gradient-based descent ascent algorithm called GT-GDA that enjoys a convergence rate of $\mathcal{O}(1/T)$. To further reduce the sample complexity, we propose two decentralized stochastic optimization algorithms called GT-SRVR and GT-SRVRI, which enhance GT-GDA by variance reduction techniques. We show that all algorithms all enjoy an $\mathcal{O}(1/T)$ convergence rate to a stationary point of the reformulated minimax problem. Moreover, the fast convergence rates of GT-SRVR and GT-SRVRI imply $\mathcal{O}(\epsilon^{-2})$ communication complexity and $\mathcal{O}(m\sqrt{n}\epsilon^{-2})$ sample complexity, where $m$ is the number of agents and $n$ is the length of trajectories. To our knowledge, this paper is the first work that achieves both $\mathcal{O}(\epsilon^{-2})$ sample complexity and $\mathcal{O}(\epsilon^{-2})$ communication complexity in decentralized policy evaluation for cooperative MARL. Our extensive experiments also corroborate the theoretical performance of our proposed decentralized policy evaluation algorithms.

YNIMG Journal 2020 Journal Article

Discovering dynamic task-modulated functional networks with specific spectral modes using MEG

  • Yongjie Zhu
  • Jia Liu
  • Chaoxiong Ye
  • Klaus Mathiak
  • Piia Astikainen
  • Tapani Ristaniemi
  • Fengyu Cong

Efficient neuronal communication between brain regions through oscillatory synchronization at certain frequencies is necessary for cognition. Such synchronized networks are transient and dynamic, established on the timescale of milliseconds in order to support ongoing cognitive operations. However, few studies characterizing dynamic electrophysiological brain networks have simultaneously accounted for temporal non-stationarity, spectral structure, and spatial properties. Here, we propose an analysis framework for characterizing the large-scale phase-coupling network dynamics during task performance using magnetoencephalography (MEG). We exploit the high spatiotemporal resolution of MEG to measure time-frequency dynamics of connectivity between parcellated brain regions, yielding data in tensor format. We then use a tensor component analysis (TCA)-based procedure to identify the spatio-temporal-spectral modes of covariation among separate regions in the human brain. We validate our pipeline using MEG data recorded during a hand movement task, extracting a transient motor network with beta-dominant spectral mode, which is significantly modulated by the movement task. Next, we apply the proposed pipeline to explore brain networks that support cognitive operations during a working memory task. The derived results demonstrate the temporal formation and dissolution of multiple phase-coupled networks with specific spectral modes, which are associated with face recognition, vision, and movement. The proposed pipeline can characterize the spectro-temporal dynamics of functional connectivity in the brain on the subsecond timescale, commensurate with that of cognitive performance.

JBHI Journal 2020 Journal Article

Improved 3D Catheter Shape Estimation Using Ultrasound Imaging for Endovascular Navigation: A Further Study

  • Fang Chen
  • Jia Liu
  • Xinran Zhang
  • Daoqiang Zhang
  • Hongen Liao

Objective: Two-dimensional fluoroscopy is the standard guidance imaging method for closed endovascular intervention. However, two-dimensional fluoroscopy lacks depth perception for the intervention catheter and causes radiation exposure for both surgeons and patients. In this paper, we extend our previous study and develop the improved three-dimensional (3D) catheter shape estimation using ultrasound imaging. In addition, we perform further quantitative evaluations of endovascular navigation. Method: First, the catheter tracking accuracy in ultrasound images is improved by adjusting the state vector and adding direction information. Then, the 3D catheter points from the catheter tracking are further optimized based on the 3D catheter shape optimization with a high-quality sample set. Finally, the estimated 3D catheter shapes from ultrasound images are overlaid with preoperative 3D tissue structures for the intuitive endovascular navigation. Results: the tracking accuracy of the catheter increased by 24. 39%, and the accuracy of the catheter shape optimization step also increased by approximately 17. 34% compared with our previous study. Furthermore, the overall error of catheter shape estimation was further validated in the catheter intervention experiment of in vitro cardiovascular tissue and in a vivo swine, and the errors were 2. 13 mm and 3. 37 mm, respectively. Conclusion: Experimental results demonstrate that the improved catheter shape estimation using ultrasound imaging is accurate and appropriate for endovascular navigation. Significance: Improved navigation reduces the radiation risk because it decreases use of X-ray imaging. In addition, this navigation method can also provide accurate 3D catheter shape information for endovascular surgery.

NeurIPS Conference 2020 Conference Paper

Overfitting Can Be Harmless for Basis Pursuit, But Only to a Degree

  • Peizhong Ju
  • Xiaojun Lin
  • Jia Liu

Recently, there have been significant interests in studying the so-called "double-descent" of the generalization error of linear regression models under the overparameterized and overfitting regime, with the hope that such analysis may provide the first step towards understanding why overparameterized deep neural networks (DNN) still generalize well. However, to date most of these studies focused on the min L2-norm solution that overfits the data. In contrast, in this paper we study the overfitting solution that minimizes the L1-norm, which is known as Basis Pursuit (BP) in the compressed sensing literature. Under a sparse true linear regression model with p i. i. d. Gaussian features, we show that for a large range of p up to a limit that grows exponentially with the number of samples n, with high probability the model error of BP is upper bounded by a value that decreases with p. To the best of our knowledge, this is the first analytical result in the literature establishing the double-descent of overfitting BP for finite n and p. Further, our results reveal significant differences between the double-descent of BP and min L2-norm solutions. Specifically, the double-descent upper-bound of BP is independent of the signal strength, and for high SNR and sparse models the descent-floor of BP can be much lower and wider than that of min L2-norm solutions.

IJCAI Conference 2020 Conference Paper

Two-stage Behavior Cloning for Spoken Dialogue System in Debt Collection

  • Zihao Wang
  • Jia Liu
  • Hengbin Cui
  • Chunxiang Jin
  • Minghui Yang
  • Yafang Wang
  • Xiaolong Li
  • Renxin Mao

With the rapid growth of internet finance and the booming of financial lending, the intelligent calling for debt collection in FinTech companies has driven increasing attention. Nowadays, the widely used intelligent calling system is based on dialogue flow, namely configuring the interaction flow with the finite-state machine. In our scenario of debt collection, the completed dialogue flow contains more than one thousand interactive paths. All the dialogue procedures are artificially specified, with extremely high maintenance costs and error-prone. To solve this problem, we propose the behavior-cloning-based collection robot framework without any dialogue flow configuration, called two-stage behavior cloning (TSBC). In the first stage, we use multi-label classification model to obtain policies that may be able to cope with the current situation according to the dialogue state; in the second stage, we score several scripts under each obtained policy to select the script with the highest score as the reply for the current state. This framework makes full use of the massive manual collection records without labeling and fully absorbs artificial wisdom and experience. We have conducted extensive experiments in both single-round and multi-round scenarios and showed the effectiveness of the proposed system. The accuracy of a single round of dialogue can be improved by 5%, and the accuracy of multiple rounds of dialogue can be increased by 3. 1%.

YNIMG Journal 2019 Journal Article

A novel training-free externally-regulated neurofeedback (ER-NF) system using phase-guided visual stimulation for alpha modulation

  • Gan Huang
  • Jia Liu
  • Linling Li
  • Li Zhang
  • Yixuan Zeng
  • Lijie Ren
  • Shiqing Ye
  • Zhiguo Zhang

The efficacy of neurofeedback is a point of great controversy, because a certain proportion of users cannot properly regulate their brain activities and thereby fail to benefit from neurofeedback. To address the neurofeedback inefficacy problem, the present study is aimed to design and implement a new neurofeedback system that can more effectively and consistently regulate users’ brain activities than the conventional way of training users to voluntarily regulate brain activities. The new neurofeedback system delivers external visual stimuli continuously at a specific alpha phase, which is real-time decoded from ongoing alpha wave, to regulate the alpha wave. Experimental results show that the proposed training-free externally-regulated neurofeedback (ER-NF) system can achieve consistent (effective in almost all sessions for almost all users), flexible (either increasing or decreasing peak alpha frequency and alpha power), and immediate (taking or losing effect immediately after stimulation is on or off) modulation effects on alpha wave. Therefore, the ER-NF system holds great potential to be able to more reliably and flexibly modulate cognition and behavior.

JBHI Journal 2019 Journal Article

Three-Dimensional Feature-Enhanced Network for Automatic Femur Segmentation

  • Fang Chen
  • Jia Liu
  • Zhe Zhao
  • Mingyu Zhu
  • Hongen Liao

Automatic femur segmentation from computed tomography volume is a crucial but challenging task for computer-aided diagnosis in orthopedic surgeries. The main obstacles are weak bone boundaries, narrowness of joint space, variations in femur density and shape, as well as diverse leg postures. In this paper, we presented a novel 3-D feature-enhanced network to address these challenges. The novelty of our approach lies in two feature enhancement modules, including the edge detection task and the multi-scale features fusion. First, the edge detection task was embedded into femur segmentation from computed tomography volume to solve the problems of narrow joint space and weak femur boundary. Crucially, a task-specific edge detector was used to optimize the performance of femur segmentation in an end-to-end trainable system. Second, the multi-scale features fusion provided both local and global contexts to handle the problems of large variations in leg postures as well as femur shape and density. The results demonstrated that accurate 3-D femur segmentation with a high Dice similarity coefficient of 96. 88% was achieved using the developed method, and the segmentation of computed tomography volume took 0. 93 s on an average.

JBHI Journal 2018 Journal Article

Clustering of Morphological Features for Identifying Femur Cavity Subtypes With Difficulties of Intramedullary Nail Implantation

  • Fang Chen
  • Zhe Zhao
  • Cong Gao
  • Jia Liu
  • Xiuyun Su
  • Jingxin Zhao
  • Peifu Tang
  • Hongen Liao

Intramedullary (IM) nail implantation is currently the standard treatment for femoral intertrochanteric fractures. However, individual differences in femur cavity bring a challenge in designing well-matched IM nails and cause difficulties in IM nail implantation. Therefore, there is an intense need to analyze femur cavities to predict difficulties in IM nail implantation to assist the design of IM nails. This study proposed a method to automatically identify subtypes of femur cavities that exhibit differences in potential difficulties in nail implantation by clustering the morphological features of femur models. The unsupervised subtype extraction method offers a scientific approach to stratify patients for designing and choosing well-matched IM nails. First, the quantitative morphological features of 422 femur cavities were extracted from computed tomography patient models. Second, 422 femur cavities were clustered into three distinct subtypes using a density peak-based k-means clustering method to provide a possible solution for the scientific design of IM nails. The effectiveness of the identified subtypes was validated by comparing subtype differences associated with IM nail implantation and the natural attributes of the patient. Quantitative evaluation of the mismatch degree and real clinical cases confirmed that the clustering results were clinically effective, with clear differences in the subtypes. Therefore, particular IM nails designed from the identified subtypes will potentially facilitate IM nail implantation and reduce complications. Compared with state-of-the-art methods, we used the largest scale dataset and unsupervised clustering to achieve subtype identification of femur cavities with clinical significance.

YNIMG Journal 2018 Journal Article

The neural network for face recognition: Insights from an fMRI study on developmental prosopagnosia

  • Yuanfang Zhao
  • Zonglei Zhen
  • Xiqin Liu
  • Yiying Song
  • Jia Liu

Face recognition is supported by collaborative work of multiple face-responsive regions in the brain. Based on findings from individuals with normal face recognition ability, a neural model has been proposed with the occipital face area (OFA), fusiform face area (FFA), and face-selective posterior superior temporal sulcus (pSTS) as the core face network (CFN) and the rest of the face-responsive regions as the extended face network (EFN). However, little is known about how these regions work collaboratively for face recognition in our daily life. Here we focused on individuals suffering developmental prosopagnosia (DP), a neurodevelopmental disorder specifically impairing face recognition, to shed light on the infrastructure of the neural model of face recognition. Specifically, we used a variant of global brain connectivity method to comprehensively explore resting-state functional connectivity (FC) among face-responsive regions in a large sample of DPs (N = 64). We found that both the FCs within the CFN and those between the CFN and EFN were largely reduced in DP. Importantly, the right OFA and FFA served as the dysconnectivity hubs within the CFN, i. e. , FCs concerning these two regions within the CFN were largely disrupted. In addition, DPs' right FFA also showed reduced FCs with the EFN. Moreover, these disrupted FCs were related to DP's behavioral deficit in face recognition, with the FCs from the FFA to the anterior temporal lobe (ATL) and pSTS the most predictive. Based on these findings, we proposed a revised neural model of face recognition demonstrating the relatedness of interactions among face-responsive regions to face recognition.

YNIMG Journal 2017 Journal Article

Comparison of fMRI analysis methods for heterogeneous BOLD responses in block design studies

  • Jia Liu
  • Ben A. Duffy
  • David Bernal-Casas
  • Zhongnan Fang
  • Jin Hyung Lee

A large number of fMRI studies have shown that the temporal dynamics of evoked BOLD responses can be highly heterogeneous. Failing to model heterogeneous responses in statistical analysis can lead to significant errors in signal detection and characterization and alter the neurobiological interpretation. However, to date it is not clear that, out of a large number of options, which methods are robust against variability in the temporal dynamics of BOLD responses in block-design studies. Here, we used rodent optogenetic fMRI data with heterogeneous BOLD responses and simulations guided by experimental data as a means to investigate different analysis methods’ performance against heterogeneous BOLD responses. Evaluations are carried out within the general linear model (GLM) framework and consist of standard basis sets as well as independent component analysis (ICA). Analyses show that, in the presence of heterogeneous BOLD responses, conventionally used GLM with a canonical basis set leads to considerable errors in the detection and characterization of BOLD responses. Our results suggest that the 3rd and 4th order gamma basis sets, the 7th to 9th order finite impulse response (FIR) basis sets, the 5th to 9th order B-spline basis sets, and the 2nd to 5th order Fourier basis sets are optimal for good balance between detection and characterization, while the 1st order Fourier basis set (coherence analysis) used in our earlier studies show good detection capability. ICA has mostly good detection and characterization capabilities, but detects a large volume of spurious activation with the control fMRI data.

IJCAI Conference 2017 Conference Paper

DRLnet: Deep Difference Representation Learning Network and An Unsupervised Optimization Framework

  • Puzhao Zhang
  • Maoguo Gong
  • Hui Zhang
  • Jia Liu

Change detection and analysis (CDA) is an important research topic in the joint interpretation of spatial-temporal remote sensing images. The core of CDA is to effectively represent the difference and measure the difference degree between bi-temporal images. In this paper, we propose a novel difference representation learning network (DRLnet) and an effective optimization framework without any supervision. Difference measurement, difference representation learning and unsupervised clustering are combined as a single model, i. e. , DRLnet, which is driven to learn clustering-friendly and discriminative difference representations (DRs) for different types of changes. Further, DRLnet is extended into a recurrent learning framework to update and reuse limited training samples and prevent the semantic gaps caused by the saltation in the number of change types from over-clustering stage to the desired one. Experimental results identify the effectiveness of the proposed framework.

IJCAI Conference 2017 Conference Paper

Modeling Hebb Learning Rule for Unsupervised Learning

  • Jia Liu
  • Maoguo Gong
  • Qiguang Miao

This paper presents to model the Hebb learning rule and proposes a neuron learning machine (NLM). Hebb learning rule describes the plasticity of the connection between presynaptic and postsynaptic neurons and it is unsupervised itself. It formulates the updating gradient of the connecting weight in artificial neural networks. In this paper, we construct an objective function via modeling the Hebb rule. We make a hypothesis to simplify the model and introduce a correlation based constraint according to the hypothesis and stability of solutions. By analysis from the perspectives of maintaining abstract information and increasing the energy based probability of observed data, we find that this biologically inspired model has the capability of learning useful features. NLM can also be stacked to learn hierarchical features and reformulated into convolutional version to extract features from 2-dimensional data. Experiments on single-layer and deep networks demonstrate the effectiveness of NLM in unsupervised feature learning.

AAAI Conference 2017 Short Paper

Neuron Learning Machine for Representation Learning

  • Jia Liu
  • Maoguo Gong
  • Qiguang Miao

This paper presents a novel neuron learning machine (NLM) which can extract hierarchical features from data. We focus on the single-layer neural network architecture and propose to model the network based on the Hebbian learning rule. Hebbian learning rule describes how synaptic weight changes with the activations of presynaptic and postsynaptic neurons. We model the learning rule as the objective function by considering the simplicity of the network and stability of solutions. We make a hypothesis and introduce a correlation based constraint according to the hypothesis. We find that this biologically inspired model has the ability of learning useful features from the perspectives of retaining abstract information. NLM can also be stacked to learn hierarchical features and reformulated into convolutional version to extract features from 2-dimensional data.

YNIMG Journal 2017 Journal Article

Sex-linked association between cortical scene selectivity and navigational ability

  • Xiang-Zhen Kong
  • Yi Huang
  • Xin Hao
  • Siyuan Hu
  • Jia Liu

Spatial navigation is a crucial ability for living. Previous studies have shown that males are better at navigation than females, but little is known about the neural basis underlying the sex differences. In this study, we investigated whether cortical scene processing in three well-established scene-selective regions was sexually different, by examining sex differences in scene selectivity and its behavioral relevance to navigation. To do this, we used functional magnetic resonance imaging (fMRI) to scan the parahippocampal place area (PPA), retrosplenial complex (RSC), and occipital place area (OPA) in a large cohort of healthy young adults viewing navigationally relevant scenes (N = 202), and correlated their neural selectivity to scenes with their self-reported navigational ability. Behaviorally, we replicated the previous finding that males were better at navigation than females. Neurally, we found that the scene selectivity in the bilateral PPA, not in the RSC or OPA, was significantly higher in males than females. Such differences could not be explained by confounding factors including brain size and fMRI data quality. Importantly, males, not females, with stronger scene selectivity in the left PPA possessed better navigational ability. This brain-behavior association could not be accounted for by non-navigational abilities (i. e. , intelligence and mental rotation ability). Overall, our study provides novel empirical evidence demonstrating sex differences in the brain activity, inviting further studies on sex differences in the neural network for spatial navigation.

YNIMG Journal 2016 Journal Article

Dissociable roles of internal feelings and face recognition ability in facial expression decoding

  • Lin Zhang
  • Yiying Song
  • Ling Liu
  • Jia Liu

The problem of emotion recognition has been tackled by researchers in both affective computing and cognitive neuroscience. While affective computing relies on analyzing visual features from facial expressions, it has been proposed that humans recognize emotions by internally simulating the emotional states conveyed by others' expressions, in addition to perceptual analysis of facial features. Here we investigated whether and how our internal feelings contributed to the ability to decode facial expressions. In two independent large samples of participants, we observed that individuals who generally experienced richer internal feelings exhibited a higher ability to decode facial expressions, and the contribution of internal feelings was independent of face recognition ability. Further, using voxel-based morphometry, we found that the gray matter volume (GMV) of bilateral superior temporal sulcus (STS) and the right inferior parietal lobule was associated with facial expression decoding through the mediating effect of internal feelings, while the GMV of bilateral STS, precuneus, and the right central opercular cortex contributed to facial expression decoding through the mediating effect of face recognition ability. In addition, the clusters in bilateral STS involved in the two components were neighboring yet separate. Our results may provide clues about the mechanism by which internal feelings, in addition to face recognition ability, serve as an important instrument for humans in facial expression decoding.

YNIMG Journal 2015 Journal Article

Direct in vivo assessment of human stem cell graft–host neural circuits

  • Blake Byers
  • Hyun Joo Lee
  • Jia Liu
  • Andrew J. Weitz
  • Peter Lin
  • Pengbo Zhang
  • Aleksandr Shcheglovitov
  • Ricardo Dolmetsch

Despite the potential of stem cell-derived neural transplants for treating intractable neurological diseases, the global effects of a transplant's electrical activity on host circuitry have never been measured directly, preventing the systematic optimization of such therapies. Here, we overcome this problem by combining optogenetics, stem cell biology, and neuroimaging to directly map stem cell-driven neural circuit formation in vivo. We engineered human induced pluripotent stem cells (iPSCs) to express channelrhodopsin-2 and transplanted resulting neurons to striatum of rats. To non-invasively visualize the function of newly formed circuits, we performed high-field functional magnetic resonance imaging (fMRI) during selective stimulation of transplanted cells. fMRI successfully detected local and remote neural activity, enabling the global graft–host neural circuit function to be assessed. These results demonstrate the potential of a novel neuroimaging-based platform that can be used to identify how a graft's electrical activity influences the brain network in vivo.

YNIMG Journal 2015 Journal Article

Extraversion mediates the relationship between structural variations in the dorsolateral prefrontal cortex and social well-being

  • Feng Kong
  • Siyuan Hu
  • Song Xue
  • Yiying Song
  • Jia Liu

Social well-being reflects the appraisal of one's circumstance and functioning in society, which is crucial for individuals' mental and physical health. However, little is known about the neural processes associated with social well-being. In this study, we used voxel-based morphometry (VBM) to identify the brain regions underlying individual differences in social well-being, as measured by the Social Well-being Scale (SWBS), in a large sample of young healthy adults. We found that social well-being was negatively correlated with gray matter volume in left mid-dorsolateral prefrontal cortex (mid-DLPFC) that is implicated in executive functioning, emotional regulation and social reasoning. The results remained significant even after controlling for the effect of socioeconomic status. Furthermore, although basic personality factors such as neuroticism, extraversion, and conscientiousness (as measured by the NEO Personality Inventory) all contributed to social well-being, only extraversion acted as a mediational mechanism underlying the association between the left mid-DLPFC volume and social well-being. Together, our findings provide the first evidence for the structural basis of individual differences in social well-being, and suggest that the personality trait of extraversion might play an important role in the acquisition and process of social well-being.

YNIMG Journal 2015 Journal Article

Neural correlates of psychological resilience and their relation to life satisfaction in a sample of healthy young adults

  • Feng Kong
  • Xu Wang
  • Siyuan Hu
  • Jia Liu

Psychological resilience refers to the ability to thrive in the face of risk and adversity, which is crucial for individuals' mental and physical health. However, its precise neural correlates are still largely unknown. Here we used resting-state functional magnetic resonance imaging (rs-fMRI) to identify the brain regions underlying this construct by correlating individuals' psychological resilience scores with the regional homogeneity (ReHo) and then examined how these resilience-related regions predicted life satisfaction in a sample of healthy young adults. We found that the ReHo in the bilateral insula, right dorsal anterior cingulate cortex (dACC) and right rostral ACC (rACC) negatively predicted individual differences in psychological resilience, revealing the critical role of the salience network (SN) in psychological resilience. Crucially, the ReHo in the dACC within the SN mediated the effects of psychological resilience on life satisfaction. In summary, these findings suggest that spontaneous activity of the human brain reflect the efficiency of psychological resilience and highlight the dACC within the SN as a neural substrate linking psychological resilience and life satisfaction.

YNIMG Journal 2015 Journal Article

Neural correlates of the happy life: The amplitude of spontaneous low frequency fluctuations predicts subjective well-being

  • Feng Kong
  • Siyuan Hu
  • Xu Wang
  • Yiying Song
  • Jia Liu

Subjective well-being is assumed to be distributed in the hedonic hotspots of subcortical and cortical structures. However, the precise neural correlates underlying this construct, especially how it is maintained during the resting state, are still largely unknown. Here, we explored the neural basis of subjective well-being by correlating the regional fractional amplitude of low frequency fluctuations (fALFF) with the self-reported subjective well-being of healthy individuals. Behaviorally, we demonstrated that subjective well-being contained two related but distinct components: cognitive and affective well-being. Neurally, we showed that the fALFF in the bilateral posterior superior temporal gyrus (pSTG), right posterior mid-cingulate cortex (pMCC), right thalamus, left postcentral gyrus (PCG), right lingual gyrus, and left planum temporale (PT) positively predicted cognitive well-being, whereas the fALFF in the bilateral superior frontal gyrus (SFG), right orbitofrontal cortex (OFC), and left inferior temporal gyrus (ITG) negatively predicted cognitive well-being. In contrast, only the fALFF in the right amygdala reliably predicted affective well-being. Furthermore, emotional intelligence partially mediated the effects of the right pSTG and thalamus on cognitive well-being, as well as the effect of the right amygdala on affective well-being. In summary, we provide the first evidence that spontaneous brain activity in multiple regions associated with sensation, social perception, cognition, and emotion contributes to cognitive well-being, whereas the spontaneous brain activity in only one emotion-related region contributes to affective well-being, suggesting that the spontaneous activity of the human brain reflect the efficiency of subjective well-being.

YNIMG Journal 2015 Journal Article

Optogenetic fMRI reveals distinct, frequency-dependent networks recruited by dorsal and intermediate hippocampus stimulations

  • Andrew J. Weitz
  • Zhongnan Fang
  • Hyun Joo Lee
  • Robert S. Fisher
  • Wesley C. Smith
  • ManKin Choy
  • Jia Liu
  • Peter Lin

Although the connectivity of hippocampal circuits has been extensively studied, the way in which these connections give rise to large-scale dynamic network activity remains unknown. Here, we used optogenetic fMRI to visualize the brain network dynamics evoked by different frequencies of stimulation of two distinct neuronal populations within dorsal and intermediate hippocampus. Stimulation of excitatory cells in intermediate hippocampus caused widespread cortical and subcortical recruitment at high frequencies, whereas stimulation in dorsal hippocampus led to activity primarily restricted to hippocampus across all frequencies tested. Sustained hippocampal responses evoked during high-frequency stimulation of either location predicted seizure-like afterdischarges in video-EEG experiments, while the widespread activation evoked by high-frequency stimulation of intermediate hippocampus predicted behavioral seizures. A negative BOLD signal observed in dentate gyrus during dorsal, but not intermediate, hippocampus stimulation is proposed to underlie the mechanism for these differences. Collectively, our results provide insight into the dynamic function of hippocampal networks and their role in seizures.

YNIMG Journal 2015 Journal Article

Quantifying interindividual variability and asymmetry of face-selective regions: A probabilistic functional atlas

  • Zonglei Zhen
  • Zetian Yang
  • Lijie Huang
  • Xiang-Zhen Kong
  • Xu Wang
  • Xiaobin Dang
  • Yangyue Huang
  • Yiying Song

Face-selective regions (FSRs) are among the most widely studied functional regions in the human brain. However, individual variability of the FSRs has not been well quantified. Here we use functional magnetic resonance imaging (fMRI) to localize the FSRs and quantify their spatial and functional variabilities in 202 healthy adults. The occipital face area (OFA), posterior and anterior fusiform face areas (pFFA and aFFA), posterior continuation of the superior temporal sulcus (pcSTS), and posterior and anterior STS (pSTS and aSTS) were delineated for each individual with a semi-automated procedure. A probabilistic atlas was constructed to characterize their interindividual variability, revealing that the FSRs were highly variable in location and extent across subjects. The variability of FSRs was further quantified on both functional (i. e. , face selectivity) and spatial (i. e. , volume, location of peak activation, and anatomical location) features. Considerable interindividual variability and rightward asymmetry were found in all FSRs on these features. Taken together, our work presents the first effort to characterize comprehensively the variability of FSRs in a large sample of healthy subjects, and invites future work on the origin of the variability and its relation to individual differences in behavioral performance. Moreover, the probabilistic functional atlas will provide an adequate spatial reference for mapping the face network.

ICRA Conference 2014 Conference Paper

High performance control of high-acceleration motions based on time-domain relay feedback technique

  • Chao Liu 0019
  • Jia Liu
  • Jianhua Wu 0005
  • Zhenhua Xiong 0001

This paper focuses on proposing an easily-implemented control method for high-acceleration point-to-point motions. The control algorithm used here to handle the disturbances consists of a model-based feedforward controller, a pole-placement PD controller and a disturbance observer. Then, a fast time-domain identification technique is implemented to give the accurate model parameters, which can be directly utilized in the control algorithm. The method avoids the complicated parameters tuning process, which would be attractive in the industrial application. Experiments are carried out on a permanent magnet linear synchronous motor (PMLSM) and the results demonstrate that the proposed method is capable of achieving high-precision and fast positioning by reducing the tracking error and overshoot.

TCS Journal 2012 Journal Article

A complete symbolic bisimulation for full applied pi calculus

  • Jia Liu
  • Huimin Lin

Symbolic characterisations of bisimilarities for the applied pi calculus proposed so far are sound but incomplete, even restricted to the finite fragment of the calculus. In this paper we present a novel approach to symbolic semantics for the applied pi calculus, leading to a notion of symbolic bisimulation which is both sound and complete with respect to the standard labelled bisimilarity. Moreover, our framework accommodates replications hence works for the full calculus.

TIST Journal 2011 Journal Article

Automatic player labeling, tracking and field registration and trajectory mapping in broadcast soccer video

  • Xiaofeng Tong
  • Jia Liu
  • Tao Wang
  • Yimin Zhang

In this article, we present a method to perform automatic player trajectories mapping based on player detection, unsupervised labeling, efficient multi-object tracking, and playfield registration in broadcast soccer videos. Player detector determines the players' positions and scales by combining the ability of dominant color based background subtraction and a boosting detector with Haar features. We first learn the dominant color with accumulate color histogram at the beginning of processing, then use the player detector to collect hundreds of player samples, and learn player appearance codebook by unsupervised clustering. In a soccer game, a player can be labeled as one of four categories: two teams, referee or outlier. The learning capability enables the method to be generalized well to different videos without any manual initialization. With the dominant color and player appearance model, we can locate and label each player. After that, we perform multi-object tracking by using Markov Chain Monte Carlo (MCMC) data association to generate player trajectories. Some data driven dynamics are proposed to improve the Markov chain's efficiency, such as label consistency, motion consistency, and track length, etc. Finally, we extract key-points and find the mapping from an image plane to the standard field model, and then map players' position and trajectories to the field. A large quantity of experimental results on FIFA World Cup 2006 videos demonstrate that this method can reach high detection and labeling precision, reliably tracking in scenes of player occlusion, moderate camera motion and pose variation, and yield promising field registration results.

v2026.09.13