Arrow Research search

Author name cluster

Jie Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

81 papers
2 author rows

Possible papers

81

TMLR Journal 2026 Journal Article

Achieving Faster than O(1/t) Convergence in General Convex Federated Learning

  • Jie Liu
  • Zuang Wang
  • Yongqiang Wang

This paper aims to achieve faster than O(1/t) convergence in federated learning for general convex loss functions. Under the independent and identical distribution (IID) condition, we show that accurate convergence to an optimal solution can be achieved in convex federated learning even when individual clients select stepsizes locally without any coordination. More importantly, this local stepsize strategy allows exploitation of the local geometry of individual clients’ loss functions, and is shown to lead to faster convergence than the case where a same universal stepsize is used for all clients. Then, when the distribution is non-IID, we employ the sharing of gradients besides the global model parameter to ensure o(1/t) convergence to an optimal solution in convex federated learning. For both algorithms, we theoretically prove that stepsizes that are much larger than existing counterparts are allowed, which leads to much faster convergence in empirical evaluations. It is worth noting that, beyond providing a general framework for federated learning with drift correction, our second algorithm’s achievement of o(1/t) convergence to the exact optimal solution under general convex loss functions has not been previously reported in the federated learning literature—except in certain restricted convex cases with additional constraints. We believe that this is significant because even after incorporating momentum, existing first-order federated learning algorithms can only ensure O(1/t) convergence for general convex loss functions when no additional assumptions on heterogeneity are imposed.

AAAI Conference 2026 Conference Paper

ALERT: Adversarial Learning Enhanced Stability-aware Routing Transformer for Adaptive Depression Detection

  • Liangyi Kang
  • Wei Hua
  • Yan Yang
  • Jie Liu
  • DAN YE

Detecting depression through social media is a complex task, as noisy user-generated content creates significant interference between persistent depressive patterns and transient emotional expressions. Two main challenges arise: First, negative mood indicators are not exclusive to depressed individuals, making it difficult to distinguish between pathological symptoms and situational emotional variations. Second, existing static models fail to adapt to diverse user expression styles and effectively filter out confounding noise from posts by non-depressed individuals. This results in conventional approaches either overfitting to superficial emotional cues or overlooking subtle long-term symptom progression. To address these issues, we propose the Adversarial Learning Enhanced Stability-aware Routing Transformer for Adaptive Depression Detection(ALERT), a novel framework integrating adaptive attention routing and adversarial learning to enhance robustness against confounding mood signals. Specifically, ALERT employs a stability-aware dynamic routing mechanism to annotate user-specific mood valence trends, providing a structured representation of affective progression over time. An adversarial learning module then leverages these mood-based representations to distinguish between expressions indicative of persistent depressive mood and variations in situational mood states, ensuring adaptability to diverse user behaviors. Experimental results on public social media datasets demonstrate that ALERT outperforms state-of-the-art methods in depression detection, effectively reducing false alarm from transient mood states and improving classification accuracy.

AAAI Conference 2026 Conference Paper

CaTFormer: Causal Temporal Transformer with Dynamic Contextual Fusion for Driving Intention Prediction

  • Sirui Wang
  • Zhou Guan
  • Bingxi Zhao
  • Tongjia Gu
  • Jie Liu

Accurate prediction of driving intention is key to enhancing the safety and interactive efficiency of human-machine co-driving systems. It serves as a cornerstone for achieving high-level autonomous driving. However, current approaches remain inadequate for accurately modeling the complex spatiotemporal interdependencies and the unpredictable variability of human driving behavior. To address these challenges, we propose CaTFormer, a causal Temporal Transformer that explicitly models causal interactions between driver behavior and environmental context for robust intention prediction. Specifically, CaTFormer introduces a novel Reciprocal Delayed Fusion (RDF) mechanism for precise temporal alignment of interior and exterior feature streams, a Counterfactual Residual Encoding (CRE) module that systematically eliminates spurious correlations to reveal authentic causal dependencies, and an innovative Feature Synthesis Network (FSN) that adaptively synthesizes these purified representations into coherent temporal representations. Experimental results demonstrate that CaTFormer attains state-of-the-art performance on the Brain4Cars dataset. It effectively captures complex causal temporal dependencies and enhances both the accuracy and transparency of driving intention prediction.

EAAI Journal 2026 Journal Article

Deconfounding enhanced image-based knowledge distillation for fault diagnosis with an application in manufacturing process

  • Jianping Zhang
  • Jie Liu

Deep learning (DL) models are widely adopted in fault diagnosis due to their powerful feature extraction capabilities, yet their high computational burden restricts deployment on edge devices. Knowledge distillation (KD) offers a lightweight solution by transferring knowledge from complex teacher models to lightweight student models. However, existing KD methods often fail to extract fault-specific features in images, as background features—highly correlated with labels—can mislead the model. To address this, we propose a deconfounding-enhanced knowledge distillation (DE-KD) method that integrates causal inference into KD to eliminate spurious correlations caused by background confounders. Specifically, a variational autoencoder (VAE) is incorporated to reconstruct the background of input images, enabling the teacher model to isolate and focus on fault-relevant features. The background reconstruction error is used to extract causal feature maps, which are then aligned with intermediate representations in the student model. The student model is trained using a multi-loss function incorporating hard labels, soft labels from the teacher, and deconfounded intermediate features. Applied to carbon accumulation diagnosis in automotive conductor rails, DE-KD achieveshigher accuracy (95. 63 %)and improved interpretability compared to state-of-the-art lightweight methods, demonstrating its effectiveness in industrial scenarios.

AAAI Conference 2026 Conference Paper

Dual-Branch Asymmetric Discrepancy Learning Based on Fake Image Pattern-Coexistence for AI-Generated Image Detection

  • Chunli Song
  • Jie Liu
  • Peiyang Wang
  • Ying Huang
  • Guixuan Zhang
  • Zhi Zeng
  • Shuwu Zhang

With the rapid advancement of generative models, high-fidelity AI-generated images have become increasingly indistinguishable from real images, posing significant challenges to traditional detection methods that rely on explicit artifacts or uniform feature learning. We hypothesize that detection ambiguity originates from pattern coexistence: synthetic images simultaneously embed (a) authentic patterns inherited from real-image distributions and (b) synthetic patterns induced by generative architectures, whereas real images maintain consistent patterns. We validate this hypothesis through SHAP-based quantitative analysis, demonstrating that synthetic images inherently exhibit a dual distribution—simultaneously containing authentic patterns and synthetic traces—while real images show a unimodal distribution. Building on this insight, this paper proposes a Dual-Branch Asymmetric Discrepancy Learning (DADL) framework. The DADL leverages multi-scale feature extraction and Asymmetric Feature Discrepancy Loss to capture and amplify such pattern differences across multiple scales. Extensive experiments on three benchmarks (AIGCDetectBenchmark, GenImage, and Chameleon) show that DADL achieves state-of-the-art performance, with particular strengths in detecting high-fidelity synthetic images from diffusion models (e.g., Midjourney, SDv1.4, SDv1.5) and enhancing generalization across diverse generative paradigms. This study not only offers an effective approach for AIGI detection but also sheds light on the intrinsic properties of synthetic images, providing a new perspective for advancing AIGI forensics.

AAAI Conference 2026 Conference Paper

Enhanced Recommendation Systems with Retrieval-Augmented Large Language Model (Abstract Reprint)

  • Chuyuan Wei
  • Ke Duan
  • Shengda Zhuo
  • Hongchun Wang
  • Shuqiang Huang
  • Jie Liu

Recommender systems have long struggled with challenges such as cold start and data sparsity, which can lead to poor recommendation performance. While previous approaches have attempted to address these issues by incorporating side information, they often introduce noise, lack flexibility for data expansion, and suffer from inconsistent data quality—factors that hinder accurate user preference inference and reduce recommendation performance. With the vast knowledge bases and advanced reasoning capabilities of large language models (LLMs), these models are particularly well-suited to supplement auxiliary information and capture implicit user intent. To address these challenges, we propose a novel framework, ER2ALM, which leverages the capabilities of LLMs enhanced by Retrieval-Augmented Generation (RAG) to improve recommendation outcomes. Our framework specifically addresses the challenges by flexibly and accurately augmenting auxiliary information and capturing users’ implicit preferences and interests. Additionally, to mitigate the risk of introducing noise, we incorporate a noise reduction strategy to ensure the reliability of the augmented information. Experimental validation on two real-world datasets demonstrates the efficacy of our approach, significantly enhancing both the accuracy and robustness of recommendations compared to state-of-the-art methods. This demonstrates the potential of our framework as a new paradigm for preference mining in recommendation systems.

EAAI Journal 2026 Journal Article

Interpretable hybrid learning for fracturing optimization in deep coalbed methane under data scarcity

  • Jie Liu
  • Cong Xiao
  • Xiaolun Yan
  • Shicheng Zhang
  • Tong Zhou
  • Lei Zou
  • Gui Cao
  • Ying Zhou

The efficient development of deep coalbed methane (CBM) faces challenges including complex geological conditions, sensitivity of fracturing parameters, and data scarcity. This study focuses on a block within the Ordos Basin and proposes a hybrid Artificial Intelligence modeling methodology integrating data augmentation, ensemble learning, and interpretability analysis. The Synthetic Minority Over-sampling Technique (SMOTE) was employed for data enhancement. Based on 17 geological and engineering parameters, a Stacked Generalization ensemble model integrating multiple algorithms including Random Forest, Support Vector Machine, and Gradient Boosting was constructed through randomized search hyperparameter optimization. Furthermore, the Shapley Additive Explanations (SHAP) method was introduced to identify dominant controlling factors, combined with Particle Swarm Optimization (PSO) to achieve collaborative optimization of fracturing parameters. Results demonstrate that geological parameters are the primary controlling factors for post-fracturing productivity. Among geological parameters, gas content, reservoir pressure, and Young's modulus show significant influence, while among engineering parameters, low-viscosity slickwater volume, pad fluid volume, high-viscosity slickwater volume, and pumping rate exhibit considerable impact. After SMOTE and Stacking integration modeling, the production prediction model achieved acceptable prediction accuracy. The optimal fracturing parameter intervals were determined as: low-viscosity slickwater volume 100–150 cubic meters (m3), pad fluid volume 200–400 m3, high-viscosity slickwater volume 100–300 m3, and pumping rate 18–20 cubic meters per minute (m3/min). This study provides an interpretable and scalable methodological framework for fracturing optimization under data-scarce conditions in deep CBM development, offering valuable references for intelligent development of unconventional oil and gas resources.

AAAI Conference 2026 Conference Paper

Learning to Generate Structured Meshes with In-Context: Toward Generalization in Mesh Generation

  • Jing Xiao
  • Xinhai Chen
  • Jiaming Peng
  • Jie Liu

Structured mesh generation serves as a crucial preprocessing step in numerical simulations and can be formulated as a mapping problem from geometry to structured mesh. Existing approaches typically establish an isolated mapping for each geometry. This geometry-specific paradigm fails to capture and leverage commonalities across geometries, inevitably requiring recomputation or costly retraining for new geometries. To overcome this limitation, we propose ICL-Mesh, a meta-learning framework based on in-context learning (ICL) for structured mesh generation. It treats learning one mapping as one task and trains a single neural network to extract commonalities across tasks and learn from in-context examples within each task, enabling rapid generalization to unseen tasks without parameter updates. Experimental results demonstrate that ICL-Mesh effectively generalizes to diverse geometries with only a few context examples, and even without examples. It also exhibits robustness to in-context example order sensitivity and can be extended to various mesh generation scenarios, including mesh refinement and coarsening.

TMLR Journal 2026 Journal Article

Leveraging Recursive Methods for Efficient Federated Learning

  • Jie Liu
  • Zuang Wang
  • Yongqiang Wang

Federated learning algorithms perform multiple local updates on clients before communicating with the parameter server to reduce communication overhead and improve overall training efficiency. However, local updates also lead to the “client-drift” problem under non-IID data, which avoids convergence to the exact optimal solution under heterogeneous data distributions. To ensure accurate convergence, existing federated-learning algorithms employ auxiliary variables to locally estimate the global gradient or the drift from the global gradient, which, however, also incurs extra communication and storage overhead. In this paper, we propose a new recursion-based federated-learning architecture that completely eliminates the need for auxiliary variables while ensuring accurate convergence under het- erogeneous data distributions. This new federated-learning architecture, called FedRecu, can significantly reduce communication and storage overhead compared with existing federated- learning algorithms with accurate convergence guarantees. More importantly, this novel ar- chitecture enables FedRecu to employ much larger stepsizes than existing federated-learning algorithms, thereby leading to much faster convergence. We provide rigorous convergence analysis of FedRecu under both convex and nonconvex loss functions, in both the determin- istic gradient case and the stochastic gradient case. In fact, our theoretical analysis shows that FedRecu ensures o(1/K) convergence to an accurate solution under general convex loss functions, which improves upon the existing achievable O(1/K) convergence rate for general convex loss functions. Numerical experiments on benchmark datasets confirm the effectiveness of the proposed algorithm

EAAI Journal 2026 Journal Article

Physics informed Dual-Layer Bidirectional Gated Recurrent Unit for Nuclear-Grade Electric Gate Valves Fault Prognostics

  • Jie Liu
  • Mian Zhang
  • Chenwei Tang
  • Jiancheng Lv
  • Yanping Huang
  • Yanshan Li
  • Chenhui Li

Nuclear-grade electric gate valves (NEGVs) are mission-critical components in nuclear power plants, characterized by widespread deployment yet prone to high failure rates. Sticking faults pose the most significant risk, often triggering unscheduled plant shutdowns and potentially resulting in severe safety incidents. While accurate fault prediction is crucial for plants safety, current prognostic investigations for NEGVs facing challenges: (1) Inadequate actual operational data, (2) Suboptimal feature selection, (3) Limited prediction accuracy. To overcome these limitations, this study introduces an integrated prognostic framework combining physics informed data augmentation (DA) with optimized feature selection and a Dual-Layer Bidirectional Gated Recurrent Unit (DL-BiGRU) architecture. The proposed DA method capitalizes on ‘segmented wave’ patterns in operating current during sticking faults to effectively describe the degradation trend. Feature selection is enhanced through a random weighting method that simultaneously evaluates feature monotonicity, correlation, and robustness. Case study validation using actual operational data demonstrates the proposed model architecture’s superior predictive capability than other deep learning models, establishing a reasonable strategy in NEGVs degradation trend prediction.

AAAI Conference 2026 Conference Paper

Stop Mixing Things Up! BISCUIT Teaches Vision-Language Models to Learn New Concepts from Images on the Spot

  • Jiahua Bao
  • Siyao Cheng
  • Jiaxing Du
  • Yuhang Jia
  • Boyang Niu
  • Zeming Lang
  • Changjiang He
  • Hao Zhang

Vision-Language Models (VLMs) have achieved impressive performance across various tasks, but often struggle to apply newly introduced visual concepts during inference. A common failure pattern is what we call Mixing Things Up: VLMs frequently confuse concept names, resulting in vague descriptions and failure to ground the concept correctly. Existing approaches mainly address person-related concepts through text prompts or tokenizer modifications. However, VLMs still miss or misinterpret untrained visual concepts, underscoring the need to learn new concepts directly from visual input, without relying on prior textual injection. To overcome these limitations, we propose BISCUIT (Basis-aligned Inference through Structured Concept Unification and Identification-aware Tuning), a two-step training method. Step I proposes a dual-stream structure-aware vision encoder that fuses RGB and edge-based embeddings within a shared basis space to enhance concept recognition. Step II enhances generation quality through identification-aware tuning, which encourages alignment between the generated text and the newly introduced visual concepts. Existing methods mainly focus on person concepts and lack comprehensive evaluation across diverse visual categories. We further propose a benchmark BiscuitVQA to evaluate VLMs performance on recognizing and applying novel image-introduced concepts across diverse concept types and task types, including real people, cartoons, animals, and symbolic content. We apply BISCUIT to LLaVA-1.5 and Qwen2.5-VL, achieving competitive results among open-source models and narrowing the gap to Gemini-2.5 and GPT-4o. Interestingly, our BISCUIT maintains strong generalization, showing minimal degradation on other downstream tasks.

JAIR Journal 2026 Journal Article

TeamTTA: Efficient Multi-Device Collaboration for Open-Set Test-Time Adaptation via Cloud Integration

  • Anqi Lu
  • Youbing Hu
  • Yun Cheng
  • Dawei Wei
  • Zhiqiang Cao
  • Jie Liu
  • Zhijun Li

Deep neural networks (DNNs) deployed on edge devices often suffer from severe performance degradation when exposed to dynamic and continually shifting environments. Test-time adaptation (TTA) has emerged as a promising solution by updating models online with incoming test data. However, edge deployment poses unique challenges: limited computational resources, latency caused by adaptation delays, and knowledge isolation across devices. The situation becomes even more complex in open-world scenarios, where the presence of unknown categories further disrupts adaptation. To overcome these limitations, we propose TeamTTA, a cloud-integrated framework designed for efficient multi-device collaboration open-set test-time adaptation. Specifically, TeamTTA aggregates reliable samples from multiple edge devices through crowdsourcing, uploads them to the cloud, and maintains a memory buffer for continual adaptation. A large vision model (LVM) in the cloud leverages its zero-shot generalization ability to filter out open-set samples and acts as a teacher model, distilling its knowledge into a replicated student edge model stored in the cloud. The adapted model parameters, or alternatively global statistics under poor network conditions, are then transmitted back to the edge devices for efficient inference. Extensive experiments on standard public TTA benchmarks, including corrupted and open-set datasets, show that TeamTTA achieves superior adaptation accuracy, robustness to distribution shifts, and communication efficiency, outperforming state-of-the-art TTA baselines. These results validate the effectiveness of integrating cloud-edge collaboration and LVM-driven knowledge distillation for real-world edge intelligence.

AAAI Conference 2026 Conference Paper

TFRank: Think-Free Reasoning Enables Practical Pointwise LLM Ranking

  • Yongqi Fan
  • Xiaoyang Chen
  • Dezhi Ye
  • Jie Liu
  • Haijin Liang
  • Jin Ma
  • Ben He
  • Yingfei Sun

Reasoning-intensive ranking models built on Large Language Models (LLMs) have made notable progress. However, existing approaches often rely on large-scale LLMs and explicit Chain-of-Thought (CoT) reasoning, resulting in high computational cost and latency that limit real-world use. To address this, we propose TFRank, an efficient pointwise reasoning ranker based on small-scale LLMs. To improve ranking performance, TFRank effectively integrates CoT data, fine-grained score supervision, and multi-task training. Furthermore, it achieves an efficient "Think-Free" reasoning capability by employing a "think-mode switch" and pointwise format constraints. Specifically, this allows the model to leverage explicit reasoning during training while delivering precise relevance scores for complex queries at inference without generating any reasoning chains. Experiments show that TFRank achieves performance comparable to models with four times more parameters on the BRIGHT benchmark, and demonstrates strong competitiveness on the BEIR benchmark. Further analysis shows that TFRank achieves an effective balance between performance and efficiency, providing a practical solution for integrating advanced reasoning into real-world systems.

EAAI Journal 2025 Journal Article

A multi-scale deep feature memory and recovery network for multi-sensor fault diagnosis in the channel missing scenario

  • Tianao Zhang
  • Li Jiang
  • Jie Liu
  • Xin Zhang
  • Qing Zhang

In the field of data-driven intelligent fault diagnosis, the monitoring data provided by single sensor are typically inadequate to reveal the complex states of large-scale equipment comprehensively. Intelligent fault diagnosis techniques based on the fusion of multi-sensor signals have achieved considerable success. Nevertheless, the primary multi-channel fault diagnosis methods based on multi-sensor signals have not yet effectively resolved the problem of sudden sensor failures. Given the challenges mentioned above, we introduce the channel missing scenario to emulate the situation where partial channels are suddenly missing during online inference. To alleviate the impact of missing channels, we propose a multi-scale deep feature memory and recovery network (MDFMR). The feature memory and channel recovery mechanisms of MDFMR can improve the robustness of the model under channel missing scenario. The experiments were conducted on two rotating machinery datasets. Experimental results demonstrated that in the most severe channel missing scenario, conventional multi-channel diagnosis methods become unreliable, while MDFMR maintains a diagnosis accuracy of over 95%.

YNIMG Journal 2025 Journal Article

Aligning with the good in urgency: The enhanced prosocial influence under high time pressure

  • Zhengjie Liu
  • Xiaobo Zhong
  • Jie Liu
  • Fang Cui

Prosocial behavior is essential for enhancing human welfare, particularly in urgent situations. This study employed computational models and fMRI to examine how prosocial influence affects helping behaviors under varying levels of time pressure. Participants were tasked with deciding whether to reduce electric shocks to strangers at their own expense, influenced by varying levels of time pressure (high or low) and social influences (prosocial or selfish choices made by others). Both results of study 1 (n = 31) and study 2 (n = 39) showed that prosocial influence significantly increased helping tendencies, especially under high time pressure. The hierarchical drift diffusion model demonstrates that under high time pressure, prosocial influence accelerates evidence accumulation toward prosocial choices, while a conflict emerges between prosocial priori information and the process of evidence accumulation under low time pressure. Neural correlates also indicated distinct activation patterns associated with prosocial influence under high and low time pressures: heightened affective-related activation in the insula and medial cingulate gyrus under high pressure, and increased activation in the valuation related caudate nucleus, with altered connectivity to the lateral prefrontal cortex under low pressure. In urgent contexts, witnessing altruistic actions of others significantly enhances helping behaviors through increased activation of empathy-related neural regions. Conversely, in non-urgent situations, the impact of prosocial influence diminishes, as evidenced by changes in neural activity. These findings underscore the critical role of social influence in fostering prosocial behavior during emergencies, highlighting the importance of immediate action in urgent contexts.

AAAI Conference 2025 Conference Paper

Both Supply and Precision: Sample Debias and Ranking Consistency Joint Learning for Large Scale Pre-Ranking System

  • Feng Gao
  • Xin Zhou
  • Yinning Shao
  • Yue Wu
  • Jiahua Gao
  • Yujian Ren
  • Fengyang Qi
  • Ruochen Deng

Cascade ranking architecture, composed of matching, pre-ranking, ranking and re-ranking stages, is usually adopted to balance the efficiency and effectiveness in real-world recommendation system (RS). As the middle stage of RS, pre-ranking aims to quickly filter out the low-quality items selected at the matching stage and then forwarding high-quality items to the ranking stage. Existing pre-ranking approaches mainly endure two problems 1) Sample Selection Bias (SSB) problem, which heavily limits the performance improvement of filtering out low-quality items owing to ignoring the data flow between stages; and 2) Ranking Consistency (RC) problem, which may cause the ranked lists of the ranking stage and previous pre-ranking stage to be inconsistent. As a result, the competitive items with high scores at the ranking stage may not be selected because of low scores at the pre-ranking stage. These both two problems may cause sub-optimal performances, but previous works usually only focus on the one of them. In this paper, we propose a novel Sample Debias and Ranking Consistency Joint Learning Framework (SDCL) to jointly alleviate SSB and RC problems. SDCL consists of two main modules including 1) Multi-Task Distillation Module (MTD), which enhances the ability of identifying high-quality items by distilling knowledge across all tasks simultaneously from the more complex ranking model which jointly trained with the pre-ranking model; and 2) Adaptive Negative Sample Learning Module (ANSL), which improves the performance of filtering out low-quality items by adaptively adjusting negative samples learning weights based on the current performance of model. SDCL seamlessly integrates two modules in an end-to-end multi-task learning framework. Evaluations on both real-world large-scale traffic logs and online A/B test demonstrate the efficacy and superiority of SDCL.

EAAI Journal 2025 Journal Article

Causality-guided fault diagnosis under visual interference in fused deposition modeling

  • Qian Li
  • Tingting Huang
  • Jie Liu
  • Shanggang Wang

Fused deposition modeling (FDM) is one of the additive manufacturing (AM) technologies widely used in various industrial fields. Several factors can affect the manufacturing process, leading to quality issues in produced parts. Fault diagnosis techniques aiming at identifying the cause of quality degradation are critical to managing the FDM process. Image-based methods are a trending topic in fault diagnosis with non-contact sensors. Visual interference may cause an unreliable correlation between the background and the root cause labels. Herein, a causality-guided model is proposed for fault diagnosis in the FDM process, which has practical value for improving the accuracy of the fault diagnosis models and helps provide reliable recommendations to the operators. From a causal inference perspective, a structural causal model (SCM) is constructed to describe the "pixel-feature-label" relationship in fault diagnosis tasks. A causality-guided network is proposed to extract features causally related to the root cause labels by eliminating the background information. It comprises three parts: a variational autoencoder (VAE) for learning the background from raw images, a convolution-based fault learning network for extracting defective features, and a feature intervention unit for implementing intervention strategy in SCM. In the experiments, the proposed method achieved a diagnostic F1-score of 92. 80 % in the classification of one normal and six abnormal scenarios. It outperformed state-of-art networks including non-causality and causality-guided ones. Furthermore, Grad-CAM visualizations reveal that the proposed model tends to focus on more compact defect areas, aiding in precise fault localization. This enhances both the interpretability and trustworthiness of the fault diagnosis results.

AAAI Conference 2025 Conference Paper

DAMPER: A Dual-Stage Medical Report Generation Framework with Coarse-Grained MeSH Alignment and Fine-Grained Hypergraph Matching

  • Xiaofei Huang
  • Wenting Chen
  • Jie Liu
  • Qisheng Lu
  • Xiaoling Luo
  • Linlin Shen

Medical report generation is crucial for clinical diagnosis and patient management, summarizing diagnoses and recommendations based on medical imaging. However, existing work often overlook the clinical pipeline involved in report writing, where physicians typically conduct an initial quick review followed by a detailed examination. Moreover, current alignment methods may lead to misaligned relationships. To address these issues, we propose DAMPER, a dual-stage framework for medical report generation that mimics the clinical pipeline of report writing in two stages. In the first stage, a MeSH-Guided Coarse-Grained Alignment (MCG) stage that aligns chest X-ray (CXR) image features with medical subject headings (MeSH) features to generate a rough keyphrase representation of the overall impression. In the second stage, a Hypergraph-Enhanced Fine-Grained Alignment (HFG) stage that constructs hypergraphs for image patches and report annotations, modeling high-order relationships within each modality and performing hypergraph matching to capture semantic correlations between image regions and textual phrases. Finally,the coarse-grained visual features, generated MeSH representations, and visual hypergraph features are fed into a report decoder to produce the final medical report. Extensive experiments on public datasets demonstrate the effectiveness of DAMPER in generating comprehensive and accurate medical reports, outperforming state-of-the-art methods across various evaluation metrics.

AAAI Conference 2025 Conference Paper

E4: Energy-Efficient DNN Inference for Edge Video Analytics via Early Exiting and DVFS

  • Ziyang Zhang
  • Yang Zhao
  • Ming-Ching Chang
  • Changyao Lin
  • Jie Liu

Deep neural network (DNN) models are increasingly popular in edge video analytic applications. However, the computeintensive nature of DNN models pose challenges for energyefficient inference on resource-constrained edge devices. Most existing solutions focus on optimizing DNN inference latency and accuracy, often overlooking energy efficiency. They also fail to account for the varying complexity of video frames, leading to sub-optimal performance in edge video analytics. In this paper, we propose an EnergyEfficient Early-Exit (E4) framework that enhances DNN inference efficiency for edge video analytics by integrating a novel early-exit mechanism with dynamic voltage and frequency scaling (DVFS) governors. It employs an attentionbased cascade module to analyze video frame diversity and automatically determine optimal DNN exit points. Additionally, E4 features a just-in-time (JIT) profiler that uses coordinate descent search to co-optimize CPU and GPU clock frequencies for each layer before the DNN exit points. Extensive evaluations demonstrate that E4 outperforms current state-of-the-art methods, achieving up to 2.8× speedup and 26% average energy saving while maintaining high accuracy.

JAIR Journal 2025 Journal Article

Enhanced Recommendation Systems with Retrieval-Augmented Large Language Model

  • Chuyuan Wei
  • Ke Duan
  • Shengda Zhuo
  • Hongchun Wang
  • Shuqiang Huang
  • Jie Liu

Recommender systems have long struggled with challenges such as cold start and data sparsity, which can lead to poor recommendation performance. While previous approaches have attempted to address these issues by incorporating side information, they often introduce noise, lack flexibility for data expansion, and suffer from inconsistent data quality—factors that hinder accurate user preference inference and reduce recommendation performance. With the vast knowledge bases and advanced reasoning capabilities of large language models (LLMs), these models are particularly well-suited to supplement auxiliary information and capture implicit user intent. To address these challenges, we propose a novel framework, ER2ALM, which leverages the capabilities of LLMs enhanced by Retrieval-Augmented Generation (RAG) to improve recommendation outcomes. Our framework specifically addresses the challenges by flexibly and accurately augmenting auxiliary information and capturing users’ implicit preferences and interests. Additionally, to mitigate the risk of introducing noise, we incorporate a noise reduction strategy to ensure the reliability of the augmented information. Experimental validation on two real-world datasets demonstrates the efficacy of our approach, significantly enhancing both the accuracy and robustness of recommendations compared to state-of-the-art methods. This demonstrates the potential of our framework as a new paradigm for preference mining in recommendation systems.

JBHI Journal 2025 Journal Article

FIND: A Framework for Iterative to Non-Iterative Distillation for Lightweight Deformable Registration

  • Yongtai Zhuo
  • Mingkang Liu
  • Jie Liu
  • Zhikai Yang
  • Rui Liu
  • Peng Xue
  • Lixu Gu

Deformable image registration is crucial for medical image analysis, yet the complexity of deep learning networks often limits their deployment on resource-limited devices. Current distillation methods in registration tasks fail to effectively transfer complex deformation handling capabilities to non-iterative lightweight networks, leading to insignificant performance improvement. To address this, we propose the Framework for Iterative to Non-iterative Distillation (FIND), which efficiently transfers these capabilities to a Non-Iterative Lightweight (NIL) network. FIND employs a dual-step process: first, using recurrent distillation to derive a high-performance non-iterative teacher assistant from an iterative network; second, using advanced feature distillation from the assistant to the lightweight network. This enables NIL to perform rapid, effective registration on resource-limited devices. Experiments across four datasets show that NIL can achieve up to 60 times faster performance on CPU and 89 times on GPU than compared deep learning methods, with superior registration accuracy improvements of up to 3. 5 points in Dice scores.

NeurIPS Conference 2025 Conference Paper

Flow-GRPO: Training Flow Matching Models via Online RL

  • Jie Liu
  • Gongye Liu
  • Jiajun Liang
  • Yangguang Li
  • Jiaheng Liu
  • Xintao Wang
  • Pengfei Wan
  • Di Zhang

We propose Flow-GRPO, the first method to integrate online policy gradient reinforcement learning (RL) into flow matching models. Our approach uses two key strategies: (1) an ODE-to-SDE conversion that transforms a deterministic Ordinary Differential Equation (ODE) into an equivalent Stochastic Differential Equation (SDE) that matches the original model's marginal distribution at all timesteps, enabling statistical sampling for RL exploration; and (2) a Denoising Reduction strategy that reduces training denoising steps while retaining the original number of inference steps, significantly improving sampling efficiency without sacrificing performance. Empirically, Flow-GRPO is effective across multiple text-to-image tasks. For compositional generation, RL-tuned SD3. 5-M generates nearly perfect object counts, spatial relations, and fine-grained attributes, increasing GenEval accuracy from $63\%$ to $95\%$. In visual text rendering, accuracy improves from $59\%$ to $92\%$, greatly enhancing text generation. Flow-GRPO also achieves substantial gains in human preference alignment. Notably, very little reward hacking occurred, meaning rewards did not increase at the cost of appreciable image quality or diversity degradation.

EAAI Journal 2025 Journal Article

Heterogeneous graph contrastive learning-based transductive health condition assessment of Francis turbine unit

  • Fengyuan Zhang
  • Jie Liu
  • Yujie Liu
  • Yuxin Li
  • Ran Duan
  • Zhidi Chen
  • Xingxing Jiang

To assess the Francis turbine unit's (FTU) degradation status from the onsite data without state labeling, a series of health benchmark model (HBM) driven health condition assessment (HCA) methods have been proposed. However, some limitation still exists, including: 1) Multiple similar HBMs are repeatedly constructed and trained to process multiple heterogeneous monitoring signals, however, there is a lack of a unified model that can handle the full task flow directly. 2) Simple linear methods are used to fuse multiple heterogeneous signals to construct performance degradation indexes (PDIs), ignoring temporal changes in the relationships within the signals and resulting in inadequate unit state representation. In this paper, a heterogeneous graph contrastive learning-based transductive health condition assessment of FTU considering multi-source monitoring signals is proposed. First, multi-source heterogeneous signals are converted into a series of heterogeneous signal graphs via the designed similarity-based edge-connection function considering the temporal representation differences. Then, an innovative graph-level heterogenous signal fusion method is proposed to represent unit conditions completely. By concatenating adjacency matrices of the heterogeneous signal graphs, multiple edge connections within the assessment period data are fused, obtaining a unique interactive graph with balanced and extended edge connections. Further, GCL-driven graph representator model is used to maximize the differences between the healthy and degraded interactive graph (IGs), and then the distances between the model readout vectors of the above two graphs is calculated as IPDIs, considering the signal correlations influenced by unit state changes. Verification experiments show that the proposed transductive HCA method effectively evaluates FTU's degradation without state labeling.

NeurIPS Conference 2025 Conference Paper

Improving Video Generation with Human Feedback

  • Jie Liu
  • Gongye Liu
  • Jiajun Liang
  • Ziyang Yuan
  • Xiaokun Liu
  • Mingwu Zheng
  • Xiele Wu
  • Qiulin Wang

Video generation has achieved significant advances through rectified flow techniques, but issues like unsmooth motion and misalignment between videos and prompts persist. In this work, we develop a systematic pipeline that harnesses human feedback to mitigate these problems and refine the video generation model. Specifically, we begin by constructing a large-scale human preference dataset focused on modern video generation models, incorporating pairwise annotations across multi-dimensions. We then introduce VideoReward, a multi-dimensional video reward model, and examine how annotations and various design choices impact its rewarding efficacy. From a unified reinforcement learning perspective aimed at maximizing reward with KL regularization, we introduce three alignment algorithms for flow-based models. These include two training-time strategies: direct preference optimization for flow (Flow-DPO) and reward weighted regression for flow (Flow-RWR), and an inference-time technique, Flow-NRG, which applies reward guidance directly to noisy videos. Experimental results indicate that VideoReward significantly outperforms existing reward models, and Flow-DPO demonstrates superior performance compared to both Flow-RWR and supervised fine-tuning methods. Additionally, Flow-NRG lets users assign custom weights to multiple objectives during inference, meeting personalized video quality needs.

AAAI Conference 2025 Conference Paper

K-hop Hypergraph Neural Network: A Comprehensive Aggregation Approach

  • Linhuang Xie
  • Shihao Gao
  • Jie Liu
  • Ming Yin
  • Taisong Jin

The powerful capability of HyperGraph Neural Networks (HGNNs) in modeling intricate, high-order relationships among multiple data samples stems primarily from their ability to aggregate both the direct neighborhood features of individual nodes and those associated with hyperedges. However, the limited scope of feature propagation in existing HGNNs significantly reduces the utilization of hypergraph information, exacerbating over-squashing and over-smoothing issues. To this end, we propose a novel K-hop HyperGraph Neural Network (KHGNN) to facilitate the interactions of distant nodes and hyperedges. Specifically, the bisection nested convolution based on HyperGINE is employed to extract features from nodes, hyperedges, and structures along all shortest paths between nodes or hyperedges, providing representations of long-distance relationships. With these comprehensive path features, nodes and hyperedges are guided to aggregate distant information while learning their complex relationships. The extensive experiments, particularly on long-range graph datasets, demonstrate that the proposed method achieves SOTA performance compared to existing HGNNs and graph neural networks.

NeurIPS Conference 2025 Conference Paper

Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need

  • Kecheng Chen
  • Pingping Zhang
  • Hui Liu
  • Jie Liu
  • Yibing Liu
  • Jiaxin Huang
  • Shiqi Wang
  • Hong Yan

We have recently witnessed that ''Intelligence" and `''Compression" are the two sides of the same coin, where the language large model (LLM) with unprecedented intelligence is a general-purpose lossless compressor for various data modalities. This attribute is particularly appealing to the lossless image compression community, given the increasing need to compress high-resolution images in the current streaming media era. Consequently, a spontaneous envision emerges: Can the compression performance of the LLM elevate lossless image compression to new heights? However, our findings indicate that the naive application of LLM-based lossless image compressors suffers from a considerable performance gap compared with existing state-of-the-art (SOTA) codecs on common benchmark datasets. In light of this, we are dedicated to fulfilling the unprecedented intelligence (compression) capacity of the LLM for lossless image compression tasks, thereby bridging the gap between theoretical and practical compression performance. Specifically, we propose P -LLM, a next-pixel prediction-based LLM, which integrates various elaborated insights and methodologies, \textit{e. g. ,} pixel-level priors, the in-context ability of LLM, and a pixel-level semantic preservation strategy, to enhance the understanding capacity of pixel sequences for better next-pixel predictions. Extensive experiments on benchmark datasets demonstrate that P-LLM can beat SOTA classical and learned codecs.

NeurIPS Conference 2025 Conference Paper

MASTER: Enhancing Large Language Model via Multi-Agent Simulated Teaching

  • Liang Yue
  • Yihong Tang
  • Kehai Chen
  • Jie Liu
  • Min Zhang

Instruction fine-tuning is crucial in NLP tasks, enhancing pretrained models' instruction-following capabilities and task-specific performance. However, obtaining high-quality fine-tuning data for large models is challenging due to data collection difficulties and high production costs. To address this, we propose MASTER, a novel data augmentation method that enriches original data through interactions among multiple agents with varying cognitive levels. We simulate three pedagogically grounded teaching scenarios, leveraging multi-agent conversations to generate high-quality teacher-student interaction data. Utilizing MASTER, we construct BOOST-QA, a fine-tuning dataset augmented from existing datasets like Orca-Math-200k, ProcQA, and OpenHermes2. 5. Experiments show that models fine-tuned with BOOST-QA perform excellently across multiple benchmarks, demonstrating strong multitask generalization. Notably, MASTER significantly improves models' reasoning abilities in complex tasks, providing valuable insights for future research.

NeurIPS Conference 2025 Conference Paper

MedChain: Bridging the Gap Between LLM Agents and Clinical Practice with Interactive Sequence

  • Jie Liu
  • Wenxuan Wang
  • Zizhan Ma
  • Guolin Huang
  • Yihang SU
  • Kao-Jung Chang
  • Haoliang Li
  • Linlin Shen

Clinical decision making (CDM) is a complex, dynamic process crucial to healthcare delivery, yet it remains a significant challenge for artificial intelligence systems. While Large Language Model (LLM)-based agents have been tested on general medical knowledge using licensing exams and knowledge question-answering tasks, their performance in the CDM in real-world scenarios is limited due to the lack of comprehensive benchmark that mirror actual medical practice. To address this gap, we present MedChain, a dataset of 12, 163 clinical cases that covers five key stages of clinical workflow. MedChain distinguishes itself from existing benchmarks with three key features of real-world clinical practice: personalization, interactivity, and sequentiality. Further, to tackle real-world CDM challenges, we also propose MedChain-Agent, an AI system that integrates a feedback mechanism and a MedCase-RAG module to learn from previous cases and adapt its responses. MedChain-Agent demonstrates remarkable adaptability in gathering information dynamically and handling sequential clinical tasks, significantly outperforming existing approaches. The relevant dataset and code will be released upon acceptance of this paper.

ICML Conference 2025 Conference Paper

PEINR: A Physics-enhanced Implicit Neural Representation for High-Fidelity Flow Field Reconstruction

  • Liming Shen
  • Liang Deng
  • Chongke Bi
  • Yu Wang
  • Xinhai Chen
  • Yueqing Wang
  • Jie Liu

Implicit neural representation (INR) has now been thrust into the limelight with its flexibility in high-fidelity flow field reconstruction tasks. However, the lack of standard benchmarking datasets and the grid independence assumption for INR-based methods hinder progress and adoption in real-world simulation scenarios. Moreover, naive adoptions of existing INR frameworks suffer from limited accuracy in capturing fine-scale structures and spatiotemporal dynamics. Tacking these issues, we first introduce HFR-Beach, a 5. 4 TB public large-scale CFD dataset with 33, 600 unsteady 2D and 3D vector fields for reconstructing high-fidelity flow fields. We further present PEINR, a physics-enhanced INR framework, to enrich the flow fields by concurrently enhancing numerical-precision and grid-resolution. Specifically, PEINR is mainly composed of physical encoding and transformer-based spatiotemporal fuser (TransSTF). Physical encoding decouples temporal and spatial components, employing Gaussian coordinate encoding and localized encoding techniques to capture the nonlinear characteristics of spatiotemporal dynamics and the stencil discretization of spatial dimensions, respectively. TransSTF fuses both spatial and temporal information via transformer for capturing long-range temporal dependencies. Qualitative and quantitative experiments and demonstrate that PEINR outperforms state-of-the-art INR-based methods in reconstruction quality.

EAAI Journal 2025 Journal Article

Research on hull form optimization at multiple speeds based on machine learning and ship model experiments

  • Jie Liu
  • Baoji Zhang
  • Lifen Hu
  • Junying Bi
  • Zheng Tian
  • Yingkai Dong

In order to improve the scientificity, efficiency and systematicness of ship form optimization, the multi-objective optimization research on the David Taylor Model Basin (DTMB) 5512 ship is carried out. First, the ship model experiment quantified the still water resistance of DTMB 5512 at six speeds at Froude number (Fr) as 0. 25–0. 40, demonstrating an almost linear resistance velocity relationship. Meanwhile, the DTMB 5512 ship is subjected to numerical simulations using the Computational Fluid Dynamics (CFD) method and the calculated results are compared with the experimental results. Then, Random Forest (RF)-based approximate models were developed for multi-speed resistance prediction, and verified its feasibility using Maximum Absolute Error (MAE). Finally, the parametric modeling method, the CFD method, and the optimization algorithm are integrated to construct a multi-objective optimization design system for ship forms. The resistance performance of the DTMB 5512 ship is optimized using the Multi-Objective Particle Swarm Optimization (MOPSO) algorithm. The results show that under the constructed hull form optimization framework, the optimized hull forms that meet the constraint conditions can be obtained. The total resistance of the obtained optimized ship at six speeds is reduced by 2. 95 %, 4. 44 %, 3. 71 %, 5. 22 %, 5. 51 % and 4. 83 % respectively. The research results indicate that the optimized hull forms with improved resistance performance can be obtained through the proposed methods, significantly enhancing the optimization efficiency. It also verifies the effectiveness of the random forest method in addressing the challenges of actual engineering optimization.

NeurIPS Conference 2025 Conference Paper

Segment Policy Optimization: Effective Segment-Level Credit Assignment in RL for Large Language Models

  • Yiran Guo
  • Lijie Xu
  • Jie Liu
  • DAN YE
  • Shuang Qiu

Enhancing the reasoning capabilities of large language models effectively using reinforcement learning (RL) remains a crucial challenge. Existing approaches primarily adopt two contrasting advantage estimation granularities: token-level methods (e. g. , PPO) aim to provide fine-grained advantage signals but suffer from inaccurate estimation due to difficulties in training an accurate critic model. On the other extreme, trajectory-level methods (e. g. , GRPO) solely rely on a coarse-grained advantage signal from the final reward, leading to imprecise credit assignment. To address these limitations, we propose Segment Policy Optimization (SPO), a novel RL framework that leverages segment-level advantage estimation at an intermediate granularity, achieving a better balance by offering more precise credit assignment than trajectory-level methods and requiring fewer estimation points than token-level methods, enabling accurate advantage estimation based on Monte Carlo (MC) without a critic model. SPO features three components with novel strategies: (1) flexible segment partition; (2) accurate segment advantage estimation; and (3) policy optimization using segment advantages, including a novel probability-mask strategy. We further instantiate SPO for two specific scenarios: (1) SPO-chain for short chain-of-thought (CoT), featuring novel cutpoint-based partition and chain-based advantage estimation, achieving $6$-$12$ percentage point improvements in accuracy over PPO and GRPO on GSM8K. (2) SPO-tree for long CoT, featuring novel tree-based advantage estimation, which significantly reduces the cost of MC estimation, achieving $7$-$11$ percentage point improvements over GRPO on MATH500 under 2K and 4K context evaluation. We make our code publicly available at https: //github. com/AIFrameResearch/SPO.

EAAI Journal 2025 Journal Article

Self-information and prediction mask enhanced blind inpainting network for dunhuang murals

  • Jiahao Meng
  • Weirong Liu
  • Changhong Shi
  • Zhijun Li
  • Jie Liu

Blind image inpainting methods based on deep learning have shown promising results in digital image inpainting of dunhuang mural images in recent years. However, current blind inpainting methods still suffer from color patches and structural confusion in the repair results caused by contamination of damaged features and sub-network interference. To address the above problems, a self-information and prediction mask enhanced blind inpainting network (SIME-BINet) for dunhuang mural images is proposed. SIME-BINet redesigns blind inpainting method of phased guidance paradigm into information enhance paradigm, which continuously optimizes enhanced information in dynamic form during training process and provides guidance for encoding process. Meanwhile, an information enhanced transformer block is designed to overcome the problem of damaged feature contamination by introducing enhanced information. Experiments show that SIME-BINet outperforms recent state-of-the-art blind inpainting methods on DhMurals1714 dataset and real damage mask. SIME-BINet offers a new paradigm for blind image inpainting based on deep learning and provides an innovative approach for inpainting of dunhuang mural images. The code, data, and pre-trained models will be made available at https: //github. com/IPCSRG/SIME-BINet after the paper is published.

YNICL Journal 2025 Journal Article

Structural and functional changes of Post-Stroke Depression: A multimodal magnetic resonance imaging study

  • Qiuhong Lu
  • Shunzu Lu
  • Xue Wang
  • Yanlan Huang
  • Jie Liu
  • Zhijian Liang

This study investigated changes in gray matter volume (GMV), white matter microstructure, and spontaneous brain activity in post-stroke depression (PSD) using multiple MRI techniques, including neurite orientation dispersion and density imaging (NODDI). Changes in GMV, neurite density index (NDI), orientation dispersion index (ODI), fraction of isotropic water (ISO), diffusion tensor imaging (DTI) parameters, and the amplitude of frequency fluctuations (ALFF) were assessed between PSD (n = 20), post-stroke without depression (n = 20), and normal control (n = 20) groups. Receiver operating characteristic (ROC) curve analysis was performed to test the classification performance of the variant parameters of each MRI modality, each single MRI modality and multiple MRI modality. Compared to patients with post-stroke without depression (non-PSD), those with PSD showed increased ODI and ISO in the widespread white matter, as well as increased ALFF in the left pallidum. No significant differences in the GMV or DTI parameters were observed between the two groups. Furthermore, the ODI of the right superior longitudinal fasciculus and NODDI showed the best classification performance for PSD at their respective comparison level (the areas under the ROC curves (AUC) = 0.917(0.000), 0.933(0.000)). The model of NODDI-derived parameters combined with non-diffusion MRI modality parameters (i.e., GMV and ALFF) showed better diagnostic performance than that of DTI-derived parameters. These findings suggest that PSD is associated with structural and functional abnormalities that may contribute to depressive symptoms. Additionally, NODDI showed its advantages in the description of structural alterations in emotion-related white matter pathways and classification performance in PSD.

NeurIPS Conference 2025 Conference Paper

UGM2N: An Unsupervised and Generalizable Mesh Movement Network via M-Uniform Loss

  • Zhichao Wang
  • Xinhai Chen
  • Qinglin Wang
  • Xiang Gao
  • Qingyang Zhang
  • Menghan Jia
  • Xiang Zhang
  • Jie Liu

Partial differential equations (PDEs) form the mathematical foundation for modeling physical systems in science and engineering, where numerical solutions demand rigorous accuracy-efficiency tradeoffs. Mesh movement techniques address this challenge by dynamically relocating mesh nodes to rapidly-varying regions, enhancing both simulation accuracy and computational efficiency. However, traditional approaches suffer from high computational complexity and geometric inflexibility, limiting their applicability, and existing supervised learning-based approaches face challenges in zero-shot generalization across diverse PDEs and mesh topologies. In this paper, we present an $\textbf{U}$nsupervised and $\textbf{G}$eneralizable $\textbf{M}$esh $\textbf{M}$ovement $\textbf{N}$etwork (UGM2N). We first introduce unsupervised mesh adaptation through localized geometric feature learning, eliminating the dependency on pre-adapted meshes. We then develop a physics-constrained loss function, M-Uniform loss, that enforces mesh equidistribution at the nodal level. Experimental results demonstrate that the proposed network exhibits equation-agnostic generalization and geometric independence in efficient mesh adaptation. It demonstrates consistent superiority over existing methods, including robust performance across diverse PDEs and mesh geometries, scalability to multi-scale resolutions and guaranteed error reduction without mesh tangling.

AAAI Conference 2024 Conference Paper

A Perspective of Q-value Estimation on Offline-to-Online Reinforcement Learning

  • Yinmin Zhang
  • Jie Liu
  • Chuming Li
  • Yazhe Niu
  • Yaodong Yang
  • Yu Liu
  • Wanli Ouyang

Offline-to-online Reinforcement Learning (O2O RL) aims to improve the performance of offline pretrained policy using only a few online samples. Built on offline RL algorithms, most O2O methods focus on the balance between RL objective and pessimism, or the utilization of offline and online samples. In this paper, from a novel perspective, we systematically study the challenges that remain in O2O RL and identify that the reason behind the slow improvement of the performance and the instability of online finetuning lies in the inaccurate Q-value estimation inherited from offline pretraining. Specifically, we demonstrate that the estimation bias and the inaccurate rank of Q-value cause a misleading signal for the policy update, making the standard offline RL algorithms, such as CQL and TD3-BC, ineffective in the online finetuning. Based on this observation, we address the problem of Q-value estimation by two techniques: (1) perturbed value update and (2) increased frequency of Q-value updates. The first technique smooths out biased Q-value estimation with sharp peaks, preventing early-stage policy exploitation of sub-optimal actions. The second one alleviates the estimation bias inherited from offline pretraining by accelerating learning. Extensive experiments on the MuJoco and Adroit environments demonstrate that the proposed method, named SO2, significantly alleviates Q-value estimation issues, and consistently improves the performance against the state-of-the-art methods by up to 83.1%.

IROS Conference 2024 Conference Paper

Adaptive Visual-Aided 4D Radar Odometry Through Transformer-Based Feature Fusion

  • Yuanfan Zhang
  • Renxiang Xiao
  • Ziyang Hong 0001
  • Liang Hu 0002
  • Jie Liu

Multimodal sensor fusion has been successfully utilized in many odometry and localization methods as it increases both estimate accuracy and robustness in application scenarios. To address the challenge of odometry under varying-weather conditions, we propose a novel visual 4D radar fusion based odometry in an unsupervised deep learning approach. In our method, we adopt transformer-based cascaded decoders to facilitate efficient feature extraction of images and radar point clouds. Considering that radars are weather-agnostic and information-rich cameras are susceptible to adverse weathers, we deliberately introduce an adaptive attention-based feature fusion mechanism, in which the attention shifts dynamically to adapt to changing weather conditions based on the amount of information content in image features. Through extensive comparative experiments, our method surpasses different state-of-the-art single-modal odometry estimation methods. Our code and trained model will be released publicly.

NeurIPS Conference 2024 Conference Paper

DDK: Distilling Domain Knowledge for Efficient Large Language Models

  • Jiaheng Liu
  • Chenchen Zhang
  • Jinyang Guo
  • Yuanxing Zhang
  • Haoran Que
  • Ken Deng
  • Zhiqi Bai
  • Jie Liu

Despite the advanced intelligence abilities of large language models (LLMs) in various applications, they still face significant computational and storage demands. Knowledge Distillation (KD) has emerged as an effective strategy to improve the performance of a smaller LLM (i. e. , the student model) by transferring knowledge from a high-performing LLM (i. e. , the teacher model). Prevailing techniques in LLM distillation typically use a black-box model API to generate high-quality pretrained and aligned datasets, or utilize white-box distillation by altering the loss function to better transfer knowledge from the teacher LLM. However, these methods ignore the knowledge differences between the student and teacher LLMs across domains. This results in excessive focus on domains with minimal performance gaps and insufficient attention to domains with large gaps, reducing overall performance. In this paper, we introduce a new LLM distillation framework called DDK, which dynamically adjusts the composition of the distillation dataset in a smooth manner according to the domain performance differences between the teacher and student models, making the distillation process more stable and effective. Extensive evaluations show that DDK significantly improves the performance of student models, outperforming both continuously pretrained baselines and existing knowledge distillation methods by a large margin.

AAAI Conference 2024 Conference Paper

Hierarchical Aligned Multimodal Learning for NER on Tweet Posts

  • Peipei Liu
  • Hong Li
  • Yimo Ren
  • Jie Liu
  • Shuaizong Si
  • Hongsong Zhu
  • Limin Sun

Mining structured knowledge from tweets using named entity recognition (NER) can be beneficial for many downstream applications such as recommendation and intention under standing. With tweet posts tending to be multimodal, multimodal named entity recognition (MNER) has attracted more attention. In this paper, we propose a novel approach, which can dynamically align the image and text sequence and achieve the multi-level cross-modal learning to augment textual word representation for MNER improvement. To be specific, our framework can be split into three main stages: the first stage focuses on intra-modality representation learning to derive the implicit global and local knowledge of each modality, the second evaluates the relevance between the text and its accompanying image and integrates different grained visual information based on the relevance, the third enforces semantic refinement via iterative cross-modal interactions and co-attention. We conduct experiments on two open datasets, and the results and detailed analysis demonstrate the advantage of our model.

AAAI Conference 2024 Conference Paper

LF-ViT: Reducing Spatial Redundancy in Vision Transformer for Efficient Image Recognition

  • Youbing Hu
  • Yun Cheng
  • Anqi Lu
  • Zhiqiang Cao
  • Dawei Wei
  • Jie Liu
  • Zhijun Li

The Vision Transformer (ViT) excels in accuracy when handling high-resolution images, yet it confronts the challenge of significant spatial redundancy, leading to increased computational and memory requirements. To address this, we present the Localization and Focus Vision Transformer (LF-ViT). This model operates by strategically curtailing computational demands without impinging on performance. In the Localization phase, a reduced-resolution image is processed; if a definitive prediction remains elusive, our pioneering Neighborhood Global Class Attention (NGCA) mechanism is triggered, effectively identifying and spotlighting class-discriminative regions based on initial findings. Subsequently, in the Focus phase, this designated region is used from the original image to enhance recognition. Uniquely, LF-ViT employs consistent parameters across both phases, ensuring seamless end-to-end optimization. Our empirical tests affirm LF-ViT's prowess: it remarkably decreases Deit-S's FLOPs by 63% and concurrently amplifies throughput twofold. Code of this project is at https://github.com/edgeai1/LF-ViT.git.

EAAI Journal 2024 Journal Article

Machine learning-driven high-fidelity ensemble surrogate modeling of Francis turbine unit based on data-model interactive simulation

  • Jian Wang
  • Jie Liu
  • Yanglong Lu
  • Haoliang Li
  • Xin Zhang

Abnormal mechanical properties of Francis turbine units (FTUs) lead to unstable output power and operation fault, and may cause catastrophic hazards. At present, computational fluid dynamics (CFD) and machine learning (ML) methods are popular in predicting FTUs' mechanical behaviors, but there are limitations as follows: 1) CFD simulations focus on rated power. The water head and active power variability are neglected, leading to sparse coverage of FTUs operation conditions. 2) Numerous data are required to generate high-fidelity prediction results, occupying vast resources with low efficiency. 3) Computation software and statistical tools may develop inherent errors in generated values, leading to a considerable deviation in final prediction results. In this study, a high-fidelity data-model interactive ensemble surrogate model for FTUs' mechanical behaviors prediction is proposed. First, to form a comprehensive operation conditions sample space, the monitoring data of various operation conditions are collected and clustered by using density-based spatial clustering of application with noise (DBSCAN). Next, to reduce the resource consumption, a small sample space is formed by Latin hypercube sampling (LHS) based on the cluster weights, then sent into the physical model for data-model interactive simulation. Subsequently, on the premise of maintaining the data characteristic, simulated and monitoring data are mixed to weaken the effect of inherent errors. Thus, an ensemble surrogate model, capable of feature extraction and regression analysis, is proposed to predict FTUs’ mechanical behaviors. The experimental results show that the proposed method obtains great prediction accuracy in various conditions, and it outperforms in resource consumption while enables the high-fidelity prediction. Finally, a comparison experiment shows that the ensemble model exhibits significantly lower relative prediction errors and converges more rapidly.

TMLR Journal 2024 Journal Article

MaskMA: Towards Zero-Shot Multi-Agent Decision Making with Mask-Based Collaborative Learning

  • Jie Liu
  • Yinmin Zhang
  • Chuming Li
  • Zhiyuan You
  • Zhanhui Zhou
  • Chao Yang
  • Yaodong Yang
  • Yu Liu

Building a single generalist agent with strong zero-shot capability has recently sparked significant advancements. However, extending this capability to multi-agent decision making scenarios presents challenges. Most current works struggle with zero-shot transfer, due to two challenges particular to the multi-agent settings: (a) a mismatch between centralized training and decentralized execution; and (b) difficulties in creating generalizable representations across diverse tasks due to varying agent numbers and action spaces. To overcome these challenges, we propose a Mask-Based collaborative learning framework for Multi-Agent decision making (MaskMA). Firstly, we randomly mask part of the units and collaboratively learn the policies of unmasked units to handle the mismatch. In addition, MaskMA integrates a generalizable action representation by dividing the action space into intrinsic actions solely related to the unit itself and interactive actions involving interactions with other units. This flexibility allows MaskMA to tackle tasks with varying agent numbers and thus different action spaces. Extensive experiments in SMAC reveal MaskMA, with a single model trained on 11 training maps, can achieve an impressive 77.8% average zero-shot win rate on 60 unseen test maps by decentralized execution, while also performing effectively on other types of downstream tasks (e.g., varied policies collaboration, ally malfunction, and ad hoc team play).

JBHI Journal 2024 Journal Article

MHD-Net: Memory-Aware Hetero-Modal Distillation Network for Thymic Epithelial Tumor Typing With Missing Pathology Modality

  • Huaqi Zhang
  • Jie Liu
  • Weifan Liu
  • Huang Chen
  • Zekuan Yu
  • Yixuan Yuan
  • Pengyu Wang
  • Jing Qin

Fusing multi-modal radiology and pathology data with complementary information can improve the accuracy of tumor typing. However, collecting pathology data is difficult since it is high-cost and sometimes only obtainable after the surgery, which limits the application of multi-modal methods in diagnosis. To address this problem, we propose comprehensively learning multi-modal radiology-pathology data in training, and only using uni-modal radiology data in testing. Concretely, a Memory-aware Hetero-modal Distillation Network (MHD-Net) is proposed, which can distill well-learned multi-modal knowledge with the assistance of memory from the teacher to the student. In the teacher, to tackle the challenge in hetero-modal feature fusion, we propose a novel spatial-differentiated hetero-modal fusion module (SHFM) that models spatial-specific tumor information correlations across modalities. As only radiology data is accessible to the student, we store pathology features in the proposed contrast-boosted typing memory module (CTMM) that achieves type-wise memory updating and stage-wise contrastive memory boosting to ensure the effectiveness and generalization of memory items. In the student, to improve the cross-modal distillation, we propose a multi-stage memory-aware distillation (MMD) scheme that reads memory-aware pathology features from CTMM to remedy missing modal-specific information. Furthermore, we construct a Radiology-Pathology Thymic Epithelial Tumor (RPTET) dataset containing paired CT and WSI images with annotations. Experiments on the RPTET and CPTAC-LUAD datasets demonstrate that MHD-Net significantly improves tumor typing and outperforms existing multi-modal methods on missing modality situations.

AAAI Conference 2024 Conference Paper

Sketch and Refine: Towards Fast and Accurate Lane Detection

  • Chao Chen
  • Jie Liu
  • Chang Zhou
  • Jie Tang
  • Gangshan Wu

Lane detection is to determine the precise location and shape of lanes on the road. Despite efforts made by current methods, it remains a challenging task due to the complexity of real-world scenarios. Existing approaches, whether proposal-based or keypoint-based, suffer from depicting lanes effectively and efficiently. Proposal-based methods detect lanes by distinguishing and regressing a collection of proposals in a streamlined top-down way, yet lack sufficient flexibility in lane representation. Keypoint-based methods, on the other hand, construct lanes flexibly from local descriptors, which typically entail complicated post-processing. In this paper, we present a “Sketch-and-Refine” paradigm that utilizes the merits of both keypoint-based and proposal-based methods. The motivation is that local directions of lanes are semantically simple and clear. At the “Sketch” stage, local directions of keypoints can be easily estimated by fast convolutional layers. Then we can build a set of lane proposals accordingly with moderate accuracy. At the “Refine” stage, we further optimize these proposals via a novel Lane Segment Association Module (LSAM), which allows adaptive lane segment adjustment. Last but not least, we propose multi-level feature integration to enrich lane feature representations more efficiently. Based on the proposed “Sketch-and-Refine” paradigm, we propose a fast yet effective lane detector dubbed “SRLane”. Experiments show that our SRLane can run at a fast speed (i.e., 278 FPS) while yielding an F1 score of 78.9%. The source code is available at: https://github.com/passerer/SRLane.

AAAI Conference 2024 Short Paper

THGFormer: Time-Aware Hypergraph Learning for Multimodal Social Media Popularity Prediction (Student Abstract)

  • Jienan Zhang
  • Jie Liu
  • Zhangtao Cheng
  • Xovee Xu
  • Fang Liu
  • Ting Zhong
  • Kunpeng Zhang

Social media popularity prediction of multimodal user-generated content (UGC) is a crucial task for many real-world applications. However, existing efforts are often limited by missing inter-instance correlations and UGC temporal patterns. To address these issues, we propose a novel time-aware hypergraph Transformer framework, THGFormer. It fully represents inter-instance and intra-instance relations by hypergraphs, captures the temporal dependencies with a time encoder, and enhances UGC's representations via a neighborhood knowledge aggregation. Extensive experiments conducted on two real-world datasets demonstrate that THGFormer outperforms state-of-the-art popularity prediction models across several settings.

NeurIPS Conference 2024 Conference Paper

Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models

  • Zhanhui Zhou
  • Zhixuan Liu
  • Jie Liu
  • Zhichen Dong
  • Chao Yang
  • Yu Qiao

Large language models are usually fine-tuned to align with human preferences. However, fine-tuning a large language model can be challenging. In this work, we introduce $\textit{weak-to-strong search}$, framing the alignment of a large language model as a test-time greedy search to maximize the log-probability difference between small tuned and untuned models while sampling from the frozen large model. This method serves both as (1) a compute-efficient model up-scaling strategy that avoids directly tuning the large model and as (2) an instance of weak-to-strong generalization that enhances a strong model with weak test-time guidance. Empirically, we demonstrate the flexibility of weak-to-strong search across different tasks. In controlled-sentiment generation and summarization, we use tuned and untuned $\texttt{gpt2}$s to improve the alignment of large models without additional training. Crucially, in a more difficult instruction-following benchmark, AlpacaEval 2. 0, we show that reusing off-the-shelf small models (e. g. , $\texttt{zephyr-7b-beta}$ and its untuned version) can improve the length-controlled win rates of both white-box and black-box large models against $\texttt{gpt-4-turbo}$ (e. g. , $34. 4\% \rightarrow 37. 9\%$ for $\texttt{Llama-3-70B-Instruct}$ and $16. 0\% \rightarrow 20. 1\%$ for $\texttt{gpt-3. 5-turbo-instruct}$), despite the small models' low win rates $\approx 10. 0\%$.

JBHI Journal 2024 Journal Article

Wearable Surface Deformation Myography (sDMG) System for Recognition of Locomotion Modes

  • Haoran Sun
  • Xiangyu Peng
  • Junlang Wang
  • Jie Liu
  • Tingting Fu
  • Chaoming He

This study designs a wearable sensing system for locomotion mode recognition using lower-limb skin surface curvature deformation caused by the morphological changes of musculotendinous complexes and soft tissues. Flexible bending sensors are embedded into stretch pants, enabling curvature deformations of specific skin segments above lower-limb muscle groups to be captured in a noncontact manner. To evaluate the performance of this system, we conducted experiments on eight able-bodied subjects completing seven common locomotive activities, including walking, running, ramp ascending/descending, stair ascending/descending, and standing. The system measured seven channels of deformation signals from two cross-sections on the shank and the thigh. The collected signals were distinguishable across different locomotion modes and exhibited consistency when monitoring steps. Using selected time-domain features and a linear discriminant analysis (LDA) classifier enabled the proposed system to continuously recognize locomotion modes with an average accuracy of 96. 5%. Furthermore, the system maintains recognition performance with 95. 7% accuracy even after removing and reapplying the sensors. Finally, we conducted comparison experiments to analyze how window length, feature selection, and the number of channels affect recognition performance, providing insights for optimization. We believe that this novel signal platform holds great potential as a valuable supplementary tool in wearable human motion detection, enriching the information diversity for motion analysis, and enabling new possibilities for further advancements and applications in fields including biomedical engineering, textiles, and computer graphics.

EAAI Journal 2023 Journal Article

A health condition assessment and prediction method of Francis turbine units using heterogeneous signal fusion and graph-driven health benchmark model

  • Fengyuan Zhang
  • Jie Liu
  • Yuxin Li
  • Yujie Liu
  • Ming-Feng Ge
  • Xingxing Jiang

To ensure the safety and efficiency of hydroelectric power generation, the health condition assessment and prediction (HCAP) of Francis turbine units (FTUs) have been widely concerned. To this end, some data-driven methods based on health benchmark model (HBM) and performance deterioration index (PDI) have been proposed, but there are still some shortcomings: 1) Only one type of signal is used for FTU health monitoring and assessment, which cannot fully represent the deterioration of the system. 2) The establishment of HBM only focuses on the time sequence dependences of signals, while ignoring the inter-correlations between signals. 3) PDI based on linear difference measurement cannot integrate heterogeneous signal representations to fully assess FTU status. In this paper, a HCAP method of FTUs using heterogeneous signal fusion and graph-driven HBM is proposed. First, the multivariate data of health status are transformed into spatial-temporal graphs by establishing connections between similar signals. Furthermore, a hybrid neural network-based HBM is designed to excavate the spatial-temporal dependence relationships existing in these graphs, and learn the mapping relationship between working condition parameters and monitoring signals. Finally, Mahalanobis distance between the heterogeneous signals predicted by HBM and the measured signals in the degraded status is calculated, and the comprehensive PDI for HCAP tasks is obtained by the designed heterogeneous signal fusion function. Verification experiments show that the proposed HCAP method effectively assesses FTU deterioration degree earlier with a higher sensitivity.

EAAI Journal 2023 Journal Article

A meta-path graph-based graph homogenization framework for machine fault diagnosis

  • Chaoying Yang
  • Jie Liu
  • Kaibo Zhou
  • Xiaohui Yuan
  • Xingxing Jiang

Graph data-driven methods have swept the field of machine fault diagnosis by merits of modeling relationships between samples. Their performance is highly affected by the constructed graphs quality. Compared to the single-sensor data, multi-sensor data can provide more information, so as to construct higher-quality graphs. However, existing graph data-driven diagnosis methods using multiple sensors still have two limitations. Firstly, heterogeneous multi-sensor data are mainly processed as homogeneous data, ignoring the heterogeneity of heterogeneous multi-sensor data. Secondly, the heterogeneous graph is often with a complex graph structure, and consumes much computational cost to learn. To overcome these limitations, A meta-path graph-based graph homogenization framework for machine fault diagnosis is proposed. Heterogeneous multi-sensor data are converted into the heterogeneous graph, modeling the heterogeneity of heterogeneous multi-sensor data. Further, instead of directly inputting the heterogeneous graph into graph deep learning model, a heterogeneous graph homogenization framework is designed to generate a meta-path graph, reducing the complexity of graph structure and improving the graph quality. Finally, a graph convolutional network is used for graph feature learning, obtaining the diagnosis results. Verification experiments show that the proposed method performs better than machine learning-based and graph deep learning-based methods. In addition, discussive experiments show that the meta-path graph is with lower complexity in graph structure and a higher clustering accuracy than single-sensor data-based K-nearest neighborhood graph.

NeurIPS Conference 2023 Conference Paper

AbdomenAtlas-8K: Annotating 8,000 CT Volumes for Multi-Organ Segmentation in Three Weeks

  • Chongyu Qu
  • Tiezheng Zhang
  • Hualin Qiao
  • Jie Liu
  • Yucheng Tang
  • Alan L. Yuille
  • Zongwei Zhou

Annotating medical images, particularly for organ segmentation, is laborious and time-consuming. For example, annotating an abdominal organ requires an estimated rate of 30-60 minutes per CT volume based on the expertise of an annotator and the size, visibility, and complexity of the organ. Therefore, publicly available datasets for multi-organ segmentation are often limited in data size and organ diversity. This paper proposes an active learning procedure to expedite the annotation process for organ segmentation and creates the largest multi-organ dataset (by far) with the spleen, liver, kidneys, stomach, gallbladder, pancreas, aorta, and IVC annotated in 8, 448 CT volumes, equating to 3. 2 million slices. The conventional annotation methods would take an experienced annotator up to 1, 600 weeks (or roughly 30. 8 years) to complete this task. In contrast, our annotation procedure has accomplished this task in three weeks (based on an 8-hour workday, five days a week) while maintaining a similar or even better annotation quality. This achievement is attributed to three unique properties of our method: (1) label bias reduction using multiple pre-trained segmentation models, (2) effective error detection in the model predictions, and (3) attention guidance for annotators to make corrections on the most salient errors. Furthermore, we summarize the taxonomy of common errors made by AI algorithms and annotators. This allows for continuous improvement of AI and annotations, significantly reducing the annotation costs required to create large-scale datasets for a wider variety of medical imaging tasks. Code and dataset are available at https: //github. com/MrGiovanni/AbdomenAtlas

AAAI Conference 2023 Conference Paper

ACE: Cooperative Multi-Agent Q-learning with Bidirectional Action-Dependency

  • Chuming Li
  • Jie Liu
  • Yinmin Zhang
  • Yuhong Wei
  • Yazhe Niu
  • Yaodong Yang
  • Yu Liu
  • Wanli Ouyang

Multi-agent reinforcement learning (MARL) suffers from the non-stationarity problem, which is the ever-changing targets at every iteration when multiple agents update their policies at the same time. Starting from first principle, in this paper, we manage to solve the non-stationarity problem by proposing bidirectional action-dependent Q-learning (ACE). Central to the development of ACE is the sequential decision making process wherein only one agent is allowed to take action at one time. Within this process, each agent maximizes its value function given the actions taken by the preceding agents at the inference stage. In the learning phase, each agent minimizes the TD error that is dependent on how the subsequent agents have reacted to their chosen action. Given the design of bidirectional dependency, ACE effectively turns a multi-agent MDP into a single-agent MDP. We implement the ACE framework by identifying the proper network representation to formulate the action dependency, so that the sequential decision process is computed implicitly in one forward pass. To validate ACE, we compare it with strong baselines on two MARL benchmarks. Empirical experiments demonstrate that ACE outperforms the state-of-the-art algorithms on Google Research Football and StarCraft Multi-Agent Challenge by a large margin. In particular, on SMAC tasks, ACE achieves 100% success rate on almost all the hard and super hard maps. We further study extensive research problems regarding ACE, including extension, generalization and practicability.

AAAI Conference 2023 Conference Paper

From Coarse to Fine: Hierarchical Pixel Integration for Lightweight Image Super-resolution

  • Jie Liu
  • Chao Chen
  • Jie Tang
  • Gangshan Wu

Image super-resolution (SR) serves as a fundamental tool for the processing and transmission of multimedia data. Recently, Transformer-based models have achieved competitive performances in image SR. They divide images into fixed-size patches and apply self-attention on these patches to model long-range dependencies among pixels. However, this architecture design is originated for high-level vision tasks, which lacks design guideline from SR knowledge. In this paper, we aim to design a new attention block whose insights are from the interpretation of Local Attribution Map (LAM) for SR networks. Specifically, LAM presents a hierarchical importance map where the most important pixels are located in a fine area of a patch and some less important pixels are spread in a coarse area of the whole image. To access pixels in the coarse area, instead of using a very large patch size, we propose a lightweight Global Pixel Access (GPA) module that applies cross-attention with the most similar patch in an image. In the fine area, we use an Intra-Patch Self-Attention (IPSA) module to model long-range pixel dependencies in a local patch, and then a spatial convolution is applied to process the finest details. In addition, a Cascaded Patch Division (CPD) strategy is proposed to enhance perceptual quality of recovered images. Extensive experiments suggest that our method outperforms state-of-the-art lightweight SR methods by a large margin. Code is available at https://github.com/passerer/HPINet.

EAAI Journal 2023 Journal Article

Graph features dynamic fusion learning driven by multi-head attention for large rotating machinery fault diagnosis with multi-sensor data

  • Xin Zhang
  • Xi Zhang
  • Jie Liu
  • Bo Wu
  • Youmin Hu

Recently, rotating machinery fault diagnosis studies based on graph neural networks (GNN) have received some satisfactory achievements. But most of them are based on the analysis of the single sensor signals, which cannot capture the comprehensive fault information, especially aiming at large rotating machineries. A few research using GNN for multi-sensor fault diagnosis only fuse multi-source features in the construction of the input graph, and the fusion effect largely depends on the manual feature selection. Graph attention network (GAT), as an emerging GNN, can give trainable weights to vertices based on the self-attention mechanism to improve the effectiveness of feature learning. And it has not yet been used in the field of multi-sensor fault diagnosis. To fill this gap and utilize GAT’s advantages, this paper presents a multi-sensor multi-head GAT (MMHGAT) model for large rotating machinery fault diagnosis. With the input of several subgraphs, the designed MMHGAT model consisting of two graph attention layers (GAL), a feature fusion process and a Softmax classifier, can dynamically fuse and mine the high-level fault characteristics during the training process. By employing the experiment on the axial flow pump, the effectiveness and superiority of the proposed method are validated.

IJCAI Conference 2023 Conference Paper

Imbalanced Node Classification Beyond Homophilic Assumption

  • Jie Liu
  • Mengting He
  • Guangtao Wang
  • Quoc Viet Hung Nguyen
  • Xuequn Shang
  • Hongzhi Yin

Imbalanced node classification widely exists in real-world networks where graph neural networks (GNNs) are usually highly inclined to majority classes and suffer from severe performance degradation on classifying minority class nodes. Various imbalanced node classification methods have been proposed recently which construct synthetic nodes and edges w. r. t. minority classes to balance the label/topology distribution. However, they are all based on homophilic assumption that nodes of the same label tend to connect despite the widely existence of heterophilic edges in real-world graphs. Thus, they uniformly aggregate features from both homophilic and heterophilic neighbors and rely on feature similarity to generate synthetic edges, which cannot be applied to imbalanced graphs in high heterophily. To address this problem, we propose a novel GraphSANN for imbalanced node classification on both homophilic and heterophilic graphs. Firstly, we propose a unified feature mixer to generate synthetic nodes with both homophilic and heterophilic interpolation in a unified way. Next, by randomly sampling edges between synthetic nodes and existing nodes as candidata edges, we design an adaptive subgraph extractor to dynamically extract the contextual subgraphs of candidate edges with flexible ranges. Finally, we develop a multi-filter subgraph encoder which constructs multiple different filter channels to discriminatively aggregate neighbors’ information along the homophilic and heterophilic edges. Extensive experiments on eight benchmark datasets demonstrate the superiority of our model for imbalanced node classificaiton on both homophilic and heterophilic graphs.

YNICL Journal 2023 Journal Article

Predicting treatment response in adolescents and young adults with major depressive episodes from fMRI using graph isomorphism network

  • Jia Duan
  • Yueying Li
  • Xiaotong Zhang
  • Shuai Dong
  • Pengfei Zhao
  • Jie Liu
  • Junjie Zheng
  • Rongxin Zhu

BACKGROUND: Major depressive episode (MDE) is the main clinical feature of mood disorders (major depressive disorder and bipolar disorder) in adolescents and young adults and accounts for most of the disease course. However, 30%-40% of MDE patients not responding to clinical first-line interventions. It is crucial to predict treatment response in the early stages and identify biomarkers associated with treatment response. Graph Isomorphism Network (GIN), a deep learning method, is promising for predicting treatment response for individual MDE patients with more powerful representation ability to capture the features of brain functional connectivity. METHODS: In this study, GIN was used to predict individual treatment response in 198 adolescents and young adults with MDE. The most discriminating regions were also identified for the treatment response prediction. RESULTS: Using GIN approach, the baseline functional connectivity could predict 79.8% responders and 67.4% non-responders to treatment (accuracy 74.24%). Furthermore, the most discriminating brain regions were mainly involved in paralimbic and subcortical areas. CONCLUSIONS: GIN has shown potential in predicting treatment response for individual patients, which may enable personalized treatment decisions. Furthermore, targeted interventions focused on modulating the activity and connectivity within paralimbic and subcortical regions could potentially improve treatment outcomes and enable personalized interventions for adolescents and young adults with MDE.

EAAI Journal 2023 Journal Article

Research on decision-level fusion method based on structural causal model in system-level fault detection and diagnosis

  • Haoyuan Pu
  • Zhi Chen
  • Jie Liu
  • Xiaohua Yang
  • Changan Ren
  • Hua Liu
  • Yifan Jian

At present, system-level fault detection and diagnosis (FDD) research often uses correlation-based machine learning methods combined with multiple heterogeneous diagnosis methods to improve the fault detection rate (FDR), that is, decision-level fusion. Since it does not take into account the causal direction of the decision relationship, it will affect the realization of the fusion objectives, and lead to the reduction of the fusion range and the decrease of the global decision on FDR. In this regard, the structural causal model (SCM), a commonly used causal model in causal science, can use the causal graph to ensure causal direction of fusion, and the structural equation can be used to achieve fusion objectives to increase FDR, which can improve this problem. In this paper, we propose seven fusion objectives according to the diagnostic advantage interval of each preliminary method, and use SCM to construct causal graph and structural equation to achieve decision-level fusion according to the proposed seven fusion objectives, thereby improving FDR. The proposed method is validated through the simulation platform Tennessee Eastman process. We choose to combine the prediction results of Linear Discriminant Analysis method and Gaussian Naive Bayes method to achieve decision-level fusion. The results show that compared with the single method and the Bayesian network decision-level fusion method, the proposed method can achieve the best results in the FDR of each single system state and average FDR, and the above indicators are significantly improved.

AILAW Journal 2023 Journal Article

Semantic matching based legal information retrieval system for COVID-19 pandemic

  • Junlin Zhu
  • Jiaye Wu
  • Xudong Luo
  • Jie Liu

Abstract Recently, the pandemic caused by COVID-19 is severe in the entire world. The prevention and control of crimes associated with COVID-19 are critical for controlling the pandemic. Therefore, to provide efficient and convenient intelligent legal knowledge services during the pandemic, we develop an intelligent system for legal information retrieval on the WeChat platform in this paper. The data source we used for training our system is “The typical cases of national procuratorial authorities handling crimes against the prevention and control of the new coronary pneumonia pandemic following the law”, which is published online by the Supreme People’s Procuratorate of the People’s Republic of China. We base our system on convolutional neural network and use the semantic matching mechanism to capture inter-sentence relationship information and make a prediction. Moreover, we introduce an auxiliary learning process to help the network better distinguish the relation between two sentences. Finally, the system uses the trained model to identify the information entered by a user and responds to the user with a reference case similar to the query case and gives the reference legal gist applicable to the query case.

PRL Workshop 2023 Workshop Paper

Theoretically Guaranteed Policy Improvement Distilled from Model-Based Planning

  • Chuming Li
  • Ruonan Jia
  • Jiawei Yao
  • Jie Liu
  • Yinmin Zhang
  • Yazhe Niu
  • Yaodong Yang
  • Yu Liu

Model-based reinforcement learning (RL) has demonstrated remarkable successes on a range of continuous control tasks due to its high sample efficiency. To save the computation cost of conducting planning online, recent practices tend to distill optimized action sequences into an RL policy during the training phase. Although the distillation can incorporate both the foresight of planning and the exploration ability of RL policies, the theoretical understanding of these methods is yet unclear. In this paper, we extend the policy improvement step of Soft Actor-Critic (SAC) by developing an approach to distill from model-based planning to the policy. We then demonstrate that such an approach of policy improvement has a theoretical guarantee of monotonic improvement and convergence to the maximum value defined in SAC. We discuss effective design choices and implement our theory as a practical algorithm---$\textit{\textbf{M}odel-based \textbf{P}lanning \textbf{D}istilled to \textbf{P}olicy (MPDP)}$---that updates the policy jointly over multiple future time steps. Extensive experiments show that MPDP achieves better sample efficiency and asymptotic performance than both model-free and model-based planning algorithms on six continuous control benchmark tasks in MuJoCo.

IJCAI Conference 2023 Conference Paper

Video Frame Interpolation with Densely Queried Bilateral Correlation

  • Chang Zhou
  • Jie Liu
  • Jie Tang
  • Gangshan Wu

Video Frame Interpolation (VFI) aims to synthesize non-existent intermediate frames between existent frames. Flow-based VFI algorithms estimate intermediate motion fields to warp the existent frames. Real-world motions' complexity and the reference frame's absence make motion estimation challenging. Many state-of-the-art approaches explicitly model the correlations between two neighboring frames for more accurate motion estimation. In common approaches, the receptive field of correlation modeling at higher resolution depends on the motion fields estimated beforehand. Such receptive field dependency makes common motion estimation approaches poor at coping with small and fast-moving objects. To better model correlations and to produce more accurate motion fields, we propose the Densely Queried Bilateral Correlation (DQBC) that gets rid of the receptive field dependency problem and thus is more friendly to small and fast-moving objects. The motion fields generated with the help of DQBC are further refined and up-sampled with context features. After the motion fields are fixed, a CNN-based SynthNet synthesizes the final interpolated frame. Experiments show that our approach enjoys higher accuracy and less inference time than the state-of-the-art. Source code is available at https: //github. com/kinoud/DQBC.

JBHI Journal 2022 Journal Article

Cross-Boosted Multi-Target Domain Adaptation for Multi-Modality Histopathology Image Translation and Segmentation

  • Huaqi Zhang
  • Jie Liu
  • Pengyu Wang
  • Zekuan Yu
  • Weifan Liu
  • Huang Chen

Recent digital pathology workflows mainly focus on mono-modality histopathology image analysis. However, they ignore the complementarity between Haematoxylin & Eosin (H&E) and Immunohistochemically (IHC) stained images, which can provide comprehensive gold standard for cancer diagnosis. To resolve this issue, we propose a cross-boosted multi-target domain adaptation pipeline for multi-modality histopathology images, which contains Cross-frequency Style-auxiliary Translation Network (CSTN) and Dual Cross-boosted Segmentation Network (DCSN). Firstly, CSTN achieves the one-to-many translation from fluorescence microscopy images to H&E and IHC images for providing source domain training data. To generate images with realistic color and texture, Cross-frequency Feature Transfer Module (CFTM) is developed to pertinently restructure and normalize high-frequency content and low-frequency style features from different domains. Then, DCSN fulfills multi-target domain adaptive segmentation, where a dual-branch encoder is introduced, and Bidirectional Cross-domain Boosting Module (BCBM) is designed to implement cross-modality information complementation through bidirectional inter-domain collaboration. Finally, we establish Multi-modality Thymus Histopathology (MThH) dataset, which is the largest publicly available H&E and IHC image benchmark. Experiments on MThH dataset and several public datasets show that the proposed pipeline outperforms state-of-the-art methods on both histopathology image translation and segmentation.

AAAI Conference 2022 Short Paper

MMAN: Metapath Based Multi-Level Graph Attention Networks for Heterogeneous Network Embedding (Student Abstract)

  • Jie Liu
  • Lingyun Song
  • Li Gao
  • Xuequn Shang

Current Heterogeneous Network Embedding (HNE) models can be roughly divided into two types, i. e. , relation-aware and metapath-aware models. However, they either fail to represent the non-pairwise relations in heterogeneous graph, or only capable of capturing local information around target node. In this paper, we propose a metapath based multilevel graph attention networks (MMAN) to jointly learn node embeddings on two substructures, i. e. , metapath based graphs and hypergraphs extracted from original heterogeneous graph. Extensive experiments on three benchmark datasets for node classification and node clustering demonstrate the superiority of MMAN over the state-of-the-art works.

AAAI Conference 2022 Conference Paper

Multi-Scale Distillation from Multiple Graph Neural Networks

  • Chunhai Zhang
  • Jie Liu
  • Kai Dang
  • WenZheng Zhang

Knowledge Distillation (KD), which is an effective model compression and acceleration technique, has been successfully applied to graph neural networks (GNNs) recently. Existing approaches utilize a single GNN model as the teacher to distill knowledge. However, we notice that GNN models with different number of layers demonstrate different classification abilities on nodes with different degrees. On the one hand, for nodes with high degrees, their local structures are dense and complex, hence more message passing is needed. Therefore, GNN models with more layers perform better. On the other hand, for nodes with low degrees, whose local structures are relatively sparse and simple, the repeated message passing can easily lead to over-smoothing. Thus, GNN models with less layers are more suitable. However, existing single-teacher GNN knowledge distillation approaches which are based on a single GNN model, are sub-optimal. To this end, we propose a novel approach to distill multi-scale knowledge, which learns from multiple GNN teacher models with different number of layers to capture the topological semantic at different scales. Instead of learning from the teacher models equally, the proposed method automatically assigns proper weights for each teacher model via an attention mechanism which enables the student to select teachers for different local structures. Extensive experiments are conducted to evaluate the proposed method on four public datasets. The experimental results demonstrate the superiority of our proposed method over state-of-the-art methods. Our code is publicly available at https: //github. com/NKU-IIPLab/MSKD.

TIST Journal 2021 Journal Article

A Comprehensive Survey of Grammatical Error Correction

  • Yu Wang
  • Yuelin Wang
  • Kai Dang
  • Jie Liu
  • Zhuo Liu

Grammatical error correction (GEC) is an important application aspect of natural language processing techniques, and GEC system is a kind of very important intelligent system that has long been explored both in academic and industrial communities. The past decade has witnessed significant progress achieved in GEC for the sake of increasing popularity of machine learning and deep learning. However, there is not a survey that untangles the large amount of research works and progress in this field. We present the first survey in GEC for a comprehensive retrospective of the literature in this area. We first give the definition of GEC task and introduce the public datasets and data annotation schema. After that, we discuss six kinds of basic approaches, six commonly applied performance boosting techniques for GEC systems, and three data augmentation methods. Since GEC is typically viewed as a sister task of Machine Translation (MT), we put more emphasis on the statistical machine translation (SMT)-based approaches and neural machine translation (NMT)-based approaches for the sake of their importance. Similarly, some performance-boosting techniques are adapted from MT and are successfully combined with GEC systems for enhancement on the final performance. More importantly, after the introduction of the evaluation in GEC, we make an in-depth analysis based on empirical results in aspects of GEC approaches and GEC systems for a clearer pattern of progress in GEC, where error type analysis and system recapitulation are clearly presented. Finally, we discuss five prospective directions for future GEC researches.

AAAI Conference 2021 Conference Paper

Bidirectional Machine Reading Comprehension for Aspect Sentiment Triplet Extraction

  • Shaowei Chen
  • Yu Wang
  • Jie Liu
  • Yuelin Wang

Aspect sentiment triplet extraction (ASTE), which aims to identify aspects from review sentences along with their corresponding opinion expressions and sentiments, is an emerging task in fine-grained opinion mining. Since ASTE consists of multiple subtasks, including opinion entity extraction, relation detection, and sentiment classification, it is critical and challenging to appropriately capture and utilize the associations among them. In this paper, we transform ASTE task into a multi-turn machine reading comprehension (MTMRC) task and propose a bidirectional MRC (BMRC) framework to address this challenge. Specifically, we devise three types of queries, including non-restrictive extraction queries, restrictive extraction queries and sentiment classification queries, to build the associations among different subtasks. Furthermore, considering that an aspect sentiment triplet can derive from either an aspect or an opinion expression, we design a bidirectional MRC structure. One direction sequentially recognizes aspects, opinion expressions, and sentiments to obtain triplets, while the other direction identifies opinion expressions first, then aspects, and at last sentiments. By making the two directions complement each other, our framework can identify triplets more comprehensively. To verify the effectiveness of our approach, we conduct extensive experiments on four benchmark datasets. The experimental results demonstrate that BMRC achieves state-of-the-art performances.

YNIMG Journal 2021 Journal Article

Development of the default-mode network during childhood and adolescence: A longitudinal resting-state fMRI study

  • Fengmei Fan
  • Xuhong Liao
  • Tianyuan Lei
  • Tengda Zhao
  • Mingrui Xia
  • Weiwei Men
  • Yanpei Wang
  • Mingming Hu

The default-mode network (DMN) is a set of functionally connected regions that play crucial roles in internal cognitive processing. Previous resting-state fMRI studies have demonstrated that the intrinsic functional organization of the DMN undergoes remarkable reconfigurations during childhood and adolescence. However, these studies have mainly focused on cross-sectional designs with small sample sizes, limiting the consistency and interpretations of the findings. Here, we used a large sample of longitudinal resting-state fMRI data comprising 305 typically developing children (6-12 years of age at baseline, 491 scans in total) and graph theoretical approaches to delineate the developmental trajectories of the functional architecture of the DMN. For each child, the DMN was constructed according to a prior parcellation with 32 brain nodes. We showed that the overall connectivity increased in strength from childhood to adolescence and became spatially similar to that in the young adult group (N = 61, 18-28 years of age). These increases were primarily located in the midline structures. Global and local network efficiency in the DMN also increased with age, indicating an enhanced capability in parallel information communication within the brain system. Based on the divergent developmental rates of nodal centrality, we identified three subclusters within the DMN, with the fastest rates in the cluster mainly comprising the anterior medial prefrontal cortex and posterior cingulate cortex. Together, our findings highlight the developmental patterns of the functional architecture in the DMN from childhood to adolescence, which has implications for the understanding of network mechanisms underlying the cognitive development of individuals.

IJCAI Conference 2020 Conference Paper

Attention as Relation: Learning Supervised Multi-head Self-Attention for Relation Extraction

  • Jie Liu
  • Shaowei Chen
  • Bingquan Wang
  • Jiaxin Zhang
  • Na Li
  • Tong Xu

Joint entity and relation extraction is critical for many natural language processing (NLP) tasks, which has attracted increasing research interest. However, it is still faced with the challenges of identifying the overlapping relation triplets along with the entire entity boundary and detecting the multi-type relations. In this paper, we propose an attention-based joint model, which mainly contains an entity extraction module and a relation detection module, to address the challenges. The key of our model is devising a supervised multi-head self-attention mechanism as the relation detection module to learn the token-level correlation for each relation type separately. With the attention mechanism, our model can effectively identify overlapping relations and flexibly predict the relation type with its corresponding intensity. To verify the effectiveness of our model, we conduct comprehensive experiments on two benchmark datasets. The experimental results demonstrate that our model achieves state-of-the-art performances.

IJCAI Conference 2020 Conference Paper

Dress like an Internet Celebrity: Fashion Retrieval in Videos

  • Hongrui Zhao
  • Jin Yu
  • Yanan Li
  • Donghui Wang
  • Jie Liu
  • Hongxia Yang
  • Fei Wu

Nowadays, both online shopping and video sharing have grown exponentially. Although internet celebrities in videos are ideal exhibition for fashion corporations to sell their products, audiences do not always know where to buy fashion products in videos, which is a cross-domain problem called video-to-shop. In this paper, we propose a novel deep neural network, called Detect, Pick, and Retrieval Network (DPRNet), to break the gap between fashion products from videos and audiences. For the video side, we have modified the traditional object detector, which automatically picks out the best object proposals for every commodity in videos without duplication, to promote the performance of the video-to-shop task. For the fashion retrieval side, a simple but effective multi-task loss network obtains new state-of-the-art results on DeepFashion. Extensive experiments conducted on a new large-scale cross-domain video-to-shop dataset shows that DPRNet is efficient and outperforms the state-of-the-art methods on video-to-shop task.

AAAI Conference 2020 Conference Paper

Understanding and Improving Proximity Graph Based Maximum Inner Product Search

  • Jie Liu
  • Xiao Yan
  • Xinyan Dai
  • Zhirong Li
  • James Cheng
  • Ming-Chang Yang

The inner-product navigable small world graph (ip-NSW) represents the state-of-the-art method for approximate maximum inner product search (MIPS) and it can achieve an order of magnitude speedup over the fastest baseline. However, to date it is still unclear where its exceptional performance comes from. In this paper, we show that there is a strong norm bias in the MIPS problem, which means that the large norm items are very likely to become the result of MIPS. Then we explain the good performance of ip-NSW as matching the norm bias of the MIPS problem — large norm items have big in-degrees in the ip-NSW proximity graph and a walk on the graph spends the majority of computation on these items, thus effectively avoids unnecessary computation on small norm items. Furthermore, we propose the ip-NSW+ algorithm, which improves ip-NSW by introducing an additional angular proximity graph. Search is first conducted on the angular graph to find the angular neighbors of a query and then the MIPS neighbors of these angular neighbors are used to initialize the candidate pool for search on the inner-product proximity graph. Experiment results show that ip-NSW+ consistently and significantly outperforms ip-NSW and provides more robust performance under different data distributions.

YNIMG Journal 2019 Journal Article

Diffusion tensor imaging shows mechanism-specific differences in injury pattern and progression in rat models of acute spinal cord injury

  • Andrew Yung
  • Stephen Mattucci
  • Barry Bohnet
  • Jie Liu
  • Caron Fournier
  • Wolfram Tetzlaff
  • Piotr Kozlowski
  • Thomas Oxland

We investigate the ability of diffusion tensor imaging (DTI) to distinguish between three experimental rat models of spinal cord injury mechanism – contusion, dislocation, and distraction. Ex vivo DTI scans were performed on cord specimens that were preserved at different time points of the acute injury (3 hr, 24 hr, and 7 days post-injury) across all three injury mechanisms. White matter was classified as abnormal if their DTI metric was substantially different from regional values measured from a set of uninjured controls, thus allowing generation of binary “white matter damage maps” which categorizes each pixel in the DTI image as “normal” or “damaged”. Damage classification was most robust using thresholds in the longitudinal diffusivity, which supports previous studies that show that longitudinal diffusivity is the most robust DTI metric in depicting damage in SCI. Furthermore, the spatial damage patterns from all subjects in the same group were consolidated into a "damage occurrence ratio map", which illustrates an average damage shape that characterizes the injury mechanism. Our analysis has yielded a dataset which highlights the differences in injury pattern due to the initial mode of mechanical injury. For example, contusion produced an initial injury that emanated radially outward from the central canal, with subsequent damage along the caudal corticospinal tract and rostral gracile fasciculus; dislocation injuries showed a high level of involvement in the lateral and ventral white matter which became less apparent by 7 days post-injury, and distraction injuries were found to be less focal and more distributed rostrocaudally. This work represents a first step in adopting the use of the primary injury mechanism as a clinical prognostic factor in SCI, which may help to inform the trialing of existing neuroprotective treatment candidates, the development of new therapies as well as personalize the management of SCI for the individual patient.

IJCAI Conference 2019 Conference Paper

Network Embedding with Dual Generation Tasks

  • Jie Liu
  • Na Li
  • Zhicheng He

We study the problem of Network Embedding (NE) for content-rich networks. NE models aim to learn efficient low-dimensional dense vectors for network vertices which are crucial to many network analysis tasks. The core problem of content-rich network embedding is to learn and integrate the semantic information conveyed by network structure and node content. In this paper, we propose a general end-to-end model, Dual GEnerative Network Embedding (DGENE), to leverage the complementary information of network structure and content. In this model, each vertex is regarded as an object with two modalities: node identity and textual content. Then we formulate two dual generation tasks. One is Node Identification (NI) which recognizes nodes’ identities given their contents. Inversely, the other one is Content Generation (CG) which generates textual contents given the nodes’ identities. We develop specific Content2Node and Node2Content models for the two tasks. Under the DGENE framework, the two dual models are learned by sharing and integrating intermediate layers, with which they mutually enhance each other. Extensive experimental results show that our model yields a significant performance gain compared to the state-of-the-art NE methods. Moreover, our model has an interesting and useful byproduct, that is, a component of our model can generate texts, which is potentially useful for many tasks.

AAAI Conference 2018 Conference Paper

A Framework for Multistream Regression With Direct Density Ratio Estimation

  • Ahsanul Haque
  • Hemeng Tao
  • Swarup Chandra
  • Jie Liu
  • Latifur Khan

Regression over a stream of data is challenging due to unbounded data size and non-stationary distribution over time. Typically, a traditional supervised regression model over a data stream is trained on data instances occurring within a short time period by assuming a stationary distribution. This model is later used to predict value of response-variable in future instances. Over time, the model may degrade in performance due to changes in data distribution among incoming data instances. Updating the model for change adaptation requires true value for every recent data instances, which is scarce in practice. To overcome this issue, recent studies have employed techniques that sample fewer instances to be used for model re-training. Yet, this may introduce sampling bias that adversely affects the model performance. In this paper, we study the regression problem over data streams in a novel setting. We consider two independent, yet related, nonstationary data streams, which are referred to as the source and the target stream. The target stream continuously generates data instances whose value of response variable is unknown. The source stream, however, continuously generates data instances along with corresponding value for the response-variable, and has a biased data distribution with respect to the target stream. We refer to the problem of using a model trained on the biased source stream to predict the response-variable’s value in data instances occurring on the target stream as Multistream Regression. In this paper, we describe a framework for multistream regression that simultaneously overcomes distribution bias and detects change in data distribution represented by the two streams over time using a Gaussian kernel model. We analyze the theoretical properties of the proposed approach and empirically evaluate it on both real-world and synthetic data sets. Importantly, our results indicate superior performance by the framework compared to other baseline regression methods.

IJCAI Conference 2018 Conference Paper

Hashtag2Vec: Learning Hashtag Representation with Relational Hierarchical Embedding Model

  • Jie Liu
  • Zhicheng He
  • Yalou Huang

Hashtags have always been important elements in many social network platforms and micro-blog services. Semantic understanding of hashtags is a critical and fundamental task for many applications on social networks, such as event analysis, theme discovery, information retrieval, etc. However, this task is challenging due to the sparsity, polysemy, and synonymy of hashtags. In this paper, we investigate the problem of hashtag embedding by combining the short text content with the various heterogeneous relations in social networks. Specifically, we first establish a network with hashtags as its nodes. Hierarchically, each of the hashtag nodes is associated with a set of tweets and each tweet contains a set of words. Then we devise an embedding model, called Hashtag2Vec, which exploits multiple relations of hashtag-hashtag, hashtag-tweet, tweet-word, and word-word relations based on the hierarchical heterogeneous network. In addition to embedding the hashtags, our proposed framework is capable of embedding the short social texts as well. Extensive experiments are conducted on two real-world datasets, and the results demonstrate the effectiveness of the proposed method.

YNIMG Journal 2018 Journal Article

The semantic system is involved in mathematical problem solving

  • Xinlin Zhou
  • Mengyi Li
  • Leinian Li
  • Yiyun Zhang
  • Jiaxin Cui
  • Jie Liu
  • Chuansheng Chen

Numerous studies have shown that the brain regions around bilateral intraparietal cortex are critical for number processing and arithmetical computation. However, the neural circuits for more advanced mathematics such as mathematical problem solving (with little routine arithmetical computation) remain unclear. Using functional magnetic resonance imaging (fMRI), this study (N = 24 undergraduate students) compared neural bases of mathematical problem solving (i. e. , number series completion, mathematical word problem solving, and geometric problem solving) and arithmetical computation. Direct subject- and item-wise comparisons revealed that mathematical problem solving typically had greater activation than arithmetical computation in all 7 regions of the semantic system (which was based on a meta-analysis of 120 functional neuroimaging studies on semantic processing). Arithmetical computation typically had greater activation in the supplementary motor area and left precentral gyrus. The results suggest that the semantic system in the brain supports mathematical problem solving.

TIST Journal 2017 Journal Article

Personalized Air Travel Prediction

  • Jie Liu
  • Bin Liu
  • Yanchi Liu
  • Huipeng Chen
  • Lina Feng
  • Hui Xiong
  • Yalou Huang

Human mobility analysis is one of the most important research problems in the field of urban computing. Existing research mainly focuses on the intra-city ground travel behavior modeling, while the inter-city air travel behavior modeling has been largely ignored. Actually, the inter-city travel analysis can be of equivalent importance and complementary to the intra-city travel analysis. Understanding massive passenger-air-travel behavior delivers intelligence for airlines’ precision marketing and related socioeconomic activities, such as airport planning, emergency management, local transportation planning, and tourism-related businesses. Moreover, it provides opportunities to study the characteristics of cities and the mutual relationships between them. However, modeling and predicting air traveler behavior is challenging due to the complex factors of the market situation and individual characteristics of customers (e.g., airlines’ market share, customer membership, and travelers’ intrinsic interests on destinations). To this end, in this article, we present a systematic study on the personalized air travel prediction problem, namely where a customer will fly to and which airline carrier to fly with, by leveraging real-world anonymized Passenger Name Record (PNR) data. Specifically, we first propose a relational travel topic model, which combines the merits of latent factor model with a neighborhood-based method, to uncover the personal travel preferences of aviation customers and the latent travel topics of air routes and airline carriers simultaneously. Then we present a multi-factor travel prediction framework, which fuses complex factors of the market situation and individual characteristics of customers, to predict airline customers’ personalized travel demands. Experimental results on two real-world PNR datasets demonstrate the effectiveness of our approach on both travel topic discovery and customer travel prediction.

YNIMG Journal 2017 Journal Article

The neural circuits for arithmetic principles

  • Jie Liu
  • Han Zhang
  • Chuansheng Chen
  • Hui Chen
  • Jiaxin Cui
  • Xinlin Zhou

Arithmetic principles are the regularities underlying arithmetic computation. Little is known about how the brain supports the processing of arithmetic principles. The current fMRI study examined neural activation and functional connectivity during the processing of verbalized arithmetic principles, as compared to numerical computation and general language processing. As expected, arithmetic principles elicited stronger activation in bilateral horizontal intraparietal sulcus and right supramarginal gyrus than did language processing, and stronger activation in left middle temporal lobe and left orbital part of inferior frontal gyrus than did computation. In contrast, computation elicited greater activation in bilateral horizontal intraparietal sulcus (extending to posterior superior parietal lobule) than did either arithmetic principles or language processing. Functional connectivity analysis with the psychophysiological interaction approach (PPI) showed that left temporal-parietal (MTG-HIPS) connectivity was stronger during the processing of arithmetic principle and language than during computation, whereas parietal-occipital connectivities were stronger during computation than during the processing of arithmetic principles and language. Additionally, the left fronto-parietal (orbital IFG-HIPS) connectivity was stronger during the processing of arithmetic principles than during computation. The results suggest that verbalized arithmetic principles engage a neural network that overlaps but is distinct from the networks for computation and language processing.

AAAI Conference 2017 Conference Paper

Topic Aware Neural Response Generation

  • Chen Xing
  • Wei Wu
  • Yu Wu
  • Jie Liu
  • Yalou Huang
  • Ming Zhou
  • Wei-Ying Ma

We consider incorporating topic information into a sequenceto-sequence framework to generate informative and interesting responses for chatbots. To this end, we propose a topic aware sequence-to-sequence (TA-Seq2Seq) model. The model utilizes topics to simulate prior human knowledge that guides them to form informative and interesting responses in conversation, and leverages topic information in generation by a joint attention mechanism and a biased generation probability. The joint attention mechanism summarizes the hidden vectors of an input message as context vectors by message attention and synthesizes topic vectors by topic attention from the topic words of the message obtained from a pre-trained LDA model, with these vectors jointly affecting the generation of words in decoding. To increase the possibility of topic words appearing in responses, the model modifies the generation probability of topic words by adding an extra probability item to bias the overall distribution. Empirical studies on both automatic evaluation metrics and human annotations show that TA-Seq2Seq can generate more informative and interesting responses, significantly outperforming state-of-theart response generation models.

YNIMG Journal 2017 Journal Article

Validating myelin water imaging with transmission electron microscopy in a rat spinal cord injury model

  • Henry Szu-Meng Chen
  • Nathan Holmes
  • Jie Liu
  • Wolfram Tetzlaff
  • Piotr Kozlowski

Myelin content is an important marker for neuropathology and MRI generated myelin water fraction (MWF) has been shown to correlate well with myelin content. However, because MWF is based on the amount of signal from myelin water, that is, the water trapped between the myelin lipid bilayers, the reading may depend heavily on myelin morphology. This is of special concern when there is a mix of intact myelin and myelin debris, as in the case of injury. To investigate what MWF measures in the presence of debris, we compared MWF to transmission electron microscopy (TEM) derived myelin fraction that measures the amount of compact appearing myelin. A rat spinal cord injury model was used with time points at normal (normal myelin), 3 weeks post-injury (myelin debris), and 8 weeks post-injury (myelin debris, partially cleared). The myelin period between normal and 3 or 8 weeks post-injury cords did not differ significantly, suggesting that as long as the bilayer structure is intact, myelin debris has the same water content as intact myelin. The MWF also correlated strongly with the TEM-derived myelin fraction, suggesting that MWF measures the amount of compact appearing myelin in both intact myelin and myelin debris. From the TEM images, it appears that as myelin degenerates, it tends to form large watery spaces within the myelin sheaths that are not classified as myelin water. The results presented in this study improve our understanding and allows for better interpretation of MWF in the presence of myelin debris.

AAAI Conference 2016 Conference Paper

Hashtag-Based Sub-Event Discovery Using Mutually Generative LDA in Twitter

  • Chen Xing
  • Yuan Wang
  • Jie Liu
  • Yalou Huang
  • Wei-Ying Ma

Sub-event discovery is an effective method for social event analysis in Twitter. It can discover sub-events from large amount of noisy event-related information in Twitter and semantically represent them. The task is challenging because tweets are short, informal and noisy. To solve this problem, we consider leveraging event-related hashtags that contain many locations, dates and concise sub-event related descriptions to enhance sub-event discovery. To this end, we propose a hashtag-based mutually generative Latent Dirichlet Allocation model(MGe-LDA). In MGe-LDA, hashtags and topics of a tweet are mutually generated by each other. The mutually generative process models the relationship between hashtags and topics of tweets, and highlights the role of hashtags as a semantic representation of the corresponding tweets. Experimental results show that MGe-LDA can significantly outperform state-of-the-art methods for sub-event discovery.

JMLR Journal 2016 Journal Article

Structure-Leveraged Methods in Breast Cancer Risk Prediction

  • Jun Fan
  • Yirong Wu
  • Ming Yuan
  • David Page
  • Jie Liu
  • Irene M. Ong
  • Peggy Peissig
  • Elizabeth Burnside

Predicting breast cancer risk has long been a goal of medical research in the pursuit of precision medicine. The goal of this study is to develop novel penalized methods to improve breast cancer risk prediction by leveraging structure information in electronic health records. We conducted a retrospective case- control study, garnering 49 mammography descriptors and 77 high- frequency/low-penetrance single-nucleotide polymorphisms (SNPs) from an existing personalized medicine data repository. Structured mammography reports and breast imaging features have long been part of a standard electronic health record (EHR), and genetic markers likely will be in the near future. Lasso and its variants are widely used approaches to integrated learning and feature selection, and our methodological contribution is to incorporate the dependence structure among the features into these approaches. More specifically, we propose a new methodology by combining group penalty and $\ell^p$ ($1\leq p\leq2$) fusion penalty to improve breast cancer risk prediction, taking into account structure information in mammography descriptors and SNPs. We demonstrate that our method provides benefits that are both statistically significant and potentially significant to people's lives. [abs] [ pdf ][ bib ] &copy JMLR 2016. ( edit, beta )

YNIMG Journal 2015 Journal Article

Central artery stiffness, baroreflex sensitivity, and brain white matter neuronal fiber integrity in older adults

  • Takashi Tarumi
  • Daan L.K. de Jong
  • David C. Zhu
  • Benjamin Y. Tseng
  • Jie Liu
  • Candace Hill
  • Jonathan Riley
  • Kyle B. Womack

Cerebral hypoperfusion elevates the risk of brain white matter (WM) lesions and cognitive impairment. Central artery stiffness impairs baroreflex, which controls systemic arterial perfusion, and may deteriorate neuronal fiber integrity of brain WM. The purpose of this study was to examine the associations among brain WM neuronal fiber integrity, baroreflex sensitivity (BRS), and central artery stiffness in older adults. Fifty-four adults (65±6years) with normal cognitive function or mild cognitive impairment (MCI) were tested. The neuronal fiber integrity of brain WM was assessed from diffusion metrics acquired by diffusion tensor imaging. BRS was measured in response to acute changes in blood pressure induced by bolus injections of vasoactive drugs. Central artery stiffness was measured by carotid–femoral pulse wave velocity (cfPWV). The WM diffusion metrics including fractional anisotropy (FA) and radial (RD) and axial (AD) diffusivities, BRS, and cfPWV were not different between the control and MCI groups. Thus, the data from both groups were combined for subsequent analyses. Across WM, fiber tracts with decreased FA and increased RD were associated with lower BRS and higher cfPWV, with many of the areas presenting spatial overlap. In particular, the BRS assessed during hypotension was strongly correlated with FA and RD when compared with hypertension. Executive function performance was associated with FA and RD in the areas that correlated with cfPWV and BRS. These findings suggest that baroreflex-mediated control of systemic arterial perfusion, especially during hypotension, may play a crucial role in maintaining neuronal fiber integrity of brain WM in older adults.

JBHI Journal 2014 Journal Article

A Comparative Study of Different Level Interpolations for Improving Spatial Resolution in Diffusion Tensor Imaging

  • Feng Yang
  • Yue-Min Zhu
  • Jian-Hua Luo
  • Marc Robini
  • Jie Liu
  • Pierre Croisille

This paper studies and evaluates the feasibility and the performance of different level interpolations for improving spatial resolution of diffusion tensor magnetic resonance imaging (DT-MRI or DTI). In particular, the following techniques are investigated: anisotropic interpolation operating on scalar gray-level images, log-Euclidean interpolation method, and the quaternion interpolation method, which operate on diffusion tensor fields. The performance is evaluated both qualitatively and quantitatively using criteria such as tensor determinant, fractional anisotropy (FA), mean diffusivity (MD), fiber length, etc. We conclude that tensor field interpolations allow avoiding undesirable swelling effect in DTI, which is not the case with scalar gray-level interpolation, and that scalar gray-level image interpolation and log-Euclidean tensor field interpolation suffer from decrease in FA and MD, which may mislead the interpretation of the clinical parameters FA and MD. In contrast, the quaternion tensor field interpolation avoids such FA and MD decrease, which suggests its use for clinical applications.

NeurIPS Conference 2013 Conference Paper

Bayesian Estimation of Latently-grouped Parameters in Undirected Graphical Models

  • Jie Liu
  • David Page

In large-scale applications of undirected graphical models, such as social networks and biological networks, similar patterns occur frequently and give rise to similar parameters. In this situation, it is beneficial to group the parameters for more efficient learning. We show that even when the grouping is unknown, we can infer these parameter groups during learning via a Bayesian approach. We impose a Dirichlet process prior on the parameters. Posterior inference usually involves calculating intractable terms, and we propose two approximation algorithms, namely a Metropolis-Hastings algorithm with auxiliary variables and a Gibbs sampling algorithm with stripped Beta approximation (Gibbs SBA). Simulations show that both algorithms outperform conventional maximum likelihood estimation (MLE). Gibbs SBA's performance is close to Gibbs sampling with exact likelihood calculation. Models learned with Gibbs_SBA also generalize better than the models learned by MLE on real-world Senate voting data.

v2026.09.13