Arrow Research search

Author name cluster

Zhen Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

36 papers
2 author rows

Possible papers

36

AAAI Conference 2026 Conference Paper

A Unified Shape-Aware Foundation Model for Time Series Classification

  • Zhen Liu
  • Yucheng Wang
  • Boyuan Li
  • Junhao Zheng
  • Emadeldeen Eldele
  • Min Wu
  • Qianli Ma

Foundation models pre-trained on large-scale source datasets are reshaping the traditional training paradigm for time series classification. However, existing time series foundation models primarily focus on forecasting tasks and often overlook classification-specific challenges, such as modeling interpretable shapelets that capture class-discriminative temporal features. To bridge this gap, we propose UniShape, a unified shape-aware foundation model designed for time series classification. UniShape incorporates a shape-aware adapter that adaptively aggregates multiscale discriminative subsequences (shapes) into class tokens, effectively selecting the most relevant subsequence scales to enhance model interpretability. Meanwhile, a prototype-based pretraining module is introduced to jointly learn instance- and shape-level representations, enabling the capture of transferable shape patterns. Pre-trained on a large-scale multi-domain time series dataset comprising 1.89 million samples, UniShape exhibits superior generalization across diverse target domains. Experiments on 128 UCR datasets and 30 additional time series datasets demonstrate that UniShape achieves state-of-the-art classification performance, with interpretability and ablation analyses further validating its effectiveness.

EAAI Journal 2026 Journal Article

Non-stationary multi-scale prediction model based on Patch Time Series Transformer for multi-step coal price forecasting

  • Shuai Ding
  • Bing Du
  • Kaidi Sun
  • Jiabo Xu
  • Huansheng Ning
  • Zhen Liu

Accurate forecasting of coal prices is essential for maintaining coal market stability and supporting carbon neutrality initiatives. However, the inherent non-stationarity and multi-scale characteristics of coal price time series pose significant challenges to effective modeling. To address these issues, we propose the Non-Stationary Multi-Scale Patch Time Series Transformer (NS-MS-PTST), a novel forecasting model that integrates Patch Time Series Transformer with Reversible Instance Normalization (ReVIN) to jointly capture fine-grained temporal structures and mitigate distribution shifts. Building upon this integration, the model further introduces a Hierarchical Reversible Normalization mechanism, which combines local-wise Reversible Patch Normalization (ReVPN) and global ReVIN to capture both fine-grained local statistics and global distributional properties. This layered design significantly enhances the model’s adaptability to complex and evolving coal price dynamics. The model also incorporates a multi-encoder architecture with a scale-wise parameter decay strategy to enable multi-scale temporal modeling while controlling complexity, along with a stepwise gated fusion strategy that dynamically balances contributions from different temporal scales. In addition, a response time metric is introduced to quantitatively evaluate the model’s recovery speed following abrupt price fluctuations, highlighting its dynamic adaptability and practical utility. Extensive experiments on real-world coal price datasets show that NS-MS-PTST achieves the lowest prediction errors across multiple standard evaluation metrics and exhibits the fastest response and recovery speed when facing abnormal price fluctuations, particularly in medium- and long-term forecasting tasks. These results confirm the model’s superior accuracy, strong generalization ability, and robustness in real-world coal market scenarios.

AAAI Conference 2026 Conference Paper

RAW-Flow: Advancing RGB-to-RAW Image Reconstruction with Deterministic Latent Flow Matching

  • Zhen Liu
  • Diedong Feng
  • Hai Jiang
  • Liaoyuan Zeng
  • Hao Wang
  • Chaoyu Feng
  • Lei Lei
  • Bing Zeng

RGB-to-RAW reconstruction, or the reverse modeling of a camera Image Signal Processing (ISP) pipeline, aims to recover high-fidelity RAW data from RGB images. Despite notable progress, existing learning-based methods typically treat this task as a direct regression objective and still struggle with detail inconsistency and color deviation, due to the ill-posed nature of inverse ISP and the inherent information loss in quantized RGB images. To address these limitations, we pioneer a generative perspective by reformulating RGB-to-RAW reconstruction as a deterministic latent transport problem and introduce a novel framework named RAW-Flow, which leverages flow matching to learn a deterministic vector field in latent space, to effectively bridge the gap between RGB and RAW representations and enable accurate reconstruction of structural details and color information. To further enhance latent transport, we introduce a cross-scale context guidance module that injects hierarchical RGB features into the flow estimation process. Moreover, we design a Dual-domain Latent Autoencoder (DLAE) with a feature alignment constraint to support the proposed latent transport framework, which jointly encodes RGB and RAW inputs while promoting stable training and high-fidelity reconstruction. Extensive experiments demonstrate that RAW-Flow outperforms state-of-the-art approaches both quantitatively and visually.

EAAI Journal 2025 Journal Article

A fuzzy decision-making method for unloading command at crushing stations based on deep learning and dynamic coal flow features slicing

  • Tongyu Cui
  • Yongtai Pan
  • Yankun Bi
  • Zhen Liu
  • Jiacheng Huang
  • Bingjia Liu

The contemporary world hosts numerous open-pit coal mines, as mining efficiency increases, the unloading process of the crushing station still relies on manual command. Adversely affected by the harsh production environment and the high-intensity fatigue of operators, incorrect unloading commands can lead to issues, such as blockages at the receiver bin discharge outlet and coal swelling in the crushing chamber, resulting in production accidents. To address this issue, this study analyzes the dynamic characteristics of the coal flow in the receiver bin, in addition, monitors the bin status in real time through the camera, combining a lightweight selective kernel network (Li-SKNet) with the average classification confidence from a small segment of consecutive video frames to assess the suitability of unloading conditions. By introducing a convolutional kernel attention mechanism, the model achieved a classification accuracy of 99. 8 %, subsequently, in order to adapt to the changes in the working conditions of the crushing station, a segmented fuzzy function is proposed to further optimize the model inference results. Finally, an intelligent unloading command system for the crushing station is established, which has been put into production and operated stably for six months. In comparison to traditional manual commanding methods for unloading, the system has approximately demonstrated a 15 % increase in production efficiency.

EAAI Journal 2025 Journal Article

Dynamic Interactive Graph Convolutional Recurrent Network for bidirectional spatiotemporal traffic flow forecasting

  • Zhen Liu
  • Shiqi Zhang
  • Yuzhuang Pian
  • Yonghong Liu

Accurate prediction of traffic inflow and outflow is essential for efficient urban mobility management and multimodal transit systems. However, existing approaches struggle with two main challenges: (i) The dynamic spatiotemporal heterogeneity that varies across different regions and times, complicating the prediction task. (ii) The asymmetric interdependence between inflows and outflows is often overlooked, leading to an inadequate representation of intricate bidirectional relationships. To address these challenges, we propose the Dynamic Interactive Graph Convolutional Recurrent Network (DIGCRN). In particular, DIGCRN incorporates an inflow and outflow feature interaction learning that utilizes an interactive gated mechanism to achieve spatiotemporal characteristic transformation, thereby capturing the asymmetric interdependence between inflows and outflows. Subsequently, a gated recurrent unit based on an adaptive graph convolutional network is employed to recursively capture spatiotemporal features. The final multi-scale convolution module realizes the fusion of inflow and outflow features at different granularity levels. Comprehensive empirical evaluations on the Hangzhou Metro and New York City Taxi datasets indicate that DIGCRN surpasses all baselines, achieving improvements of up to 3. 46% in mean absolute error (MAE) and 7. 28% in root mean square error (RMSE) compared to the best-performing baseline models. The code is available at https: //github. com/LiuZhen1234567/DIGCRN.

AAAI Conference 2025 Conference Paper

FlowPolicy: Enabling Fast and Robust 3D Flow-Based Policy via Consistency Flow Matching for Robot Manipulation

  • Qinglun Zhang
  • Zhen Liu
  • Haoqiang Fan
  • Guanghui Liu
  • Bing Zeng
  • Shuaicheng Liu

Robots can acquire complex manipulation skills by learning policies from expert demonstrations, which is often known as vision-based imitation learning. Generating policies based on diffusion and flow matching models has been shown to be effective, particularly in robotic manipulation tasks. However, recursion-based approaches are inference inefficient in working from noise distributions to policy distributions, posing a challenging trade-off between efficiency and quality. This motivates us to propose FlowPolicy, a novel framework for fast policy generation based on consistency flow matching and 3D vision. Our approach refines the flow dynamics by normalizing the self-consistency of the velocity field, enabling the model to derive task execution policies in a single inference step. Specifically, FlowPolicy conditions on the observed 3D point cloud, where consistency flow matching directly defines straight-line flows from different time states to the same action space, while simultaneously constraining their velocity values, that is, we approximate the trajectories from noise to robot actions by normalizing the self-consistency of the velocity field within the action space, thus improving the inference efficiency. We validate the effectiveness of FlowPolicy in Adroit and Metaworld, demonstrating a 7× increase in inference speed while maintaining competitive average success rates compared to state-of-the-art methods.

NeurIPS Conference 2025 Conference Paper

Luminance-Aware Statistical Quantization: Unsupervised Hierarchical Learning for Illumination Enhancement

  • Derong Kong
  • Zhixiong Yang
  • Shengxi Li
  • Shuaifeng Zhi
  • Li Liu
  • Zhen Liu
  • Jingyuan Xia

Low-light image enhancement (LLIE) faces persistent challenges in balancing reconstruction fidelity with cross-scenario generalization. While existing methods predominantly focus on deterministic pixel-level mappings between paired low/normal-light images, they often neglect the continuous physical process of luminance transitions in real-world environments, leading to performance drop when normal-light references are unavailable. Inspired by empirical analysis of natural luminance dynamics revealing power-law distributed intensity transitions, this paper introduces Luminance-Aware Statistical Quantification (LASQ), a novel framework that reformulates LLIE as a statistical sampling process over hierarchical luminance distributions. Our LASQ re-conceptualizes luminance transition as a power-law distribution in intensity coordinate space that can be approximated by stratified power functions, therefore, replacing deterministic mappings with probabilistic sampling over continuous luminance layers. A diffusion forward process is designed to autonomously discover optimal transition paths between luminance layers, achieving unsupervised distribution emulation without normal-light references. In this way, it considerably improves the performance in practical situations, enabling more adaptable and versatile light restoration. This framework is also readily applicable to cases with normal-light references, where it achieves superior performance on domain-specific datasets alongside better generalization-ability across non-reference datasets. The code is available at: https: //github. com/XYLGroup/LASQ.

NeurIPS Conference 2025 Conference Paper

Nabla-R2D3: Effective and Efficient 3D Diffusion Alignment with 2D Rewards

  • Qingming LIU
  • Zhen Liu
  • Dinghuai Zhang
  • Kui Jia

Generating high-quality and photorealistic 3D assets remains a longstanding challenge in 3D vision and computer graphics. Although state-of-the-art generative models, such as diffusion models, have made significant progress in 3D generation, they often fall short of human-designed content due to limited ability to follow instructions, align with human preferences, or produce realistic textures, geometries, and physical attributes. In this paper, we introduce Nabla-R2D3, a highly effective and sample-efficient reinforcement learning alignment framework for 3D-native diffusion models using 2D rewards. Built upon the recently proposed Nabla-GFlowNet method for reward finetuning, our Nabla-R2D3 enables effective adaptation of 3D diffusion models through pure 2D reward feedback. Extensive experiments show that, unlike naive finetuning baselines which either fail to converge or suffer from overfitting, Nabla-R2D3 consistently achieves higher rewards and reduced prior forgetting within few finetuning steps.

EAAI Journal 2025 Journal Article

SkeletonDETR: A novel multimodal fusion based object detection framework for chemical safety applications

  • Yudi Tang
  • Bing Wang
  • Wangli He
  • Feng Qian
  • Zhen Liu

Regarding the object detection algorithm in field operations in chemical plants, how to accurately detect objects carried and used by construction workers has become a crucial challenge in safety monitoring. Current object detection algorithms usually perform well for large objects but can easily ignore small objects, particularly under partial occlusion. Moreover, existing methods fail to recognize the significance of workers’ pose information during the construction process, which provides significant benefits for the detection task, especially when the targets for detection are closely associated with the construction workers. To solve this problem, we proposed a novel multimodal fusion based object detection framework, which can effectively use human pose information to improve the detection effect of small targets and occlusions. Furthermore, we propose a multimodal sampling module to fully utilize the features of different modalities to enhance the encoder’s ability to aggregate features. Compared with the baseline model, our proposed method achieves an 8. 3% improvement in small object detection performance. Comprehensive experiments demonstrate that our proposed method outperforms existing efficient models, especially in field operation in chemical plants.

NeurIPS Conference 2025 Conference Paper

Value Gradient Guidance for Flow Matching Alignment

  • Zhen Liu
  • Tim Xiao
  • Carles Domingo i Enrich
  • Weiyang Liu
  • Dinghuai Zhang

While methods exist for aligning flow matching models––a popular and effective class of generative models––with human preferences, existing approaches fail to achieve both adaptation efficiency and probabilistically sound prior preservation. In this work, we leverage the theory of optimal control and propose VGG-Flow, a gradient matching–based method for finetuning pretrained flow matching models. The key idea in this algorithm is that the optimal difference between the finetuned velocity field and the pretrained one should be matched with the gradient field of a value function. This method not only incorporates first-order information from the reward model but also benefits from heuristic initialization of the value function to enable fast adaptation. Empirically, we show on a popular text-to-image flow matching model, Stable Diffusion 3, that our method can finetune flow matching models under limited computational budgets while achieving effective and prior-preserving alignment.

TCS Journal 2024 Journal Article

A secure hierarchical deterministic wallet with stealth address from lattices

  • Xin Yin
  • Zhen Liu

The concept of Hierarchical Deterministic Wallet (HDW) was introduced by Wuille in Bitcoin Improvement Proposal 32 (BIP32). HDW enables an individual/organization to generate cryptographic keys and subsequently ease the key management problems (e. g. , backup and recovery). Since the first HDW algorithm in 2012, HDW has gradually shown its fit for many promising use cases, such as Bitcoin-like cryptocurrencies, global key revocation in FIDO2 standard. In order to achieve all the features (i. e. , deterministic derivation, master public key and hierarchy) and the security (i. e. , safety of cryptocurrencies and privacy protection of users) requirements for HDW, Yin et al. (ESORICS 2022) conceptualized Hierarchical Deterministic Wallet supporting Stealth Address (HDWSA), and gave a provably secure construction from the standard Computational Diffie-Hellman Assumption. Unfortunately, the construction is not quantum-resistant. In this work, we propose the first HDWSA construction from lattices to fill this gap, we provide the security proof for the construction in the random oracle model (ROM) based on hard problems over lattices. Compared with existing works, to the best of our knowledge, our construction not only captures all the HDW features and security properties, but also provides the potential quantum resistance.

EAAI Journal 2024 Journal Article

Boosting cluster tree with reciprocal nearest neighbors scoring

  • Wen-Bo Xie
  • Zhen Liu
  • Bin Chen
  • Jaideep Srivastava

Clustering plays a pivotal role in knowledge processing, knowledge bases, and expert systems, enabling AI systems to acquire knowledge effectively. Hierarchical clustering, in particular, offers an intelligent approach to represent knowledge hierarchically by transforming raw data into one/multiple tree-shaped components. However, a notable difficulty arises when attempting to pinpoint appropriate representative points within lower levels of the cluster tree. These points are of paramount importance, as they serve as the roots for subsequent aggregation within the upper levels of the cluster tree. Traditional hierarchical clustering algorithms have relied on rudimentary techniques to select these representative points, which may not provide an adequate representation. Consequently, the resulting cluster tree often falls short in terms of empirical performance. To address this shortcoming, we proposed an innovative hierarchical clustering algorithm in this paper. The proposed algorithm is designed to efficiently identify the representative point within each sub-minimum-spanning-tree during the construction of the cluster tree, achieved by topology-based scoring the reciprocal nearest data points. Rigorous testing on UCI datasets has demonstrated the superior clustering accuracy (measured by Rand Index and Normalized Mutual Information) of our proposed algorithm compared to other benchmark algorithms. Further analysis reveals that our algorithm boasts a O ( n log n ) time-complexity and a O ( log n ) space-complexity, indicating its scalability and efficiency in handling large-scale data with minimal time and storage costs. Importantly, our algorithm’s ability to process up to two million data points on a standard personal computer underscores its cost-effectiveness.

AAAI Conference 2024 Conference Paper

Diffusion Language-Shapelets for Semi-supervised Time-Series Classification

  • Zhen Liu
  • Wenbin Pei
  • Disen Lan
  • Qianli Ma

Semi-supervised time-series classification could effectively alleviate the issue of lacking labeled data. However, existing approaches usually ignore model interpretability, making it difficult for humans to understand the principles behind the predictions of a model. Shapelets are a set of discriminative subsequences that show high interpretability in time series classification tasks. Shapelet learning-based methods have demonstrated promising classification performance. Unfortunately, without enough labeled data, the shapelets learned by existing methods are often poorly discriminative, and even dissimilar to any subsequence of the original time series. To address this issue, we propose the Diffusion Language-Shapelets model (DiffShape) for semi-supervised time series classification. In DiffShape, a self-supervised diffusion learning mechanism is designed, which uses real subsequences as a condition. This helps to increase the similarity between the learned shapelets and real subsequences by using a large amount of unlabeled data. Furthermore, we introduce a contrastive language-shapelets learning strategy that improves the discriminability of the learned shapelets by incorporating the natural language descriptions of the time series. Experiments have been conducted on the UCR time series archive, and the results reveal that the proposed DiffShape method achieves state-of-the-art performance and exhibits superior interpretability over baselines.

EAAI Journal 2024 Journal Article

Intelligent analysis method for the global vertical displacement field of foundation pits in dense karst cave areas

  • Jin Liao
  • Chunxiu Lin
  • Chunhui Lan
  • Yongtao Wu
  • Zhen Liu
  • Cuiying Zhou

Prior research on excavation in dense karst cave foundation pits has primarily concentrated on evaluating the localized spatio-temporal influence and isolated geological factors. Nonetheless, this approach oversimplifies modeling conditions, thereby limiting its ability to provide a comprehensive understanding of the vertical displacement field. Consequently, this oversimplification can inflate the safety factor and increase project costs. Therefore, we propose a feedforward neural network (FNN), updated with the loop nested optimal iterative method (LNOIM), which incorporates the spatiotemporal characteristics of monitoring points and geological factors to analyze the engineering sensitivity of karst caves. Ultimately, the global foundation pit vertical displacement field was obtained. Our method has been demonstrated to be effective in a foundation pit in South China ((P value) P > 0. 050, Cohen's d < 0. 200). Furthermore, it has been validated in other cases ((Root Mean Square Error) RMSE = 1. 576–2. 916). This work provides a new perspective on the accurate reflection of the global vertical displacement state of a foundation pit. Additionally, it enhances the ability to sensitively identify caves in dense karst cave areas, thereby improving the safety of foundation pit works.

NeurIPS Conference 2024 Conference Paper

Knowledge-Empowered Dynamic Graph Network for Irregularly Sampled Medical Time Series

  • Yicheng Luo
  • Zhen Liu
  • Linghao Wang
  • Junhao Zheng
  • Binquan Wu
  • Qianli Ma

Irregularly Sampled Medical Time Series (ISMTS) are commonly found in the healthcare domain, where different variables exhibit unique temporal patterns while interrelated. However, many existing methods fail to efficiently consider the differences and correlations among medical variables together, leading to inadequate capture of fine-grained features at the variable level in ISMTS. We propose Knowledge-Empowered Dynamic Graph Network (KEDGN), a graph neural network empowered by variables' textual medical knowledge, aiming to model variable-specific temporal dependencies and inter-variable dependencies in ISMTS. Specifically, we leverage a pre-trained language model to extract semantic representations for each variable from their textual descriptions of medical properties, forming an overall semantic view among variables from a medical perspective. Based on this, we allocate variable-specific parameter spaces to capture variable-specific temporal patterns and generate a complete variable graph to measure medical correlations among variables. Additionally, we employ a density-aware mechanism to dynamically adjust the variable graph at different timestamps, adapting to the time-varying correlations among variables in ISMTS. The variable-specific parameter spaces and dynamic graphs are injected into the graph convolutional recurrent network to capture intra-variable and inter-variable dependencies in ISMTS together. Experiment results on four healthcare datasets demonstrate that KEDGN significantly outperforms existing methods.

TMLR Journal 2024 Journal Article

Large Language Models Synergize with Automated Machine Learning

  • Jinglue Xu
  • Jialong Li
  • Zhen Liu
  • NAV Suryanarayanan
  • Guoyuan Zhou
  • Jia Guo
  • Hitoshi Iba
  • Kenji Tei

Recently, program synthesis driven by large language models (LLMs) has become increasingly popular. However, program synthesis for machine learning (ML) tasks still poses significant challenges. This paper explores a novel form of program synthesis, targeting ML programs, by combining LLMs and automated machine learning (autoML). Specifically, our goal is to fully automate the generation and optimization of the code of the entire ML workflow, from data preparation to modeling and post-processing, utilizing only textual descriptions of the ML tasks. To manage the length and diversity of ML programs, we propose to break each ML program into smaller, manageable parts. Each part is generated separately by the LLM, with careful consideration of their compatibilities. To ensure compatibilities, we design a testing technique for ML programs. Unlike traditional program synthesis, which typically relies on binary evaluations (i.e., correct or incorrect), evaluating ML programs necessitates more than just binary judgments. Our approach automates the numerical evaluation and optimization of these programs, selecting the best candidates through autoML techniques. In experiments across various ML tasks, our method outperforms existing methods in 10 out of 12 tasks for generating ML programs. In addition, autoML significantly improves the performance of the generated ML programs. In experiments, given the textual task description, our method, Text-to-ML, generates the complete and optimized ML program in a fully autonomous process. The implementation of our method is available at https://github.com/JLX0/llm-automl.

AAAI Conference 2024 Conference Paper

Uncertainty-Aware Yield Prediction with Multimodal Molecular Features

  • Jiayuan Chen
  • Kehan Guo
  • Zhen Liu
  • Olexandr Isayev
  • Xiangliang Zhang

Predicting chemical reaction yields is pivotal for efficient chemical synthesis, an area that focuses on the creation of novel compounds for diverse uses. Yield prediction demands accurate representations of reactions for forecasting practical transformation rates. Yet, the uncertainty issues broadcasting in real-world situations prohibit current models to excel in this task owing to the high sensitivity of yield activities and the uncertainty in yield measurements. Existing models often utilize single-modal feature representations, such as molecular fingerprints, SMILES sequences, or molecular graphs, which is not sufficient to capture the complex interactions and dynamic behavior of molecules in reactions. In this paper, we present an advanced Uncertainty-Aware Multimodal model (UAM) to tackle these challenges. Our approach seamlessly integrates data sources from multiple modalities by encompassing sequence representations, molecular graphs, and expert-defined chemical reaction features for a comprehensive representation of reactions. Additionally, we address both the model and data-based uncertainty, refining the model's predictive capability. Extensive experiments on three datasets, including two high throughput experiment (HTE) datasets and one chemist-constructed Amide coupling reaction dataset, demonstrate that UAM outperforms the state-of-the-art methods. The code and used datasets are available at https://github.com/jychen229/Multimodal-reaction-yield-prediction.

TMLR Journal 2023 Journal Article

Continual Learning by Modeling Intra-Class Variation

  • Longhui Yu
  • Tianyang Hu
  • Lanqing Hong
  • Zhen Liu
  • Adrian Weller
  • Weiyang Liu

It has been observed that neural networks perform poorly when the data or tasks are presented sequentially. Unlike humans, neural networks suffer greatly from catastrophic forgetting, making it impossible to perform life-long learning. To address this issue, memory-based continual learning has been actively studied and stands out as one of the best-performing methods. We examine memory-based continual learning and identify that large variation in the representation space is crucial for avoiding catastrophic forgetting. Motivated by this, we propose to diversify representations by using two types of perturbations: model-agnostic variation (i.e., the variation is generated without the knowledge of the learned neural network) and model-based variation (i.e., the variation is conditioned on the learned neural network). We demonstrate that enlarging representational variation serves as a general principle to improve continual learning. Finally, we perform empirical studies which demonstrate that our method, as a simple plug-and-play component, can consistently improve a number of memory-based continual learning methods by a large margin.

NeurIPS Conference 2023 Conference Paper

Controlling Text-to-Image Diffusion by Orthogonal Finetuning

  • Zeju Qiu
  • Weiyang Liu
  • Haiwen Feng
  • Yuxuan Xue
  • Yao Feng
  • Zhen Liu
  • Dan Zhang
  • Adrian Weller

Large text-to-image diffusion models have impressive capabilities in generating photorealistic images from text prompts. How to effectively guide or control these powerful models to perform different downstream tasks becomes an important open problem. To tackle this challenge, we introduce a principled finetuning method -- Orthogonal Finetuning (OFT), for adapting text-to-image diffusion models to downstream tasks. Unlike existing methods, OFT can provably preserve hyperspherical energy which characterizes the pairwise neuron relationship on the unit hypersphere. We find that this property is crucial for preserving the semantic generation ability of text-to-image diffusion models. To improve finetuning stability, we further propose Constrained Orthogonal Finetuning (COFT) which imposes an additional radius constraint to the hypersphere. Specifically, we consider two important finetuning text-to-image tasks: subject-driven generation where the goal is to generate subject-specific images given a few images of a subject and a text prompt, and controllable generation where the goal is to enable the model to take in additional control signals. We empirically show that our OFT framework outperforms existing methods in generation quality and convergence speed.

IJCAI Conference 2023 Conference Paper

CTW: Confident Time-Warping for Time-Series Label-Noise Learning

  • Peitian Ma
  • Zhen Liu
  • Junhao Zheng
  • Linghao Wang
  • Qianli Ma

Noisy labels seriously degrade the generalization ability of Deep Neural Networks (DNNs) in various classification tasks. Existing studies on label-noise learning mainly focus on computer vision, while time series also suffer from the same issue. Directly applying the methods from computer vision to time series may reduce the temporal dependency due to different data characteristics. How to make use of the properties of time series to enable DNNs to learn robust representations in the presence of noisy labels has not been fully explored. To this end, this paper proposes a method that expands the distribution of Confident instances by Time-Warping (CTW) to learn robust representations of time series. Specifically, since applying the augmentation method to all data may introduce extra mislabeled data, we select confident instances to implement Time-Warping. In addition, we normalize the distribution of the training loss of each class to eliminate the model's selection preference for instances of different classes, alleviating the class imbalance caused by sample selection. Extensive experimental results show that CTW achieves state-of-the-art performance on the UCR datasets when dealing with different types of noise. Besides, the t-SNE visualization of our method verifies that augmenting confident data improves the generalization ability. Our code is available at https: //github. com/qianlima-lab/CTW.

AAAI Conference 2023 Conference Paper

Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA

  • Yongxin Zhu
  • Zhen Liu
  • Yukang Liang
  • Xin Li
  • Hao Liu
  • Changcun Bao
  • Linli Xu

In this paper, we propose a novel multi-modal framework for Scene Text Visual Question Answering (STVQA), which requires models to read scene text in images for question answering. Apart from text or visual objects, which could exist independently, scene text naturally links text and visual modalities together by conveying linguistic semantics while being a visual object in an image simultaneously. Different to conventional STVQA models which take the linguistic semantics and visual semantics in scene text as two separate features, in this paper, we propose a paradigm of "Locate Then Generate" (LTG), which explicitly unifies this two semantics with the spatial bounding box as a bridge connecting them. Specifically, at first, LTG locates the region in an image that may contain the answer words with an answer location module (ALM) consisting of a region proposal network and a language refinement network, both of which can transform to each other with one-to-one mapping via the scene text bounding box. Next, given the answer words selected by ALM, LTG generates a readable answer sequence with an answer generation module (AGM) based on a pre-trained language model. As a benefit of the explicit alignment of the visual and linguistic semantics, even without any scene text based pre-training tasks, LTG can boost the absolute accuracy by +6.06% and +6.92% on the TextVQA dataset and the ST-VQA dataset respectively, compared with a non-pre-training baseline. We further demonstrate that LTG effectively unifies visual and text modalities through the spatial bounding box connection, which is underappreciated in previous methods.

NeurIPS Conference 2023 Conference Paper

Scale-teaching: Robust Multi-scale Training for Time Series Classification with Noisy Labels

  • Zhen Liu
  • ma peitian
  • Dongliang Chen
  • Wenbin Pei
  • Qianli Ma

Deep Neural Networks (DNNs) have been criticized because they easily overfit noisy (incorrect) labels. To improve the robustness of DNNs, existing methods for image data regard samples with small training losses as correctly labeled data (small-loss criterion). Nevertheless, time series' discriminative patterns are easily distorted by external noises (i. e. , frequency perturbations) during the recording process. This results in training losses of some time series samples that do not meet the small-loss criterion. Therefore, this paper proposes a deep learning paradigm called Scale-teaching to cope with time series noisy labels. Specifically, we design a fine-to-coarse cross-scale fusion mechanism for learning discriminative patterns by utilizing time series at different scales to train multiple DNNs simultaneously. Meanwhile, each network is trained in a cross-teaching manner by using complementary information from different scales to select small-loss samples as clean labels. For unselected large-loss samples, we introduce multi-scale embedding graph learning via label propagation to correct their labels by using selected clean samples. Experiments on multiple benchmark time series datasets demonstrate the superiority of the proposed Scale-teaching paradigm over state-of-the-art methods in terms of effectiveness and robustness.

AAAI Conference 2023 Conference Paper

Temporal-Frequency Co-training for Time Series Semi-supervised Learning

  • Zhen Liu
  • Qianli Ma
  • Peitian Ma
  • Linghao Wang

Semi-supervised learning (SSL) has been actively studied due to its ability to alleviate the reliance of deep learning models on labeled data. Although existing SSL methods based on pseudo-labeling strategies have made great progress, they rarely consider time-series data's intrinsic properties (e.g., temporal dependence). Learning representations by mining the inherent properties of time series has recently gained much attention. Nonetheless, how to utilize feature representations to design SSL paradigms for time series has not been explored. To this end, we propose a Time Series SSL framework via Temporal-Frequency Co-training (TS-TFC), leveraging the complementary information from two distinct views for unlabeled data learning. In particular, TS-TFC employs time-domain and frequency-domain views to train two deep neural networks simultaneously, and each view's pseudo-labels generated by label propagation in the representation space are adopted to guide the training of the other view's classifier. To enhance the discriminative of representations between categories, we propose a temporal-frequency supervised contrastive learning module, which integrates the learning difficulty of categories to improve the quality of pseudo-labels. Through co-training the pseudo-labels obtained from temporal-frequency representations, the complementary information in the two distinct views is exploited to enable the model to better learn the distribution of categories. Extensive experiments on 106 UCR datasets show that TS-TFC outperforms state-of-the-art methods, demonstrating the effectiveness and robustness of our proposed model.

TMLR Journal 2023 Journal Article

Using Representation Expressiveness and Learnability to Evaluate Self-Supervised Learning Methods

  • Yuchen Lu
  • Zhen Liu
  • Aristide Baratin
  • Romain Laroche
  • Aaron Courville
  • Alessandro Sordoni

We address the problem of evaluating the quality of self-supervised learning (SSL) models without access to supervised labels, while being agnostic to the architecture, learning algorithm or data manipulation used during training. We argue that representations can be evaluated through the lens of expressiveness and learnability. We propose to use the Intrinsic Dimension (ID) to assess expressiveness and introduce Cluster Learnability (CL) to assess learnability. CL is measured in terms of the performance of a KNN classifier trained to predict labels obtained by clustering the representations with K-means. We thus combine CL and ID into a single predictor – CLID. Through a large-scale empirical study with a diverse family of SSL algorithms, we find that CLID better correlates with in-distribution model performance than other competing recent evaluation schemes. We also benchmark CLID on out-of-domain generalization, where CLID serves as a predictor of the transfer performance of SSL models on several visual classification tasks, yielding improvements with respect to the competing baselines.

TCS Journal 2022 Journal Article

Purchase preferences-based air passenger choice behavior analysis from sales transaction data

  • Xinghua Li
  • Suixiang Gao
  • Wenguo Yang
  • Yu Si
  • Zhen Liu

Travel providers such as airlines are becoming more and more interested in understanding how passengers choose among alternative products, especially the purchasing preferences of passengers. Getting information of air passenger choice behavior helps them better display and adapt their offer. Discrete choice models are appealing for airline revenue management (RM). In this paper, we apply latent class multinomial logit model (LC-MNL) to passenger choice behavior. The analysis based on actual sales transaction data reveals the purchase preferences of different passenger types. According to the distribution of the market, we divide passengers into three groups: low-price oriented, high-price oriented and no specific price preference. The low-price oriented passengers only choose products from the set which consists of the lowest price cabin classes of each flight, while the high-price oriented passengers do the opposite. Considering that the passenger types in the transaction sales data are unknown, the latent class passenger choice model can better represent their heterogeneous purchasing preference. An improved EM algorithm is applied to solve the LC-MNL. In the improved EM algorithm, an indicator function containing both the type of passengers and the first choice information in period t is devised, the iterative process of the EM algorithm is more effective consequently. The proposed model and algorithm are evaluated on actual aviation sales transaction data in China. Experimental results show that the passenger choice behavior analysis based on the specific purchasing preferences performs well on actual aviation sales transaction data.

NeurIPS Conference 2021 Conference Paper

A Consciousness-Inspired Planning Agent for Model-Based Reinforcement Learning

  • Mingde Zhao
  • Zhen Liu
  • Sitao Luan
  • Shuyuan Zhang
  • Doina Precup
  • Yoshua Bengio

We present an end-to-end, model-based deep reinforcement learning agent which dynamically attends to relevant parts of its state during planning. The agent uses a bottleneck mechanism over a set-based representation to force the number of entities to which the agent attends at each planning step to be small. In experiments, we investigate the bottleneck mechanism with several sets of customized environments featuring different challenges. We consistently observe that the design allows the planning agents to generalize their learned task-solving abilities in compatible unseen environments by attending to the relevant objects, leading to better out-of-distribution generalization performance.

NeurIPS Conference 2021 Conference Paper

Iterative Teaching by Label Synthesis

  • Weiyang Liu
  • Zhen Liu
  • Hanchen Wang
  • Liam Paull
  • Bernhard Schölkopf
  • Adrian Weller

In this paper, we consider the problem of iterative machine teaching, where a teacher provides examples sequentially based on the current iterative learner. In contrast to previous methods that have to scan over the entire pool and select teaching examples from it in each iteration, we propose a label synthesis teaching framework where the teacher randomly selects input teaching examples (e. g. , images) and then synthesizes suitable outputs (e. g. , labels) for them. We show that this framework can avoid costly example selection while still provably achieving exponential teachability. We propose multiple novel teaching algorithms in this framework. Finally, we empirically demonstrate the value of our framework.

NeurIPS Conference 2019 Conference Paper

Exponential Family Estimation via Adversarial Dynamics Embedding

  • Bo Dai
  • Zhen Liu
  • Hanjun Dai
  • Niao He
  • Arthur Gretton
  • Le Song
  • Dale Schuurmans

We present an efficient algorithm for maximum likelihood estimation (MLE) of exponential family models, with a general parametrization of the energy function that includes neural networks. We exploit the primal-dual view of the MLE with a kinetics augmented model to obtain an estimate associated with an adversarial dual sampler. To represent this sampler, we introduce a novel neural architecture, dynamics embedding, that generalizes Hamiltonian Monte-Carlo (HMC). The proposed approach inherits the flexibility of HMC while enabling tractable entropy estimation for the augmented model. By learning both a dual sampler and the primal model simultaneously, and sharing parameters between them, we obviate the requirement to design a separate sampling procedure once the model has been trained, leading to more effective learning. We show that many existing estimators, such as contrastive divergence, pseudo/composite-likelihood, score matching, minimum Stein discrepancy estimator, non-local contrastive objectives, noise-contrastive estimation, and minimum probability flow, are special cases of the proposed approach, each expressed by a different (fixed) dual sampler. An empirical investigation shows that adapting the sampler during MLE can significantly improve on state-of-the-art estimators.

NeurIPS Conference 2019 Conference Paper

Neural Similarity Learning

  • Weiyang Liu
  • Zhen Liu
  • James Rehg
  • Le Song

Inner product-based convolution has been the founding stone of convolutional neural networks (CNNs), enabling end-to-end learning of visual representation. By generalizing inner product with a bilinear matrix, we propose the neural similarity which serves as a learnable parametric similarity measure for CNNs. Neural similarity naturally generalizes the convolution and enhances flexibility. Further, we consider the neural similarity learning (NSL) in order to learn the neural similarity adaptively from training data. Specifically, we propose two different ways of learning the neural similarity: static NSL and dynamic NSL. Interestingly, dynamic neural similarity makes the CNN become a dynamic inference network. By regularizing the bilinear matrix, NSL can be viewed as learning the shape of kernel and the similarity measure simultaneously. We further justify the effectiveness of NSL with a theoretical viewpoint. Most importantly, NSL shows promising performance in visual recognition and few-shot learning, validating the superiority of NSL over the inner product-based convolution counterparts.

NeurIPS Conference 2018 Conference Paper

Coupled Variational Bayes via Optimization Embedding

  • Bo Dai
  • Hanjun Dai
  • Niao He
  • Weiyang Liu
  • Zhen Liu
  • Jianshu Chen
  • Lin Xiao
  • Le Song

Variational inference plays a vital role in learning graphical models, especially on large-scale datasets. Much of its success depends on a proper choice of auxiliary distribution class for posterior approximation. However, how to pursue an auxiliary distribution class that achieves both good approximation ability and computation efficiency remains a core challenge. In this paper, we proposed coupled variational Bayes which exploits the primal-dual view of the ELBO with the variational distribution class generated by an optimization procedure, which is termed optimization embedding. This flexible function class couples the variational distribution with the original parameters in the graphical models, allowing end-to-end learning of the graphical models by back-propagation through the variational distribution. Theoretically, we establish an interesting connection to gradient flow and demonstrate the extreme flexibility of this implicit distribution family in the limit sense. Empirically, we demonstrate the effectiveness of the proposed method on multiple graphical models with either continuous or discrete latent variables comparing to state-of-the-art methods.

NeurIPS Conference 2018 Conference Paper

Learning towards Minimum Hyperspherical Energy

  • Weiyang Liu
  • Rongmei Lin
  • Zhen Liu
  • Lixin Liu
  • Zhiding Yu
  • Bo Dai
  • Le Song

Neural networks are a powerful class of nonlinear functions that can be trained end-to-end on various applications. While the over-parametrization nature in many neural networks renders the ability to fit complex functions and the strong representation power to handle challenging tasks, it also leads to highly correlated neurons that can hurt the generalization ability and incur unnecessary computation cost. As a result, how to regularize the network to avoid undesired representation redundancy becomes an important issue. To this end, we draw inspiration from a well-known problem in physics -- Thomson problem, where one seeks to find a state that distributes N electrons on a unit sphere as evenly as possible with minimum potential energy. In light of this intuition, we reduce the redundancy regularization problem to generic energy minimization, and propose a minimum hyperspherical energy (MHE) objective as generic regularization for neural networks. We also propose a few novel variants of MHE, and provide some insights from a theoretical point of view. Finally, we apply neural networks with MHE regularization to several challenging tasks. Extensive experiments demonstrate the effectiveness of our intuition, by showing the superior performance with MHE regularization.

EAAI Journal 2018 Journal Article

Sparse Self-Represented Network Map: A fast representative-based clustering method for large dataset and data stream

  • Zhen Liu
  • Qiuhua Zheng
  • Zhongping Ji
  • Weihua Zhao

The demand of fast clustering increases rapidly as we keep collecting tremendously large amount of data in the last decade. In this paper, we propose a nonparametric and representative-based Sparse Self-Represented Network Map for fast clustering on large dataset. Each node in the network generates a heat map for the dataset by receiving stimulations from data within its Accepting Field. We developed a weight adjusting method to learn and summarize the clustering pattern of the data. Such learned map is used for computing clustering results, by breaking weak links and finding connected components Rather than employing an iterative process to find local minima, our network passes the dataset only once and is able to capture the global pattern of the dataset as well as detecting natural number of clusters. As a nonparametric method, we propose Sparse Dynamic Instantiation to avoid the curse of dimensionality, namely a node or a link is instantiated only when stimulated by input data. As a result, the overall complexity is linear to the data dimension. Our algorithm is tested on synthetic and real datasets and compare with popular clustering algorithms (K-means + +, Expectation–Maximization, Mean-Shift and StreamKM + + ) as well as state-of-art clustering algorithm (Affinity Propagation and Density Peak). We also applied our clustering algorithm to mobile location clustering, building a Visual Dictionary for image recognition, and clustering data streams. Our experiments indicate that our algorithm can be a better alternative for all compared popular clustering algorithms especially when efficiency is the primary consideration, namely we drastically improve time and space complexity but retain equal level of accuracy.

ICRA Conference 2017 Conference Paper

Motion planning with graph-based trajectories and Gaussian process inference

  • Eric Huang
  • Mustafa Mukadam
  • Zhen Liu
  • Byron Boots

Motion planning as trajectory optimization requires generating trajectories that minimize a desired objective function or performance metric. Finding a globally optimal solution is often intractable in practice: despite the existence of fast motion planning algorithms, most are prone to local minima, which may require re-solving the problem multiple times with different initializations. In this work we provide a novel motion planning algorithm, GPMP-GRAPH, that considers a graph-based initialization that simultaneously explores multiple homotopy classes, helping to contend with the local minima problem. Drawing on previous work to represent continuous-time trajectories as samples from a Gaussian process (GP) and formulating the motion planning problem as inference on a factor graph, we construct a graph of interconnected states such that each path through the graph is a valid trajectory and efficient inference can be performed on the collective factor graph. We perform a variety of benchmarks and show that our approach allows the evaluation of an exponential number of trajectories within a fraction of the computational time required to evaluate them one at a time, yielding a more thorough exploration of the solution space and a higher success rate.

TCS Journal 2000 Journal Article

Dynamic scheduling of parallel computations

  • Zhen Liu

Structures of parallel programs are usually represented by task graphs in the scheduling literature. Such graphs are sometimes obtained at compile time. In many other cases, however, they can be determined only at run time. In this paper, we consider the scheduling of parallel computations whose task graphs are generated at run time. We analyze the case where the task graph is a random out-tree. When the number of offspring of a task has a geometric distribution whose parameter is decreasing and convex in the level, then the breadth-first policy stochastically minimizes the makespan. If, however, this parameter is increasing and concave, then the depth-first policy stochastically minimizes the makespan.

TCS Journal 1996 Journal Article

Scheduling UET-UCT series-parallel graphs on two processors

  • Lucian Finta
  • Zhen Liu
  • Ioannis Mills
  • Evripidis Bampis

The scheduling of task graphs on two identical processors is considered. It is assumed that tasks have unit-execution-time, and arcs are associated with unit-communication-time delays. The problem is to assign the tasks to the two processors and schedule their execution in order to minimize the makespan. A quadratic algorithm is proposed to compute an optimal schedule for a class of series-parallel graphs, called SP1 graphs, which includes in particular in-forests and out-forests.

v2026.09.13