Arrow Research search

Author name cluster

Tong Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

25 papers
2 author rows

Possible papers

25

EAAI Journal 2025 Journal Article

A lightweight segmentation model based on dilated multi-scale residual attention U-Net for brain tumor segmentation

  • Lihong Zhang
  • Yuzhuo Li
  • Yingbo Liang
  • Tong Liu
  • Wenwu Zhang
  • Junding Sun

To address the limited computational capacity of current clinical medical devices, which hampers the effective use of complex high-performance segmentation algorithms, this paper proposes a new lightweight brain tumor segmentation algorithm: Lightweight Dilated Multi-Scale Residual Attention U-Net (LDMRA U-Net). This model is based on the previously proposed Dilated Multi-Scale Residual Attention U-Net (DMRA U-Net). The overall network architecture incorporates the proposed Lightweight Channel Aggregation (LCA) mechanism to reduce the number of channels, along with a multi-level decoder aggregation strategy to minimize key information loss. Additionally, to enhance more direct information transfer between different levels of the encoder-decoder, the proposed Cross Enhanced Attention (CEA) module is incorporated into the skip connection part, improving segmentation performance. We evaluated our proposed method on the Brain Tumor Segmentation Challenge (BraTS) 2019 dataset, and the experimental results show that the number of parameters in LDMRA U-Net is only about 7. 41 % of that in DMRA U-Net, with Dice Similarity Coefficients (DSC) of 0. 8882, 0. 8685, and 0. 8622 in the Whole Tumor (WT), Tumor Core (TC), and Enhancing Tumor (ET) regions, respectively. LDMRA U-Net demonstrates significant performance improvements compared to existing lightweight methods. Our method helps to improve brain tumor segmentation accuracy with very small parameter sizes, and has the potential to be applied in a clinical setting.

EAAI Journal 2025 Journal Article

An end-to-end autonomous learning approach for waterway images analysis of inland river

  • Zechen Li
  • Bingjie Chen
  • Tong Liu
  • Zhonglin Zuo
  • Shan Liang

With the ongoing trend of implementing intelligent and remote traffic control on inland waterways, relying solely on Automatic Identification System as a data source is insufficient to meet diverse demands. Surveillance Video (SV) has attracted attention because of its cheap and real-time information compensation characteristics. However, existing SV process models still face many challenges, including the lack of exclusive data sets, the traditional detection/classification single-function model scarcity providing the whole surface waterway elements, and the model’s limited adaptability to diverse engineering scenarios. To tackle these challenges, this paper proposes a comprehensive solution for engineering applications, from data set establishment to multi-functional model design as well as model self-updating. Specifically, we propose a data augmentation strategy based on Auto-Augment to reduce the data annotation cost. Furthermore, we improve a multi-function model based on Transformer architecture by simultaneously achieving target detection and panoramic segmentation to obtain complete surface waterway elements. Moreover, inspired by knowledge distillation and generative adversarial ideas, we propose an automatic SV data labeling method and an asynchronous end-to-end autonomous learning strategy, ensuring that the model has sustainable learning ability when transferring to new application scenarios. In a regular test, our multi-functional model’s panoramic segmentation and target detection capability can achieve 0. 7675 Intersection over Union and 0. 9412 mean Average Precision (mAP), and the target detection task can achieve 0. 6530 mAP in the deplorable vision case. This result reaches the engineering level that can be deployed in practice.

ICRA Conference 2025 Conference Paper

AVD2: Accident Video Diffusion for Accident Video Description

  • Cheng Li
  • Keyuan Zhou
  • Tong Liu
  • Yu Wang
  • Mingqiao Zhuang
  • Huan-ang Gao
  • Bu Jin
  • Hao Zhao 0002

Traffic accidents present complex challenges for autonomous driving, often featuring unpredictable scenarios that hinder accurate system interpretation and responses. Nonetheless, prevailing methodologies fall short in elucidating the causes of accidents and proposing preventive measures due to the paucity of training data specific to accident scenarios. In this work, we introduce AVD2 (Accident Video Diffusion for Accident Video Description), a novel framework that enhances accident scene understanding by generating accident videos that aligned with detailed natural language descriptions and reasoning, resulting in the contributed EMM-AU (Enhanced Multi-Modal Accident Video Understanding) dataset. Empirical results reveal that the integration of the EMM-AU dataset establishes state-of-the-art performance across both automated metrics and human evaluations, markedly advancing the domains of accident analysis and prevention. Project resources are available at https://an-answer-tree.github.io

ICML Conference 2025 Conference Paper

LipsNet++: Unifying Filter and Controller into a Policy Network

  • Xujie Song
  • Liangfa Chen
  • Tong Liu
  • Wenxuan Wang 0004
  • Yinuo Wang
  • Shentao Qin
  • Yinsong Ma
  • Jingliang Duan

Deep reinforcement learning (RL) is effective for decision-making and control tasks like autonomous driving and embodied AI. However, RL policies often suffer from the action fluctuation problem in real-world applications, resulting in severe actuator wear, safety risk, and performance degradation. This paper identifies the two fundamental causes of action fluctuation: observation noise and policy non-smoothness. We propose LipsNet++, a novel policy network with Fourier filter layer and Lipschitz controller layer to separately address both causes. The filter layer incorporates a trainable filter matrix that automatically extracts important frequencies while suppressing noise frequencies in the observations. The controller layer introduces a Jacobian regularization technique to achieve a low Lipschitz constant, ensuring smooth fitting of a policy function. These two layers function analogously to the filter and controller in classical control theory, suggesting that filtering and control capabilities can be seamlessly integrated into a single policy network. Both simulated and real-world experiments demonstrate that LipsNet++ achieves the state-of-the-art noise robustness and action smoothness. The code and videos are publicly available at https: //xjsong99. github. io/LipsNet_v2.

TIST Journal 2025 Journal Article

Modeling the Chaotic Semantic States of Generative Artificial Intelligence (AI): A Quantum Mechanics Analogy Approach

  • Tong Liu
  • Timothy R. McIntosh
  • Teo Susnjak
  • Paul Watters
  • Malka N. Halgamuge

Generative AI models have revolutionized intelligent systems by enabling machines to produce human-like content across diverse domains. However, their outputs often exhibit unpredictability due to complex and opaque internal semantic states, posing challenges for reliability in real-world applications. In this article, we introduce the AI Uncertainty Principle, a novel theoretical framework inspired by quantum mechanics, to model and quantify the inherent unpredictability in generative AI outputs. By drawing parallels with the uncertainty principle and superposition, we formalize the tradeoff between the precision of internal semantic states and output variability. Through comprehensive experiments involving state-of-the-art models and a variety of prompt designs, we analyze how factors such as specificity, complexity, tone, and style influence model behavior. Our results demonstrate that carefully engineered prompts can significantly enhance output predictability and consistency, while excessive complexity or irrelevant information can increase uncertainty. We also show that ensemble techniques, such as Sigma-weighted aggregation across models and prompt variations, effectively improve reliability. Our findings have profound implications for the development of intelligent systems, emphasizing the critical role of prompt engineering and theoretical modeling in creating AI technologies that perceive, reason, and act predictably in the real world.

EAAI Journal 2025 Journal Article

Multimodality based deep learning method for cancer-related T-cell receptor sequence prediction

  • Junjiang Liu
  • Shusen Zhou
  • Mujun Zang
  • Chanjuan Liu
  • Tong Liu
  • Qingjun Wang

T-cell receptor sequences (TCR-seq) are closely related to cancers, and in particular, cancer-related TCR-seq are crucial in cancer diagnosis and treatment. Current prediction methods for cancer-related TCR-seq often focus solely on the sequence structure, neglecting its spatial structure. Therefore, we propose a multimodal deep learning method based on parallel and residual structures (MDPR) for the detection of cancer-related TCR-seq. MDPR can effectively integrate the spatial and sequence structure of TCR-seq for accurately identifying cancer-related sequences. First, we introduce a TCR-seq encoding method based on atomic three-dimensional spatial coordinates, allowing for more effective extraction of the spatial structural features of TCR-seq. Second, we use high-dimensional word vectors instead of the amino acid feature vectors traditionally used by other researchers. Third, we pretrain the spatial feature extraction module and then conduct joint training with the sequence feature extraction module. This approach allows the model to better consider the relationship between the two modalities, thereby improving prediction accuracy. Finally, MDPR achieved an area under the curve (AUC) of 0. 971 after ten rounds of three-fold cross-validation on the dataset. The AUC of MDPR is 5% higher than that of the previous best method. In short, we propose an artificial intelligence method called MDPR, and apply it to the biomedical field. MDPR can be obtained from https: //github. com/biomg/MDPR.

ICLR Conference 2025 Conference Paper

ODE-based Smoothing Neural Network for Reinforcement Learning Tasks

  • Yinuo Wang
  • Wenxuan Wang 0004
  • Xujie Song
  • Tong Liu
  • Yuming Yin
  • Liangfa Chen
  • Likun Wang
  • Jingliang Duan

The smoothness of control actions is a significant challenge faced by deep reinforcement learning (RL) techniques in solving optimal control problems. Existing RL-trained policies tend to produce non-smooth actions due to high-frequency input noise and unconstrained Lipschitz constants in neural networks. This article presents a Smooth ODE (SmODE) network capable of simultaneously addressing both causes of unsmooth control actions, thereby enhancing policy performance and robustness under noise condition. We first design a smooth ODE neuron with first-order low-pass filtering expression, which can dynamically filter out high frequency noises of hidden state by a learnable state-based system time constant. Additionally, we construct a state-based mapping function, $g$, and theoretically demonstrate its capacity to control the ODE neuron's Lipschitz constant. Then, based on the above neuronal structure design, we further advanced the SmODE network serving as RL policy approximators. This network is compatible with most existing RL algorithms, offering improved adaptability compared to prior approaches. Various experiments show that our SmODE network demonstrates superior anti-interference capabilities and smoother action outputs than the multi-layer perception and smooth network architectures like LipsNet.

IJCAI Conference 2025 Conference Paper

PerfSeer: An Efficient and Accurate Deep Learning Models Performance Predictor

  • Xinlong Zhao
  • Jiande Sun
  • Jia Zhang
  • Tong Liu
  • Ke Liu

Predicting the performance of deep learning (DL) models, such as execution time and resource utilization, is crucial for Neural Architecture Search (NAS), DL cluster schedulers, and other technologies that advance deep learning. The representation of a model is the foundation for its performance prediction. However, existing methods cannot comprehensively represent diverse model configurations, resulting in unsatisfactory accuracy. To address this, we represent a model as a graph that includes the topology, along with node, edge, and global features, all of which are crucial for effectively capturing the performance of the model. Based on this representation, we propose PerfSeer, a novel predictor that uses a Graph Neural Network (GNN)-based performance prediction model, SeerNet. SeerNet fully leverages the topology and various features, while incorporating optimizations such as Synergistic Max-Mean aggregation (SynMM) and Global-Node Perspective Boost (GNPB) to more effectively capture the critical performance information, enabling it to predict the performance of models accurately. Furthermore, SeerNet can be extended to SeerNet-Multi by using Project Conflicting Gradients (PCGrad), enabling efficient simultaneous prediction of multiple performance metrics without significantly affecting accuracy. We constructed a dataset containing performance metrics for 53k+ model configurations, including execution time, memory usage, and Streaming Multiprocessor (SM) utilization during both training and inference. The evaluation results show that PerfSeer outperforms nn-Meter, Brp-NAS, and DIPPM.

EAAI Journal 2025 Journal Article

Resilient kernel-based unsupervised multi-view feature selection via compact binary hashing

  • Rongyao Hu
  • Mengmeng Zhan
  • Jiangzhang Gan
  • Li Li
  • Fei Ye
  • Tong Liu

Multi-view feature selection across diverse views identifying a compact subset of the most informative feature across various data views without relying on labeled information. While most of the solutions are limited to linear multi-view data or utilize weakly-supervised single-label learning to assist in feature selection, leading to the loss of valuable semantic information, especially when dealing with complex real-world multi-view datasets. To overcome these limitations, we introduce a novel Resilient Kernel-based Unsupervised Multi-view Feature Selection via compact Binary Hashing (RKUMBH), which aims to search a robust and consistent graph representation across views, leveraging binary hashing codes to guide feature selection. Specifically, we first standardize the dimensionality of multi-view data by using non-linear kernel mapping. Then, we explore consistent graph structures across different views by fusing individual similarity graph of each view under a self-representation guidance. Moreover, the low-rank constraints are used to preserve the primary structures and patterns embedding within the data, and an unsupervised hashing feature selection framework is conducted to generate reliable hashing codes across views. Additionally, we design a customized iterative optimization method to solve the unified model. Extensive experiments on six public multi-view datasets demonstrate that our proposed method obtains state-of-the-art results compared to existing works for both clustering and feature selection tasks.

EAAI Journal 2024 Journal Article

An efficient astronomical seeing forecasting method by random convolutional Kernel transformation

  • Weijian Ni
  • Chengqin Zhang
  • Tong Liu
  • Qingtian Zeng
  • Lingzhe Xu
  • Huaiqing Wang

Astronomical seeing is the primary indicator of the observation quality of large-scale optical telescopes. Forecasting astronomical seeing is an essential and challenging task in the operation of large optical telescopes, as it can help observers more efficiently arrange observation tasks as well as providing hints for mechanism optimization of optical telescopes. In this paper, we propose an efficient data-driven astronomical seeing forecasting method by leveraging the environment monitoring data from multiple sensors deployed with optical telescopes. The basic idea is to use a large number of random yet fixed convolution kernels to extract features from the original data, and then train a simple tree-based predictor based on these features. In particular, the raw monitoring data is transformed by differential operations, enabling a more explicit representation of spatial and temporal features; then an elaborate convolution mechanism, including the convolution kernel and the pooling operator, is designed to extract a variety set of features of the input time series, which serve as the foundation for constructing the time series forecasting model. In contrast to existing deep learning models, there is no need to calculate and back-propagate the gradients during the training procedure, making the proposed method much more efficient. We collect real-world monitoring data from Large Sky Area Multi-Object Fiber Spectroscopy Telescope (LAMOST), one of the world’s largest optical reflecting telescopes, and conduct extensive empirical evaluations. The experimental results demonstrate that the proposed method can achieve state-of-the-art prediction accuracy, i. e. , 0. 1125 in MAE, 0. 0252 in MSE, and 0. 0322 in MAPE, outperforming most deep-learning-based forecasting models. Moreover, it surpasses both machine learning and deep learning approaches in terms of training efficiency with a significant margin.

NeurIPS Conference 2024 Conference Paper

Diffusion Actor-Critic with Entropy Regulator

  • Yinuo Wang
  • Likun Wang
  • Yuxuan Jiang
  • Wenjun Zou
  • Tong Liu
  • Xujie Song
  • Wenxuan Wang
  • Liming Xiao

Reinforcement learning (RL) has proven highly effective in addressing complex decision-making and control tasks. However, in most traditional RL algorithms, the policy is typically parameterized as a diagonal Gaussian distribution with learned mean and variance, which constrains their capability to acquire complex policies. In response to this problem, we propose an online RL algorithm termed diffusion actor-critic with entropy regulator (DACER). This algorithm conceptualizes the reverse process of the diffusion model as a novel policy function and leverages the capability of the diffusion model to fit multimodal distributions, thereby enhancing the representational capacity of the policy. Since the distribution of the diffusion policy lacks an analytical expression, its entropy cannot be determined analytically. To mitigate this, we propose a method to estimate the entropy of the diffusion policy utilizing Gaussian mixture model. Building on the estimated entropy, we can learn a parameter $\alpha$ that modulates the degree of exploration and exploitation. Parameter $\alpha$ will be employed to adaptively regulate the variance of the added noise, which is applied to the action output by the diffusion model. Experimental trials on MuJoCo benchmarks and a multimodal task demonstrate that the DACER algorithm achieves state-of-the-art (SOTA) performance in most MuJoCo control tasks while exhibiting a stronger representational capacity of the diffusion policy.

EAAI Journal 2024 Journal Article

Graph attention network with convolutional layer for predicting gene regulations from single-cell ribonucleic acid sequence data

  • Junjiang Liu
  • Shusen Zhou
  • Jing Ma
  • Mujun Zang
  • Chanjuan Liu
  • Tong Liu
  • Qingjun Wang

Reconstructing gene regulatory networks (GRNs) is an important task to reveal the regulatory relationship between genes and understand the mechanism of intracellular gene expression regulation. With the development of single-cell ribonucleic acid sequencing (scRNA-seq) technology, researchers begin to attempt to infer GRN within cells. In this paper, we propose graph attention network with convolutional layer (GATCL) to infer the latent interactions between transcription factors (TFs) and target genes in GRN. Firstly, GATCL uses graph attention network (GAT), which can effectively extract information about genes and TFs. Secondly, we combine multi-head attention layer with one-head attention layer and propose a new method of using convolution instead of weight matrix. Thirdly, we use the exponential linear unit (ELU) activation function to replace the leaky rectified linear unit (LReLU) commonly used in GAT, which further improves the accuracy of GATCL. The AUROC of our method on seven scRNA-seq datasets with four types of ground-truth networks reached an average of 0. 827, which is higher than other state-of-the-art models. GATCL applies a supervised deep learning algorithm to solve the problems existing in the inference of GRN from scRNA-seq in the engineering field, and the effectiveness of this method is verified by a large number of experiments.

ICRA Conference 2024 Conference Paper

Robust Collaborative Perception against Temporal Information Disturbance

  • Xunjie He
  • Yiming Li
  • Te Cui
  • Meiling Wang
  • Tong Liu
  • Yufeng Yue

Collaborative perception facilitates a more comprehensive representation of the environment by leveraging complementary information shared among various agents and sensors. However, practical applications often encounter information disturbance which includes perception packet loss and time delays, and a comprehensive framework that can simultaneously address such issues is absent. In addition, the feature extraction process prior to fusion is not sufficient, as it lacks exploration of the local semantics and context dependencies of individual features. To enhance both accuracy and robustness, this paper introduces a novel framework named Robust Collaborative Perception against Temporal Information Disturbance, which predicts perception information when disturbance occurs. Specifically, the Historical Frame Prediction (HFP) module is introduced to make compensation for information loss with temporal association excavation of historical features. Based on the predicted features generated by the HFP module, the Pyramid Attention Integration (PAI) module is introduced to augment local semantics and incorporate global long-range dependencies through multi-scale window attention. Compared with existing methods on the publicly available dataset OPV2V, our approach exhibits superior performance and expanded robustness in the 3D object detection task. The code will be publicly available at https://github.com/hexunjie/Ro-temd.

NeurIPS Conference 2023 Conference Paper

Active Negative Loss Functions for Learning with Noisy Labels

  • Xichen Ye
  • Xiaoqiang Li
  • Songmin Dai
  • Tong Liu
  • Yan Sun
  • Weiqin Tong

Robust loss functions are essential for training deep neural networks in the presence of noisy labels. Some robust loss functions use Mean Absolute Error (MAE) as its necessary component. For example, the recently proposed Active Passive Loss (APL) uses MAE as its passive loss function. However, MAE treats every sample equally, slows down the convergence and can make training difficult. In this work, we propose a new class of theoretically robust passive loss functions different from MAE, namely Normalized Negative Loss Functions (NNLFs), which focus more on memorized clean samples. By replacing the MAE in APL with our proposed NNLFs, we improve APL and propose a new framework called Active Negative Loss (ANL). Experimental results on benchmark and real-world datasets demonstrate that the new set of loss functions created by our ANL framework can outperform state-of-the-art methods. The code is available athttps: //github. com/Virusdoll/Active-Negative-Loss.

AIIM Journal 2023 Journal Article

Automated atrial fibrillation and ventricular fibrillation recognition using a multi-angle dual-channel fusion network

  • Weiyi Yang
  • Di Wang
  • Wei Fan
  • Gong Zhang
  • Chunying Li
  • Tong Liu

Atrial fibrillation (AFIB) and ventricular fibrillation (VFIB) are two common cardiovascular diseases that cause numerous deaths worldwide. Medical staff usually adopt long-term ECGs as a tool to diagnose AFIB and VFIB. However, since ECG changes are occasionally subtle and similar, visual observation of ECG changes is challenging. To address this issue, we proposed a multi-angle dual-channel fusion network (MDF-Net) to automatically recognize AFIB and VFIB heartbeats in this work. MDF-Net can be seen as the fusion of a task-related component analysis (TRCA)-principal component analysis (PCA) network (TRPC-Net), a canonical correlation analysis (CCA)-PCA network (CPC-Net), and the linear support vector machine-weighted softmax with average (LS-WSA) method. TRPC-Net and CPC-Net are employed to extract deep task-related and correlation features, respectively, from two-lead ECGs, by which multi-angle feature-level information fusion is realized. Since the convolution kernels of the above methods can be directly extracted through TRCA, CCA and PCA technologies, their training time is faster than that of convolutional neural networks. Finally, LS-WSA is employed to fuse the above features at the decision level, by which the classification results are obtained. In distinguishing AFIB and VFIB heartbeats, the proposed method achieved accuracies of 99. 39 % and 97. 17 % in intra- and inter-patient experiments, respectively. In addition, this method performed well on noisy data and extremely imbalanced data, in which abnormal heatbeats are much less than normal heartbeats. Our proposed method has the potential to be used as a diagnostic tool in the clinic.

IJCAI Conference 2023 Conference Paper

Cross-Domain Facial Expression Recognition via Disentangling Identity Representation

  • Tong Liu
  • Jing Li
  • Jia Wu
  • Lefei Zhang
  • Shanshan Zhao
  • Jun Chang
  • Jun Wan

Most existing cross-domain facial expression recognition (FER) works require target domain data to assist the model in analyzing distribution shifts to overcome negative effects. However, it is often hard to obtain expression images of the target domain in practical applications. Moreover, existing methods suffer from the interference of identity information, thus limiting the discriminative ability of the expression features. We exploit the idea of domain generalization (DG) and propose a representation disentanglement model to address the above problems. Specifically, we learn three independent potential subspaces corresponding to the domain, expression, and identity information from facial images. Meanwhile, the extracted expression and identity features are recovered as Fourier phase information reconstructed images, thereby ensuring that the high-level semantics of images remain unchanged after disentangling the domain information. Our proposed method can disentangle expression features from expression-irrelevant ones (i. e. , identity and domain features). Therefore, the learned expression features exhibit sufficient domain invariance and discriminative ability. We conduct experiments with different settings on multiple benchmark datasets, and the results show that our method achieves superior performance compared with state-of-the-art methods.

AAAI Conference 2023 Conference Paper

GradPU: Positive-Unlabeled Learning via Gradient Penalty and Positive Upweighting

  • Songmin Dai
  • Xiaoqiang Li
  • Yue Zhou
  • Xichen Ye
  • Tong Liu

Positive-unlabeled learning is an essential problem in many real-world applications with only labeled positive and unlabeled data, especially when the negative samples are difficult to identify. Most existing positive-unlabeled learning methods will inevitably overfit the positive class to some extent due to the existence of unidentified positive samples. This paper first analyzes the overfitting problem and proposes to bound the generalization errors via Wasserstein distances. Based on that, we develop a simple yet effective positive-unlabeled learning method, GradPU, which consists of two key ingredients: A gradient-based regularizer that penalizes the gradient norms in the interpolated data region, which improves the generalization of positive class; An unnormalized upweighting mechanism that assigns larger weights to those positive samples that are hard, not-well-fitted and less frequently labeled. It enforces the training error of each positive sample to be small and increases the robustness to the labeling bias. We evaluate our proposed GradPU on three datasets: MNIST, FashionMNIST, and CIFAR10. The results demonstrate that GradPU achieves state-of-the-art performance on both unbiased and biased positive labeling scenarios.

IJCAI Conference 2023 Conference Paper

Hierarchical Semantic Contrast for Weakly Supervised Semantic Segmentation

  • Yuanchen Wu
  • Xiaoqiang Li
  • Songmin Dai
  • Jide Li
  • Tong Liu
  • Shaorong Xie

Weakly supervised semantic segmentation (WSSS) with image-level annotations has achieved great processes through class activation map (CAM). Since vanilla CAMs are hardly served as guidance to bridge the gap between full and weak supervision, recent studies explore semantic representations to make CAM fit for WSSS and demonstrate encouraging results. However, they generally exploit single-level semantics, which may hamper the model to learn a comprehensive semantic structure. Motivated by the prior that each image has multiple levels of semantics, we propose hierarchical semantic contrast (HSC) to ameliorate the above problem. It conducts semantic contrast from coarse-grained to fine-grained perspective, including ROI level, class level, and pixel level, making the model learn a better object pattern understanding. To further improve CAM quality, building upon HSC, we explore consistency regularization of cross supervision and develop momentum prototype learning to utilize abundant semantics across different images. Extensive studies manifest that our plug-and-play learning paradigm, HSC, can significantly boost CAM quality on both non-saliency-guided and saliency-guided baselines, and establish new state-of-the-art WSSS performance on PASCAL VOC 2012 dataset. Code is available at https: //github. com/Wu0409/HSC_WSSS.

JAIR Journal 2023 Journal Article

On the Evaluation of (Meta-)solver Approaches

  • Roberto Amadini
  • Maurizio Gabbrielli
  • Tong Liu
  • Jacopo Mauro

Meta-solver approaches exploit many individual solvers to potentially build a better solver. To assess the performance of meta-solvers, one can adopt the metrics typically used for individual solvers (e.g., runtime or solution quality) or employ more specific evaluation metrics (e.g., by measuring how close the meta-solver gets to its virtual best performance). In this paper, based on some recently published works, we provide an overview of different performance metrics for evaluating (meta-)solvers by exposing their strengths and weaknesses.

AAAI Conference 2023 Conference Paper

Robust Temporal Smoothness in Multi-Task Learning

  • Menghui Zhou
  • Yu Zhang
  • Yun Yang
  • Tong Liu
  • Po Yang

Multi-task learning models based on temporal smoothness assumption, in which each time point of a sequence of time points concerns a task of prediction, assume the adjacent tasks are similar to each other. However, the effect of outliers is not taken into account. In this paper, we show that even only one outlier task will destroy the performance of the entire model. To solve this problem, we propose two Robust Temporal Smoothness (RoTS) frameworks. Compared with the existing models based on temporal relation, our methods not only chase the temporal smoothness information but identify outlier tasks, however, without increasing the computational complexity. Detailed theoretical analyses are presented to evaluate the performance of our methods. Experimental results on synthetic and real-life datasets demonstrate the effectiveness of our frameworks. We also discuss several potential specific applications and extensions of our RoTS frameworks.

IJCAI Conference 2022 Conference Paper

sunny-as2: Enhancing SUNNY for Algorithm Selection (Extended Abstract)

  • Tong Liu
  • Roberto Amadini
  • Maurizio Gabbrielli
  • Jacopo Mauro

SUNNY is a k-nearest neighbors based Algorithm Selection (AS) approach that schedules and runs a number of solvers for a given unforeseen problem. In this work we present sunny-as2, an enhancement of SUNNY for generic AS scenarios that advances the original approach with wrapper-based feature selection, neighborhood-size configuration and a greedy approach to speed-up the training phase. Empirical evidence shows that sunny-as2 is competitive w. r. t. state-of-the-art AS approaches.

AAAI Conference 2021 Conference Paper

KGDet: Keypoint-Guided Fashion Detection

  • Shenhan Qian
  • Dongze Lian
  • Binqiang Zhao
  • Tong Liu
  • Bohui Zhu
  • Hai Li
  • Shenghua Gao

Locating and classifying clothes, usually referred to as clothing detection, is a fundamental task in fashion analysis. Motivated by the strong structural characteristics of clothes, we pursue a detection method enhanced by clothing keypoints, which is a compact and effective representation of structures. To incorporate the keypoint cues into clothing detection, we design a simple yet effective Keypoint-Guided clothing Detector, named KGDet. Such a detector can fully utilize information provided by keypoints with the following two aspects: i) integrating local features around keypoints to benefit both classification and regression; ii) generating accurate bounding boxes from keypoints. To effectively incorporate local features, two alternative modules are proposed. One is a multi-column keypoint-encoding-based feature aggregation module; the other is a keypoint-selection-based feature aggregation module. With either of the above modules as a bridge, a cascade strategy is introduced to refine detection performance progressively. Thanks to the keypoints, our KGDet obtains superior performance on the DeepFashion2 dataset and the FLD dataset with high efficiency.

JAIR Journal 2021 Journal Article

sunny-as2: Enhancing SUNNY for Algorithm Selection

  • Tong Liu
  • Roberto Amadini
  • Maurizio Gabbrielli
  • Jacopo Mauro

SUNNY is an Algorithm Selection (AS) technique originally tailored for Constraint Programming (CP). SUNNY is based on the k-nearest neighbors algorithm and enables one to schedule, from a portfolio of solvers, a subset of solvers to be run on a given CP problem. This approach has proved to be effective for CP problems. In 2015, the ASlib benchmarks were released for comparing AS systems coming from disparate fields (e.g., ASP, QBF, and SAT) and SUNNY was extended to deal with generic AS problems. This led to the development of sunny-as, a prototypical algorithm selector based on SUNNY for ASlib scenarios. A major improvement of sunny-as, called sunny-as2, was then submitted to the Open Algorithm Selection Challenge (OASC) in 2017, where it turned out to be the best approach for the runtime minimization of decision problems. In this work we present the technical advancements of sunny-as2, by detailing through several empirical evaluations and by providing new insights. Its current version, built on the top of the preliminary version submitted to OASC, is able to outperform sunny-as and other state-of-the-art AS methods, including those who did not attend the challenge.

JBHI Journal 2020 Journal Article

MR-Forest: A Deep Decision Framework for False Positive Reduction in Pulmonary Nodule Detection

  • Hongbo Zhu
  • Hai Zhao
  • Chunhe Song
  • Zijian Bian
  • Yuanguo Bi
  • Tong Liu
  • Xuan He
  • Dongxiang Yang

With the development of deep learning methods such as convolutional neural network (CNN), the accuracy of automated pulmonary nodule detection has been greatly improved. However, the high computational and storage costs of the large-scale network have been a potential concern for the future widespread clinical application. In this paper, an alternative Multi-ringed (MR)-Forest framework, against the resource-consuming neural networks (NN)-based architectures, has been proposed for false positive reduction in pulmonary nodule detection, which consists of three steps. First, a novel multi-ringed scanning method is used to extract the order ring facets (ORFs) from the surface voxels of the volumetric nodule models; Second, Mesh-LBP and mapping deformation are employed to estimate the texture and shape features. By sliding and resampling the multi-ringed ORFs, feature volumes with different lengths are generated. Finally, the outputs of multi-level are cascaded to predict the candidate class. On 1034 scans merging the dataset from the Affiliated Hospital of Liaoning University of Traditional Chinese Medicine (AH-LUTCM) and the LUNA16 Challenge dataset, our framework performs enough competitiveness than state-of-the-art in false positive reduction task (CPM score of 0. 865). Experimental results demonstrate that MR-Forest is a successful solution to satisfy both resource-consuming and effectiveness for automated pulmonary nodule detection. The proposed MR-forest is a general architecture for 3D target detection, it can be easily extended in many other medical imaging analysis tasks, where the growth trend of the targeting object is approximated as a spheroidal expansion.

EAAI Journal 2015 Journal Article

Finger-vein pattern restoration with Direction-Variance-Boundary Constraint Search

  • Tong Liu
  • Jianbin Xie
  • Wei Yan
  • Peiqin Li
  • Huanzhang Lu

Finger-vein verification is an emerging biometrics technology. Its first task is extracting finger-vein patterns. Although existing algorithms can extract most finger-vein patterns robustly, some branch of these patterns always breaks, which leads to adverse effects for features extraction and matching. In this paper, a Direction-Variance-Boundary Constraint Search (DVBCS) model is presented to restore the broken finger-vein patterns. At the beginning, endpoints of broken finger-vein branches are located. Then, a direction constraint for searching candidate point set is demonstrated. Following the second stage, an optimal target point is selected from the candidate point set according to a minimum within-cluster variance criterion. Eventually, the boundary constraint and variance constraint are introduced as the termination conditions. Experimental results illustrate that, while maintaining low segmentation error, the proposed method can restore above 10% lost target points. Moreover, the equal error rate of finger-vein recognition is reduced from 0. 57% to 0. 29% when using the proposed method to restore finger-vein patterns.

v2026.09.13