Arrow Research search

Author name cluster

Jiaqi Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

21 papers
2 author rows

Possible papers

21

AAAI Conference 2026 Conference Paper

ARCHE: A Novel Task to Evaluate LLMs on Latent Reasoning Chain Extraction

  • Pengze Li
  • Jiaqi Liu
  • Junchi Yu
  • Lihao Liu
  • Mingyu Ding
  • Wanli Ouyang
  • Shixiang Tang
  • Xi Chen

Large language models (LLMs) are increasingly used in scientific domains. While they can produce reasoning-like content via methods such as chain-of-thought prompting, these outputs are typically unstructured and informal, obscuring whether models truly understand the fundamental reasoning paradigms that underpin scientific inference. To address this, we introduce a novel task named Latent Reasoning Chain Extraction (ARCHE), in which models must decompose complex reasoning arguments into combinations of standard reasoning paradigms in the form of a Reasoning Logic Tree (RLT). In an RLT, all reasoning steps are explicitly categorized as one of three variants of Peirce’s fundamental inference modes: deduction, induction, or abduction. To facilitate this task, we release ARCHE Bench, a new benchmark derived from 70 Nature Communications articles, including more than 1,900 references and 38,000 viewpoints. We propose two logic-aware evaluation metrics: Entity Coverage (EC) for content completeness and Reasoning Edge Accuracy (REA) for step-by-step logical validity. Evaluations on 10 leading LLMs on ARCHE Bench reveal that models exhibit a trade-off between REA and EC, and none are yet able to extract a complete and standard reasoning chain. These findings highlight a substantial gap between the abilities of current reasoning models and the rigor required for scientific argumentation.

AAMAS Conference 2026 Conference Paper

Bridging Expertise and Data: Multi-Label Disease Detection via Causal Learning and Decision Fusion

  • Xin Zhang
  • Minhui Zhang
  • Jiaqi Liu
  • Zhiwen Yu
  • Bin Guo

Recent multi-label disease detection methods exploit disease causality and disease–image feature interactions, but causal learning is ofteninaccurateandcomputationallycostly. Meanwhile, human–AI collaboration in diagnosis can outperform either clinicians or models alone. We propose a framework that combines expert-guided causal learning with Bayesian human–AI decision fusion. First, we learn an expert causal matrix via a GCN from expert labels and authoritative medical knowledge, and use it to regularize inter-disease causallearning. Second, weconvertper-labelprobabilitiesintojoint label-set probabilities and fuse them with expert decisions using a Bayesian scheme. Experiments on three medical datasets show that our method outperforms state-of-the-art multi-label disease detection models by up to 13. 18%.

AAMAS Conference 2026 Conference Paper

CentaurMD: Confidence-Aware Human-AI Decision Fusion for Multi-Label Disease Diagnosis via Label-Specific MoE

  • Youcheng Zhang
  • Hui Wang
  • Jiaqi Liu
  • Yao Zhang
  • Zhiwen Yu
  • Bin Guo

Multi-label disease diagnosis is prevalent in clinical applications, such as chest X-rays that may indicate multiple coexisting diseases. Despite advances in AI, current models remain insufficient for reliably addressing such complexity. Human–AI synergy thus emerges as both a necessary and promising approach, motivating our focus on effective decision fusion for multi-label disease diagnosis. There are two challenges. Confidence, a key factor in decision fusion, is often unrecorded in human annotations, making its estimation nontrivial. Moreover, label-specific variations in human and model expertise must be considered to achieve effective fusion. To address these challenges, we propose CentaurMD, a confidence-aware human–AI decision fusion framework based on label-specific Mixture-of-Experts (MoE). We first present a novel multi-label confusion matrix construction method that employs maximum entropy modeling to capture label correlations, enabling more accurate confidence estimation and weight allocation. Then, we develop a label-specific MoE module with dedicated gating networks and thresholds, which dynamically adjust expert weights using information extracted from the confusion matrix via a Transformer encoder. Extensive experiments on three real-world clinical datasets demonstrate that our method reduces Hamming loss by 39. 14% and improves MMR (missed-misdiagnosis reduction) by 17. 38%, achieving substantial diagnostic improvements.

AAAI Conference 2026 Conference Paper

Dynamic Cognitive Planning for Cognitive-Functional Dialogue: A Case Study in Emotional Support Conversation

  • Jiaqi Liu
  • Yankun Yang
  • Jiakang Xu
  • Zhongqiang Du
  • Wenbin Jiang

Cognitive-functional dialogues, such as those for persuasion, consultation, and question-answering, are prevalent throughout human social interaction. The core difference between these dialogues and casual chat lies in their objective: to guide a person's cognitive and psychological state toward a predetermined one. Existing conversational technologies perform poorly in handling such dialogues. The fundamental reason is that the transformation of human cognitive psychology follows specific patterns, yet existing technologies neither account for these patterns nor possess cognitive guidance planning based on them. This deficiency makes it difficult for dialogues to achieve their intended cognitive-functional goals effectively. To address this, we propose a dynamic cognitive planning method (DyCoP). By modeling the long-term evolution of a user's cognitive psychology during the dialogue process, this method dynamically generates dialogue guidance plans that align with the principles of cognitive-psychological evolution. This allows for the generation of appropriate dialogue responses based on prior user psychology and the immediate conversational context, thereby achieving cognitive-functional goals more efficiently and accurately. Simultaneously, we constructed an evaluation framework for cognitive-functional dialogues and constructed a richly annotated emotional support conversation dataset. Comprehensive automatic and human evaluations show that our proposed DyCoP method demonstrates significant advantages over existing baseline models.

EAAI Journal 2026 Journal Article

G-LFFN: A Global-Local Feature Fusion Network Leveraging Transformer-Encoder and Contrastive Learning for Multimodal Sentiment Analysis

  • Cong Liu
  • Yong Wang
  • Jing Yang
  • Xiaohui Tao
  • Jiaqi Liu

Due to the varieties of sentiment expressions, multimodal sentiment analysis for social media requires a comprehensive fusion of image and textual information. However, most of the previous studies have only modeled the inter-modal local or global interactions, ignoring inter-modal global and local co-influences, resulting in insufficient fusion of sentiment information. In addition, the introduction of multiple features may generate more sentiment-irrelevant information, thus leading to a weaker sentiment association of the fusion features. To solve the above issues, we propose a global-local feature fusion network model leveraging transformer-encoder and contrastive learning. Firstly, considering inter-modal global and local co-influences, the model extracts global and local features in the image. Secondly, we propose a cross-modal synchronous fusion transformer-encoder and its simplified version to capture inter-modal global and local consistent features, and combine it with soft self-attention to further enhance inter-modal interaction. On this basis, we utilize multiple contrastive learning to enhance the interactions among multiple fusion features and improve the sentiment associations of multimodal fusion features to assist the final sentiment analysis. Extensive experiments on three public multimodal datasets show that our model can adequately capture inter-modal global-local information interactions and effectively improve sentiment associations, thus demonstrating its validity and superiority.

AAAI Conference 2026 Conference Paper

SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts

  • Jiaqi Liu
  • Ronghao Fu
  • Lang Sun
  • Haoran Liu
  • Xiao Yang
  • Weipeng Zhang
  • Xu Na
  • Zhuoran Duan

The emergence of large vision-language models (VLMs) has significantly enhanced the efficiency and flexibility of geospatial interpretation. However, general-purpose VLMs remain suboptimal for remote sensing (RS) tasks. Existing geospatial VLMs typically adopt a unified modeling strategy and struggle to differentiate between task types and interpretation granularities, limiting their ability to balance local detail perception and global contextual understanding. In this paper, we present SkyMoE, a Mixture-of-Experts (MoE) vision-language model tailored for multimodal, multi-task RS interpretation. SkyMoE employs an adaptive router that generates task- and granularity-aware routing instructions, enabling specialized large language model experts to handle diverse sub-tasks. To further promote expert decoupling and granularity sensitivity, we introduce a context-disentangled augmentation strategy that creates contrastive pairs between local and global features, guiding experts toward level-specific representation learning. We also construct MGRS-Bench, a comprehensive benchmark covering multiple RS interpretation tasks and granularity levels, to evaluate generalization in complex scenarios. Extensive experiments on 21 public datasets demonstrate that SkyMoE achieves state-of-the-art performance across tasks, validating its adaptability, scalability, and superior multi-granularity understanding in remote sensing.

IJCAI Conference 2025 Conference Paper

ActiveHAI: Active Collection Based Human-AI Diagnosis with Limited Expert Predictions

  • Xuehan Zhao
  • Jiaqi Liu
  • Xin Zhang
  • Zhiwen Yu
  • Bin Guo

Recent studies indicate that human-AI collaboration performs better than either alone, particularly in medical diagnosis. Beyond collaboration methods that focus on assigning tasks to humans or AI, like deferral, combining human and AI decisions with their confidence scores is emerging as a promising strategy. Due to high cognitive load, doctors often struggle to provide confidence assessments, necessitating explicit human uncertainty evaluation through a limited number of additional expert predictions. There are two challenges. (1) how to actively collect limited yet representative expert predictions? (2) how to accurately evaluate human uncertainty with limited expert predictions? To address the challenges, we propose ActiveHAI, an active human-AI diagnosis method that reduces expert costs through a median-window sampling strategy that actively selects representative samples near the estimated median; and evaluate expert confidence through an evaluator module that integrates sample features and expert predictions, converting them into probability distributions. Experiments on three real-world datasets show that ActiveHAI surpasses doctor and other human-AI methods by 16. 3% and 3. 6% in accuracy, respectively. Furthermore, ActiveHAI reaches 97. 2% relative accuracy, even with just eight expert predictions per class.

NeurIPS Conference 2025 Conference Paper

BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset

  • Zhiheng Xi
  • Guanyu Li
  • Yutao Fan
  • Honglin Guo
  • Yufang Liu
  • Xiaoran Fan
  • Jiaqi Liu
  • Wangmeng Zuo

In this paper, we introduce BMMR, a large-scale bilingual, multimodal, multi-disciplinary reasoning dataset for the community to develop and evaluate large multimodal models (LMMs). BMMR comprises 100k university-level questions drawn from 300 UNESCO-defined subjects, spanning diverse formats—multiple-choice, fill-in-the-blank, and open-ended QA—and sourced from both print and digital media such as books, exams, and quizzes. All data are curated and filtered via a human-in-the-loop, automated, and scalable framework, and each instance is paired with a high-quality reasoning path. The dataset is organized into two parts: BMMR-Eval that comprises 20k high-quality instances to comprehensively assess LMMs’ knowledge and reasoning across multiple disciplines in both Chinese and English; and BMMR-Train that contains 80k instances to support further research and development, extending the current focus on mathematical reasoning to diverse disciplines and domains. In addition, we propose the process-based multi-discipline BMMR-Verifier for accurate and fine-grained evaluation of LMMs’ reasoning. Extensive experiments reveal that (i) even SOTA models leave substantial headroom on BMMR-Eval; (ii) reasoning models exhibit discipline bias and outperform LMMs only on specific subjects; (iii) open-source models still trail their proprietary counterparts; and (iv) fine-tuning on BMMR-Train narrows this gap. Additionally, we conduct reasoning-chain analyses using BMMR-Verifier and other in-depth studies, uncovering the challenges LMMs currently face in multidisciplinary reasoning. We will release the data and models, and we believe our work can offers valuable insights and contributions to the community.

EAAI Journal 2025 Journal Article

Enhanced feedback analysis of vertical load reliability parameters for airplane landing gear using an improved generative adversarial network and explainable artificial intelligence techniques

  • Weihuang Pan
  • Yunwen Feng
  • Cheng Lu
  • Jiaqi Liu
  • Jingcui Liang

Effective feedback analysis of critical equipment data is essential for improving performance and optimizing design parameters in aviation systems. This study presents a novel framework that integrates an improved generative adversarial network (GAN) with explainable artificial intelligence techniques (XAI) to evaluate the reliability of the vertical load for airplane landing gear. By utilizing limited data from the Quick Access Recorder (QAR), the improved GAN generates extensive synthetic data to expand the dataset and strengthen the analysis. Each parameter's importance and influence on vertical load reliability are then evaluated through the Shapley Additive Explanations (SHAP) method, a key approach in XAI. Validation using landing gear data from a typical civil airplane demonstrates the effectiveness of this method and confirms the viability of explainable artificial intelligence for parametric feedback analysis. The results highlight the impact of each parameter on vertical load reliability, providing valuable insights to support enhanced design and operational efficiency of landing gear.

EAAI Journal 2025 Journal Article

Few-shot augmentation based on variational auto-generative adversarial network with moving losses: Application to the variable stiffness prediction in composites

  • Zhicen Song
  • Yunwen Feng
  • Cheng Lu
  • Jiaqi Liu

The dataset of carbon fiber reinforced plastics (CFRP) materials presents characteristics of high-dimensional, low-rank, and sparse, which pose difficulties in the combination of mechanical modeling. In this paper, a Variational Auto-Generative Adversarial Network (VAGAN) with moving losses is proposed as a data augmentation method, which extends the size of the CFRP dataset covering components, processes, elasticity, and strengths factors, and increases the information conveyed in the surrogate modeling. A compression strength prediction model for CFRP laminates was constructed by combining the component, process, and mechanical tensor with a neural network optimized by the search algorithm. Combined with the data augmentation strategy, not only was the amount of data expanded, but the prediction accuracy was also significantly improved. The allowable value of compression strength is analyzed and calculated by the predicted values, which brings direct benefits in simplifying the test.

AAAI Conference 2025 Conference Paper

Revisiting Multimodal Fusion for 3D Anomaly Detection from an Architectural Perspective

  • Kaifang Long
  • Guoyang Xie
  • Lianbo Ma
  • Jiaqi Liu
  • Zhichao Lu

Existing efforts to boost multimodal fusion of 3D anomaly detection (3D-AD) primarily concentrate on devising more effective multimodal fusion strategies. However, little attention was devoted to analyzing the role of multimodal fusion architecture (topology) design in contributing to 3D-AD. In this paper, we aim to bridge this gap and present a systematic study on the impact of multimodal fusion architecture design on 3D-AD. This work considers the multimodal fusion architecture design at the intra-module fusion level, i.e., independent modality-specific modules, involving early, middle or late multimodal features with specific fusion operations, and also at the inter-module fusion level, i.e., the strategies to fuse those modules. In both cases, we first derive insights through theoretically and experimentally exploring how architectural designs influence 3D-AD. Then, we extend SOTA neural architecture search (NAS) paradigm and propose 3D-ADNAS to simultaneously search across multimodal fusion strategies and modality-specific modules for the first time. Extensive experiments show that 3D-ADNAS obtains consistent improvements in 3D-AD across various model capacities in terms of accuracy, frame rate, and memory usage, and it exhibits great potential in dealing with few-shot 3D-AD tasks.

IJCAI Conference 2025 Conference Paper

Towards Generalizable Neural Simulators: Addressing Distribution Shifts Induced by Environmental and Temporal Variations

  • Jiaqi Liu
  • Jiaxu Cui
  • Shiang Sun
  • Yizhu Zhao
  • Bo Yang

With advancements in deep learning, neural simulators have become increasingly important for improving the efficiency and effectiveness of simulating complex dynamical systems in various scientific and technological fields. This paper presents a novel neural simulator called Context-informed Polymorphic Neural ODE Processes (CoPoNDP), aimed at addressing the challenges of modeling dynamical systems encountering concurrent environmental and temporal distribution shifts, which are common in real-world scenarios. CoPoNDP employs a context-driven neural stochastic process governed by a combination of basic differential equations in a time-sensitive manner to adaptively modulate the evolution of system states. This allows for flexible adaptation to changing temporal dynamics and generalization across different environments. Extensive experiments conducted on dynamical systems from ecology, chemistry, physics, and energy demonstrate that by effectively utilizing contextual information, CoPoNDP outperforms the state-of-the-art models in handling joint distribution shifts. It also shows robustness in sparse and noisy settings, making it a promising approach for modeling dynamical systems in complex real-world applications.

AAAI Conference 2024 Conference Paper

Unsupervised Continual Anomaly Detection with Contrastively-Learned Prompt

  • Jiaqi Liu
  • Kai Wu
  • Qiang Nie
  • Ying Chen
  • Bin-Bin Gao
  • Yong Liu
  • Jinbao Wang
  • Chengjie Wang

Unsupervised Anomaly Detection (UAD) with incremental training is crucial in industrial manufacturing, as unpredictable defects make obtaining sufficient labeled data infeasible. However, continual learning methods primarily rely on supervised annotations, while the application in UAD is limited due to the absence of supervision. Current UAD methods train separate models for different classes sequentially, leading to catastrophic forgetting and a heavy computational burden. To address this issue, we introduce a novel Unsupervised Continual Anomaly Detection framework called UCAD, which equips the UAD with continual learning capability through contrastively-learned prompts. In the proposed UCAD, we design a Continual Prompting Module (CPM) by utilizing a concise key-prompt-knowledge memory bank to guide task-invariant 'anomaly' model predictions using task-specific 'normal' knowledge. Moreover, Structure-based Contrastive Learning (SCL) is designed with the Segment Anything Model (SAM) to improve prompt learning and anomaly segmentation results. Specifically, by treating SAM's masks as structure, we draw features within the same mask closer and push others apart for general feature representations. We conduct comprehensive experiments and set the benchmark on unsupervised continual anomaly detection and segmentation, demonstrating that our method is significantly better than anomaly detection methods, even with rehearsal training. The code will be available at https://github.com/shirowalker/UCAD.

NeurIPS Conference 2023 Conference Paper

Certified Minimax Unlearning with Generalization Rates and Deletion Capacity

  • Jiaqi Liu
  • Jian Lou
  • Zhan Qin
  • Kui Ren

We study the problem of $(\epsilon, \delta)$-certified machine unlearning for minimax models. Most of the existing works focus on unlearning from standard statistical learning models that have a single variable and their unlearning steps hinge on the direct Hessian-based conventional Newton update. We develop a new $(\epsilon, \delta)$-certified machine unlearning algorithm for minimax models. It proposes a minimax unlearning step consisting of a total Hessian-based complete Newton update and the Gaussian mechanism borrowed from differential privacy. To obtain the unlearning certification, our method injects calibrated Gaussian noises by carefully analyzing the ''sensitivity'' of the minimax unlearning step (i. e. , the closeness between the minimax unlearning variables and the retraining-from-scratch variables). We derive the generalization rates in terms of population strong and weak primal-dual risk for three different cases of loss functions, i. e. , (strongly-)convex-(strongly-)concave losses. We also provide the deletion capacity to guarantee that a desired population risk can be maintained as long as the number of deleted samples does not exceed the derived amount. With training samples $n$ and model dimension $d$, it yields the order $\mathcal O(n/d^{1/4})$, which shows a strict gap over the baseline method of differentially private minimax learning that has $\mathcal O(n/d^{1/2})$. In addition, our rates of generalization and deletion capacity match the state-of-the-art rates derived previously for standard statistical learning models.

NeurIPS Conference 2023 Conference Paper

Real3D-AD: A Dataset of Point Cloud Anomaly Detection

  • Jiaqi Liu
  • Guoyang Xie
  • Ruitao Chen
  • Xinpeng Li
  • Jinbao Wang
  • Yong Liu
  • Chengjie Wang
  • Feng Zheng

High-precision point cloud anomaly detection is the gold standard for identifying the defects of advancing machining and precision manufacturing. Despite some methodological advances in this area, the scarcity of datasets and the lack of a systematic benchmark hinder its development. We introduce Real3D-AD, a challenging high-precision point cloud anomaly detection dataset, addressing the limitations in the field. With 1, 254 high-resolution 3D items (from forty thousand to millions of points for each item), Real3D-AD is the largest dataset for high-precision 3D industrial anomaly detection to date. Real3D-AD surpasses existing 3D anomaly detection datasets available in terms of point cloud resolution (0. 0010mm-0. 0015mm), $360^{\circ}$ degree coverage and perfect prototype. Additionally, we present a comprehensive benchmark for Real3D-AD, revealing the absence of baseline methods for high-precision point cloud anomaly detection. To address this, we propose Reg3D-AD, a registration-based 3D anomaly detection method incorporating a novel feature memory bank that preserves local and global representations. Extensive experiments on the Real3D-AD dataset highlight the effectiveness of Reg3D-AD. For reproducibility and accessibility, we provide the Real3D-AD dataset, benchmark source code, and Reg3D-AD on our website: https: //github. com/M-3LAB/Real3D-AD.

IROS Conference 2023 Conference Paper

Visual-Kinematics Graph Learning for Procedure-Agnostic Instrument Tip Segmentation in Robotic Surgeries

  • Jiaqi Liu
  • Yonghao Long 0001
  • Kai Chen 0028
  • Cheuk Hei Leung
  • Zerui Wang
  • Qi Dou 0001

Accurate segmentation of surgical instrument tip is an important task for enabling downstream applications in robotic surgery, such as surgical skill assessment, tool-tissue interaction and deformation modeling, as well as surgical autonomy. However, this task is very challenging due to the small sizes of surgical instrument tips, and significant variance of surgical scenes across different procedures. Although much effort has been made on visual-based methods, existing segmentation models still suffer from low robustness thus not usable in practice. Fortunately, kinematics data from the robotic system can provide reliable prior for instrument location, which is consistent regardless of different surgery types. To make use of such multi-modal information, we propose a novel visual-kinematics graph learning framework to accurately segment the instrument tip given various surgical procedures. Specifically, a graph learning framework is proposed to encode relational features of instrument parts from both image and kinematics. Next, a cross-modal contrastive loss is designed to incorporate robust geometric prior from kinematics to image for tip segmentation. We have conducted experiments on a private paired visual-kinematics dataset including multiple procedures, i. e. , prostatectomy, total mesorectal excision, fundoplication and distal gastrectomy on cadaver, and distal gastrectomy on porcine. The leave-one-procedure-out cross validation demon-strated that our proposed multi-modal segmentation method significantly outperformed current image-based state-of-the-art approaches, exceeding averagely 11. 2% on Dice.

TIST Journal 2022 Journal Article

An Efficient Learning Framework for Federated XGBoost Using Secret Sharing and Distributed Optimization

  • Lunchen Xie
  • Jiaqi Liu
  • Songtao Lu
  • Tsung-Hui Chang
  • Qingjiang Shi

XGBoost is one of the most widely used machine learning models in the industry due to its superior learning accuracy and efficiency. Targeting at data isolation issues in the big data problems, it is crucial to deploy a secure and efficient federated XGBoost (FedXGB) model. Existing FedXGB models either have data leakage issues or are only applicable to the two-party setting with heavy communication and computation overheads. In this article, a lossless multi-party federated XGB learning framework is proposed with a security guarantee, which reshapes the XGBoost’s split criterion calculation process under a secret sharing setting and solves the leaf weight calculation problem by leveraging distributed optimization. Remarkably, a thorough analysis of model security is provided as well, and multiple numerical results showcase the superiority of the proposed FedXGB compared with the state-of-the-art models on benchmark datasets.

EAAI Journal 2022 Journal Article

ConvPatchTrans: A script identification network with global and local semantics deeply integrated

  • Ke Yang
  • Jizheng Yi
  • Aibin Chen
  • Jiaqi Liu
  • Wenjie Chen
  • Ze Jin

Optical Character Recognition (OCR) system serves the need of reading text from images. Script identification that identifies the language of the text in the image is an important part of OCR technology and an indispensable role in the stability and accuracy of the OCR system. The most challenging for script identification is the interference caused by similarities between texts in different languages. In this paper, a two-branch network named ConvPatchTrans is designed to process global and local semantic features separately, focusing on the text and each word in a picture. The ConvPatchTrans extracts feature from different stages of the Visual Geometry Group network (VGGNet) as global and local semantics. For the global branch, the linear classifier is recommended. For the local branch, text image data is converted to image sequence data. Then, multi-layers convolution-enhanced Transformer (MCET) is proposed to bring about the deep fusion of sequence. Finally, the global and local branches are fused by an adaptive weighted fusion method to get the best result. In order to verify the effectiveness of our proposed method, four public script identification datasets are used for comparative experiments. Our method has obtained the highest values among currently published methods on the CVSI2015 and MLE2E datasets, which are 98. 90% and 97. 50%, respectively. At the same time, satisfactory results are also obtained on the other two datasets.

TIST Journal 2022 Journal Article

DeepExpress: Heterogeneous and Coupled Sequence Modeling for Express Delivery Prediction

  • Siyuan Ren
  • Bin Guo
  • Longbing Cao
  • Ke Li
  • Jiaqi Liu
  • Zhiwen Yu

The prediction of express delivery sequence, i.e., modeling and estimating the volumes of daily incoming and outgoing parcels for delivery, is critical for online business, logistics, and positive customer experience, and specifically for resource allocation optimization and promotional activity arrangement. A precise estimate of consumer delivery requests has to involve sequential factors such as shopping behaviors, weather conditions, events, business campaigns, and their couplings. Despite that various methods have integrated external features to enhance the effects, extant works fail to address complex feature-sequence couplings in the following aspects: weaken the inter-dependencies when processing heterogeneous data and ignore the cumulative and evolving situation of coupling relationships. To address these issues, we propose DeepExpress—a deep-learning-based express delivery sequence prediction model, which extends the classic seq2seq framework to learn feature-sequence couplings. DeepExpress leverages an express delivery seq2seq learning, a carefully designed heterogeneous feature representation, and a novel joint training attention mechanism to adaptively handle heterogeneity issues and capture feature-sequence couplings for accurate prediction. Experimental results on real-world data demonstrate that the proposed method outperforms both shallow and deep baseline models.

TIST Journal 2022 Journal Article

Dynamic Probabilistic Graphical Model for Progressive Fake News Detection on Social Media Platform

  • Ke Li
  • Bin Guo
  • Jiaqi Liu
  • Jiangtao Wang
  • Haoyang Ren
  • Fei Yi
  • Zhiwen Yu

Recently, fake news has been readily spread by massive amounts of users in social media, and automatic fake news detection has become necessary. The existing works need to prepare the overall data to perform detection, losing important information about the dynamic evolution of crowd opinions, and usually neglect the issue of uneven arrival of data in the real world. To address these issues, in this article, we focus on a kind of approach for fake news detection, namely progressive detection, which can be achieved by the dynamic Probabilistic Graphical Model. Based on the observation on real-world datasets, we adaptively improve the Kalman Filter to the Labeled Variable Dimension Kalman Filter (LVDKF) that learns two universal patterns from true and fake news, respectively, which can capture the temporal information of time-series data that arrive unevenly. It can take sequential data as input, distill the dynamic evolution knowledge regarding a post, and utilize crowd wisdom from users’ responses to achieve progressive detection. Then we derive the formulas using the Forward, Backward, and EM Algorithm, and we design a dynamic detection algorithm using Bayes’ theorem. Finally, we design experimental scenarios simulating progressive detection and evaluate LVDKF on two public datasets. It outperforms the baseline methods in these experimental scenarios, which indicates that it is adequate for progressive detection.

ICRA Conference 2021 Conference Paper

Discriminative Asymmetric Learning for Efficient Surgical Instrument Parsing

  • Jiaqi Liu
  • Yu Qiao 0003
  • Jie Yang 0002
  • Guang-Zhong Yang
  • Yun Gu

Semantic segmentation of surgical instruments provides essential priors for autonomous surgery. This task is however challenging since the fine-structure of surgical instruments requires the accurate segmentation of detailed regions in images. As the visual guidance for autonomous surgery, the algorithm should also be real-time and friendly to embedded systems. In this paper, a discriminative asymmetric learning framework is proposed to balance the efficiency and effectiveness of surgical instrument segmentation. Two convolutional neural networks with specific designs are deployed to extract the detail and semantic features of instruments. To reduce the redundancy of visual representation, the aggregator-discriminator mechanism is proposed to distinguish the features learned from different levels. Experiments demonstrate that the proposed method contributes to competitive segmentation accuracy and a higher efficiency compared to existing methods.

v2026.09.13