Arrow Research search

Author name cluster

Changsheng Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

31 papers
2 author rows

Possible papers

31

AAAI Conference 2026 Conference Paper

PSA-MF: Personality-Sentiment Aligned Multi-Level Fusion for Multimodal Sentiment Analysis

  • Heng Xie
  • Kang Zhu
  • Zhengqi Wen
  • Jianhua Tao
  • Xuefei Liu
  • Ruibo Fu
  • Changsheng Li

Multimodal sentiment analysis (MSA) is a research field that recognizes human sentiments by combining textual, visual, and audio modalities. The main challenge lies in integrating sentiment-related information from different modalities, which typically arises during the unimodal feature extraction phase and the multimodal feature fusion phase. Existing methods extract only shallow information from unimodal features during the extraction phase, neglecting sentimental differences across different personalities. During the fusion phase, they directly merge the feature information from each modality without considering differences at the feature level. This ultimately affects the model's recognition performance. To address this problem, we propose a personality-sentiment aligned multi-level fusion framework. We introduce personality traits during the feature extraction phase and propose a novel personality-sentiment alignment method to obtain personalized sentiment embeddings from the textual modality for the first time. In the fusion phase, we introduce a novel multi-level fusion method. This method gradually integrates sentimental information from textual, visual, and audio modalities through multimodal pre-fusion and a multi-level enhanced fusion strategy. Our method has been evaluated through multiple experiments on two commonly used datasets, achieving state-of-the-art results.

IROS Conference 2025 Conference Paper

Development of a Novel Miniaturized Dexterous Manipulator with Variable Stiffness for NOTES

  • Rong Cong
  • Xipeng Wu
  • Chao Qian 0015
  • Kaijie Zhang
  • Xingguang Duan
  • Changsheng Li

Natural Orifice Transluminal Endoscopic Surgery (NOTES) holds great promise due to its ability to eliminate external incisions, reduce trauma, and accelerate recovery. However, the adoption of NOTES is hindered by the limited capabilities of existing instruments, particularly in achieving the required balance between compact size, dexterity, and load capacity. This paper introduces a novel robotic manipulator designed for NOTES, featuring a 5 mm diameter and 7 degrees of freedom (DoF). The manipulator incorporates an innovative 3-PRS flexible parallel mechanism combined with a continuum parallel structure, achieving enhanced dexterity and variable stiffness functionality within a miniaturized design. A kinematic and variable stiffness analysis is performed, and experimental validation demonstrates its bending performance and stiffness modulation. Additionally, the feasibility and practicality of the robotic system are confirmed through a peg-transfer experiment, proving its potential for real-world surgical applications. This research offers a viable solution for enhancing the performance of NOTES instruments.

IROS Conference 2025 Conference Paper

Fast Real-Time Neural Network-Based Kinematics Solving of the Cosserat Rod Model for a Parallel Continuum Surgical Manipulator

  • Xipeng Wu
  • Chao Qian 0015
  • Jinpeng Diao
  • Xingguang Duan
  • Changsheng Li

The parallel continuum mechanism offers distinct advantages in the design of surgical manipulators, including enhanced stiffness, improved precision, and a simplified structure compared to traditional Tendon-driven systems. Conventional kinematic models based on constant-curvature assumptions are often inadequate for accurately capturing the complex bending behaviors of this mechanism. In contrast, the Cosserat rod theory provides a rigorous framework for precise kinematic modeling of flexible structures. However, its computational complexity results in slow solving speeds, particularly when dealing with spatial points that are widely separated. This paper focuses on a miniaturized parallel continuum manipulator and employs the Cosserat rod model for kinematic modeling, combined with a neural network-based inverse kinematics solver to achieve rapid real-time computation. To expedite inverse kinematics, a multilayer perceptron is trained on 5, 000 samples generated from the Cosserat rod model, yielding the average absolute error of 0. 046mm and the average relative error of 0. 41% in predicting rod lengths. Experimental validation demonstrates that the neural network solver reduces computation time to about 0. 16ms compared to 700–3100ms for conventional numerical methods, underscoring its potential for enhancing the precision and responsiveness of surgical systems in minimally invasive procedures.

NeurIPS Conference 2025 Conference Paper

Fira: Can We Achieve Full-rank Training of LLMs Under Low-rank Constraint?

  • Xi Chen
  • Kaituo Feng
  • Changsheng Li
  • Xunhao Lai
  • Xiangyu Yue
  • Ye Yuan
  • Guoren Wang

Low-rank training has emerged as a promising approach for reducing memory usage in training Large Language Models (LLMs). Previous methods either rely on decomposing weight matrices (e. g. , LoRA), or seek to decompose gradient matrices (e. g. , GaLore) to ensure reduced memory consumption. However, both of them constrain the training in a low-rank subspace, thus inevitably leading to sub-optimal performance. To resolve this, we propose a new plug-and-play training framework for LLMs called Fira, as the first attempt to consistently preserve the low-rank constraint for memory efficiency, while achieving full-rank training (i. e. , training with full-rank gradients of full-rank weights) to avoid inferior outcomes. First, we observe an interesting phenomenon during LLM training: the scaling impact of adaptive optimizers (e. g. , Adam) on the gradient norm remains similar from low-rank to full-rank training. In light of this, we propose a \textit{norm-based scaling} method, which utilizes the scaling impact of low-rank optimizers as substitutes for that of original full-rank optimizers to achieve this goal. Moreover, we find that there are potential loss spikes during training. To address this, we further put forward a norm-growth limiter to smooth the gradient. Extensive experiments on the pre-training and fine-tuning of LLMs show that Fira outperforms both LoRA and GaLore. Notably, for pre-training LLaMA 7B, our Fira uses $8\times$ smaller memory of optimizer states than Galore, yet outperforms it by a large margin.

JBHI Journal 2025 Journal Article

Self-Supervised Monocular Depth Estimation for Endoscopic Imaging

  • Changsheng Li
  • Xue Li
  • Kaifeng Wang
  • Wenxin Chen
  • Qingyao Liu
  • Xingguang Duan

Endoscopy holds a pivotal role in the early detection and treatment of diverse diseases, with artificial intelligence (AI)-assisted methods increasingly gaining prominence in disease screening. Among them, the depth estimation from endoscopic sequences is crucial for a spectrum of AI-assisted surgical techniques. However, the development of endoscopic depth estimation algorithms presents a formidable challenge due to the unique environmental intricacies and constraints within the dataset. This paper proposes a self-supervised depth estimation network to comprehensively explore the brightness changes in endoscopic images, and fuse different features at multiple levels to achieve an accurate prediction of endoscopic depth. First, a FlowNet is designed to evaluate the brightness changes of adjacent frames by calculating the multi-scale structural similarity. Second, a feature fusion module is presented to capture multi-scale contextual information. Experiments show that the average accuracy of the algorithm is 97. 03% in the Stereo Correspondence and Reconstruction of Endoscopic Data (SCARED dataset). Based on the training parameters of the SCARED dataset, the algorithm achieves superior performance on the other two datasets (EndoSLAM and KVASIR dataset), indicating that the algorithm has good generalization performance.

NeurIPS Conference 2025 Conference Paper

Towards Multi-Table Learning: A Novel Paradigm for Complementarity Quantification and Integration

  • Zhang Junyu
  • Lizhong Ding
  • Ye Yuan
  • Xingcan Li
  • Pengqi Li
  • Tihang Xi
  • Guoren Wang
  • Changsheng Li

Multi-table data integrate various entities and attributes, with potential interconnections between them. However, existing tabular learning methods often struggle to describe and leverage the underlying complementarity across distinct tables. To address this limitation, we propose the first unified paradigm for multi-table learning that systematically quantifies and integrates complementary information across tables. Specifically, we introduce a metric called complementarity strength (CS), which captures inter-table complementarity by incorporating relevance, similarity, and informativeness. For the first time, we systematically formulate the paradigm towards multi-table learning by establishing formal definitions of tasks and loss functions. Correspondingly, we present a network for multi-table learning that combines Adaptive Table encoder and Cross table Attention mechanism (ATCA-Net), achieving the simultaneous integration of complementary information from distinct tables. Extensive experiments show that ATCA-Net effectively leverages complementary information and that the CS metric accurately quantifies the richness of complementarity across multiple tables. To the best of our knowledge, this is the first work to establish theoretical and practical foundations for multi-table learning.

NeurIPS Conference 2024 Conference Paper

Cross-Device Collaborative Test-Time Adaptation

  • Guohao Chen
  • Shuaicheng Niu
  • Deyu Chen
  • Shuhai Zhang
  • Changsheng Li
  • Yuanqing Li
  • Mingkui Tan

In this paper, we propose test-time Collaborative Lifelong Adaptation (CoLA), which is a general paradigm that can be incorporated with existing advanced TTA methods to boost the adaptation performance and efficiency in a multi-device collaborative manner. Specifically, we maintain and store a set of device-shared domain knowledge vectors, which accumulates the knowledge learned from all devices during their lifelong adaptation process. Based on this, CoLA conducts two collaboration strategies for devices with different computational resources and latency demands. 1) Knowledge reprogramming learning strategy jointly learns new domain-specific model parameters and a reweighting term to reprogram existing shared domain knowledge vectors, termed adaptation on principal agents. 2) Similarity-based knowledge aggregation strategy solely aggregates the knowledge stored in shared domain vectors according to domain similarities in an optimization-free manner, termed adaptation on follower agents. Experiments verify that CoLA is simple but effective, which boosts the efficiency of TTA and demonstrates remarkable superiority in collaborative, lifelong, and single-domain TTA scenarios, e. g. , on follower agents, we enhance accuracy by over 30\% on ImageNet-C while maintaining nearly the same efficiency as standard inference. The source code is available at https: //github. com/Cascol-Chen/COLA.

ICRA Conference 2024 Conference Paper

Excitation Trajectory Optimization for Dynamic Parameter Identification Using Virtual Constraints in Hands-on Robotic System

  • Huanyu Tian
  • Martin Huber
  • Christopher E. Mower
  • Zhe Han
  • Changsheng Li
  • Xingguang Duan
  • Christos Bergeles

This paper proposes a novel, more computationally efficient method for optimizing robot excitation trajectories for dynamic parameter identification, emphasizing self-collision avoidance. This addresses the system identification challenges for getting high-quality training data associated with co-manipulated robotic arms that can be equipped with a variety of tools, a common scenario in industrial but also clinical and research contexts. Utilizing the Unified Robotics Description Format (URDF) to implement a symbolic Python implementation of the Recursive Newton-Euler Algorithm (RNEA), the approach aids in dynamically estimating parameters such as inertia using regression analyses on data from real robots. The excitation trajectory was evaluated and achieved on par criteria when compared to state-of-the-art reported results which didn’t consider self-collision and tool calibrations. Furthermore, physical Human-Robot Interaction (pHRI) admittance control experiments were conducted in a surgical context to evaluate the derived inverse dynamics model showing a 30. 1% workload reduction by the NASA TLX questionnaire.

IROS Conference 2024 Conference Paper

FDNet: Feature Decoupling Framework for Trajectory Prediction

  • Yuhang Li 0007
  • Changsheng Li
  • Baoyu Fan
  • Rongqing Li
  • Ziyue Zhang
  • Dongchun Ren
  • Ye Yuan 0001
  • Guoren Wang

Trajectory prediction plays a significant role in autonomous driving, with current challenges primarily focused on capturing complex interactions in traffic scenes. Previous methods usually directly encode non-interactive and interactive information together, and then decode them for trajectory prediction. However, given the complexity inherent property in the trajectory generation process (e. g. , the generation of trajectory points are influenced by the interactions among multiple moving agents, as well as the interactions between agents and the static environment), previous approaches fail to precisely capture separate variations of the trajectory generation process. In this paper, we propose a general and plug-and-play feature decoupling framework for trajectory prediction called FDNet, which can learn the interactive and non-interactive factors in the latent space to capture separate variations of the trajectory generation process. At its core, FDNet is comprised of a Non-interactive Feature Extraction Module to extract non-interactive features, and an Interactive Feature Decoupling Module to decouple interactive features. Extensive experiments conducted on Argoverse and nuScenes demonstrate that FDNet significantly improves the performance of existing methods.

NeurIPS Conference 2024 Conference Paper

Hierarchical Object-Aware Dual-Level Contrastive Learning for Domain Generalized Stereo Matching

  • Yikun Miao
  • Meiqing Wu
  • Siew-Kei Lam
  • Changsheng Li
  • Thambipillai Srikanthan

Stereo matching algorithms that leverage end-to-end convolutional neural networks have recently demonstrated notable advancements in performance. However, a common issue is their susceptibility to domain shifts, hindering their ability in generalizing to diverse, unseen realistic domains. We argue that existing stereo matching networks overlook the importance of extracting semantically and structurally meaningful features. To address this gap, we propose an effective hierarchical object-aware dual-level contrastive learning (HODC) framework for domain generalized stereo matching. Our framework guides the model in extracting features that support semantically and structurally driven matching by segmenting objects at different scales and enhances correspondence between intra- and inter-scale regions from the left feature map to the right using dual-level contrastive loss. HODC can be integrated with existing stereo matching models in the training stage, requiring no modifications to the architecture. Remarkably, using only synthetic datasets for training, HODC achieves state-of-the-art generalization performance with various existing stereo matching network architectures, across multiple realistic datasets.

ICML Conference 2024 Conference Paper

Keypoint-based Progressive Chain-of-Thought Distillation for LLMs

  • Kaituo Feng
  • Changsheng Li
  • Xiaolu Zhang
  • Jun Zhou 0011
  • Ye Yuan 0001
  • Guoren Wang

Chain-of-thought distillation is a powerful technique for transferring reasoning abilities from large language models (LLMs) to smaller student models. Previous methods typically require the student to mimic the step-by-step rationale produced by LLMs, often facing the following challenges: (i) Tokens within a rationale vary in significance, and treating them equally may fail to accurately mimic keypoint tokens, leading to reasoning errors. (ii) They usually distill knowledge by consistently predicting all the steps in a rationale, which falls short in distinguishing the learning order of step generation. This diverges from the human cognitive progression of starting with easy tasks and advancing to harder ones, resulting in sub-optimal outcomes. To this end, we propose a unified framework, called KPOD, to address these issues. Specifically, we propose a token weighting module utilizing mask learning to encourage accurate mimicry of keypoint tokens by the student during distillation. Besides, we develop an in-rationale progressive distillation strategy, starting with training the student to generate the final reasoning steps and gradually extending to cover the entire rationale. To accomplish this, a weighted token generation loss is proposed to assess step reasoning difficulty, and a value function is devised to schedule the progressive distillation by considering both step difficulty and question diversity. Extensive experiments on four reasoning benchmarks illustrate our KPOD outperforms previous methods by a large margin.

NeurIPS Conference 2024 Conference Paper

LaKD: Length-agnostic Knowledge Distillation for Trajectory Prediction with Any Length Observations

  • Yuhang Li
  • Changsheng Li
  • Ruilin Lv
  • Rongqing Li
  • Ye Yuan
  • Guoren Wang

Trajectory prediction is a crucial technology to help systems avoid traffic accidents, ensuring safe autonomous driving. Previous methods typically use a fixed-length and sufficiently long trajectory of an agent as observations to predict its future trajectory. However, in real-world scenarios, we often lack the time to gather enough trajectory points before making predictions, e. g. , when a car suddenly appears due to an obstruction, the system must make immediate predictions to prevent a collision. This poses a new challenge for trajectory prediction systems, requiring them to be capable of making accurate predictions based on observed trajectories of arbitrary lengths, leading to the failure of existing methods. In this paper, we propose a Length-agnostic Knowledge Distillation framework, named LaKD, which can make accurate trajectory predictions, regardless of the length of observed data. Specifically, considering the fact that long trajectories, containing richer temporal information but potentially additional interference, may perform better or worse than short trajectories, we devise a dynamic length-agnostic knowledge distillation mechanism for exchanging information among trajectories of arbitrary lengths, dynamically determining the transfer direction based on prediction performance. In contrast to traditional knowledge distillation, LaKD employs a unique model that simultaneously serves as both the teacher and the student, potentially causing knowledge collision during the distillation process. Therefore, we design a dynamic soft-masking mechanism, where we first calculate the importance of neuron units and then apply soft-masking to them, so as to safeguard critical units from disruption during the knowledge distillation process. In essence, LaKD is a general and principled framework that can be naturally compatible with existing trajectory prediction models of different architectures. Extensive experiments on three benchmark datasets, Argoverse 1, nuScenes and Argoverse 2, demonstrate the effectiveness of our approach.

NeurIPS Conference 2023 Conference Paper

BCDiff: Bidirectional Consistent Diffusion for Instantaneous Trajectory Prediction

  • Rongqing Li
  • Changsheng Li
  • Dongchun Ren
  • Guangyi Chen
  • Ye Yuan
  • Guoren Wang

The objective of pedestrian trajectory prediction is to estimate the future paths of pedestrians by leveraging historical observations, which plays a vital role in ensuring the safety of self-driving vehicles and navigation robots. Previous works usually rely on a sufficient amount of observation time to accurately predict future trajectories. However, there are many real-world situations where the model lacks sufficient time to observe, such as when pedestrians abruptly emerge from blind spots, resulting in inaccurate predictions and even safety risks. Therefore, it is necessary to perform trajectory prediction based on instantaneous observations, which has rarely been studied before. In this paper, we propose a Bi-directional Consistent Diffusion framework tailored for instantaneous trajectory prediction, named BCDiff. At its heart, we develop two coupled diffusion models by designing a mutual guidance mechanism which can bidirectionally and consistently generate unobserved historical trajectories and future trajectories step-by-step, to utilize the complementary information between them. Specifically, at each step, the predicted unobserved historical trajectories and limited observed trajectories guide one diffusion model to generate future trajectories, while the predicted future trajectories and observed trajectories guide the other diffusion model to predict unobserved historical trajectories. Given the presence of relatively high noise in the generated trajectories during the initial steps, we introduce a gating mechanism to learn the weights between the predicted trajectories and the limited observed trajectories for automatically balancing their contributions. By means of this iterative and mutually guided generation process, both the future and unobserved historical trajectories undergo continuous refinement, ultimately leading to accurate predictions. Essentially, BCDiff is an encoder-free framework that can be compatible with existing trajectory prediction models in principle. Experiments show that our proposed BCDiff significantly improves the accuracy of instantaneous trajectory prediction on the ETH/UCY and Stanford Drone datasets, compared to related approaches.

ICML Conference 2023 Conference Paper

Detecting Adversarial Data by Probing Multiple Perturbations Using Expected Perturbation Score

  • Shuhai Zhang
  • Feng Liu 0003
  • Jiahao Yang
  • Yifan Yang
  • Changsheng Li
  • Bo Han 0003
  • Mingkui Tan

Adversarial detection aims to determine whether a given sample is an adversarial one based on the discrepancy between natural and adversarial distributions. Unfortunately, estimating or comparing two data distributions is extremely difficult, especially in high-dimension spaces. Recently, the gradient of log probability density (a. k. a. , score) w. r. t. the sample is used as an alternative statistic to compute. However, we find that the score is sensitive in identifying adversarial samples due to insufficient information with one sample only. In this paper, we propose a new statistic called expected perturbation score (EPS), which is essentially the expected score of a sample after various perturbations. Specifically, to obtain adequate information regarding one sample, we perturb it by adding various noises to capture its multi-view observations. We theoretically prove that EPS is a proper statistic to compute the discrepancy between two samples under mild conditions. In practice, we can use a pre-trained diffusion model to estimate EPS for each sample. Last, we pro- pose an EPS-based adversarial detection (EPS- AD) method, in which we develop EPS-based maximum mean discrepancy (MMD) as a metric to measure the discrepancy between the test sample and natural samples. We also prove that the EPS-based MMD between natural and adversarial samples is larger than that among natural samples. Extensive experiments show the superior adversarial detection performance of our EPS-AD.

ICLR Conference 2023 Conference Paper

Towards Open Temporal Graph Neural Networks

  • Kaituo Feng
  • Changsheng Li
  • Xiaolu Zhang
  • Jun Zhou 0011

Graph neural networks (GNNs) for temporal graphs have recently attracted increasing attentions, where a common assumption is that the class set for nodes is closed. However, in real-world scenarios, it often faces the open set problem with the dynamically increased class set as the time passes by. This will bring two big challenges to the existing dynamic GNN methods: (i) How to dynamically propagate appropriate information in an open temporal graph, where new class nodes are often linked to old class nodes. This case will lead to a sharp contradiction. This is because typical GNNs are prone to make the embeddings of connected nodes become similar, while we expect the embeddings of these two interactive nodes to be distinguishable since they belong to different classes. (ii) How to avoid catastrophic knowledge forgetting over old classes when learning new classes occurred in temporal graphs. In this paper, we propose a general and principled learning approach for open temporal graphs, called OTGNet, with the goal of addressing the above two challenges. We assume the knowledge of a node can be disentangled into class-relevant and class-agnostic one, and thus explore a new message passing mechanism by extending the information bottleneck principle to only propagate class-agnostic knowledge between nodes of different classes, avoiding aggregating conflictive information. Moreover, we devise a strategy to select both important and diverse triad sub-graph structures for effective class-incremental learning. Extensive experiments on three real-world datasets of different domains demonstrate the superiority of our method, compared to the baselines.

IROS Conference 2022 Conference Paper

GESRsim: Gastrointestinal Endoscopic Surgical Robot Simulator

  • Huxin Gao
  • Zedong Zhang
  • Changsheng Li
  • Xiao Xiao 0006
  • Liang Qiu 0002
  • Xiaoxiao Yang
  • Ruoyi Hao
  • Xiuli Zuo

Robot-assisted gastrointestinal endoscopic surgery (GES) as a kind of natural orifice transluminal endoscopic surgery (NOTES) is the next-generation minimally invasive surgery (MIS). Besides, rendering certain autonomy to a Gas-trointestinal Endoscopic Surgical Robot (GESR) is promising but highly challenging. Therefore, to accelerate the development and augment the autonomy of GESR, we use CoppeliaSim to develop the first robotic simulator for the GESR system (GESRsim) based on our previous design. The GESRsim provides several 3D models and kinematics of our designed manipulators and endoscopic snake bone. Additionally, we build several scenes for robotic GES training and then utilize different programming interfaces to perform teleoperation. Furthermore, several advanced control algorithms, including visual servoing (VS) and deep reinforcement learning (DRL), are implemented to verify the performance of the GESRsim.

AAAI Conference 2021 Conference Paper

LRSC: Learning Representations for Subspace Clustering

  • Changsheng Li
  • Chen Yang
  • Bo Liu
  • Ye Yuan
  • Guoren Wang

Deep learning based subspace clustering methods have attracted increasing attention in recent years, where a basic theme is to non-linearly map data into a latent space, and then uncover subspace structures based upon the data selfexpressiveness property. However, almost all existing deep subspace clustering methods only rely on target domain data, and always resort to shallow neural networks for modeling data, leaving huge room to design more effective representation learning mechanisms tailored for subspace clustering. In this paper, we propose a novel subspace clustering framework through learning precise sample representations. In contrast to previous approaches, the proposed method aims to leverage external data through constructing lots of relevant tasks to guide the training of the encoder, motivated by the idea of meta-learning. Considering limited networks layers of current deep subspace clustering models, we intend to distill knowledge from a deeper network trained on the external data, and transfer it into the shallower model. To reach the above two goals, we propose a new loss function to realize them in a unified framework. Moreover, we propose to construct a new auxiliary task for self-supervised training of the model, such that the representation ability of the model can be further improved. Extensive experiments are performed on four publicly available datasets, and experimental results clearly demonstrate the efficacy of our method, compared to state-of-the-art methods.

ICRA Conference 2021 Conference Paper

Magnetically-Connected Modular Reconfigurable Mini-robotic System with Bilateral Isokinematic Mapping and Fast On-site Assembly towards Minimally Invasive Procedures

  • Xiao Xiao 0006
  • Shilei Xu
  • Changsheng Li
  • Xiaoyi Gu
  • Huxin Gao
  • Max Q. -H. Meng
  • Hongliang Ren 0001

This paper presents a modular and reconfigurable mini-robotic system with 5 degrees of freedom (DoFs) towards minimally invasive surgery (MIS). The mini-robotic system consists of two modules, a 2-DoFs rotational end-effector, and a 3-DoFs positioning platform. The 2-DoFs rotational end-effector is based on a spring-spherical joint mechanism, whose rotation is controlled by Bowden-cable. The 3-DoFs positioning platform is based on the linear Delta parallel mechanism. Magnetic spherical joints are adopted to replace the traditional spherical joint. The magnetic joint connections enable fast assembling and disassemble of the end platform and kinematic chains. Different surgical instruments can be installed without changing the driver and control system. A flexible shaft actuates the 3-DoFs positioning platform to arrange the motors away from the manipulator side. Based on these structure characteristics, the 3-DoFs positioning platform’s size is dramatically reduced. The outer diameter of the current prototype is 32. 5 mm. The single-axis positioning accuracy of the 3-DoFs positioning platform is within -1 mm to 0. 85 mm. Three axes tracking experiments are also carried out, with the positioning errors of ± 1. 2 mm for cylindrical curves and -1. 5 mm to 2 mm for spherical helix curves. Static and dynamic load capabilities are also tested. Finally, the feasibility of the proposed system is demonstrated.

AAAI Conference 2021 Conference Paper

Unsupervised Active Learning via Subspace Learning

  • Changsheng Li
  • Kaihang Mao
  • Lingyan Liang
  • Dongchun Ren
  • Wei Zhang
  • Ye Yuan
  • Guoren Wang

Unsupervised active learning has been an active research topic in machine learning community, with the purpose of choosing representative samples to be labelled in an unsupervised manner. Previous works usually take the minimization of data reconstruction loss as the criterion to select representative samples, by which the original inputs can be better approximated. However, data are often drawn from low-dimensional subspaces embedded in an arbitrary highdimensional space in many scenarios, thus it might severely bring in noise if attempting to precisely reconstruct all entries of one observation, leading to a suboptimal solution. In view of this, this paper proposes a novel unsupervised Active Learning model via Subspace Learning, called ALSL. In contrast to previous approaches, ALSL aims to discover low-rank structures of data, and then perform sample selection based on the learnt low-rank representations. To this end, we devise two different strategies and propose two corresponding formulations to select samples with and under low-rank sample representations, respectively. Since the proposed formulations involve several non-smooth regularization terms, we develop a simple but effective optimization procedure to solve them. Extensive experiments are performed on five publicly available datasets, and experimental results demonstrate the proposed first formulation achieves comparable performance with the state-of-the-arts, while the second formulation significantly outperforms them, achieving a 13% improvement over the second best baseline at most.

AAAI Conference 2020 Conference Paper

Multi-Instance Multi-Label Action Recognition and Localization Based on Spatio-Temporal Pre-Trimming for Untrimmed Videos

  • Xiao-Yu Zhang
  • Haichao Shi
  • Changsheng Li
  • Peng Li

Weakly supervised action recognition and localization for untrimmed videos is a challenging problem with extensive applications. The overwhelming irrelevant background contents in untrimmed videos severely hamper effective identification of actions of interest. In this paper, we propose a novel multi-instance multi-label modeling network based on spatio-temporal pre-trimming to recognize actions and locate corresponding frames in untrimmed videos. Motivated by the fact that person is the key factor in a human action, we spatially and temporally segment each untrimmed video into person-centric clips with pose estimation and tracking techniques. Given the bag-of-instances structure associated with video-level labels, action recognition is naturally formulated as a multi-instance multi-label learning problem. The network is optimized iteratively with selective coarse-to-fine pre-trimming based on instance-label activation. After convergence, temporal localization is further achieved with localglobal temporal class activation map. Extensive experiments are conducted on two benchmark datasets, i. e. THUMOS14 and ActivityNet1. 3, and experimental results clearly corroborate the efficacy of our method when compared with the stateof-the-arts.

IJCAI Conference 2020 Conference Paper

On Deep Unsupervised Active Learning

  • Changsheng Li
  • Handong Ma
  • Zhao Kang
  • Ye Yuan
  • Xiao-Yu Zhang
  • Guoren Wang

Unsupervised active learning has attracted increasing attention in recent years, where its goal is to select representative samples in an unsupervised setting for human annotating. Most existing works are based on shallow linear models by assuming that each sample can be well approximated by the span (i. e. , the set of all linear combinations) of certain selected samples, and then take these selected samples as representative ones to label. However, in practice, the data do not necessarily conform to linear models, and how to model nonlinearity of data often becomes the key point to success. In this paper, we present a novel Deep neural network framework for Unsupervised Active Learning, called DUAL. DUAL can explicitly learn a nonlinear embedding to map each input into a latent space through an encoder-decoder architecture, and introduce a selection block to select representative samples in the the learnt latent space. In the selection block, DUAL considers to simultaneously preserve the whole input patterns as well as the cluster structure of data. Extensive experiments are performed on six publicly available datasets, and experimental results clearly demonstrate the efficacy of our method, compared with state-of-the-arts.

AAAI Conference 2020 Conference Paper

Region Focus Network for Joint Optic Disc and Cup Segmentation

  • Ge Li
  • Changsheng Li
  • Chan Zeng
  • Peng Gao
  • Guotong Xie

Glaucoma is one of the three leading causes of blindness in the world and is predicted to affect around 80 million people by 2020. The optic cup (OC) to optic disc (OD) ratio (CDR) in fundus images plays a pivotal role in the screening and diagnosis of glaucoma. Existing methods usually crop the optic disc region first, and subsequently perform segmentation in this region. However, these approaches come up with high complexities due to the separate operations. To remedy this issue, we propose a Region Focus Network (RF-Net) that innovatively integrates detection and multi-class segmentation into a unified architecture for end-to-end joint optic disc and cup segmentation with global optimization. The key idea of our method is designing a novel multi-class mask branch which generates a high-quality segmentation in the detected region for both disc and cup. To bridge the connection between the backbone and multi-class mask branch, a Fusion Feature Pooling (FFP) structure is presented to extract features from each level of the pyramid network and fuse them into a final feature representation for segmentation. Extensive experimental results on the REFUGE-2018 challenge dataset and the Drishti-GS dataset show that the proposed method achieves the best performance, compared with competitive approaches reported in the literature and the official leaderboard. Our code will be released soon.

AAAI Conference 2019 Conference Paper

Learning Transferable Self-Attentive Representations for Action Recognition in Untrimmed Videos with Weak Supervision

  • Xiao-Yu Zhang
  • Haichao Shi
  • Changsheng Li
  • Kai Zheng
  • Xiaobin Zhu
  • Lixin Duan

Action recognition in videos has attracted a lot of attention in the past decade. In order to learn robust models, previous methods usually assume videos are trimmed as short sequences and require ground-truth annotations of each video frame/sequence, which is quite costly and time-consuming. In this paper, given only video-level annotations, we propose a novel weakly supervised framework to simultaneously locate action frames as well as recognize actions in untrimmed videos. Our proposed framework consists of two major components. First, for action frame localization, we take advantage of the self-attention mechanism to weight each frame, such that the influence of background frames can be effectively eliminated. Second, considering that there are trimmed videos publicly available and also they contain useful information to leverage, we present an additional module to transfer the knowledge from trimmed videos for improving the classification performance in untrimmed ones. Extensive experiments are conducted on two benchmark datasets (i. e. , THUMOS14 and ActivityNet1. 3), and experimental results clearly corroborate the efficacy of our method.

AAAI Conference 2019 Conference Paper

Residual Invertible Spatio-Temporal Network for Video Super-Resolution

  • Xiaobin Zhu
  • Zhuangzi Li
  • Xiao-Yu Zhang
  • Changsheng Li
  • Yaqi Liu
  • Ziyu Xue

Video super-resolution is a challenging task, which has attracted great attention in research and industry communities. In this paper, we propose a novel end-to-end architecture, called Residual Invertible Spatio-Temporal Network (RISTN) for video super-resolution. The RISTN can sufficiently exploit the spatial information from low-resolution to high-resolution, and effectively models the temporal consistency from consecutive video frames. Compared with existing recurrent convolutional network based approaches, RISTN is much deeper but more efficient. It consists of three major components: In the spatial component, a lightweight residual invertible block is designed to reduce information loss during feature transformation and provide robust feature representations. In the temporal component, a novel recurrent convolutional model with residual dense connections is proposed to construct deeper network and avoid feature degradation. In the reconstruction component, a new fusion method based on the sparse strategy is proposed to integrate the spatial and temporal features. Experiments on public benchmark datasets demonstrate that RISTN outperforms the state-ofthe-art methods.

AAAI Conference 2019 Conference Paper

Similarity Learning via Kernel Preserving Embedding

  • Zhao Kang
  • Yiwei Lu
  • Yuanzhang Su
  • Changsheng Li
  • Zenglin Xu

Data similarity is a key concept in many data-driven applications. Many algorithms are sensitive to similarity measures. To tackle this fundamental problem, automatically learning of similarity information from data via self-expression has been developed and successfully applied in various models, such as low-rank representation, sparse subspace learning, semisupervised learning. However, it just tries to reconstruct the original data and some valuable information, e. g. , the manifold structure, is largely ignored. In this paper, we argue that it is beneficial to preserve the overall relations when we extract similarity information. Specifically, we propose a novel similarity learning framework by minimizing the reconstruction error of kernel matrices, rather than the reconstruction error of original data adopted by existing work. Taking the clustering task as an example to evaluate our method, we observe considerable improvements compared to other state-ofthe-art methods. More importantly, our proposed framework is very general and provides a novel and fundamental building block for many other similarity-based tasks. Besides, our proposed kernel preserving opens up a large number of possibilities to embed high-dimensional data into low-dimensional space.

IJCAI Conference 2018 Conference Paper

Improving Maximum Likelihood Estimation of Temporal Point Process via Discriminative and Adversarial Learning

  • Junchi Yan
  • Xin Liu
  • Liangliang Shi
  • Changsheng Li
  • Hongyuan Zha

Point process is an expressive tool in learning temporal event sequence which is ubiquitous in real-world applications. Traditional predictive models are based on maximum likelihood estimation (MLE). This paper aims to improve MLE by discriminative and adversarial learning. The initial model is learned by MLE explaining the joint distribution of the occurred event history. Then it is refined by devising a gradient based learning procedure with two complementary recipes: i) mean square error (MSE) that directly reflects the prediction accuracy of the model; ii) adversarial classification loss which induces the Wasserstein distance loss. The hope is that the adversarial loss can add sharpness to the smooth effect inherently caused by the MSE loss. The method is generic and compatible with different differentiable parametric forms of the intensity function. Empirical results via a variant of the Hawkes processes demonstrate its effectiveness of our method.

IJCAI Conference 2017 Conference Paper

Autoencoder Regularized Network For Driving Style Representation Learning

  • Weishan Dong
  • Ting Yuan
  • Kai Yang
  • Changsheng Li
  • Shilei Zhang

In this paper, we study learning generalized driving style representations from automobile GPS trip data. We propose a novel Autoencoder Regularized deep neural Network (ARNet) and a trip encoding framework trip2vec to learn drivers' driving styles directly from GPS records, by combining supervised and unsupervised feature learning in a unified architecture. Experiments on a challenging driver number estimation problem and the driver identification problem show that ARNet can learn a good generalized driving style representation: It significantly outperforms existing methods and alternative architectures by reaching the least estimation error on average (0. 68, less than one driver) and the highest identification accuracy (by at least 3% improvement) compared with traditional supervised learning methods.

AAAI Conference 2017 Conference Paper

Self-Paced Multi-Task Learning

  • Changsheng Li
  • Junchi Yan
  • Fan Wei
  • Weishan Dong
  • Qingshan Liu
  • Hongyuan Zha

Multi-task learning is a paradigm, where multiple tasks are jointly learnt. Previous multi-task learning models usually treat all tasks and instances per task equally during learning. Inspired by the fact that humans often learn from easy concepts to hard ones in the cognitive process, in this paper, we propose a novel multi-task learning framework that attempts to learn the tasks by simultaneously taking into consideration the complexities of both tasks and instances per task. We propose a novel formulation by presenting a new task-oriented regularizer that can jointly prioritize tasks and instances. Thus it can be interpreted as a self-paced learner for multi-task learning. An efficient block coordinate descent algorithm is developed to solve the proposed objective function, and the convergence of the algorithm can be guaranteed. Experimental results on the toy and real-world datasets demonstrate the effectiveness of the proposed approach, compared to the state-of-the-arts.

IJCAI Conference 2016 Conference Paper

Modeling Contagious Merger and Acquisition via Point Processes with a Profile Regression Prior

  • Junchi Yan
  • Shuai Xiao
  • Changsheng Li
  • Bo Jin
  • Xiangfeng Wang
  • Bin Ke
  • Xiaokang Yang
  • Hongyuan Zha

Merger and Acquisition (M&A) has been a critical practice about corporate restructuring. Previous studies are mostly devoted to evaluating the suitability of M&A between a pair of investor and target company, or a target company for its propensity of being acquired. This paper focuses on the dual problem of predicting an investor's prospective M&A based on its activities and firmographics. We propose to use a mutually-exciting point process with a regression prior to quantify the investor's M&A behavior. Our model is motivated by the so-called contagious 'wave-like' M&A phenomenon, which has been well-recognized by the economics and management communities. A tailored model learning algorithm is devised that incorporates both static profile covariates and past M&A activities. Results on CrunchBase suggest the superiority of our model. The collected dataset and code will be released together with the paper.

IJCAI Conference 2016 Conference Paper

On Modeling and Predicting Individual Paper Citation Count over Time

  • Shuai Xiao
  • Junchi Yan
  • Changsheng Li
  • Bo Jin
  • Xiangfeng Wang
  • Xiaokang Yang
  • Stephen M. Chu
  • Hongyuan Zha

Evaluating a scientist's past and future potential impact is key in decision making concerning with recruitment and funding, and is increasingly linked to publication citation count. Meanwhile, timely identifying those valuable work with great potential before they receive wide recognition and become highly cited Abstracts is both useful for readers and authors in many regards. We propose a method for predicting the citation counts of individual publications, over an arbitrary time period. Our approach explores paper-specific covariates, and a point process model to account for the aging effect and triggering role of recent citations, through which Abstracts lose and gain their popularity, respectively. Empirical results on the Microsoft Academic Graph data suggests that our model can be useful for both prediction and interpretability.

v2026.09.13