Arrow Research search

Author name cluster

Xin Zhou

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

22 papers
2 author rows

Possible papers

22

AAAI Conference 2026 Conference Paper

TransFR: Transferable Federated Recommendation with Adapter Tuning on Pre-trained Language Models

  • Honglei Zhang
  • Zhiwei Li
  • Haoxuan Li
  • Xin Zhou
  • Jie Zhang
  • Yidong Li

Federated recommendations (FRs), facilitating multiple local clients to collectively learn a global model without disclosing user private data, have emerged as a prevalent on-device service. In conventional FRs, a dominant paradigm is to utilize discrete identities to represent clients and items, which are then mapped to domain-specific embeddings to participate in model training. Despite considerable performance, we reveal three inherent limitations that can not be ignored in federated settings, i.e., non-transferability across domains, ineffectiveness in cold-start settings, and potential privacy violations during federated training. To this end, we propose a transferable federated recommendation model, TransFR, which delicately incorporates the general capabilities empowered by pre-trained models and the personalized abilities by fine-tuning local private data. Specifically, it first learns domain-agnostic representations of items by exploiting pre-trained models with public textual corpora. To tailor for FR tasks, we further introduce efficient federated adapter-tuning and post-adaptation personalization, which facilitate personalized adapters for each client by fitting local private data. We theoretically prove the advantages of incorporating adapter tuning in FRs regarding both effectiveness and privacy. Through extensive experiments, we show that our TransFR surpasses state-of-the-art FRs on transferability.

AAAI Conference 2026 Conference Paper

USE: A Unified Model for Universal Sound Separation and Extraction

  • Hongyu Wang
  • Chenda Li
  • Xin Zhou
  • Shuai Wang
  • Yanmin Qian

Sound separation (SS) and target sound extraction (TSE) are fundamental techniques for addressing complex acoustic scenarios. While existing SS methods struggle with determining the unknown number of sound sources, TSE approaches require precisely specified clues to achieve optimal performance. This paper proposes a unified framework that synergistically combines SS and TSE to overcome their individual limitations. Our architecture employs two complementary components: 1) An Encoder-Decoder Attractor (EDA) network that automatically infers both the source count and corresponding acoustic clues for SS, and 2) A multi-modal fusion network that precisely interprets diverse user-provided clues (acoustic, semantic, or visual) for TSE. Through joint training with cross-task consistency constraints, we establish a unified latent space that bridges both paradigms. During inference, the system adaptively operates in either fully autonomous SS mode or clue-driven TSE mode. Experiments demonstrate remarkable performance in both tasks, with notable improvements of 1.4 dB SDR improvement in SS compared to baseline and 86% TSE accuracy.

AAAI Conference 2025 Conference Paper

Both Supply and Precision: Sample Debias and Ranking Consistency Joint Learning for Large Scale Pre-Ranking System

  • Feng Gao
  • Xin Zhou
  • Yinning Shao
  • Yue Wu
  • Jiahua Gao
  • Yujian Ren
  • Fengyang Qi
  • Ruochen Deng

Cascade ranking architecture, composed of matching, pre-ranking, ranking and re-ranking stages, is usually adopted to balance the efficiency and effectiveness in real-world recommendation system (RS). As the middle stage of RS, pre-ranking aims to quickly filter out the low-quality items selected at the matching stage and then forwarding high-quality items to the ranking stage. Existing pre-ranking approaches mainly endure two problems 1) Sample Selection Bias (SSB) problem, which heavily limits the performance improvement of filtering out low-quality items owing to ignoring the data flow between stages; and 2) Ranking Consistency (RC) problem, which may cause the ranked lists of the ranking stage and previous pre-ranking stage to be inconsistent. As a result, the competitive items with high scores at the ranking stage may not be selected because of low scores at the pre-ranking stage. These both two problems may cause sub-optimal performances, but previous works usually only focus on the one of them. In this paper, we propose a novel Sample Debias and Ranking Consistency Joint Learning Framework (SDCL) to jointly alleviate SSB and RC problems. SDCL consists of two main modules including 1) Multi-Task Distillation Module (MTD), which enhances the ability of identifying high-quality items by distilling knowledge across all tasks simultaneously from the more complex ranking model which jointly trained with the pre-ranking model; and 2) Adaptive Negative Sample Learning Module (ANSL), which improves the performance of filtering out low-quality items by adaptively adjusting negative samples learning weights based on the current performance of model. SDCL seamlessly integrates two modules in an end-to-end multi-task learning framework. Evaluations on both real-world large-scale traffic logs and online A/B test demonstrate the efficacy and superiority of SDCL.

NeurIPS Conference 2025 Conference Paper

More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models

  • Hongkai Lin
  • Dingkang Liang
  • Mingyang Du
  • Xin Zhou
  • Xiang Bai

Generative depth estimation methods leverage the rich visual priors stored in pretrained text-to-image diffusion models, demonstrating astonishing zero-shot capability. However, parameter updates during training lead to catastrophic degradation in the image generation capability of the pretrained model. We introduce MERGE, a unified model for image generation and depth estimation, starting from a fixed-parameters pretrained text-to-image model. MERGE demonstrates that the pretrained text-to-image model can do more than image generation but also expand to depth estimation effortlessly. Specifically, MERGE introduces a plug-and-play framework that enables seamless switching between image generation and depth estimation modes through simple and pluggable converters. Meanwhile, we propose a Group Reuse Mechanism to encourage parameter reuse and improve the utilization of the additional learnable parameter. MERGE unleashes the powerful depth estimation capability of the pretrained text-to-image model while preserving its original image generation ability. Compared to other unified models for image generation and depth estimation, MERGE achieves state-of-the-art performance across multiple depth estimation benchmarks. The code and model will be made available.

JBHI Journal 2025 Journal Article

Multi-Modal Disease Prediction With Hierarchical Self-Supervised Learning

  • Zhe Qu
  • Taihua Chen
  • Xin Zhou
  • Fanglin Zhu
  • Wei Guo
  • Yonghui Xu
  • Yixin Zhang
  • Lizhen Cui

The proliferation of healthcare data sources, including diverse imaging modalities and biochemical measurements, has created unprecedented opportunities for comprehensive disease prediction. Multi-modal clinical data, encompassing medical imaging reports, biochemical assays, and longitudinal clinical records, provides a rich foundation for developing sophisticated diagnostic models. Graph Neural Networks (GNNs) have emerged as a leading methodological framework, distinguished by their capacity to model complex inter-patient relationships and capture community structures within patient data. Despite their promise, current GNN-based approaches exhibit limitations in handling noisy, low-quality data and often impose overly restrictive graph smoothness constraints. These limitations can obscure patient-specific variations and compromise model robustness. To overcome these challenges, we propose HierSSL ( Hier archical S elf- S upervised L earning), a novel multi-modal disease prediction framework that enhances representational learning through dual-scale self-supervision mechanisms operating at both local and global levels. HierSSL's architecture specifically addresses two critical aspects: 1) the capture of local inter-modality dependencies and global community patterns, and 2) the optimization of multi-modal feature integration through an innovative combination of feature consistency constraints and graph contrastive learning. Empirical evaluation across two distinct disease prediction datasets demonstrates that HierSSL achieves statistically significant performance improvements compared to state-of-the-art methods, highlighting its efficacy in robust multi-modal data integration for disease prediction tasks.

TMLR Journal 2025 Journal Article

Remembering to Be Fair Again: Reproducing Non-Markovian Fairness in Sequential Decision Making

  • Domonkos Nagy
  • Lohithsai Yadala Chanchu
  • Krystof Bobek
  • Xin Zhou
  • Jacobus Smit

Ensuring long-term fairness in sequential decision-making is a key challenge in machine learning. Alamdari et al. (2024) introduced FairQCM, a reinforcement learning algorithm that enforces fairness in non-Markovian settings via memory augmentations and counterfactual reasoning. We reproduce and extend their findings by validating their claims and introducing novel enhancements. We confirm that FairQCM outperforms standard baselines in fairness enforcement and sample efficiency across different environments. However, alternative fairness metrics (Egalitarian, Gini) yield mixed results, and counterfactual memories show limited impact on fairness improvement. Further, we introduce a realistic COVID-19 vaccine allocation environment based on SEIR, a popular compartmental model of epidemiology. To accommodate continuous action spaces, we develop FairSCM, which integrates counterfactual memories into a Soft Actor-Critic framework. Our results reinforce that counterfactual memories provide little fairness benefit and, in fact, hurt performance, especially in complex, dynamic settings. The original code, modified to be 70% more efficient, and our extensions are available on GitHub: https://github.com/bozo22/remembering-to-be-fair-again.

IROS Conference 2025 Conference Paper

UniTac-NV: A Unified Tactile Representation For Non-Vision-Based Tactile Sensors *

  • Jian Hou 0019
  • Xin Zhou
  • Qihan Yang
  • Adam J. Spiers

Generalizable algorithms for tactile sensing remain underexplored, primarily due to the diversity of sensor modalities. Recently, many methods for cross-sensor transfer between optical (vision-based) tactile sensors have been investigated, yet little work focus on non-optical tactile sensors. To address this gap, we propose an encoder-decoder architecture to unify tactile data across non-vision-based sensors. By leveraging sensor-specific encoders, the framework creates a latent space that is sensor-agnostic, enabling cross-sensor data transfer with low errors and direct use in downstream applications. We leverage this network to unify tactile data from two commercial tactile sensors: the Xela uSkin uSPa 46 and the Contactile PapillArray. Both were mounted on a UR5e robotic arm, performing force-controlled pressing sequences against distinct object shapes (circular, square, and hexagonal prisms) and two materials (rigid PLA and flexible TPU). Another more complex unseen object was also included to investigate the model’s generalization capabilities. We show that alignment in latent space can be implicitly learned from joint autoencoder training with matching contacts collected via different sensors. We further demonstrate the practical utility of our approach through contact geometry estimation, where downstream models trained on one sensor’s latent representation can be directly applied to another without retraining.

ICRA Conference 2025 Conference Paper

Variable-Friction In-Hand Manipulation for Arbitrary Objects via Diffusion-Based Imitation Learning

  • Qiyang Yan
  • Zihan Ding
  • Xin Zhou
  • Adam J. Spiers

Dexterous in-hand manipulation (IHM) for arbitrary objects is challenging due to the rich and subtle contact process. Variable-friction manipulation is an alternative approach to dexterity, previously demonstrating robust and versatile 2D IHM capabilities with only two single-joint fingers. However, the hard-coded manipulation methods for variable friction hands are restricted to regular polygon objects and limited target poses, as well as requiring the policy to be tailored for each object. This paper proposes an end-to-end learning-based manipulation method to achieve arbitrary object manipulation for any target pose on real hardware, with minimal engineering efforts and data collection. The method features a diffusion policy-based imitation learning method with cotraining from simulation and a small amount of real-world data. With the proposed framework, arbitrary objects including polygons and non-polygons can be precisely manipulated to reach arbitrary goal poses within 2 hours of training on an A100 GPU and only 1 hour of real-world data collection. The precision is higher than previous customized object-specific policies, achieving an average success rate of 71. 3 % with average pose error being 2. 676 mm and 1. 902°. Code and videos can be found at: https://sites.google.com/view/vf-ihm-il/home.

NeurIPS Conference 2024 Conference Paper

A Unified Framework for 3D Scene Understanding

  • Wei Xu
  • Chunsheng Shi
  • Sifan Tu
  • Xin Zhou
  • Dingkang Liang
  • Xiang Bai

We propose UniSeg3D, a unified 3D scene understanding framework that achieves panoptic, semantic, instance, interactive, referring, and open-vocabulary segmentation tasks within a single model. Most previous 3D segmentation approaches are typically tailored to a specific task, limiting their understanding of 3D scenes to a task-specific perspective. In contrast, the proposed method unifies six tasks into unified representations processed by the same Transformer. It facilitates inter-task knowledge sharing, thereby promoting comprehensive 3D scene understanding. To take advantage of multi-task unification, we enhance performance by establishing explicit inter-task associations. Specifically, we design knowledge distillation and contrastive learning to transfer task-specific knowledge across different tasks. Experiments on three benchmarks, including ScanNet20, ScanRefer, and ScanNet200, demonstrate that the UniSeg3D consistently outperforms current SOTA methods, even those specialized for individual tasks. We hope UniSeg3D can serve as a solid unified baseline and inspire future work. Code and models are available at \url{https: //dk-liang. github. io/UniSeg3D/}.

AAAI Conference 2024 Conference Paper

Dual-View Whitening on Pre-trained Text Embeddings for Sequential Recommendation

  • Lingzi Zhang
  • Xin Zhou
  • Zhiwei Zeng
  • Zhiqi Shen

Recent advances in sequential recommendation models have demonstrated the efficacy of integrating pre-trained text embeddings with item ID embeddings to achieve superior performance. However, our study takes a unique perspective by exclusively focusing on the untapped potential of text embeddings, obviating the need for ID embeddings. We begin by implementing a pre-processing strategy known as whitening, which effectively transforms the anisotropic semantic space of pre-trained text embeddings into an isotropic Gaussian distribution. Comprehensive experiments reveal that applying whitening to pre-trained text embeddings in sequential recommendation models significantly enhances performance. Yet, a full whitening operation might break the potential manifold of items with similar text semantics. To retain the original semantics while benefiting from the isotropy of the whitened text features, we propose a Dual-view Whitening method for Sequential Recommendation (DWSRec), which leverages both fully whitened and relaxed whitened item representations as dual views for effective recommendations. We further examine the advantages of our approach through both empirical and theoretical analyses. Experiments on three public benchmark datasets show that DWSRec outperforms state-of-the-art methods for sequential recommendation.

AIIM Journal 2024 Journal Article

MSEF-Net: Multi-scale edge fusion network for lumbosacral plexus segmentation with MR image

  • Junyong Zhao
  • Liang Sun
  • Zhi Sun
  • Xin Zhou
  • Haipeng Si
  • Daoqiang Zhang

Nerve damage of spine areas is a common cause of disability and paralysis. The lumbosacral plexus segmentation from magnetic resonance imaging (MRI) scans plays an important role in many computer-aided diagnoses and surgery of spinal nerve lesions. Due to the complex structure and low contrast of the lumbosacral plexus, it is difficult to delineate the regions of edges accurately. To address this issue, we propose a Multi-Scale Edge Fusion Network (MSEF-Net) to fully enhance the edge feature in the encoder and adaptively fuse multi-scale features in the decoder. Specifically, to highlight the edge structure feature, we propose an edge feature fusion module (EFFM) by combining the Sobel operator edge detection and the edge-guided attention module (EAM), respectively. To adaptively fuse the multi-scale feature map in the decoder, we introduce an adaptive multi-scale fusion module (AMSF). Our proposed MSEF-Net method was evaluated on the collected spinal MRI dataset with 89 patients (a total of 2848 MR images). Experimental results demonstrate that our MSEF-Net is effective for lumbosacral plexus segmentation with MR images, when compared with several state-of-the-art segmentation methods.

NeurIPS Conference 2024 Conference Paper

PointMamba: A Simple State Space Model for Point Cloud Analysis

  • Dingkang Liang
  • Xin Zhou
  • Wei Xu
  • Xingkui Zhu
  • Zhikang Zou
  • Xiaoqing Ye
  • Xiao Tan
  • Xiang Bai

Transformers have become one of the foundational architectures in point cloud analysis tasks due to their excellent global modeling ability. However, the attention mechanism has quadratic complexity, making the design of a linear complexity method with global modeling appealing. In this paper, we propose PointMamba, transferring the success of Mamba, a recent representative state space model (SSM), from NLP to point cloud analysis tasks. Unlike traditional Transformers, PointMamba employs a linear complexity algorithm, presenting global modeling capacity while significantly reducing computational costs. Specifically, our method leverages space-filling curves for effective point tokenization and adopts an extremely simple, non-hierarchical Mamba encoder as the backbone. Comprehensive evaluations demonstrate that PointMamba achieves superior performance across multiple datasets while significantly reducing GPU memory usage and FLOPs. This work underscores the potential of SSMs in 3D vision-related tasks and presents a simple yet effective Mamba-based baseline for future research. The code is available at https: //github. com/LMD0311/PointMamba.

TMLR Journal 2024 Journal Article

Sequential Best-Arm Identification with Application to P300 Speller

  • Xin Zhou
  • Botao Hao
  • Tor Lattimore
  • Jian Kang
  • Lexin Li

A brain-computer interface (BCI) is an advanced technology that facilitates direct communication between the human brain and a computer system, by enabling individuals to interact with devices using only their thoughts. The P300 speller is a primary type of BCI system, which allows users to spell words without using a physical keyboard, but instead by capturing and interpreting brain electroencephalogram (EEG) signals under different stimulus presentation paradigms. Traditional non-adaptive presentation paradigms, however, treat each word selection as an isolated event, resulting in a lengthy learning process. To enhance efficiency, we cast the problem as a sequence of best-arm identification tasks within the context of multi-armed bandits, where each task corresponds to the interaction between the user and the system for a single character or word. Leveraging large language models, we utilize the prior knowledge learned from previous tasks to inform and facilitate subsequent tasks. We propose a sequential top-two Thompson sampling algorithm under two scenarios: the fixed-confidence setting and the fixed-budget setting. We study the theoretical property of the proposed algorithm, and demonstrate its substantial empirical improvement through both simulations as well as the data generated from a P300 speller simulator that was built upon the real BCI experiments.

ICLR Conference 2023 Conference Paper

Do We Really Need Complicated Model Architectures For Temporal Networks?

  • Weilin Cong
  • Si Zhang
  • Jian Kang 0008
  • Baichuan Yuan
  • Hao Wu
  • Xin Zhou
  • Hanghang Tong
  • Mehrdad Mahdavi

Recurrent neural network (RNN) and self-attention mechanism (SAM) are the de facto methods to extract spatial-temporal information for temporal graph learning. Interestingly, we found that although both RNN and SAM could lead to a good performance, in practice neither of them is always necessary. In this paper, we propose GraphMixer, a conceptually and technically simple architecture that consists of three components: (1) a link-encoder that is only based on multi-layer perceptrons (MLP) to summarize the information from temporal links, (2) a node-encoder that is only based on neighbor mean-pooling to summarize node information, and (3) an MLP-based link classifier that performs link prediction based on the outputs of the encoders. Despite its simplicity, GraphMixer attains an outstanding performance on temporal link prediction benchmarks with faster convergence and better generalization performance. These results motivate us to rethink the importance of simpler model architecture.

IROS Conference 2023 Conference Paper

InstaGrasp: An Entirely 3D Printed Adaptive Gripper with TPU Soft Elements and Minimal Assembly Time

  • Xin Zhou
  • Adam J. Spiers

Fabricating existing and popular open-source adaptive robotic grippers commonly involves using multiple professional machines, purchasing a wide range of parts, and tedious, time-consuming assembly processes. This poses a significant barrier to entry for some robotics researchers and drives others to opt for expensive commercial alternatives. To provide both parties with an easier and cheaper (under £100) solution, we propose a novel adaptive gripper design where every component (with the exception of actuators and the screws that come packaged with them) can be fabricated on a hobby-grade 3D printer, via a combination of inexpensive and readily available PLA and TPU filaments. This approach means that the gripper's tendons, flexure joints and finger pads are now printed, as a replacement for traditional string-tendons and molded urethane flexures / pads. A push-fit systems results in an assembly time of under 10 minutes. The gripper design is also highly modular and requires only a few minutes to replace any part, leading to extremely user-friendly maintenance and part modifications. An extensive stress test has shown a level of durability more than suitable for research, whilst grasping experiments (with perturbations) using items from the YCB object set has also proven its mechanical adaptability to be highly satisfactory.

JBHI Journal 2023 Journal Article

Interactive Skin Wound Segmentation Based on Feature Augment Networks

  • Pengfei Zhang
  • Xinjian Chen
  • Ziting Yin
  • Xin Zhou
  • Qingxin Jiang
  • Weifang Zhu
  • Dehui Xiang
  • Yun Tang

Skin wound segmentation in photographs allows non-invasive analysis of wounds that supports dermatological diagnosis and treatment. In this paper, we propose a novel feature augment network (FANet) to achieve automatic segmentation of skin wounds, and design an interactive feature augment network (IFANet) to provide interactive adjustment on the automatic segmentation results. The FANet contains the edge feature augment (EFA) module and the spatial relationship feature augment (SFA) module, which can make full use of the notable edge information and the spatial relationship information be-tween the wound and the skin. The IFANet, with FANet as the backbone, takes the user interactions and the initial result as inputs, and outputs the refined segmentation result. The pro-posed networks were tested on a dataset composed of miscellaneous skin wound images, and a public foot ulcer segmentation challenge dataset. The results indicate that the FANet gives good segmentation results while the IFANet can effectively improve them based on simple marking. Comprehensive comparative experiments show that our proposed networks outperform some other existing automatic or interactive segmentation methods, respectively.

ICRA Conference 2023 Conference Paper

Tactile Identification of Object Shapes via In-Hand Manipulation with A Minimalistic Barometric Tactile Sensor Array

  • Xin Zhou
  • Adam J. Spiers

With the goal of providing an alternative to optical and other tactile sensors, we set out to stress test the object shape identification capabilities of barometric tactile arrays in robotic manipulation tasks. These sensors are superior to optical devices in terms of form factor, ease of fabrication, and data reading/processing speeds, but lack the necessary spatial resolution to identify surface shapes via a single contact. To compensate, we utilize in-hand-manipulation, specifically in- hand-rolling to identify object shapes via a spatiotemporal approach. To increase task difficulty, we only use three neighboring barometric sensors and designed strict experiment requirements with the purpose of creating a set of extremely confusable test objects. The E- TRoll robotic hand, equipped with a barometric tactile array on one finger, was used to roll test objects within its grasp, taking just under 3. 4 seconds for data collection under the fastest tested speed setting, compared to 33 seconds in our previous work. We also designed and implemented a feature extraction algorithm, based and improved upon our recently published algorithm. This captures enough information from the collected spatiotemporal data samples for successful classification with only 13 features. Finally, a bagged tree classification algorithm was trained and optimized with data from 1, 164 trials of rolling 9 prismatic test objects, leading to a five-fold cross validation accuracy of 90. 5% for identifying the 9 object classes.

IROS Conference 2022 Conference Paper

E-TRoll: Tactile Sensing and Classification via A Simple Robotic Gripper for Extended Rolling Manipulations

  • Xin Zhou
  • Adam J. Spiers

Robotic tactile sensing provides a method of recognizing objects and their properties where vision fails. Prior work on tactile perception in robotic manipulation has frequently focused on exploratory procedures (EPs). However, the also-human-inspired technique of in-hand-manipulation can glean rich data in a fraction of the time of EPs. We propose a simple 3-DOF robotic hand design, optimized for object rolling tasks via a variable-width palm and associated control system. This system dynamically adjusts the distance between the finger bases in response to object behavior. Compared to fixed finger bases, this technique significantly increases the area of the object that is exposed to finger-mounted tactile arrays during a single rolling motion (an increase of over 60% was observed for a cylinder with a 30-millimeter diameter). In addition, this paper presents a feature extraction algorithm for the collected spatiotemporal dataset, which focuses on object corner identification, analysis, and compact representation. This technique drastically reduces the dimensionality of each data sample from $\boldsymbol{10\times 1500}$ time series data to 80 features, which was further reduced by Principal Component Analysis (PCA) to 22 components. An ensemble subspace k-nearest neighbors (KNN) classification model was trained with 90 observations on rolling three different geometric objects, resulting in a three-fold cross-validation accuracy of 95. 6% for object shape recognition.

IJCAI Conference 2022 Conference Paper

Searching for Optimal Subword Tokenization in Cross-domain NER

  • Ruotian Ma
  • Yiding Tan
  • Xin Zhou
  • Xuanting Chen
  • Di Liang
  • Sirui Wang
  • Wei Wu
  • Tao Gui

Input distribution shift is one of the vital problems in unsupervised domain adaptation (UDA). The most popular UDA approaches focus on domain-invariant representation learning, trying to align the features from different domains into a similar feature distribution. However, these approaches ignore the direct alignment of input word distributions between domains, which is a vital factor in word-level classification tasks such as cross-domain NER. In this work, we shed new light on cross-domain NER by introducing a subword-level solution, X-Piece, for input word-level distribution shift in NER. Specifically, we re-tokenize the input words of the source domain to approach the target subword distribution, which is formulated and solved as an optimal transport problem. As this approach focuses on the input level, it can also be combined with previous DIRL methods for further improvement. Experimental results show the effectiveness of the proposed method based on BERT-tagger on four benchmark NER datasets. Also, the proposed method is proved to benefit DIRL methods such as DANN.

ICRA Conference 2021 Conference Paper

No Need for Interactions: Robust Model-Based Imitation Learning using Neural ODE

  • HaoChih Lin
  • Baopu Li
  • Xin Zhou
  • Jiankun Wang 0001
  • Max Q. -H. Meng

Interactions with either environments or expert policies during training are needed for most of the current imitation learning (IL) algorithms. For IL problems with no interactions, a typical approach is Behavior Cloning (BC). However, BC-like methods tend to be affected by distribution shift. To mitigate this problem, we come up with a Robust Model-Based Imitation Learning (RMBIL) framework that casts imitation learning as an end-to-end differentiable nonlinear closed-loop tracking problem. RMBIL applies Neural ODE to learn a precise multi-step dynamics and a robust tracking controller via Nonlinear Dynamics Inversion (NDI) algorithm. Then, the learned NDI controller will be combined with a trajectory generator, a conditional VAE, to imitate an expert’s behavior. Theoretical derivation shows that the controller network can approximate an NDI when minimizing the training loss of Neural ODE. Experiments on Mujoco tasks also demonstrate that RMBIL is competitive to the state-of-the-art generative adversarial method (GAIL) and achieves at least 30% performance gain over BC in uneven surfaces.

YNIMG Journal 2016 Journal Article

Direct detection of optogenetically evoked oscillatory neuronal electrical activity in rats using SLOE sequence

  • Yuhui Chai
  • Guoqiang Bi
  • Liping Wang
  • Fuqiang Xu
  • Ruiqi Wu
  • Xin Zhou
  • Bensheng Qiu
  • Hao Lei

The direct detection of neuronal electrical activity is one of the most challenging goals in non-BOLD fMRI research. Previous work has demonstrated its feasibility in phantom and cell culture studies, but attempts in in vivo studies remain few and far between. Most recent in vivo studies used T2*-weighted sequences to directly detect neuronal electrical activity evoked by sensory stimulus. As neuronal electrical signal is usually comprised of a series of spectrally distributed oscillatory waveforms rather than being a direct current, it is most likely to be detected using oscillatory current sensitive sequences. In this study, we explored the potential of using the spin-lock oscillatory excitation (SLOE) sequence with spiral readout to directly detect optogenetically evoked oscillatory neuronal electrical activity, whose main spectral component can be manipulated artificially to match the resonance frequency of spin-lock RF field. In addition, experiments using the stimulus-induced rotary saturation (SIRS) sequence with spiral readout were also performed. Electrophysiological recording and MRI data acquisition were conducted on separate animals. Robust optogenetically evoked oscillatory LFP signals were observed and significant BOLD signals were acquired with the GE-EPI sequence before and after the whole SLOE and SIRS acquisitions, but no significant neuronal current MRI (ncMRI) signal changes were detected. These results indicate that the sensitivity of oscillatory current sensitive sequences needs to be further improved for direct detection of neuronal electrical activity.

IROS Conference 2006 Conference Paper

An Autonomous Robotic Fish for Mobile Sensing

  • Xiaobo Tan
  • Drew Kim
  • Nathan Usher
  • Dan Laboy
  • Joel Jackson
  • Azra Kapetanovic
  • Jason Rapai
  • Benjamin Sabadus

In this paper an innovative approach to robotics education is reported, where hands-on learning is integrated with cutting-edge research in the development of an autonomous, biomimetic robotic fish. The project aims to develop an energy-efficient, noiseless, untethered swimming robot for mobile sensing purposes. The robot is propelled by an ionic polymer-metal composite (IPMC) actuator and equipped with a GPS receiver, a ZigBee wireless communication module, a microcontroller, and a temperature sensor for autonomous navigation, control, and sensing. The two phases of the development are described, emphasizing both the technical approaches and the learning paradigms. The developed robotic fish will be further used as an educational kit for K-12 students and as a research tool for investigating multi-robot collaborative sensing

v2026.09.13