Arrow Research search

Author name cluster

Jian Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

126 papers
2 author rows

Possible papers

126

EAAI Journal 2026 Journal Article

A hybrid couple-task surrogate operator with Fourier space–time encoding and multi-field attention for solving three-dimensional reservoir seepage equations in heterogeneous media

  • Ren-Yao Lin
  • Tao Song
  • Jian Li

Reservoir seepage modeling involves solving high-dimensional, multi-physics coupled partial differential equations, where both efficient computation and stable quality are critical for engineering applications. Neural operators offer higher computational efficiency than numerical solvers. However, in multi-physics coupling prediction scenarios, existing methods struggle to characterize complex spatiotemporal correlations and capture the coupling evolution characteristics between multiple fluid fields, while also facing challenges such as high computational and storage costs. To address these issues, this paper proposes a coupled task surrogate operator network combining Fourier spatiotemporal encoding and a coupled attention mechanism, aiming to reliably solve the three-dimensional multi-field coupled reservoir seepage equations. This operator first uses Fourier transform to encode spatiotemporal information to enhance the representation ability of data across spatiotemporal scales. Second, it utilizes a coupled task architecture to share spatiotemporal features to learn the potential correlations between different physical fields and the simultaneous spatial domain. Finally, an adaptive multi-field coupled attention mechanism is employed to enhance the ability to capture crucial correlation information and spatiotemporal changes between multiple physical fields in reservoir fluids. Experiments validate the operator’s effectiveness in terms of Fourier encoding, coupled task, and applicability to multiple scenarios. The results show that the proposed model more accurately describes the nonlinear behavior of reservoir fluids: Fourier coding enhancement reduces prediction error by over 95. 0% compared to the baseline model; combining coupled attention and a multi-field prediction framework not only improves the quality of multi-field predictions but also reduces inference time by 70. 9% and memory consumption by 22. 0%; under various types of Gaussian stochastic initial state fields, the operator exhibits stronger robustness and stability compared to the representative benchmark models, reducing multi-field errors by an average of 70. 0%. This study provides a new approach with engineering application potential for efficient and robust data-driven simulation of three-dimensional reservoir seepage problems.

EAAI Journal 2026 Journal Article

A robust and interpretable framework for sports activity recognition based on wearable sensor signals and image representations

  • Jian Li
  • Yibo Fan
  • Junhui Gong
  • Junyi Chen
  • Ruoyu Chen
  • Wenyan Zhang
  • Yuliang Zhao

Human activity recognition (HAR) using wearable sensors has advanced rapidly, improving the precision of complex movement identification. However, existing methods rely on single-modal time-series features, limiting spatiotemporal representation and global dependency capture. This hinders dynamic characterization and reduces recognition accuracy. To overcome these limitations, this paper proposes a novel multimodal fusion deep learning approach to enhance complex action recognition. First, we adopt a multimodal input strategy that integrates time-series data and Gramian Angular Difference Field (GADF) images to comprehensively capture the spatiotemporal characteristics of motion data. Second, we design a dual-stream feature fusion network, where Bidirectional Gated Recurrent Unit (BiGRU) combined with a multi-head self-attention mechanism (MSA) is employed to extract time-series features, while Efficient Channel Attention (ECA) and residual block are utilized to enhance image feature representation, effectively leveraging the complementary information across modalities. Finally, we introduce an interpretable analysis method based on submodule optimization, enabling cross-modal attribution analysis to identify key regions contributing to model decisions for both time-series and image features. Experimental results demonstrate that the proposed method achieves an accuracy of 96. 88% in a 16-class sports activity recognition task, significantly outperforming traditional machine learning methods and existing deep learning models. This study provides an effective solution for complex action recognition and lays a technological foundation for real-time motion monitoring in wearable smart devices and broader HAR applications.

JBHI Journal 2026 Journal Article

CFRAFN: A Cross-Feature Residual Attention Fusion Network for Major Depressive Disorder Prediction Using Clinical Voice Recordings

  • Rumo Pan
  • Sidu Feng
  • Yi Sun
  • Jinqiu Xu
  • Tianzhang Zhai
  • Xiaochun Wu
  • Liangliang Tan
  • Yonggui Yuan

Major depressive disorder (MDD) is a prevalent mental disorder with a significant burden on individuals and society, and timely identification and intervention are essential for effective management. Voice data have been used as behavioral indicators of MDD, offering valuable insights into an individual's mental state. In this study, we collected voice data from 221 patients diagnosed with MDD at the inpatient ward of the Department of Psychiatry and Psychosomatics, Zhongda Hospital, Southeast University, alongside 113 healthy controls, to construct the Chinese depressive voice dataset. We proposed the cross-feature residual attention fusion network (CFRAFN), which leverages extended Geneva minimalistic acoustic parameter set features along with high-dimensional embeddings extracted from the pretrained VGGish model to effectively capture MDD-associated phonetic patterns. Specifically, CFRAFN utilizes differentiated residual blocks to maintain training stability in deep hierarchical structure. Furthermore, the self-attention fusion strategy dynamically weighted the significance of each feature modality, ensuring effective feature integration and consequently improving MDD prediction accuracy. Experimental results demonstrated that CFRAFN achieved an excellent predictive performance with an area under the receiver operating characteristic curve of 0. 924 in an independent test set, and significantly outperformed 11 baseline models across 5-fold cross-validation.

AAAI Conference 2026 Conference Paper

Deep Incomplete Multi-View Clustering via Hierarchical Imputation and Alignment

  • Yiming Du
  • Ziyu Wang
  • Jian Li
  • Rui Ning
  • Lusi Li

Incomplete multi-view clustering (IMVC) aims to discover shared cluster structures from multi-view data with partial observations. The core challenges lie in accurately imputing missing views without introducing bias, while maintaining semantic consistency across views and compactness within clusters. To address these challenges, we propose DIMVC-HIA, a novel deep IMVC framework that integrates hierarchical imputation and alignment with four key components: (1) view-specific autoencoders for latent feature extraction, coupled with a view-shared clustering predictor to produce soft cluster assignments; (2) a hierarchical imputation module that first estimates missing cluster assignments based on cross-view contrastive similarity, and then reconstructs missing features using intra-view, intra-cluster statistics; (3) an energy-based semantic alignment module, which promotes intra-cluster compactness by minimizing energy variance around low-energy cluster anchors; and (4) a contrastive assignment alignment module, which enhances cross-view consistency and encourages confident, well-separated cluster predictions. Experiments on benchmarks demonstrate that our framework achieves superior performance under varying levels of missingness.

EAAI Journal 2026 Journal Article

Detection and localization of false data injection attacks based on multi-scale feature fusion and attention enhancement network in smart grid

  • Jian Li
  • Hanting Lu
  • Qingyu Su

This study proposes a novel framework based on the multi-scale feature fusion and attention enhancement network (MSFF-AEN) for detecting and localizing false data injection attacks (FDIAs) in smart grid. The model innovatively designs improved residual block with convolutional block attention module (CBAM) after the second convolutional layer, reducing early noise interference, and enhancing interpretability. It also incorporates a bidirectional long short-term memory network (BiLSTM) and multi-head attention (MHA) to capture temporal features and global dependencies respectively. Additionally, hierarchical feature fusion (HFF) with learnable weights optimizes and integrates multi-scale features, thereby enhancing feature representation and model interpretability. Experimental results on the IEEE 14-bus and IEEE 118-bus systems show that the proposed model outperforms existing conventional models and deep learning methods across multiple evaluation metrics, including accuracy, precision, recall, and F1-score. Particularly, the model performs exceptionally well on the large-scale IEEE 118-bus power system, achieving an accuracy of 98. 73%, precision of 98. 48%, recall of 97. 45%, and F1-score of 97. 95%. Furthermore, the model demonstrates strong robustness to various Gaussian noise conditions, maintaining high localization accuracy.

AAAI Conference 2026 Conference Paper

Gentle Manipulation Policy Learning via Demonstrations from VLM Planned Atomic Skills

  • Jiayu Zhou
  • Qiwei Wu
  • Jian Li
  • Zhe Chen
  • Xiaogang Xiong
  • Renjing Xu

Autonomous execution of long-horizon, contact-rich manipulation tasks traditionally requires extensive real-world data and expert engineering, posing significant cost and scalability challenges. This paper proposes a novel framework integrating hierarchical semantic decomposition, reinforcement learning (RL), visual language models (VLMs), and knowledge distillation to overcome these limitations. Complex tasks are decomposed into atomic skills, with RL-trained policies for each primitive exclusively in simulation. Crucially, our RL formulation incorporates explicit force constraints to prevent object damage during delicate interactions. VLMs perform high-level task decomposition and skill planning, generating diverse expert demonstrations. These are distilled into a unified policy via Visual-Tactile Diffusion Policy for end-to-end execution. We conduct comprehensive ablation studies exploring different VLM-based task planners to identify optimal demonstration generation pipelines, and systematically compare imitation learning algorithms for skill distillation. Extensive simulation experiments and physical deployment validate that our approach achieves policy learning for long-horizon manipulation without costly human demonstrations, while the VLM-guided atomic skill framework enables scalable generalization to diverse tasks.

AAAI Conference 2026 Conference Paper

Kronos: A Foundation Model for the Language of Financial Markets

  • Yu Shi
  • Zongliang Fu
  • Shuo Chen
  • Bohan Zhao
  • Wei Xu
  • Changshui Zhang
  • Jian Li

The success of large-scale pre-training paradigm, exemplified by Large Language Models (LLMs), has inspired the development of Time Series Foundation Models (TSFMs). However, their application to financial candlestick (K-line) data remains limited, often underperforming non-pre-trained architectures. Moreover, existing TSFMs often overlook crucial downstream tasks such as volatility prediction and synthetic data generation. To address these limitations, we propose Kronos, a unified, scalable pre-training framework tailored to financial K-line modeling. Kronos introduces a specialized tokenizer that discretizes continuous market information into token sequences, preserving both price dynamics and trade activity patterns. We pre-train Kronos using an autoregressive objective on a massive, multi-market corpus of over 12 billion K-line records from 45 global exchanges, enabling it to learn nuanced temporal and cross-asset representations. Kronos excels in a zero-shot setting across a diverse set of financial tasks. On benchmark datasets, Kronos boosts price series forecasting RankIC by 93% over the leading TSFM and 87% over the best non-pre-trained baseline. It also achieves a 9% lower MAE in volatility forecasting and a 22% improvement in generative fidelity for synthetic K-line sequences. These results establish Kronos as a robust, versatile foundation model for end-to-end financial time series analysis.

AAAI Conference 2026 Conference Paper

L2V-CoT: Cross-Modal Transfer of Chain-of-Thought Reasoning via Latent Intervention

  • Yu-Liang Zhan
  • Xinyu Tang
  • Han Wan
  • Jian Li
  • Jirong Wen
  • Hao Sun

Recently, Chain-of-Thought (CoT) reasoning has significantly enhanced the capabilities of large language models (LLMs), but Vision–Language Models (VLMs) still struggle with multi-step reasoning tasks due to limited multimodal reasoning data. To bridge this gap, researchers have explored methods to transfer CoT reasoning from LLMs to VLMs. However, existing approaches either need high training costs or require architectural alignment. In this paper, we use Linear Artificial Tomography (LAT) to empirically show that LLMs and VLMs share similar low-frequency latent representations of CoT reasoning despite architectural differences. Based on this insight, we propose L2V-CoT, a novel training-free latent intervention approach that transfers CoT reasoning from LLMs to VLMs. L2V-CoT extracts and resamples low-frequency CoT representations from LLMs in the frequency domain, enabling dimension matching and latent injection into VLMs during inference to enhance reasoning capabilities. Extensive experiments demonstrate that our approach consistently outperforms training-free baselines and even surpasses supervised methods.

AAAI Conference 2026 Conference Paper

LLM-Oriented Token-Adaptive Knowledge Distillation

  • Xurong Xie
  • Zhucun Xue
  • Jiafu Wu
  • Jian Li
  • Yabiao Wang
  • Xiaobin Hu
  • Yong Liu
  • Jiangning Zhang

Knowledge Distillation (KD) is a key technique for compressing Large-scale Language Models (LLMs), but prevailing logit-based methods employ static strategies misaligned with the student’s dynamic learning process. By treating all tokens indiscriminately with a fixed temperature, these methods result in suboptimal knowledge transfer. To address this, we propose LLM-oriented token-Adaptive Knowledge Distillation (AdaKD), a framework that adapts the distillation process to each token’s real-time learning state. AdaKD consists of two synergistic modules driven by a unified token difficulty metric. First, the Loss-driven Adaptive Token Focusing (LATF) module dynamically concentrates distillation on valuable tokens by monitoring the student’s learning stability. Second, Inverse Difficulty Temperature Scaling (IDTS) introduces a counterintuitive token-level temperature: low for difficult tokens to target error correction, and high for easy tokens to learn the teacher’s smooth output distribution for better generalization. As a plug-and-play framework, AdaKD consistently improves performance across diverse distillation methods, model architectures, and benchmarks.

AAAI Conference 2026 Conference Paper

Navigating the Alpha Jungle: An LLM-Powered MCTS Framework for Formulaic Alpha Factor Mining

  • Yu Shi
  • Yitong Duan
  • Jian Li

Alpha factor mining is pivotal in quantitative investment for identifying predictive signals from complex financial data. While traditional formulaic alpha mining relies on human expertise, contemporary automated methods, such as those based on genetic programming or reinforcement learning, often struggle with search inefficiency or yield alpha factors that are difficult to interpret. This paper introduces a novel framework that integrates Large Language Models (LLMs) with Monte Carlo Tree Search (MCTS) to overcome these limitations. Our framework leverages the LLM's instruction-following and reasoning capability to iteratively generate and refine symbolic alpha formulas within an MCTS-driven exploration. A key innovation is the guidance of MCTS exploration by rich, quantitative feedback from financial backtesting of each candidate factor, enabling efficient navigation of the vast search space. Furthermore, a frequent subtree avoidance mechanism is introduced to enhance search diversity and prevent formulaic homogenization, further improving performance. Experimental results on real-world stock market data demonstrate that our LLM-based framework outperforms existing methods by mining alphas with superior predictive accuracy and trading performance. The resulting formulas are also more amenable to human interpretation, establishing a more effective and efficient paradigm for formulaic alpha mining.

AAAI Conference 2026 Conference Paper

Predicting Video Slot Attention Queries from Random Slot-Feature Pairs

  • Rongzhen Zhao
  • Jian Li
  • Juho Kannala
  • Joni Pajarinen

Unsupervised video Object-Centric Learning (OCL) is promising as it enables object-level scene representation and understanding as we humans do. Mainstream video OCL methods adopt a recurrent architecture: An aggregator aggregates current video frame into object features, termed slots, under some queries; A transitioner transits current slots to queries for the next frame. This is an effective architecture but all existing implementations both (i1) neglect to incorporate next frame features, the most informative source for query prediction, and (i2) fail to learn transition dynamics, the knowledge essential for query prediction. To address these issues, we propose Random Slot-Feature pair for learning Query prediction (RandSF.Q): (t1) We design a new transitioner to incorporate both slots and features, which provides more information for query prediction; (t2) We train the transitioner to predict queries from slot-feature pairs randomly sampled from available recurrences, which drives it to learn transition dynamics. Experiments on scene representation demonstrate that our method surpass existing video OCL methods significantly, e.g., up to 10 points on object discovery, setting new state-of-the-art. Such superiority also benefits downstream tasks like scene understanding.

AAAI Conference 2026 Conference Paper

RadarMP: Motion Perception for 4D mmWave Radar in Autonomous Driving

  • Ruiqi Cheng
  • Huijun Di
  • Jian Li
  • Feng Liu
  • Wei Liang

Accurate 3D scene motion perception significantly enhances the safety and reliability of an autonomous driving system. Benefiting from its all-weather operational capability and unique perceptual properties, 4D mmWave radar has emerged as an essential component in advanced autonomous driving. However, sparse and noisy radar points often lead to imprecise motion perception, leaving autonomous vehicles with limited sensing capabilities when optical sensors degrade under adverse weather conditions. In this paper, we propose RadarMP, a novel method for precise 3D scene motion perception using low-level radar echo signals from two consecutive frames. Unlike existing methods that separate radar target detection and motion estimation, RadarMP jointly models both tasks in a unified architecture, enabling consistent radar point cloud generation and pointwise 3D scene flow prediction. Tailored to radar characteristics, we design specialized self-supervised loss functions guided by Doppler shifts and echo intensity, effectively supervising spatial and motion consistency without explicit annotations. Extensive experiments on the public dataset demonstrate that RadarMP achieves reliable motion perception across diverse weather and illumination conditions, outperforming radar-based decoupled motion perception pipelines and enhancing perception capabilities for full-scenario autonomous driving systems.

YNIMG Journal 2026 Journal Article

State-Level Brain Dynamics Reveal Neural Correlates of Negative-Mode Rigidity in Non-suicidal Self-injury

  • Jian Li
  • Jingying Lai
  • Fengmin Ni
  • Ya Xie
  • Enze Tang
  • Yibing Tang
  • Chun Wang

Non-suicidal self-injury (NSSI) is marked by persistent bias toward negatively valenced or salient internal experiences and difficulty disengaging from them once activated, yet the neural dynamics that support this clinical rigidity remain poorly understood. Intrinsic brain activity normally cycles among recurrent large-scale states, and alterations in how these states are occupied or transitioned may provide a neural analogue of negatively biased internal modes. Using resting-state fMRI from 160 patients with NSSI and 50 psychiatric controls, we applied Hidden Markov Modeling to characterize latent brain states and their temporal properties. NSSI was associated with disproportionate engagement of a recurrent ventral attention-related state, reduced differentiation between states, and greater variability within states. Greater dominance of this state was linked to more severe emotion-regulation difficulties at baseline and showed prognostic relevance for subsequent improvement in NSSI behaviors over three months. These findings indicate that NSSI involves a biased tendency to settle into salience- and attention-related brain states, highlighting attentional deficits as a clinically relevant feature of this condition. When considered alongside prior evidence showing heightened variability in connectivity strength and network topology, the results point to convergent disruptions in neural flexibility across multiple organizational levels in NSSI and underscore large-scale neural dynamics as a potentially informative target for future mechanistic and translational research.

AAAI Conference 2026 Conference Paper

UniMo: Unified Motion Generation and Understanding with Chain of Thought

  • Guocun Wang
  • Kenkun Liu
  • Jing Lin
  • Guorui Song
  • Jian Li
  • Xiaoguang Han

Existing 3D human motion generation and understanding methods often exhibit limited interpretability, restricting effective mutual enhancement between these inherently related tasks. While current unified frameworks based on large language models (LLMs) leverage linguistic priors, they frequently encounter challenges in semantic alignment and task coherence. Moreover, the next-token prediction paradigm in LLMs is ill-suited for motion sequences, causing cumulative prediction errors. To address these limitations, we propose UniMo, a novel framework that integrates motion-language information and interpretable chain of thought (CoT) reasoning into the LLM via supervised fine-tuning (SFT). We further introduce reinforcement learning with Group Relative Policy Optimization (GRPO) as a post-training strategy that optimizes over groups of tokens to enforce structural correctness and semantic alignment, mitigating cumulative errors in motion token prediction. Extensive experiments demonstrate that UniMo significantly outperforms existing unified and task-specific models, achieving state-of-the-art performance in both motion generation and understanding.

AAAI Conference 2025 Conference Paper

Adapting to Non-Stationary Environments: Multi-Armed Bandit Enhanced Retrieval-Augmented Generation on Knowledge Graphs

  • Xiaqiang Tang
  • Jian Li
  • Nan Du
  • Sihong Xie

Despite the superior performance of Large language models on many NLP tasks, they still face significant limitations in memorizing extensive world knowledge. Recent studies have demonstrated that leveraging the Retrieval-Augmented Generation (RAG) framework, combined with Knowledge Graphs that encapsulate extensive factual data in a structured format, robustly enhances the reasoning capabilities of LLMs. However, deploying such systems in real-world scenarios presents challenges: the continuous evolution of non-stationary environments may lead to performance degradation and user satisfaction requires a careful balance of performance and responsiveness. To address these challenges, we introduce a Multi-objective Multi-Armed Bandit enhanced RAG framework, supported by multiple retrieval methods with diverse capabilities under rich and evolving retrieval contexts in practice. Within this framework, each retrieval method is treated as a distinct "arm''. The system utilizes real-time user feedback to adapt to dynamic environments, by selecting the appropriate retrieval method based on input queries and the historical multi-objective performance of each arm. Extensive experiments conducted on two benchmark KGQA datasets demonstrate that our method significantly outperforms baseline methods in non-stationary settings while achieving state-of-the-art performance in station environments.

EAAI Journal 2025 Journal Article

Adaptive Deformable Convolutional Neural Network Framework for depression-related behavioral analysis in mice

  • Jian Li
  • Ziyi Li
  • Peng Shan
  • Xiaoyong Lyu
  • Yu Tian
  • Chen Du
  • Ying Wang
  • Yuliang Zhao

The use of approximately 1 billion laboratory animals annually in research highlights the urgent need for advanced methods to analyze behavioral dynamics, particularly in mice. Capturing subtle and prolonged behavioral changes, such as those observed in long-term depression studies, poses a significant challenge. To address this, we propose an Adaptive Deformable Convolutional Neural Network Framework for depression-related behavioral analysis in mice. By integrating DeepLabCut (DLC) with deformable convolutional networks (DCN) and convolutional block attention module (CBAM), the framework captures subtle and prolonged behavioral changes with high precision. Adaptive image deformation encodes joint movements into image representations, enabling robust analysis of spatial and temporal patterns. In depression modeling experiment, the framework achieved over 80% classification accuracy, demonstrating its scalability and efficiency. This non-invasive, automated solution represents a transformative advancement in behavioral analysis, offering a reliable tool for long-term studies in animal models.

FOCS Conference 2025 Conference Paper

Adaptivity Gaps for Stochastic Probing with Subadditive Functions

  • Jian Li
  • Yinchen Liu
  • Yiran Zhang

In this paper, we study the stochastic probing problem under a general monotone norm objective. We are given a ground set $U=[n]=\{1, 2, \ldots, n\}$, where each element i is associated with an independent nonnegative random variable $X_{i}$ (with a known distribution). We may probe these elements adaptively, and upon probing an element i, its value $X_{i}$ is realized. The sequence of probed elements must satisfy a prefix-closed feasibility constraint $\mathcal{F}$, such as a matroid, an orienteering constraint, or any other downward-closed constraint. We also have a monotone norm function $f: \mathbb{R}_{\geq 0}^{n} \rightarrow \mathbb{R}_{\geq 0}$. Let $P \subseteq U$ be the set of probed elements. Then the reward is $f\left(X_{P}\right)$, where $X_{P}$ is an n-dimensional vector whose i-th coordinate equals the realized value of $X_{i}$ if $i \in P$ (i. e. , element i is probed), and 0 otherwise. Our objective is to design a probing strategy that maximizes the expected reward $\mathbb{E}\left[f\left(X_{P}\right)\right]$. We study the adaptivity gap of the problem, defined as the ratio between the expected reward of an optimal adaptive strategy and that of an optimal non-adaptive strategy. A small adaptivity gap allows us to focus on designing nonadaptive strategies, which are typically simpler to represent and analyze. Establishing tight adaptivity gaps is a central challenge in stochastic combinatorial optimization and has been studied extensively for stochastic probing problems with various objective functions. In this paper, we resolve a central open problem in this line of research, posed in [1], [2], by proving that the adaptivity gap for stochastic probing with general monotone norms is bounded by $O\left(\log ^{2} n\right)$. With a refined analysis, we can further strengthen the bound to $O(\log r \log n / \log \log n)$ where r is the maximum length of a sequence in the feasibility constraint ($2 \leq r \leq n$). As a by-product, we obtain an asymptotically tight adaptivity gap $\Theta(\log n / \log \log n)$ for Bernoulli stochastic probing with binary-XOS objectives, matching the lower bound in [1]. We also obtain an $O\left(\log ^{3} n\right)$ upper bound for Bernoulli stochastic probing with general subadditive objectives. Furthermore, for monotone symmetric norms, we prove that the adaptivity gap can be bounded by $O(1)$, answering an open question posed in [3] and improving upon their $O(\log n)$ upper bound. Index Terms-stochastic probing, adaptivity gap, subadditive objective

EAAI Journal 2025 Journal Article

An extended multi-criteria group decision-making method based on preference ranking under Z-number environments

  • Jian Li
  • Yuanyuan Xiang
  • Honggang Peng
  • Jianqiang Wang

Compared with traditional fuzzy numbers, using Z-numbers to illustrate fuzzy events offers two key advantages. It makes fuzzy events more intuitive for decision-makers, and the second component of Z-numbers acts as a measure of the reliability of the first component. While significant progress has been made on Z-numbers, from the theoretical and practical perspectives, some gaps remain. For example, scholars have seldom focused on the likelihood of Z-numbers, most existing decision-making methods with Z-numbers rarely consider the consensus-reaching processes, and little has been reported on the superior ordering methods in the Z-number environment. To overcome these limitations, first, the likelihood of Z-numbers is defined in combination with the preference ranking organization method for enrichment evaluations (PROMETHEE) type V preference function. Second, the ordering rules for PROMETHEE are discussed. Then, a procedure of the feedback-adjustment method is introduced to help the consensus level of group-alternative ranking reach the threshold. On these bases, an extended PROMETHEE multi-criteria group decision-making method with Z-numbers is proposed. Finally, to verify the feasibility and effectiveness of the proposed method, we examined an intelligent medical-diagnostic-system selection problem and conducted a comparison analysis. We applied the Z-number PROMETHEE approach to a challenging case study requiring a dual-data-driven application. Furthermore, the study suggests future directions for improving the proposed framework in other related contexts.

AAAI Conference 2025 Conference Paper

Block-Based Multi-Scale Image Rescaling

  • Jian Li
  • Siwang Zhou

Image rescaling (IR) seeks to determine the optimal low-resolution (LR) representation of a high-resolution (HR) image to reconstruct a high-quality super-resolution (SR) image. Typically, HR images with resolutions exceeding 2K possess rich information that is unevenly distributed across the image. Traditional image rescaling methods often fall short because they focus solely on the overall scaling rate, ignoring the varying amounts of information in different parts of the image. To address this limitation, we propose a Block-Based Multi-Scale Image Rescaling Framework (BBMR), tailored for IR tasks involving HR images of 2K resolution and higher. BBMR consists of two main components: the Downscaling Module and the Upscaling Module. In the Downscaling Module, the HR image is segmented into sub-blocks of equal size, with each sub-block receiving a dynamically allocated scaling rate while maintaining a constant overall scaling rate. For the Upscaling Module, we introduce the Joint Super-Resolution method (JointSR), which performs SR on these sub-blocks with varying scaling rates and effectively eliminates blocking artifacts. Experimental results demonstrate that BBMR significantly enhances the SR image quality in the of 2K and 4K test dataset compared to initial network image rescaling methods.

NeurIPS Conference 2025 Conference Paper

Constrained Feedback Learning for Non-Stationary Multi-Armed Bandits

  • Shaoang Li
  • Jian Li

Non-stationary multi-armed bandits (nsMAB) enable agents to adapt to changing environments by incorporating mechanisms to detect and respond to shifts in reward distributions, making them well-suited for dynamic settings. However, existing approaches typically assume that reward feedback is available at every round—an assumption that overlooks many real-world scenarios where feedback is limited. In this paper, we take a significant step forward by introducing a new model of *constrained feedback in non-stationary multi-armed bandits* (ConFee-nsMAB), where the availability of reward feedback is restricted. We propose the first prior-free algorithm—that is, one that does not require prior knowledge of the degree of non-stationarity—that achieves near-optimal dynamic regret in this setting. Specifically, our algorithm attains a dynamic regret of $\tilde {\mathcal{O}}({K^{1/3} V_T^{1/3} T }/{ B^{1/3}})$, where $T$ is the number of rounds, $K$ is the number of arms, $B$ is the query budget, and $V_T$ is the variation budget capturing the degree of non-stationarity.

AAAI Conference 2025 Conference Paper

Decentralized Federated Learning with Model Caching on Mobile Agents

  • Xiaoyu Wang
  • Guojun Xiong
  • Houwei Cao
  • Jian Li
  • Yong Liu

Federated Learning (FL) trains a shared model using data and computation power on distributed agents coordinated by a central server. Decentralized FL (DFL) utilizes local model exchange and aggregation between agents to reduce the communication and computation overheads on the central server. However, when agents are mobile, the communication opportunity between agents can be sporadic, largely hindering the convergence and accuracy of DFL. In this paper, we propose Cached Decentralized Federated Learning (Cached-DFL) to investigate delay-tolerant model spreading and aggregation enabled by model caching on mobile agents. Each agent stores not only its own model, but also models of agents encountered in the recent past. When two agents meet, they exchange their own models as well as the cached models. Local model aggregation utilizes all models stored in the cache. We theoretically analyze the convergence of Cached-DFL, explicitly taking into account the model staleness introduced by caching. We design and compare different model caching algorithms for different DFL and mobility scenarios. We conduct detailed case studies in a vehicular network to systematically investigate the interplay between agent mobility, cache staleness, and model convergence. In our experiments, Cached-DFL converges quickly, and significantly outperforms DFL without caching.

IROS Conference 2025 Conference Paper

Edge-Guided Lighting Adaptation: Real-Time Detection of Transparent Objects for Cell Culture Robot

  • Qingze Huang
  • Peng Wang
  • Xiangyan Zhang
  • Jian Li
  • Shimin Wei

In robot-assisted cell culture tasks, fluctuations in lighting conditions can result in blurred boundaries, intensified reflections, and pronounced refractions of transparent objects. These optical phenomena collectively escalate the complexity of image processing and target recognition. To address these challenges, this paper takes a dual-strategy approach. Firstly, it utilizes the Unity platform to construct a synthetic dataset (STTO-9k) containing 9, 000 images of six types of transparent objects, providing abundant training samples for the detection and recognition of transparent objects. Secondly, it proposes an improved YOLOv8 visual detection algorithm (YOLO-Edge-Guided Lighting Adaptation, YL-EGLA). The algorithm realizes feature fusion by dynamically extracting the high-dimensional features of the input through the self-attention mechanism combined with the enhanced edge features extracted by the edge detection operator, and is equipped with adaptive image enhancement module to ensure stable detection under different lighting conditions. Algorithm comparison results demonstrate that the YL-EGLA can be fully trained on the synthetic dataset and directly applied to real-world scenarios without additional fine-tuning. Furthermore, physical experiments further validate the efficiency and practicality of this algorithm in transparent object manipulation, fully showcasing its significant value in practical applications.

AAAI Conference 2025 Conference Paper

FactorGCL: A Hypergraph-Based Factor Model with Temporal Residual Contrastive Learning for Stock Returns Prediction

  • Yitong Duan
  • Weiran Wang
  • Jian Li

As a fundamental method in economics and finance, the factor model has been extensively utilized in quantitative investment. In recent years, there has been a paradigm shift from traditional linear models with expert-designed factors to more flexible nonlinear machine learning-based models with data-driven factors, aiming to enhance the effectiveness of these factor models. However, due to the low signal-to-noise ratio in market data, mining effective factors in data-driven models remains challenging. In this work, we propose a hypergraph-based factor model with temporal residual contrastive learning (FactorGCL) that employs a hypergraph structure to better capture high-order nonlinear relationships among stock returns and factors. To mine hidden factors that supplement human-designed prior factors for predicting stock returns, we design a cascading residual hypergraph architecture, in which the hidden factors are extracted from the residual information after removing the influence of prior factors. Additionally, we propose a temporal residual contrastive learning method to guide the extraction of effective and comprehensive hidden factors by contrasting stock-specific residual information over different time periods. Our extensive experiments on real stock market data demonstrate that FactorGCL not only outperforms existing state-of-the-art methods but also mines effective hidden factors for predicting stock returns.

ICLR Conference 2025 Conference Paper

Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural Networks

  • Binghui Li
  • Zhixuan Pan
  • Kaifeng Lyu
  • Jian Li

In this work, we investigate a particular implicit bias in gradient descent training, which we term “Feature Averaging,” and argue that it is one of the principal factors contributing to the non-robustness of deep neural networks. We show that, even when multiple discriminative features are present in the input data, neural networks trained by gradient descent tend to rely on an average (or a certain combination) of these features for classification, rather than distinguishing and leveraging each feature individually. Specifically, we provide a detailed theoretical analysis of the training dynamics of two-layer ReLU networks on a binary classification task, where the data distribution consists of multiple clusters with mutually orthogonal centers. We rigorously prove that gradient descent biases the network towards feature averaging, where the weights of each hidden neuron represent an average of the cluster centers (each corresponding to a distinct feature), thereby making the network vulnerable to input perturbations aligned with the negative direction of the averaged features. On the positive side, we demonstrate that this vulnerability can be mitigated through more granular supervision. In particular, we prove that a two-layer ReLU network can achieve optimal robustness when trained to classify individual features rather than merely the original binary classes. Finally, we validate our theoretical findings with experiments on synthetic datasets, MNIST, and CIFAR-10, and confirm the prevalence of feature averaging and its impact on adversarial robustness. We hope these theoretical and empirical insights deepen the understanding of how gradient descent shapes feature learning and adversarial robustness, and how more detailed supervision can enhance robustness.

IROS Conference 2025 Conference Paper

MCTrack: A Unified 3D Multi-Object Tracking Framework for Autonomous Driving

  • Xiyang Wang 0002
  • Shouzheng Qi
  • Jieyou Zhao
  • Hangning Zhou
  • Siyu Zhang 0002
  • Guoan Wang
  • Kai Tu
  • Songlin Guo

This paper introduces MCTrack, a new 3D multi-object tracking method that achieves performance across KITTI, nuScenes, and Waymo datasets. Addressing the gap in existing tracking paradigms, which often perform well on specific datasets but lack generalizability, MCTrack offers a unified solution. Additionally, we have standardized the format of perceptual results across various datasets, termed BaseVersion, facilitating researchers in the field of MOT) to concentrate on the core algorithmic development without the undue burden of data preprocessing. Finally, recognizing the limitations of current evaluation metrics, we introduce a novel set of metrics designed to evaluate the output of motion information, including velocity and acceleration, which are essential for subsequent tasks. The source codes of the proposed method are available at this link: https://github.com/megvii-research/MCTrack

IROS Conference 2025 Conference Paper

Micro-UAV with Ant-Inspired Bistable Gripper for Adaptive Perching and Wildlife Detection

  • Yuan Liu
  • Yadong Mo
  • Xuexiu Liang
  • Yongkang Jiang
  • Jian Li
  • Shimin Wei

With the global ecological environment facing continuous deterioration, effective monitoring of arboreal birds in complex canopy environments remains challenging due to limitations of conventional drones in endurance, size, and habitat disturbance. To address these challenges, this paper presents an ant-inspired micro quadrotor UAV equipped with a lightweight bistable gripper system mimicking the mandibular morphology of leafcutter ants. The design integrates shape memory alloy (SMA)-driven actuation and thermoplastic polyurethane (TPU)-based adaptive grippers, enabling rapid deformation (71 ms switching time) and energy-efficient operation (zero power consumption during perching). Experimental results demonstrate exceptional adaptability in grasping irregular objects (e. g. , branches, pen caps) with an 8: 1 payload-to-weight ratio. Field tests confirm stable navigation through dense foliage and reliable perching at heights exceeding 5 meters. The system’s compact dimensions (7 cm diameter, 70. 5 g weight) and biomimetic approach offer a non-invasive solution for prolonged wildlife observation. This work advances bistable actuator design by combining bio-inspired structural optimization with rapid energy transition principles, showing potential in agile robotics and environmental sensing.

AAAI Conference 2025 Conference Paper

Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints

  • Ming Dai
  • Jian Li
  • Jiedong Zhuang
  • Xian Zhang
  • Wankou Yang

Multi-task visual grounding involves the simultaneous execution of localization and segmentation in images based on textual expressions. The majority of advanced methods predominantly focus on transformer-based multimodal fusion, aiming to extract robust multimodal representations. However, ambiguity between referring expression comprehension (REC) and referring image segmentation (RIS) is error-prone, leading to inconsistencies between multi-task predictions. Besides, insufficient multimodal understanding directly contributes to biased target perception. To overcome these challenges, we propose a Coarse-to-fine Consistency Constraints Visual Grounding architecture (C3VG), which integrates implicit and explicit modeling approaches within a two-stage framework. Initially, query and pixel decoders are employed to generate preliminary detection and segmentation outputs, a process referred to as the Rough Semantic Perception (RSP) stage. These coarse predictions are subsequently refined through the proposed Mask-guided Interaction Module (MIM) and a novel explicit bidirectional consistency constraint loss to ensure consistent representations across tasks, which we term the Refined Consistency Interaction (RCI) stage. Furthermore, to address the challenge of insufficient multimodal understanding, we leverage pre-trained models based on visual-linguistic fusion representations. Empirical evaluations on the RefCOCO, RefCOCO+, and RefCOCOg datasets demonstrate the efficacy and soundness of C3VG, which significantly outperforms state-of-the-art REC and RIS methods by a substantial margin.

ICLR Conference 2025 Conference Paper

On the Linear Speedup of Personalized Federated Reinforcement Learning with Shared Representations

  • Guojun Xiong
  • Shufan Wang
  • Daniel Jiang
  • Jian Li

Federated reinforcement learning (FedRL) enables multiple agents to collaboratively learn a policy without needing to share the local trajectories collected during agent-environment interactions. However, in practice, the environments faced by different agents are often heterogeneous, but since existing FedRL algorithms learn a single policy across all agents, this may lead to poor performance. In this paper, we introduce a personalized FedRL framework (PFedRL) by taking advantage of possibly shared common structure among agents in heterogeneous environments. Specifically, we develop a class of PFedRL algorithms named PFedRL-Rep that learns (1) a shared feature representation collaboratively among all agents, and (2) an agent-specific weight vector personalized to its local environment. We analyze the convergence of PFedTD-Rep, a particular instance of the framework with temporal difference (TD) learning and linear representations. To the best of our knowledge, we are the first to prove a linear convergence speedup with respect to the number of agents in the PFedRL setting. To achieve this, we show that PFedTD-Rep is an example of federated two-timescale stochastic approximation with Markovian noise. Experimental results demonstrate that PFedTD-Rep, along with an extension to the control setting based on deep Q-networks (DQN), not only improve learning in heterogeneous settings, but also provide better generalization to new environments.

ICML Conference 2025 Conference Paper

Provably Efficient Exploration in Inverse Constrained Reinforcement Learning

  • Bo Yue
  • Jian Li
  • Guiliang Liu

Optimizing objective functions subject to constraints is fundamental in many real-world applications. However, these constraints are often not readily defined and must be inferred from expert agent behaviors, a problem known as Inverse Constraint Inference. Inverse Constrained Reinforcement Learning (ICRL) is a common solver for recovering feasible constraints in complex environments, relying on training samples collected from interactive environments. However, the efficacy and efficiency of current sampling strategies remain unclear. We propose a strategic exploration framework for sampling with guaranteed efficiency to bridge this gap. By defining the feasible cost set for ICRL problems, we analyze how estimation errors in transition dynamics and the expert policy influence the feasibility of inferred constraints. Based on this analysis, we introduce two exploratory algorithms to achieve efficient constraint inference via 1) dynamically reducing the bounded aggregate error of cost estimations or 2) strategically constraining the exploration policy around plausibly optimal ones. Both algorithms are theoretically grounded with tractable sample complexity, and their performance is validated empirically across various environments.

ICLR Conference 2025 Conference Paper

Rethinking Diffusion Posterior Sampling: From Conditional Score Estimator to Maximizing a Posterior

  • Tongda Xu
  • Xiyan Cai
  • Xinjie Zhang
  • Xingtong Ge
  • Dailan He
  • Ming Sun
  • Jingjing Liu
  • Ya-Qin Zhang

Recent advancements in diffusion models have been leveraged to address inverse problems without additional training, and Diffusion Posterior Sampling (DPS) (Chung et al., 2022a) is among the most popular approaches. Previous analyses suggest that DPS accomplishes posterior sampling by approximating the conditional score. While in this paper, we demonstrate that the conditional score approximation employed by DPS is not as effective as previously assumed, but rather aligns more closely with the principle of maximizing a posterior (MAP). This assertion is substantiated through an examination of DPS on 512$\times$512 ImageNet images, revealing that: 1) DPS’s conditional score estimation significantly diverges from the score of a well-trained conditional diffusion model and is even inferior to the unconditional score; 2) The mean of DPS’s conditional score estimation deviates significantly from zero, rendering it an invalid score estimation; 3) DPS generates high-quality samples with significantly lower diversity. In light of the above findings, we posit that DPS more closely resembles MAP than a conditional score estimator, and accordingly propose the following enhancements to DPS: 1) we explicitly maximize the posterior through multi-step gradient ascent and projection; 2) we utilize a light-weighted conditional score estimator trained with only 100 images and 8 GPU hours. Extensive experimental results indicate that these proposed improvements significantly enhance DPS's performance. The source code for these improvements is provided in https://github.com/tongdaxu/Rethinking-Diffusion-Posterior-Sampling-From-Conditional-Score-Estimator-to-Maximizing-a-Posterior.

YNIMG Journal 2025 Journal Article

Shared subcortical arousal systems across sensory modalities during transient modulation of attention

  • Aya Khalaf
  • Erick Lopez
  • Jian Li
  • Andreas Horn
  • Brian L. Edlow
  • Hal Blumenfeld

Subcortical arousal systems are known to play a key role in controlling sustained changes in attention and conscious awareness. Recent studies indicate that these systems have a major influence on short-term dynamic modulation of visual attention, but their role across sensory modalities is not fully understood. In this study, we investigated shared subcortical arousal systems across sensory modalities during transient changes in attention using block and event-related fMRI paradigms. We analyzed massive publicly available fMRI datasets collected while 1561 participants performed visual, auditory, tactile, and taste perception tasks. Our analyses revealed a shared circuit of subcortical arousal systems exhibiting early transient increases in activity in midbrain reticular formation and central thalamus across perceptual modalities, as well as less consistent increases in pons, hypothalamus, basal forebrain, and basal ganglia. Identifying these networks is critical for understanding mechanisms of normal attention and consciousness and may help facilitate subcortical targeting for therapeutic neuromodulation.

YNIMG Journal 2025 Journal Article

Subthalamic nucleus stimulation at high and low frequencies engages different brain networks to enhance gait performance in Parkinson's disease

  • Yin Jiang
  • Hutao Xie
  • Yutong Bai
  • Quan Zhang
  • Yu Diao
  • Houyou Fan
  • Xin Zhang
  • Hua Zhang

BACKGROUND: Subthalamic nucleus (STN) deep brain stimulation (DBS) is used to treat Parkinson's disease (PD), yet neither high-frequency stimulation (HFS) nor low frequency stimulation (LFS) fully resolves gait issues. Previous studies indicate that STN-DBS modulates motor-related brain networks. Given that PD patients with gait disturbances exhibit cognitive deficits-and considering the extensive projections between the STN and cerebral cortex-we hypothesized that varying STN stimulation frequencies may improve gait by modulating distinct brain networks. METHODS: We collected gait data, cortical electrophysiological signals, and resting-state fMRI from 44 PD patients and 32 healthy controls. Multi-network cortical activity and functional connectivity were c ompared under three conditions: DBS OFF, HFS, and LFS. Additionally, the connectivity values were correlated to the gait behaviors and clinical assessment scores. RESULTS: We found that: (1) HFS improved both motor and gait performance, while LFS enhanced gait but may not be optimal for long-term use; (2) STN-DBS induced widespread modulation across sensorimotor, frontoparietal, salience, dorsal attention, and default mode networks. HFS improved motor and gait functions via network modulation related to motor control, whereas LFS may enhance gait by boosting executive-related cortical activities and connections; (3) Relative to healthy controls, PD exhibited widespread reductions in functional connectivity, with DBS modulation trending toward normalization. CONCLUSIONS: These results reveal distinct brain network responses to different STN-DBS frequencies in PD, offering a theoretical basis for optimizing DBS treatment for gait impairments. These findings provide critical insights for tailoring DBS parameters to maximize both motor and cognitive benefits in PD patients.

NeurIPS Conference 2025 Conference Paper

The Lighthouse of Language: Enhancing LLM Agents via Critique-Guided Improvement

  • Ruihan Yang
  • Fanghua Ye
  • Jian Li
  • Siyu Yuan
  • Yikai Zhang
  • Zhaopeng Tu
  • Xiaolong Li
  • Deqing Yang

Large language models (LLMs) have recently transformed from text-based assistants to autonomous agents capable of planning, reasoning, and iteratively improving their actions. While numerical reward signals and verifiers can effectively rank candidate actions, they often provide limited contextual guidance. In contrast, natural language feedback better aligns with the generative capabilities of LLMs, providing richer and more actionable suggestions. However, parsing and implementing this feedback effectively can be challenging for LLM-based agents. In this work, we introduce Critique-Guided Improvement (CGI), a novel two-player framework, comprising an actor model that explores an environment and a critic model that generates detailed nature language feedback. By training the critic to produce fine-grained assessments and actionable revisions, and the actor to utilize these critiques, our approach promotes more robust exploration of alternative strategies while avoiding local optima. Experiments in three interactive environments show that CGI outperforms existing baselines by a substantial margin. Notably, even a small critic model surpasses GPT-4 in feedback quality. The resulting actor achieves state-of-the-art performance, demonstrating the power of explicit iterative guidance to enhance decision-making in LLM-based agents.

ICLR Conference 2025 Conference Paper

Understanding Constraint Inference in Safety-Critical Inverse Reinforcement Learning

  • Bo Yue
  • Shufan Wang
  • Ashish Gaurav
  • Jian Li
  • Pascal Poupart
  • Guiliang Liu

In practical applications, the underlying constraint knowledge is often unknown and difficult to specify. To address this issue, recent advances in Inverse Constrained Reinforcement Learning (ICRL) have focused on inferring these constraints from expert demonstrations. However, the ICRL approach typically characterizes constraint learning as a tri-level optimization problem, which is inherently complex due to its interdependent variables and multiple layers of optimization. Considering these challenges, a critical question arises: *Can we implicitly embed constraint signals into reward functions and effectively solve this problem using a classic reward inference algorithm?* The resulting method, known as Inverse Reward Correction (IRC), merits investigation. In this work, we conduct a theoretical analysis comparing the sample complexities of both solvers. Our findings confirm that the IRC solver achieves lower sample complexity than its ICRL counterpart. Nevertheless, this reduction in complexity comes at the expense of generalizability. Specifically, in the target environment, the reward correction terms may fail to guarantee the safety of the resulting policy, whereas this issue can be effectively mitigated by transferring the constraints via the ICRL solver. Advancing our inquiry, we investigate conditions under which the ICRL solver ensures $\epsilon$-optimality when transferring to new environments. Empirical results across various environments validate our theoretical findings, underscoring the nuanced trade-offs between complexity reduction and generalizability in safety-critical applications.

NeurIPS Conference 2025 Conference Paper

Understanding LLM Behaviors via Compression: Data Generation, Knowledge Acquisition and Scaling Laws

  • Zhixuan Pan
  • Shaowen Wang
  • Liao Pengfei
  • Jian Li

Large Language Models (LLMs) have demonstrated remarkable capabilities across numerous tasks, yet principled explanations for their underlying mechanisms and several phenomena, such as scaling laws, hallucinations, and related behaviors, remain elusive. In this work, we revisit the classical relationship between compression and prediction, grounded in Kolmogorov complexity and Shannon information theory, to provide deeper insights into LLM behaviors. By leveraging the Kolmogorov Structure Function and interpreting LLM compression as a two-part coding process, we offer a detailed view of how LLMs acquire and store information across increasing model and data scales -- from pervasive syntactic patterns to progressively rarer knowledge elements. Motivated by this theoretical perspective and natural assumptions inspired by Heap’s and Zipf’s laws, we introduce a simplified yet representative hierarchical data-generation framework called the Syntax-Knowledge model. Under the Bayesian setting, we show that prediction and compression within this model naturally lead to diverse learning and scaling behaviors of LLMs. In particular, our theoretical analysis offers intuitive and principled explanations for both data and model scaling laws, the dynamics of knowledge acquisition during training and fine-tuning, factual knowledge hallucinations in LLMs. The experimental results validate our theoretical predictions.

NeurIPS Conference 2025 Conference Paper

VITA-Audio: Fast Interleaved Audio-Text Token Generation for Efficient Large Speech-Language Model

  • Zuwei Long
  • Yunhang Shen
  • Chaoyou Fu
  • Heting Gao
  • Lijiang Li
  • Peixian Chen
  • Mengdan Zhang
  • Hang Shao

With the growing requirement for natural human-computer interaction, speech-based systems receive increasing attention as speech is one of the most common forms of daily communication. However, the existing speech models still experience high latency when generating the first audio token during streaming, which poses a significant bottleneck for deployment. To address this issue, we propose VITA-Audio, an end-to-end large speech model with fast audio-text token generation. Specifically, we introduce a lightweight Multiple Cross-modal Token Prediction (MCTP) module that efficiently generates multiple audio tokens within a single model forward pass, which not only accelerates the inference but also significantly reduces the latency for generating the first audio in streaming scenarios. In addition, a four-stage progressive training strategy is explored to achieve model acceleration with minimal loss of speech quality. To our knowledge, VITA-Audio is the first multi-modal large language model capable of generating audio output during the first forward pass, enabling real-time conversational capabilities with minimal latency. VITA-Audio is fully reproducible and is trained on open-source data only. Experimental results demonstrate that our model achieves an inference speedup of 3~5x at the 7B parameter scale, but also significantly outperforms open-source models of similar model size on multiple benchmarks for automatic speech recognition (ASR), text-to-speech (TTS), and spoken question answering (SQA) tasks.

IROS Conference 2024 Conference Paper

6-DoF Grasp Detection in Clutter with Enhanced Receptive Field and Graspable Balance Sampling

  • Hanwen Wang
  • Ying Zhang
  • Yunlong Wang
  • Jian Li

6-DoF grasp detection of small-scale grasps is crucial for robots to perform specific tasks. This paper focuses on enhancing the recognition capability of small-scale grasping, aiming to improve the overall accuracy of grasping prediction results and the generalization ability of the network. We propose an enhanced receptive field method that includes a multi-radii cylinder grouping module and a passive attention module. This method enhances the receptive field area within the graspable space and strengthens the learning of graspable features. Additionally, we design a graspable balance sampling module based on a 3D segmentation network, which enables the network to focus on features of small objects, thereby improving the recognition capability of small-scale grasping. Our network achieves state-of-the-art performance on the GraspNet-1Billion dataset, with an overall improvement of approximately 10% in average precision@k (AP). Furthermore, we deployed our grasp detection model on pybullet grasping platform and in real-world scenarios, which validates the effectiveness of our method.

AAAI Conference 2024 Conference Paper

Cheaper and Faster: Distributed Deep Reinforcement Learning with Serverless Computing

  • Hanfei Yu
  • Jian Li
  • Yang Hua
  • Xu Yuan
  • Hao Wang

Deep reinforcement learning (DRL) has gained immense success in many applications, including gaming AI, robotics, and system scheduling. Distributed algorithms and architectures have been vastly proposed (e.g., actor-learner architecture) to accelerate DRL training with large-scale server-based clusters. However, training on-policy algorithms with the actor-learner architecture unavoidably induces resource wasting due to synchronization between learners and actors, thus resulting in significantly extra billing. As a promising alternative, serverless computing naturally fits on-policy synchronization and alleviates resource wasting in distributed DRL training with pay-as-you-go pricing. Yet, none has leveraged serverless computing to facilitate DRL training. This paper proposes MinionsRL, the first serverless distributed DRL training framework that aims to accelerate DRL training- and cost-efficiency with dynamic actor scaling. We prototype MinionsRL on top of Microsoft Azure Container Instances and evaluate it with popular DRL tasks from OpenAI Gym. Extensive experiments show that MinionsRL reduces total training time by up to 52% and training cost by 86% compared to latest solutions.

AAAI Conference 2024 Conference Paper

Convolutional Spectral Kernel Learning with Generalization Guarantees (Abstract Reprint)

  • Jian Li
  • Yong Liu
  • Weiping Wang

Kernel methods are powerful tools to capture nonlinear patterns behind given data but often lead to poor performance on complicated tasks compared to convolutional neural networks. The reason is that kernel methods are still shallow and fully connected models, failing to reveal hierarchical features and local interdependencies. In this paper, to acquire hierarchical and local knowledge, we incorporate kernel methods with deep architectures and convolutional operators in a spectral kernel learning framework. Based on the inverse Fourier transform and Rademacher complexity theory, we provide the generalization error bounds for the proposed model and prove that under suitable initialization, deeper networks lead to tighter error bounds. Inspired by theoretical findings, we finally completed the convolutional spectral kernel network (CSKN) with two additional regularizers and an initialization strategy. Extensive ablation results validate the effectiveness of non-stationary spectral kernel, multiple layers, additional regularizers, and the convolutional filters, which coincide with our theoretical findings. We further devise a VGG-type 8-layers CSKN, and it outperforms the existing kernel-based networks and popular CNN models on the medium-sized image classification tasks.

AAAI Conference 2024 Conference Paper

DePRL: Achieving Linear Convergence Speedup in Personalized Decentralized Learning with Shared Representations

  • Guojun Xiong
  • Gang Yan
  • Shiqiang Wang
  • Jian Li

Decentralized learning has emerged as an alternative method to the popular parameter-server framework which suffers from high communication burden, single-point failure and scalability issues due to the need of a central server. However, most existing works focus on a single shared model for all workers regardless of the data heterogeneity problem, rendering the resulting model performing poorly on individual workers. In this work, we propose a novel personalized decentralized learning algorithm named DePRL via shared representations. Our algorithm relies on ideas from representation learning theory to learn a low-dimensional global representation collaboratively among all workers in a fully decentralized manner, as well as a user-specific low-dimensional local head leading to a personalized solution for each worker. We show that DePRL achieves, for the first time, a provable \textit{linear speedup for convergence} with general non-linear representations (i.e., the convergence rate is improved linearly with respect to the number of workers). Experimental results support our theoretical findings showing the superiority of our method in data heterogeneous environments.

AAAI Conference 2024 Conference Paper

FedNS: A Fast Sketching Newton-Type Algorithm for Federated Learning

  • Jian Li
  • Yong Liu
  • Weiping Wang

Recent Newton-type federated learning algorithms have demonstrated linear convergence with respect to the communication rounds. However, communicating Hessian matrices is often unfeasible due to their quadratic communication complexity. In this paper, we introduce a novel approach to tackle this issue while still achieving fast convergence rates. Our proposed method, named as Federated Newton Sketch methods (FedNS), approximates the centralized Newton's method by communicating the sketched square-root Hessian instead of the exact Hessian. To enhance communication efficiency, we reduce the sketch size to match the effective dimension of the Hessian matrix. We provide convergence analysis based on statistical learning for the federated Newton sketch approaches. Specifically, our approaches reach super-linear convergence rates w.r.t. the communication rounds for the first time. We validate the effectiveness of our algorithms through various experiments, which coincide with our theoretical findings.

NeurIPS Conference 2024 Conference Paper

Fetch and Forge: Efficient Dataset Condensation for Object Detection

  • Ding Qi
  • Jian Li
  • Jinlong Peng
  • Bo Zhao
  • Shuguang Dou
  • Jialin Li
  • Jiangning Zhang
  • Yabiao Wang

Dataset condensation (DC) is an emerging technique capable of creating compact synthetic datasets from large originals while maintaining considerable performance. It is crucial for accelerating network training and reducing data storage requirements. However, current research on DC mainly focuses on image classification, with less exploration of object detection. This is primarily due to two challenges: (i) the multitasking nature of object detection complicates the condensation process, and (ii) Object detection datasets are characterized by large-scale and high-resolution data, which are difficult for existing DC methods to handle. As a remedy, we propose DCOD, the first dataset condensation framework for object detection. It operates in two stages: Fetch and Forge, initially storing key localization and classification information into model parameters, and then reconstructing synthetic images via model inversion. For the complex of multiple objects in an image, we propose Foreground Background Decoupling to centrally update the foreground of multiple instances and Incremental PatchExpand to further enhance the diversity of foregrounds. Extensive experiments on various detection datasets demonstrate the superiority of DCOD. Even at an extremely low compression rate of 1\%, we achieve 46. 4\% and 24. 7\% $\text{AP}_{50}$ on the VOC and COCO, respectively, significantly reducing detector training duration.

AAAI Conference 2024 Conference Paper

High-Dimensional Analysis for Generalized Nonlinear Regression: From Asymptotics to Algorithm

  • Jian Li
  • Yong Liu
  • Weiping Wang

Overparameterization often leads to benign overfitting, where deep neural networks can be trained to overfit the training data but still generalize well on unseen data. However, it lacks a generalized asymptotic framework for nonlinear regressions and connections to conventional complexity notions. In this paper, we propose a generalized high-dimensional analysis for nonlinear regression models, including various nonlinear feature mapping methods and subsampling. Specifically, we first provide an implicit regularization parameter and asymptotic equivalents related to a classical complexity notion, i.e., effective dimension. We then present a high-dimensional analysis for nonlinear ridge regression and extend it to ridgeless regression in the under-parameterized and over-parameterized regimes, respectively. We find that the limiting risks decrease with the effective dimension. Motivated by these theoretical findings, we propose an algorithm, namely RFRed, to improve generalization ability. Finally, we validate our theoretical findings and the proposed algorithm through several experiments.

EAAI Journal 2024 Journal Article

HIWANet: A high imperceptibility watermarking attack network

  • Chunpeng Wang
  • Xinying Li
  • Zhiqiu Xia
  • Qi Li
  • Hao Zhang
  • Jian Li
  • Bing Han
  • Bin Ma

Digital image watermarking technology has made a significant contribution to the copyright protection of digital images. In recent years, researchers have focused on designing various watermarking algorithms to enhance resistance against different forms of attacks. However, the evolution of watermarking attack technology has been sluggish, thus impeding possible advancements in digital copyright protection. Existing watermarking attack methods exhibit a notable drawback, causing substantial deterioration in visual quality and undermining the practical utility of attacked images. In this paper, we propose a high imperceptibility watermarking attack network, named HIWANet, based on deep neural networks. To enhance the watermarking attack ability, a feature extraction module (FEM) is aimed to better capture watermark information features, and a watermarking attack module (WAM) is constructed to learn high-level abstract features of the images. In addition, to ensure the imperceptibility of the watermarking attack, an asymmetric loss function is designed to maintain the quality of the attacked watermarked image. In the experiments, we randomly select 2000 color images from the PASCAL VOC2012 database as the dataset, with the training and test sets containing 1000 distinct images. Compared to traditional watermarking attack methods, our HIWANet achieves a significant increase in bit error rate (improved by 242%), indicating a higher attack ability. Meanwhile, it brings more than 26% improvement in the attack imperceptibility. Furthermore, our HIWANet also offers significant advantages compared to deep learning-based watermarking attack methods.

IJCAI Conference 2024 Conference Paper

IMM: An Imitative Reinforcement Learning Approach with Predictive Representation Learning for Automatic Market Making

  • Hui Niu
  • Siyuan Li
  • Jiahao Zheng
  • Zhouchi Lin
  • Bo An
  • Jian Li
  • Jian Guo

Market making (MM) via Reinforcement Learning (RL) has attracted significant attention in financial trading. Most existing RL-based MM methods focus on optimizing single-price level strategies which fail at frequent order cancellations and loss of queue priority. By comparison, strategies involving multiple price levels align better with actual trading scenarios. However, given the complexity that multi-price level RL strategies involve a comprehensive trading action space, the challenge of effectively training RL persists. Inspired by the effective workflow of professional human market makers, we propose Imitative Market Maker (IMM), a novel RL framework leveraging knowledge from both suboptimal signal-based experts and direct policy interactions. Our framework starts with introducing effective state and action formulations that well encode information about multiprice level orders. Furthermore, IMM integrates a representation learning unit capable of capturing both short- and long-term market trends to mitigate adverse selection risk. Subsequently, IMM designs an expert strategy based on predictive signals, and trains the agent through the integration of RL and imitation learning techniques to achieve efficient learning. Extensive experimental results on four real-world market datasets demonstrate the superiority of IMM against current RL-based MM strategies.

NeurIPS Conference 2024 Conference Paper

LoRA-GA: Low-Rank Adaptation with Gradient Approximation

  • Shaowen Wang
  • Linxi Yu
  • Jian Li

Fine-tuning large-scale pretrained models is prohibitively expensive in terms of computational and memory costs. LoRA, as one of the most popular Parameter-Efficient Fine-Tuning (PEFT) methods, offers a cost-effective alternative by fine-tuning an auxiliary low-rank model that has significantly fewer parameters. Although LoRA reduces the computational and memory requirements significantly at each iteration, extensive empirical evidence indicates that it converges at a considerably slower rate compared to full fine-tuning, ultimately leading to increased overall compute and often worse test performance. In our paper, we perform an in-depth investigation of the initialization method of LoRA and show that careful initialization (without any change of the architecture and the training algorithm) can significantly enhance both efficiency and performance. In particular, we introduce a novel initialization method, LoRA-GA (Low Rank Adaptation with Gradient Approximation), which aligns the gradients of low-rank matrix product with those of full fine-tuning at the first step. Our extensive experiments demonstrate that LoRA-GA achieves a convergence rate comparable to that of full fine-tuning (hence being significantly faster than vanilla LoRA as well as various recent improvements) while simultaneously attaining comparable or even better performance. For example, on the subset of the GLUE dataset with T5-Base, LoRA-GA outperforms LoRA by 5. 69% on average. On larger models such as Llama 2-7B, LoRA-GA shows performance improvements of 0. 34, 11. 52%, and 5. 05% on MTbench, GSM8k, and Human-eval, respectively. Additionally, we observe up to 2-4 times convergence speed improvement compared to vanilla LoRA, validating its effectiveness in accelerating convergence and enhancing model performance.

IJCAI Conference 2024 Conference Paper

MacMic: Executing Iceberg Orders via Hierarchical Reinforcement Learning

  • Hui Niu
  • Siyuan Li
  • Jian Li

In recent years, there has been a growing interest in applying reinforcement learning (RL) techniques to order execution owing to RL’s strong sequential decision-making ability. However, realistic order execution tasks usually involve a large fine-grained action space and a long trading duration. The former hinders the RL agents from efficient exploration. The latter increases the task complexity, since the agent must capture price advantages throughout the day as well as micro changes within a few seconds on the limited order books. In addressing these challenges, we propose MacMic, a novel Hierarchical RL-based order execution approach that captures market patterns and executes orders from different temporal scales. MacMic employs a high-level agent to split the parent order into smaller slices at coarse-grained time steps. Then a low-level agent is adopted to execute these slices by placing fixed-size sub-orders at a continuous time. Besides, to balance the multifaceted objectives of the two tasks, MacMic pretrains a causal stacking hidden Markov model (SHMM) to obtain both effective macro-level and micro-level market states. Comprehensive experimental results on 200 stocks across the US and China A-share markets validate the effectiveness of the proposed method.

AAAI Conference 2024 Conference Paper

Online Restless Multi-Armed Bandits with Long-Term Fairness Constraints

  • Shufan Wang
  • Guojun Xiong
  • Jian Li

Restless multi-armed bandits (RMAB) have been widely used to model sequential decision making problems with constraints. The decision maker (DM) aims to maximize the expected total reward over an infinite horizon under an “instantaneous activation constraint” that at most B arms can be activated at any decision epoch, where the state of each arm evolves stochastically according to a Markov decision process (MDP). However, this basic model fails to provide any fairness guarantee among arms. In this paper, we introduce RMAB-F, a new RMAB model with “long-term fairness constraints”, where the objective now is to maximize the longterm reward while a minimum long-term activation fraction for each arm must be satisfied. For the online RMAB-F setting (i.e., the underlying MDPs associated with each arm are unknown to the DM), we develop a novel reinforcement learning (RL) algorithm named Fair-UCRL. We prove that Fair-UCRL ensures probabilistic sublinear bounds on both the reward regret and the fairness violation regret. Compared with off-the-shelf RL methods, our Fair-UCRL is much more computationally efficient since it contains a novel exploitation that leverages a low-complexity index policy for making decisions. Experimental results further demonstrate the effectiveness of our Fair-UCRL.

EAAI Journal 2024 Journal Article

Secondary restoration of islanded alternating current microgrids under a neural inverse optimal control

  • Jian Li
  • Cong Cai
  • Qingyu Su

A novel control strategy for secondary restoration of an islanded microgrid is proposed, focusing on restoring the frequency and voltage magnitude of an inverter-based distributed generator. A neural inverse optimal controller is integrated into the secondary control layer with a higher-order neural network trained using an extended Kalman filter (EKF). A practical and effective design strategy is provided. Unlike traditional approaches that rely on accurate mathematical models, the combination of inverse optimal control and neural networks does not require an accurate model and is more relevant to real-world engineering scenarios. To enhance the secondary control, the EKF optimization parameters are used in conjunction with the inverse optimal control to achieve more accurate and effective repair results. The synergistic effect improves control performance and ensures superior secondary recovery. Real-time validation is performed through rigorous simulations on the StarSim hardware-in-the-loop experimental platform.

IJCAI Conference 2024 Conference Paper

Trade When Opportunity Comes: Price Movement Forecasting via Locality-Aware Attention and Iterative Refinement Labeling

  • Liang Zeng
  • Lei Wang
  • Hui Niu
  • Ruchen Zhang
  • Ling Wang
  • Jian Li

Price movement forecasting, aimed at predicting financial asset trends based on current market information, has achieved promising advancements through machine learning (ML) methods. Most existing ML methods, however, struggle with the extremely low signal-to-noise ratio and stochastic nature of financial data, often mistaking noises for real trading signals without careful selection of potentially profitable samples. To address this issue, we propose LARA, a novel price movement forecasting framework with two main components: Locality-Aware Attention (LA-Attention) and Iterative Refinement Labeling (RA-Labeling). (1) LA-Attention, enhanced by metric learning techniques, automatically extracts the potentially profitable samples through masked attention scheme and task-specific distance metrics. (2) RA-Labeling further iteratively refines the noisy labels of potentially profitable samples, and combines the learned predictors robust to the unseen and noisy samples. In a set of experiments on three real-world financial markets: stocks, cryptocurrencies, and ETFs, LARA significantly outperforms several machine learning based methods on the Qlib quantitative investment platform. Extensive ablation studies confirm LARA's superior ability in capturing more reliable trading opportunities.

IROS Conference 2024 Conference Paper

Visual Loop Closure Detection with Thorough Temporal and Spatial Context Exploitation

  • Jiaxin Li
  • Zan Wang
  • Huijun Di
  • Jian Li
  • Wei Liang 0008

Despite advancements in visual Simultaneous Localization and Mapping (SLAM), prevailing visual Loop Closure Detection (LCD) methods primarily rely on computationally intensive image similarity comparisons, neglecting temporal-spatial context during long-term exploration. To address this issue, we propose TOSA, a novel visual LCD algorithm harnessing TempOral and SpAtial context for efficient LCD. Specifically, as the agent explores through time, our approach recurrently updates a latent feature incorporating historical information via a Long Short-Term Memory (LSTM) module. Upon receiving a query frame, TOSA seamlessly fuses the latent feature with the query feature to predict the candidates’ distribution, thus averting intensive similarity computation. Additionally, TOSA integrates a temporal-spatial convolution for candidate refinement by thoroughly exploiting the temporal consistency and spatial correlation to enhance selected candidates, further boosting the performance. Extensive experiments across four standard datasets showcase the superiority of our method over existing state-of-the-art techniques, demonstrating the effectiveness of utilizing rich temporal-spatial contexts.

AAAI Conference 2023 Conference Paper

AEC-GAN: Adversarial Error Correction GANs for Auto-Regressive Long Time-Series Generation

  • Lei Wang
  • Liang Zeng
  • Jian Li

Large-scale high-quality data is critical for training modern deep neural networks. However, data acquisition can be costly or time-consuming for many time-series applications, thus researchers turn to generative models for generating synthetic time-series data. In particular, recent generative adversarial networks (GANs) have achieved remarkable success in time-series generation. Despite their success, existing GAN models typically generate the sequences in an auto-regressive manner, and we empirically observe that they suffer from severe distribution shifts and bias amplification, especially when generating long sequences. To resolve this problem, we propose Adversarial Error Correction GAN (AEC-GAN), which is capable of dynamically correcting the bias in the past generated data to alleviate the risk of distribution shifts and thus can generate high-quality long sequences. AEC-GAN contains two main innovations: (1) We develop an error correction module to mitigate the bias. In the training phase, we adversarially perturb the realistic time-series data and then optimize this module to reconstruct the original data. In the generation phase, this module can act as an efficient regulator to detect and mitigate the bias. (2) We propose an augmentation method to facilitate GAN's training by introducing adversarial examples. Thus, AEC-GAN can generate high-quality sequences of arbitrary lengths, and the synthetic data can be readily applied to downstream tasks to boost their performance. We conduct extensive experiments on six widely used datasets and three state-of-the-art time-series forecasting models to evaluate the quality of our synthetic time-series data in different lengths and downstream tasks. Both the qualitative and quantitative experimental results demonstrate the superior performance of AEC-GAN over other deep generative models for time-series generation.

AAAI Conference 2023 Conference Paper

Decentralized Stochastic Multi-Player Multi-Armed Walking Bandits

  • Guojun Xiong
  • Jian Li

Multi-player multi-armed bandit is an increasingly relevant decision-making problem, motivated by applications to cognitive radio systems. Most research for this problem focuses exclusively on the settings that players have full access to all arms and receive no reward when pulling the same arm. Hence all players solve the same bandit problem with the goal of maximizing their cumulative reward. However, these settings neglect several important factors in many real-world applications, where players have limited access to a dynamic local subset of arms (i.e., an arm could sometimes be ``walking'' and not accessible to the player). To this end, this paper proposes a multi-player multi-armed walking bandits model, aiming to address aforementioned modeling issues. The goal now is to maximize the reward, however, players can only pull arms from the local subset and only collect a full reward if no other players pull the same arm. We adopt Upper Confidence Bound (UCB) to deal with the exploration-exploitation tradeoff and employ distributed optimization techniques to properly handle collisions. By carefully integrating these two techniques, we propose a decentralized algorithm with near-optimal guarantee on the regret, and can be easily implemented to obtain competitive empirical performance.

AAAI Conference 2023 Conference Paper

DeFL: Defending against Model Poisoning Attacks in Federated Learning via Critical Learning Periods Awareness

  • Gang Yan
  • Hao Wang
  • Xu Yuan
  • Jian Li

Federated learning (FL) is known to be susceptible to model poisoning attacks in which malicious clients hamper the accuracy of the global model by sending manipulated model updates to the central server during the FL training process. Existing defenses mainly focus on Byzantine-robust FL aggregations, and largely ignore the impact of the underlying deep neural network (DNN) that is used to FL training. Inspired by recent findings on critical learning periods (CLP) in DNNs, where small gradient errors have irrecoverable impact on the final model accuracy, we propose a new defense, called a CLP-aware defense against poisoning of FL (DeFL). The key idea of DeFL is to measure fine-grained differences between DNN model updates via an easy-to-compute federated gradient norm vector (FGNV) metric. Using FGNV, DeFL simultaneously detects malicious clients and identifies CLP, which in turn is leveraged to guide the adaptive removal of detected malicious clients from aggregation. As a result, DeFL not only mitigates model poisoning attacks on the global model but also is robust to detection errors. Our extensive experiments on three benchmark datasets demonstrate that DeFL produces significant performance gain over conventional defenses against state-of-the-art model poisoning attacks.

NeurIPS Conference 2023 Conference Paper

Finite-Time Analysis of Whittle Index based Q-Learning for Restless Multi-Armed Bandits with Neural Network Function Approximation

  • Guojun Xiong
  • Jian Li

Whittle index policy is a heuristic to the intractable restless multi-armed bandits (RMAB) problem. Although it is provably asymptotically optimal, finding Whittle indices remains difficult. In this paper, we present Neural-Q-Whittle, a Whittle index based Q-learning algorithm for RMAB with neural network function approximation, which is an example of nonlinear two-timescale stochastic approximation with Q-function values updated on a faster timescale and Whittle indices on a slower timescale. Despite the empirical success of deep Q-learning, the non-asymptotic convergence rate of Neural-Q-Whittle, which couples neural networks with two-timescale Q-learning largely remains unclear. This paper provides a finite-time analysis of Neural-Q-Whittle, where data are generated from a Markov chain, and Q-function is approximated by a ReLU neural network. Our analysis leverages a Lyapunov drift approach to capture the evolution of two coupled parameters, and the nonlinearity in value function approximation further requires us to characterize the approximation error. Combing these provide Neural-Q-Whittle with $\mathcal{O}(1/k^{2/3})$ convergence rate, where $k$ is the number of iterations.

NeurIPS Conference 2023 Conference Paper

GLIME: General, Stable and Local LIME Explanation

  • Zeren Tan
  • Yang Tian
  • Jian Li

As black-box machine learning models become more complex and are applied in high-stakes settings, the need for providing explanations for their predictions becomes crucial. Although Local Interpretable Model-agnostic Explanations (LIME) \cite{ribeiro2016should} is a widely adopted method for understanding model behavior, it suffers from instability with respect to random seeds \cite{zafar2019dlime, shankaranarayana2019alime, bansal2020sam} and exhibits low local fidelity (i. e. , how the explanation explains model's local behaviors) \cite{rahnama2019study, laugel2018defining}. Our study demonstrates that this instability is caused by small sample weights, resulting in the dominance of regularization and slow convergence. Additionally, LIME's sampling approach is non-local and biased towards the reference, leading to diminished local fidelity and instability to references. To address these challenges, we propose \textsc{Glime}, an enhanced framework that extends LIME and unifies several previous methods. Within the \textsc{Glime} framework, we derive an equivalent formulation of LIME that achieves significantly faster convergence and improved stability. By employing a local and unbiased sampling distribution, \textsc{Glime} generates explanations with higher local fidelity compared to LIME, while being independent of the reference choice. Moreover, \textsc{Glime} offers users the flexibility to choose sampling distribution based on their specific scenarios.

EAAI Journal 2023 Journal Article

Identification and classification for multiple cyber attacks in power grids based on the deep capsule CNN

  • Guangdou Zhang
  • Jian Li
  • Olusola Bamisile
  • Yankai Xing
  • Di Cao
  • Qi Huang

Cyber-attacks have become one of the main threats to the security, reliability, and economic operation of power systems. Detection and classification of multiple cyber-attacks pose challenges for ensuring the stability and security of power systems. To address this issue, this study proposes an automatic identification and classification method for multiple cyber-attacks based on the deep capsule convolution neural network. Spatial correlations among different nodes and temporal features from history operation status in the transmitted data packets are extracted by the convolution neural network. Capsules in the proposed structure have important implications for maintaining the topological consistency contained in the measurement matrix. Furthermore, the proposed method is model-free and avoids the impact of system parameters uncertainties on detection performance. Multiple kinds of typical cyber-attacks, including false data injection attacks, replay attacks, denial of service attacks, time-delay attacks, and deception attacks, are considered and modeled in this paper. Numerical results on the IEEE 39-bus test system show that the proposed method can achieve 99. 97% detection accuracy on single cyber-attacks and 96. 25% detection accuracy on multiple cyber-attacks. Comparative results illustrate the proposed method outperforms than traditional neural networks. This approach provides a solution for the problem of multiple cyber-attack detection and classification.

YNIMG Journal 2023 Journal Article

Identification of overlapping and interacting networks reveals intrinsic spatiotemporal organization of the human brain

  • Jian Li
  • Yijun Liu
  • Jessica L. Wisnowski
  • Richard M. Leahy

The human brain is a complex network that exhibits dynamic fluctuations in activity across space and time. Depending on the analysis method, canonical brain networks identified from resting-state fMRI (rs-fMRI) are typically constrained to be either orthogonal or statistically independent in their spatial and/or temporal domains. We avoid imposing these potentially unnatural constraints through the combination of a temporal synchronization process ("BrainSync") and a three-way tensor decomposition method ("NASCAR") to jointly analyze rs-fMRI data from multiple subjects. The resulting set of interacting networks comprises minimally constrained spatiotemporal distributions, each representing one component of functionally coherent activity across the brain. We show that these networks can be clustered into six distinct functional categories and naturally form a representative functional network atlas for a healthy population. This functional network atlas could help explore group and individual differences in neurocognitive function, as we demonstrate in the context of ADHD and IQ prediction.

AAAI Conference 2023 Conference Paper

ImGCL: Revisiting Graph Contrastive Learning on Imbalanced Node Classification

  • Liang Zeng
  • Lanqing Li
  • Ziqi Gao
  • Peilin Zhao
  • Jian Li

Graph contrastive learning (GCL) has attracted a surge of attention due to its superior performance for learning node/graph representations without labels. However, in practice, the underlying class distribution of unlabeled nodes for the given graph is usually imbalanced. This highly imbalanced class distribution inevitably deteriorates the quality of learned node representations in GCL. Indeed, we empirically find that most state-of-the-art GCL methods cannot obtain discriminative representations and exhibit poor performance on imbalanced node classification. Motivated by this observation, we propose a principled GCL framework on Imbalanced node classification (ImGCL), which automatically and adaptively balances the representations learned from GCL without labels. Specifically, we first introduce the online clustering based progressively balanced sampling (PBS) method with theoretical rationale, which balances the training sets based on pseudo-labels obtained from learned representations in GCL. We then develop the node centrality based PBS method to better preserve the intrinsic structure of graphs, by upweighting the important nodes of the given graph. Extensive experiments on multiple imbalanced graph datasets and imbalanced settings demonstrate the effectiveness of our proposed framework, which significantly improves the performance of the recent state-of-the-art GCL methods. Further experimental ablations and analyses show that the ImGCL framework consistently improves the representation quality of nodes in under-represented (tail) classes.

ICML Conference 2023 Conference Paper

Online Restless Bandits with Unobserved States

  • Bowen Jiang
  • Bo Jiang 0003
  • Jian Li
  • Tao Lin 0001
  • Xinbing Wang
  • Chenghu Zhou

We study the online restless bandit problem, where each arm evolves according to a Markov chain independently, and the reward of pulling an arm depends on both the current state of the corresponding Markov chain and the pulled arm. The agent (decision maker) does not know the transition functions and reward functions, and cannot observe the states of arms even after pulling. The goal is to sequentially choose which arms to pull so as to maximize the expected cumulative rewards collected. In this paper, we propose TSEETC, a learning algorithm based on Thompson Sampling with Episodic Explore-Then-Commit. The algorithm proceeds in episodes of increasing length and each episode is divided into exploration and exploitation phases. During the exploration phase, samples of action-reward pairs are collected in a round-robin fashion and utilized to update the posterior distribution as a mixture of Dirichlet distributions. At the beginning of the exploitation phase, TSEETC generates a sample from the posterior distribution as true parameters. It then follows the optimal policy for the sampled model for the rest of the episode. We establish the Bayesian regret bound $\tilde {\mathcal{O}}(\sqrt{T})$ for TSEETC, where $T$ is the time horizon. We show through simulations that TSEETC outperforms existing algorithms in regret.

JMLR Journal 2023 Journal Article

Optimal Convergence Rates for Distributed Nystroem Approximation

  • Jian Li
  • Yong Liu
  • Weiping Wang

The distributed kernel ridge regression (DKRR) has shown great potential in processing complicated tasks. However, DKRR only made use of the local samples that failed to capture the global characteristics. Besides, the existing optimal learning guarantees were provided in expectation and only pertain to the attainable case that the target regression lies exactly in the kernel space. In this paper, we propose distributed learning with globally-shared Nystroem centers (DNystroem), which utilizes global information across the local clients. We also study the statistical properties of DNystroem in expectation and in probability, respectively, and obtain several state-of-the-art results with the minimax optimal learning rates. Note that, the optimal convergence rates for DNystroem pertain to the non-attainable case, while the statistical results allow more partitions and require fewer Nystroem centers. Finally, we conduct experiments on several real-world datasets to validate the effectiveness of the proposed algorithm, and the empirical results coincide with our theoretical findings. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2023. ( edit, beta )

IROS Conference 2023 Conference Paper

Robotic Kinematic Calibration with Only Position Data and Consideration of Non-Geometric Errors Using POE-Based Model and Gaussian Mixture Models

  • Xiao Luo 0005
  • Yitian Xian
  • Mancheong Lei
  • Jian Li
  • Ke Xie 0007
  • Limin Zou
  • Zheng Li 0012

Kinematic calibration is crucial to improve the positioning accuracy of serial robots. This paper proposes a novel algorithm for robotic kinematic calibration based on an augmented product of exponentials (POE)-based kinematic model using Gaussian mixture models (GMMs) with only position data. In this algorithm, non-geometric errors that cannot be fitted by varying the parameters within the traditional robot model are also considered and compensated. This approach involving a three-stage calibration process which is used to identify the kinematic model parameters and to train the GMMs will be presented in this paper. Finally, this algorithm will be applied to two serial robots for simulation and experimental validation. The effectiveness of the proposed algorithm is verified from both results and significant improvement on error reduction from 26 % to 96% can be observed through the comparison with other existing approaches.

AAAI Conference 2023 Conference Paper

Symphony in the Latent Space: Provably Integrating High-Dimensional Techniques with Non-linear Machine Learning Models

  • Qiong Wu
  • Jian Li
  • Zhenming Liu
  • Yanhua Li
  • Mihai Cucuringu

This paper revisits building machine learning algorithms that involve interactions between entities, such as those between financial assets in an actively managed portfolio, or interactions between users in a social network. Our goal is to forecast the future evolution of ensembles of multivariate time series in such applications (e.g., the future return of a financial asset or the future popularity of a Twitter account). Designing ML algorithms for such systems requires addressing the challenges of high-dimensional interactions and non-linearity. Existing approaches usually adopt an ad-hoc approach to integrating high-dimensional techniques into non-linear models and recent studies have shown these approaches have questionable efficacy in time-evolving interacting systems. To this end, we propose a novel framework, which we dub as the additive influence model. Under our modeling assumption, we show that it is possible to decouple the learning of high-dimensional interactions from the learning of non-linear feature interactions. To learn the high-dimensional interactions, we leverage kernel-based techniques, with provable guarantees, to embed the entities in a low-dimensional latent space. To learn the non-linear feature-response interactions, we generalize prominent machine learning techniques, including designing a new statistically sound non-parametric method and an ensemble learning algorithm optimized for vector regressions. Extensive experiments on two common applications demonstrate that our new algorithms deliver significantly stronger forecasting power compared to standard and recently proposed methods.

IJCAI Conference 2023 Conference Paper

Towards Generalizable Reinforcement Learning for Trade Execution

  • Chuheng Zhang
  • Yitong Duan
  • Xiaoyu Chen
  • Jianyu Chen
  • Jian Li
  • Li Zhao

Optimized trade execution is to sell (or buy) a given amount of assets in a given time with the lowest possible trading cost. Recently, reinforcement learning (RL) has been applied to optimized trade execution to learn smarter policies from market data. However, we find that many existing RL methods exhibit considerable overfitting which prevents them from real deployment. In this paper, we provide an extensive study on the overfitting problem in optimized trade execution. First, we model the optimized trade execution as offline RL with dynamic context (ORDC), where the context represents market variables that cannot be influenced by the trading policy and are collected in an offline manner. Under this framework, we derive the generalization bound and find that the overfitting issue is caused by large context space and limited context samples in the offline setting. Accordingly, we propose to learn compact representations for context to address the overfitting problem, either by leveraging prior knowledge or in an end-to-end manner. To evaluate our algorithms, we also implement a carefully designed simulator based on historical limit order book (LOB) data to provide a high-fidelity benchmark for different algorithms. Our experiments on the high-fidelity simulator demonstrate that our algorithms can effectively alleviate overfitting and achieve better performance.

IJCAI Conference 2023 Conference Paper

Towards Sharp Analysis for Distributed Learning with Random Features

  • Jian Li
  • Yong Liu

In recent studies, the generalization properties for distributed learning and random features assumed the existence of the target concept over the hypothesis space. However, this strict condition is not applicable to the more common non-attainable case. In this paper, using refined proof techniques, we first extend the optimal rates for distributed learning with random features to the non-attainable case. Then, we reduce the number of required random features via data-dependent generating strategy, and improve the allowed number of partitions with additional unlabeled data. Theoretical analysis shows these techniques remarkably reduce computational cost while preserving the optimal generalization accuracy under standard assumptions. Finally, we conduct several experiments on both simulated and real-world datasets, and the empirical results validate our theoretical findings.

IJCAI Conference 2023 Conference Paper

Unbiased Gradient Boosting Decision Tree with Unbiased Feature Importance

  • Zheyu Zhang
  • Tianping Zhang
  • Jian Li

Gradient Boosting Decision Tree (GBDT) has achieved remarkable success in a wide variety of applications. The split finding algorithm, which determines the tree construction process, is one of the most crucial components of GBDT. However, the split finding algorithm has long been criticized for its bias towards features with a large number of potential splits. This bias introduces severe interpretability and overfitting issues in GBDT. To this end, we provide a fine-grained analysis of bias in GBDT and demonstrate that the bias originates from 1) the systematic bias in the gain estimation of each split and 2) the bias in the split finding algorithm resulting from the use of the same data to evaluate the split improvement and determine the best split. Based on the analysis, we propose unbiased gain, a new unbiased measurement of gain importance using out-of-bag samples. Moreover, we incorporate the unbiased property into the split finding algorithm and develop UnbiasedGBM to solve the overfitting issue of GBDT. We assess the performance of UnbiasedGBM and unbiased gain in a large-scale empirical study comprising 60 datasets and show that: 1) UnbiasedGBM exhibits better performance than popular GBDT implementations such as LightGBM, XGBoost, and Catboost on average on the 60 datasets and 2) unbiased gain achieves better average performance in feature selection than popular feature importance methods.

NeurIPS Conference 2022 Conference Paper

Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of Stability

  • Zixuan Wang
  • Zhouzi Li
  • Jian Li

Recent findings demonstrate that modern neural networks trained by full-batch gradient descent typically enter a regime called Edge of Stability (EOS). In this regime, the sharpness, i. e. , the maximum Hessian eigenvalue, first increases to the value 2/(step size) (the progressive sharpening phase) and then oscillates around this value (the EOS phase). This paper aims to analyze the GD dynamics and the sharpness along the optimization trajectory. Our analysis naturally divides the GD trajectory into four phases depending on the change in the sharpness value. We empirically identify the norm of output layer weight as an interesting indicator of the sharpness dynamics. Based on this empirical observation, we attempt to theoretically and empirically explain the dynamics of various key quantities that lead to the change of the sharpness in each phase of EOS. Moreover, based on certain assumptions, we provide a theoretical proof of the sharpness behavior in the EOS regime in two-layer fully-connected linear neural networks. We also discuss some other empirical findings and the limitation of our theoretical results.

AIJ Journal 2022 Journal Article

Convolutional spectral kernel learning with generalization guarantees

  • Jian Li
  • Yong Liu
  • Weiping Wang

Kernel methods are powerful tools to capture nonlinear patterns behind given data but often lead to poor performance on complicated tasks compared to convolutional neural networks. The reason is that kernel methods are still shallow and fully connected models, failing to reveal hierarchical features and local interdependencies. In this paper, to acquire hierarchical and local knowledge, we incorporate kernel methods with deep architectures and convolutional operators in a spectral kernel learning framework. Based on the inverse Fourier transform and Rademacher complexity theory, we provide the generalization error bounds for the proposed model and prove that under suitable initialization, deeper networks lead to tighter error bounds. Inspired by theoretical findings, we finally completed the convolutional spectral kernel network (CSKN) with two additional regularizers and an initialization strategy. Extensive ablation results validate the effectiveness of non-stationary spectral kernel, multiple layers, additional regularizers, and the convolutional filters, which coincide with our theoretical findings. We further devise a VGG-type 8-layers CSKN, and it outperforms the existing kernel-based networks and popular CNN models on the medium-sized image classification tasks.

AAAI Conference 2022 Conference Paper

FactorVAE: A Probabilistic Dynamic Factor Model Based on Variational Autoencoder for Predicting Cross-Sectional Stock Returns

  • Yitong Duan
  • Lei Wang
  • Qizhong Zhang
  • Jian Li

As an asset pricing model in economics and finance, factor model has been widely used in quantitative investment. Towards building more effective factor models, recent years have witnessed the paradigm shift from linear models to more flexible nonlinear data-driven machine learning models. However, due to low signal-to-noise ratio of the financial data, it is quite challenging to learn effective factor models. In this paper, we propose a novel factor model, FactorVAE, as a probabilistic model with inherent randomness for noise modeling. Essentially, our model integrates the dynamic factor model (DFM) with the variational autoencoder (VAE) in machine learning, and we propose a prior-posterior learning method based on VAE, which can effectively guide the learning of model by approximating an optimal posterior factor model with future information. Particularly, considering that risk modeling is important for the noisy stock data, Factor- VAE can estimate the variances from the distribution over the latent space of VAE, in addition to predicting returns. The experiments on the real stock market data demonstrate the effectiveness of FactorVAE, which outperforms various baseline methods.

NeurIPS Conference 2022 Conference Paper

Generalization Bounds for Gradient Methods via Discrete and Continuous Prior

  • Xuanyuan Luo
  • Bei Luo
  • Jian Li

Proving algorithm-dependent generalization error bounds for gradient-type optimization methods has attracted significant attention recently in learning theory. However, most existing trajectory-based analyses require either restrictive assumptions on the learning rate (e. g. , fast decreasing learning rate), or continuous injected noise (such as the Gaussian noise in Langevin dynamics). In this paper, we introduce a new discrete data-dependent prior to the PAC-Bayesian framework, and prove a high probability generalization bound of order $O(\frac{1}{n}\cdot \sum_{t=1}^T(\gamma_t/\varepsilon_t)^2\left\|{\mathrm{g}_t}\right\|^2)$ for Floored GD (i. e. a version of gradient descent with precision level $\varepsilon_t$), where $n$ is the number of training samples, $\gamma_t$ is the learning rate at step $t$, $\mathrm{g}_t$ is roughly the difference of the gradient computed using all samples and that using only prior samples. $\left\|{\mathrm{g}_t}\right\|$ is upper bounded by and and typical much smaller than the gradient norm $\left\|{\nabla f(W_t)}\right\|$. We remark that our bound holds for nonconvex and nonsmooth scenarios. Moreover, our theoretical results provide numerically favorable upper bounds of testing errors (e. g. , $0. 037$ on MNIST). Using similar technique, we can also obtain new generalization bounds for a certain variant of SGD. Furthermore, we study the generalization bounds for gradient Langevin Dynamics (GLD). Using the same framework with a carefully constructed continuous prior, we show a new high probability generalization bound of order $O(\frac{1}{n} + \frac{L^2}{n^2}\sum_{t=1}^T(\gamma_t/\sigma_t)^2)$ for GLD. The new $1/n^2$ rate is due to the concentration of the difference between the gradient of training samples and that of the prior.

NeurIPS Conference 2022 Conference Paper

Learning Infinite-Horizon Average-Reward Restless Multi-Action Bandits via Index Awareness

  • Guojun Xiong
  • Shufan Wang
  • Jian Li

We consider the online restless bandits with average-reward and multiple actions, where the state of each arm evolves according to a Markov decision process (MDP), and the reward of pulling an arm depends on both the current state of the corresponding MDP and the action taken. Since finding the optimal control is typically intractable for restless bandits, existing learning algorithms are often computationally expensive or with a regret bound that is exponential in the number of arms and states. In this paper, we advocate \textit{index-aware reinforcement learning} (RL) solutions to design RL algorithms operating on a much smaller dimensional subspace by exploiting the inherent structure in restless bandits. Specifically, we first propose novel index policies to address dimensionality concerns, which are provably optimal. We then leverage the indices to develop two low-complexity index-aware RL algorithms, namely, (i) GM-R2MAB, which has access to a generative model; and (ii) UC-R2MAB, which learns the model using an upper confidence style online exploitation method. We prove that both algorithms achieve a sub-linear regret that is only polynomial in the number of arms and states. A key differentiator between our algorithms and existing ones stems from the fact that our RL algorithms contain a novel exploitation that leverages our proposed provably optimal index policies for decision-makings.

NeurIPS Conference 2022 Conference Paper

Learning to Break the Loop: Analyzing and Mitigating Repetitions for Neural Text Generation

  • Jin Xu
  • Xiaojiang Liu
  • Jianhao Yan
  • Deng Cai
  • Huayang Li
  • Jian Li

While large-scale neural language models, such as GPT2 and BART, have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms (\textit{e. g. }, greedy search). This phenomenon is counter-intuitive since there are few consecutive sentence-level repetitions in the human corpus (e. g. , 0. 02\% in Wikitext-103). To investigate the underlying reasons for generating consecutive sentence-level repetitions, we study the relationship between the probability of repetitive tokens and their previous repetitions in context. Through our quantitative experiments, we find that 1) Models have a preference to repeat the previous sentence; 2) The sentence-level repetitions have a \textit{self-reinforcement effect}: the more times a sentence is repeated in the context, the higher the probability of continuing to generate that sentence; 3) The sentences with higher initial probabilities usually have a stronger self-reinforcement effect. Motivated by our findings, we propose a simple and effective training method \textbf{DITTO} (Pseu\underline{D}o-Repet\underline{IT}ion Penaliza\underline{T}i\underline{O}n), where the model learns to penalize probabilities of sentence-level repetitions from synthetic repetitive data. Although our method is motivated by mitigating repetitions, our experiments show that DITTO not only mitigates the repetition issue without sacrificing perplexity, but also achieves better generation quality. Extensive experiments on open-ended text generation (Wikitext-103) and text summarization (CNN/DailyMail) demonstrate the generality and effectiveness of our method.

EAAI Journal 2022 Journal Article

Light-field image watermarking based on geranion polar harmonic Fourier moments

  • Chunpeng Wang
  • Qinghua Zhang
  • Bin Ma
  • Zhiqiu Xia
  • Jian Li
  • Ting Luo
  • Qi Li

Light-field images provide a spatial and angular description of the light, and they can capture rich visual information in the natural world. In recent years, the research on light-field imaging has achieved many fruitful results. As more and more people are now focusing their attention to light-field imaging, the copyright protection of light-field images has become an urgent problem that needs to be addressed. Digital watermarking can effectively protect the copyright ownership of light-field images. Currently, there exists no light-field image watermarking scheme that can resist various geometric attacks. In this work, relying on geranion theory and polar harmonic Fourier moments (PHFMs), geranion polar harmonic Fourier moments (GPHFMs) are constructed and utilized for light-field image watermarking. The geranion represents a hypercomplex number containing one real part and thirty-one imaginary parts, and the imaginary parts of geranion can be used to encode multiple image color components while maintaining the correlation between each component. GPHFMs are stable image features, with good image reconstruction ability and stability. Essentially, the light-field image watermarking based on GPHFMs offers good imperceptibility and robustness, can resist various attacks, and effectively solves the problem of current watermarking schemes related to their inability of resisting geometric attacks. Furthermore, it is verified through experiments that the proposed scheme is more robust than the previously reported schemes.

I&C Journal 2022 Journal Article

Optimal in-place suffix sorting

  • Zhize Li
  • Jian Li
  • Hongwei Huo

The suffix array is a fundamental data structure for many applications that involve string searching and data compression. We obtain the first in-place suffix array construction algorithms that are optimal both in time and space for (read-only) integer alphabets. We make the following contributions: 1. For integer alphabets, we obtain the first linear time suffix sorting algorithm which uses only O ( 1 ) workspace. The input string may be modified during the execution of the algorithm, but should be restored upon termination of the algorithm. 2. We strengthen the first result by providing the first in-place linear time algorithm for read-only integer alphabets with | Σ | = O ( n ) (i. e. , the input string cannot be modified). This algorithm settles the open problem posed by Franceschini and Muthukrishnan in ICALP 2007. 3. Besides, for the read-only general alphabets (i. e. , only comparisons are allowed), we present an optimal in-place O ( n log ⁡ n ) time suffix sorting algorithm.

AAAI Conference 2022 Conference Paper

Reinforcement Learning Augmented Asymptotically Optimal Index Policy for Finite-Horizon Restless Bandits

  • Guojun Xiong
  • Jian Li
  • Rahul Singh

We study a finite-horizon restless multi-armed bandit problem with multiple actions, dubbed as R(MA)2 B. The state of each arm evolves according to a controlled Markov decision process (MDP), and the reward of pulling an arm depends on both the current state and action of the corresponding MDP. Since finding the optimal policy is typically intractable, we propose a computationally appealing index policy entitled Occupancy-Measured-Reward Index Policy for the finite-horizon R(MA)2 B. Our index policy is well-defined without the requirement of indexability condition and is provably asymptotically optimal. We then adopt a learning perspective where the system parameters are unknown, and propose R(MA)2 B-UCB, a generative model based reinforcement learning augmented algorithm that can fully exploit the structure of Occupancy-Measured-Reward Index Policy. Compared to existing algorithms, R(MA)2 B-UCB performs close to offline optimum, as well as achieves a sub-linear regret and a low computational complexity all at once. Experimental results show that R(MA)2 B-UCB outperforms existing algorithms in both regret and running time.

IJCAI Conference 2022 Conference Paper

Ridgeless Regression with Random Features

  • Jian Li
  • Yong Liu
  • Yingying Zhang

Recent theoretical studies illustrated that kernel ridgeless regression can guarantee good generalization ability without an explicit regularization. In this paper, we investigate the statistical properties of ridgeless regression with random features and stochastic gradient descent. We explore the effect of factors in the stochastic gradient and random features, respectively. Specifically, random features error exhibits the double-descent curve. Motivated by the theoretical findings, we propose a tunable kernel algorithm that optimizes the spectral density of kernel during training. Our work bridges the interpolation theory and practical algorithm.

AAAI Conference 2022 Conference Paper

SCSNet: An Efficient Paradigm for Learning Simultaneously Image Colorization and Super-resolution

  • Jiangning Zhang
  • Chao Xu
  • Jian Li
  • Yue Han
  • Yabiao Wang
  • Ying Tai
  • Yong Liu

In the practical application of restoring low-resolution grayscale images, we generally need to run three separate processes of image colorization, super-resolution, and dowssampling operation for the target device. However, this pipeline is redundant and inefficient for the independent processes, and some inner features could have been shared. Therefore, we present an efficient paradigm to perform Simultaneously Image Colorization and Super-resolution (SCS) and propose an end-to-end SCSNet to achieve this goal. The proposed method consists of two parts: colorization branch for learning color information that employs the proposed plug-and-play Pyramid Valve Cross Attention (PV- CAttn) module to aggregate feature maps between source and reference images; and super-resolution branch for integrating color and texture information to predict target images, which uses the designed Continuous Pixel Mapping (CPM) module to predict high-resolution images at continuous magnification. Furthermore, our SCSNet supports both automatic and referential modes that is more flexible for practical application. Abundant experiments demonstrate the superiority of our method for generating authentic images over state-of-theart methods, e. g. , averagely decreasing FID by 1. 8↓ and 5. 1 ↓ compared with current best scores for automatic and referential modes, respectively, while owning fewer parameters (more than ×2↓) and faster running speed (more than ×3↑).

AAAI Conference 2022 Conference Paper

Seizing Critical Learning Periods in Federated Learning

  • Gang Yan
  • Hao Wang
  • Jian Li

Federated learning (FL) is a popular technique to train machine learning (ML) models with decentralized data. Extensive works have studied the performance of the global model; however, it is still unclear how the training process affects the final test accuracy. Exacerbating this problem is the fact that FL executions differ significantly from traditional ML with heterogeneous data characteristics across clients, involving more hyperparameters. In this work, we show that the final test accuracy of FL is dramatically affected by the early phase of the training process, i. e. , FL exhibits critical learning periods, in which small gradient errors irrecoverably impact the final test accuracy. To further explain this phenomenon, we generalize the trace of Fisher Information Matrix (FIM) to FL and define a new notion called FedFIM, a quantity reflecting the local curvature of each client from the beginning of training in FL. Our findings suggest that the initial learning phase plays a critical role in understanding the FL performance. This is in contrast to many existing works which generally do not connect the final accuracy of FL to the early phase training. Finally, seizing critical learning periods in FL is of independent interest and could be useful for other problems such as the choices of hyperparameters including but not limited to the number of client selected per round, batch size, so as to improve the performance of FL training and testing.

JMLR Journal 2022 Journal Article

Simple and Optimal Stochastic Gradient Methods for Nonsmooth Nonconvex Optimization

  • Zhize Li
  • Jian Li

We propose and analyze several stochastic gradient algorithms for finding stationary points or local minimum in nonconvex, possibly with nonsmooth regularizer, finite-sum and online optimization problems. First, we propose a simple proximal stochastic gradient algorithm based on variance reduction called ProxSVRG+. We provide a clean and tight analysis of ProxSVRG+, which shows that it outperforms the deterministic proximal gradient descent (ProxGD) for a wide range of minibatch sizes, hence solves an open problem proposed in Reddi et al. 2016. Also, ProxSVRG+ uses much less proximal oracle calls than ProxSVRG (Reddi et al. 2016) and extends to the online setting by avoiding full gradient computations. Then, we further propose an optimal algorithm, called SSRGD, based on ARAH (Nguyen et al. 2017) and show that SSRGD further improves the gradient complexity of ProxSVRG+ and achieves the the optimal upper bound, matching the known lower bound. Moreover, we show that both ProxSVRG+ and SSRGD enjoy automatic adaptation with local structure of the objective function such as the Polyak-Lojasiewicz (PL) condition for nonconvex functions in the finite-sum case, i.e., we prove that both of them can automatically switch to faster global linear convergence without any restart performed in prior work. Finally, we focus on the more challenging problem of finding an $(\epsilon, \delta)$-local minimum instead of just finding an $\epsilon$-approximate (first-order) stationary point (which may be some bad unstable saddle points). We show that SSRGD can find an $(\epsilon, \delta)$-local minimum by simply adding some random perturbations. Our algorithm is almost as simple as its counterpart for finding stationary points, and achieves similar optimal rates. [abs] [ pdf ][ bib ] &copy JMLR 2022. ( edit, beta )

NeurIPS Conference 2021 Conference Paper

Analogous to Evolutionary Algorithm: Designing a Unified Sequence Model

  • Jiangning Zhang
  • Chao Xu
  • Jian Li
  • Wenzhou Chen
  • Yabiao Wang
  • Ying Tai
  • Shuo Chen
  • Chengjie Wang

Inspired by biological evolution, we explain the rationality of Vision Transformer by analogy with the proven practical Evolutionary Algorithm (EA) and derive that both of them have consistent mathematical representation. Analogous to the dynamic local population in EA, we improve the existing transformer structure and propose a more efficient EAT model, and design task-related heads to deal with different tasks more flexibly. Moreover, we introduce the spatial-filling curve into the current vision transformer to sequence image data into a uniform sequential format. Thus we can design a unified EAT framework to address multi-modal tasks, separating the network architecture from the data format adaptation. Our approach achieves state-of-the-art results on the ImageNet classification task compared with recent vision transformer works while having smaller parameters and greater throughput. We further conduct multi-modal tasks to demonstrate the superiority of the unified EAT, \eg, Text-Based Image Retrieval, and our approach improves the rank-1 by +3. 7 points over the baseline on the CSS dataset.

AAAI Conference 2021 Conference Paper

Exploration by Maximizing Renyi Entropy for Reward-Free RL Framework

  • Chuheng Zhang
  • Yuanying Cai
  • Longbo Huang
  • Jian Li

Exploration is essential for reinforcement learning (RL). To face the challenges of exploration, we consider a reward-free RL framework that completely separates exploration from exploitation and brings new challenges for exploration algorithms. In the exploration phase, the agent learns an exploratory policy by interacting with a reward-free environment and collects a dataset of transitions by executing the policy. In the planning phase, the agent computes a good policy for any reward function based on the dataset without further interacting with the environment. This framework is suitable for the meta RL setting where there are many reward functions of interest. In the exploration phase, we propose to maximize the Rényi entropy over the state-action space and justify this objective theoretically. The success of using Rényi entropy as the objective results from its encouragement to explore the hard-to-reach state-actions. We further deduce a policy gradient formulation for this objective and design a practical exploration algorithm that can deal with complex environments. In the planning phase, we solve for good policies given arbitrary reward functions using a batch RL algorithm. Empirically, we show that our exploration algorithm is effective and sample efficient, and results in superior policies for arbitrary reward functions in the planning phase.

AAAI Conference 2021 Short Paper

Improving the Morphology and Control Policy of Self-reconfiguring Modular Robots in Dynamic Environment (Student Abstract)

  • Dong Xue
  • Qiang Lu
  • Jian Li

Ideally, self-reconfiguration modular robots (SMR) can change their morphology and perform actions related to a specific task in any scene. However, most SMRs only adapt to several specific scenes because their morphology and control policies are designed or trained based on these scenes. Once SMRs meet an unknown scene, especially multiply unknown scenes (called dynamic environment), these policies will be useless. Although some of these policies import evolutionary algorithms to enhance the ability of SMR to explore unknown scenes, they are very time-consuming. The reason is that individual fitness depends on the interaction between SMR and these scenes. We propose a two-stage reconfiguration algorithm (TSRA) without any prior knowledge to address the time-consuming problem. In the two stages, the reconfiguration methods use the evolutionary algorithm (GA) to simultaneously generate scene-fitted morphology and actions. The first stage method uses the estimation neural network to evaluate the individual fitness to run faster and can recommend better policies to the second stage. The second stage method obtains this fitness from scenes and updates the neural network to approximate these scenes. Through experiments, TSRA can find better morphology and control policies than the other two canonical algorithms — GA and GEM-RL.

YNIMG Journal 2021 Journal Article

Mapping the subcortical connectivity of the human default mode network

  • Jian Li
  • William H. Curley
  • Bastien Guerin
  • Darin D. Dougherty
  • Adrian V. Dalca
  • Bruce Fischl
  • Andreas Horn
  • Brian L. Edlow

The default mode network (DMN) mediates self-awareness and introspection, core components of human consciousness. Therapies to restore consciousness in patients with severe brain injuries have historically targeted subcortical sites in the brainstem, thalamus, hypothalamus, basal forebrain, and basal ganglia, with the goal of reactivating cortical DMN nodes. However, the subcortical connectivity of the DMN has not been fully mapped, and optimal subcortical targets for therapeutic neuromodulation of consciousness have not been identified. In this work, we created a comprehensive map of DMN subcortical connectivity by combining high-resolution functional and structural datasets with advanced signal processing methods. We analyzed 7 Tesla resting-state functional MRI (rs-fMRI) data from 168 healthy volunteers acquired in the Human Connectome Project. The rs-fMRI blood-oxygen-level-dependent (BOLD) data were temporally synchronized across subjects using the BrainSync algorithm. Cortical and subcortical DMN nodes were jointly analyzed and identified at the group level by applying a novel Nadam-Accelerated SCAlable and Robust (NASCAR) tensor decomposition method to the synchronized dataset. The subcortical connectivity map was then overlaid on a 7 Tesla 100 µm ex vivo MRI dataset for neuroanatomic analysis using automated segmentation of nuclei within the brainstem, thalamus, hypothalamus, basal forebrain, and basal ganglia. We further compared the NASCAR subcortical connectivity map with its counterpart generated from canonical seed-based correlation analyses. The NASCAR method revealed that BOLD signal in the central lateral nucleus of the thalamus and ventral tegmental area of the midbrain is strongly correlated with that of the DMN. In an exploratory analysis, additional subcortical sites in the median and dorsal raphe, lateral hypothalamus, and caudate nuclei were correlated with the cortical DMN. We also found that the putamen and globus pallidus are negatively correlated (i.e., anti-correlated) with the DMN, providing rs-fMRI evidence for the mesocircuit hypothesis of human consciousness, whereby a striatopallidal feedback system modulates anterior forebrain function via disinhibition of the central thalamus. Seed-based analyses yielded similar subcortical DMN connectivity, but the NASCAR result showed stronger contrast and better spatial alignment with dopamine immunostaining data. The DMN subcortical connectivity map identified here advances understanding of the subcortical regions that contribute to human consciousness and can be used to inform the selection of therapeutic targets in clinical trials for patients with disorders of consciousness.

IROS Conference 2021 Conference Paper

Model Adaptation through Hypothesis Transfer with Gradual Knowledge Distillation

  • Song Tang 0001
  • Yuji Shi
  • Zhiyuan Ma 0001
  • Jian Li
  • Jianzhi Lyu
  • Qingdu Li
  • Jianwei Zhang 0001

The ability to adapt their perception to changing environments is a core characterization of intelligent robots. At present, Unsupervised Domain Adaptation (UDA) methods are used to address this problem where the adaptation task is formulated as a transfer problem from a well-described scenario (source domain) to a new scenario (target domain). In order to implement the domain adaptation, these methods require access to the source data for achieving the distribution matching between both domains. However, in many real-world applications, the source data is inaccessible and only a source model pre-trained on the source domain is available during the transfer process. Therefore, the traditional UDA methods cannot support the challenging setting. This paper developed a new hypothesis transfer method to achieve model adaptation with gradual knowledge distillation. Specifically, we first prepare a source model through training a deep network on the labeled source domain by supervised learning. Then, we transfer the source model to the unlabeled target domain by self-training. To implement gradual knowledge distillation, we sliced the self-training into several epochs and then used the soft pseudo-labels from the latest epoch to guide the current epoch. In this process, the soft labels were generated by a semantic fusion on a proposed geometry of the neighborhood. To regulate the self-training, we developed a new objective constructed on the neighborhood. Experiments on three benchmarks have confirmed the state-of-the-art results of our method.

IROS Conference 2021 Conference Paper

Multifunctional Robotic Glove with Active-Passive Training Modes for Hand Rehabilitation and Assistance

  • Yongkang Jiang
  • Diansheng Chen
  • Junlin Ma
  • Zhe Liu 0032
  • Yazhe Luo
  • Jian Li
  • Yingtian Li

Soft robotic gloves have shown great advantages in assisting individuals with hand pathologies to perform continuous exercises to restore their hand functions, which could considerably accelerate the rehabilitation process and reduce the costs. However, single rehabilitation mode, difficulty in achieving multiple degrees-of-freedom (DoF) motion, and the lack of high-fidelity feedback still challenge the development of soft robotic gloves. In this paper, we propose a novel design of a robotic glove based on soft-rigid hybrid joint actuators and minimal clutches. We first introduce structures and working principles of the proposed bending joint actuator in detail and then characterize the single joint actuator. Furthermore, we present a performance evaluation of the whole robotic glove in both active and passive modes. Preliminary experimental results showed that (1) in the active training mode, the tested human hand’s muscle effort needed to conduct gross finger flexion increased from 11. 16% to 42. 60% of the maximum value when the air pressure inside the minimal clutches changed from 0 kPa to 200 kPa; (2) in the passive mode, the 10-DoF robotic glove could assist the tested hand to perform various training exercises and grasp various objects with different hand postures. This paper focuses on the integrated design of multi-DoF structures and variable stiffness mechanisms, which will have an impact on the development of multifunctional soft robots and wearable devices.

EAAI Journal 2021 Journal Article

Octonion continuous orthogonal moments and their applications in color stereoscopic image reconstruction and zero-watermarking

  • Chunpeng Wang
  • Qixian Hao
  • Bin Ma
  • Xiaoming Wu
  • Jian Li
  • Zhiqiu Xia
  • Hongling Gao

Continuous orthogonal moments (COMs) are a type of effective image features widely used in various fields of image processing. However, most of the existing COMs are used for processing flat images and are not suitable for color stereoscopic images. For this reason, this paper first proposes an octonion theory applicable to color stereoscopic images, all color components of color stereoscopic images are coded by using the imaginary part of octonion, and all color components are processed as a whole, and the internal relations among all components are preserved. Then this paper combines the octonion theory with COMs to propose the octonion continuous orthogonal moments (OCOMs). The OCOMs fully reflect and retain the specific correlations between the left- and right-view components of color stereoscopic images, and provide good image description capability. Experimental results show that OCOMs have strong stability and good reconstruction performance when processing color stereoscopic images. Compared with other zero-watermarking methods, the zero-watermarking method embedded by OCOMs has stronger robustness.

YNIMG Journal 2021 Journal Article

Robust brain network identification from multi-subject asynchronous fMRI data

  • Jian Li
  • Jessica L. Wisnowski
  • Anand A. Joshi
  • Richard M. Leahy

We describe a novel method for robust identification of common brain networks and their corresponding temporal dynamics across subjects from asynchronous functional MRI (fMRI) using tensor decomposition. We first temporally align asynchronous fMRI data using the orthogonal BrainSync transform, allowing us to study common brain networks across sessions and subjects. We then map the synchronized fMRI data into a 3D tensor (vertices × time × subject/session). Finally, we apply Nesterov-accelerated adaptive moment estimation (Nadam) within a scalable and robust sequential Canonical Polyadic (CP) decomposition framework to identify a low rank tensor approximation to the data. As a result of CP tensor decomposition, we successfully identified twelve known brain networks with their corresponding temporal dynamics from 40 subjects using the Human Connectome Project's language task fMRI data without any prior information regarding the specific task designs. Seven of these networks show distinct subjects' responses to the language task with differing temporal dynamics; two show sub-components of the default mode network that exhibit deactivation during the tasks; the remaining three components reflect non-task-related activities. We compare results to those found using group independent component analysis (ICA) and canonical ICA. Bootstrap analysis demonstrates increased robustness of networks found using the CP tensor approach relative to ICA-based methods.

UAI Conference 2021 Conference Paper

Simple combinatorial algorithms for combinatorial bandits: corruptions and approximations

  • Haike Xu
  • Jian Li

We consider the stochastic combinatorial semi-bandit problem with adversarial corruptions. We provide a simple combinatorial algorithm that can achieve a regret of $\tilde{O}\left(C+d^2K/\Delta_{min}\right)$ where $C$ is the total amount of corruptions, $d$ is the maximal number of arms one can play in each round, $K$ is the number of arms. If one selects only one arm in each round, we achieves a regret of $\tilde{O}\left(C+\sum_{\Delta_i>0}(1/\Delta_i)\right)$. Our algorithm is combinatorial and improves on the previous combinatorial algorithm by [Gupta et al. , COLT2019] (their bound is $\tilde{O}\left(KC+\sum_{\Delta_i>0}(1/\Delta_i)\right)$), and almost matches the best known bounds obtained by [Zimmert et al. , ICML2019] and [Zimmert and Seldin, AISTATS2019] (up to logarithmic factor). Note that the algorithms in [Zimmert et al. , ICML2019] and [Zimmert and Seldin, AISTATS2019] require one to solve complex convex programs while our algorithm is combinatorial, very easy to implement, requires weaker assumptions and has very low oracle complexity and running time. We also study the setting where we only get access to an approximation oracle for the stochastic combinatorial semi-bandit problem. Our algorithm achieves an (approximation) regret bound of $\tilde{O}\left(d\sqrt{KT}\right)$. Our algorithm is very simple, only worse than the best known regret bound by $\sqrt{d}$, and has much lower oracle complexity than previous work.

TCS Journal 2020 Journal Article

Approximation algorithms for the connected sensor cover problem

  • Lingxiao Huang
  • Jian Li
  • Qicai Shi

We study the minimum connected sensor cover problem ( MIN - CSC ) and the budgeted connected sensor cover ( Budgeted - CSC ) problem, both motivated by important applications (e. g. , reduce the communication cost among sensors) in wireless sensor networks. In both problems, we are given a set of sensors and a set of target points in the Euclidean plane. In MIN - CSC, our goal is to find a set of sensors of minimum cardinality, such that all target points are covered, and all sensors can communicate with each other (i. e. , the communication graph is connected). We obtain a constant factor approximation algorithm, assuming that the ratio between the sensor radius and communication radius is bounded. In Budgeted - CSC problem, our goal is to choose a set of B sensors, such that the number of targets covered by the chosen sensors is maximized and the communication graph is connected. We also obtain a constant approximation under the same assumption.

AAAI Conference 2020 Conference Paper

Automated Spectral Kernel Learning

  • Jian Li
  • Yong Liu
  • Weiping Wang

The generalization performance of kernel methods is largely determined by the kernel, but spectral representations of stationary kernels are both input-independent and outputindependent, which limits their applications on complicated tasks. In this paper, we propose an efficient learning framework that incorporates the process of finding suitable kernels and model training. Using non-stationary spectral kernels and backpropagation w. r. t. the objective, we obtain favorable spectral representations that depends on both inputs and outputs. Further, based on Rademacher complexity, we derive data-dependent generalization error bounds, where we investigate the effect of those factors and introduce regularization terms to improve the performance. Extensive experimental results validate the effectiveness of the proposed algorithm and coincide with our theoretical findings.

AAAI Conference 2020 Conference Paper

Automatic Verification of Liveness Properties in the Situation Calculus

  • Jian Li
  • Yongmei Liu

In dynamic systems, liveness properties concern whether something good will eventually happen. Examples of liveness properties are termination of programs and goal achievability. In this paper, we consider the following theorem-proving problem: given an action theory and a goal, check whether the goal is achievable in every model of the action theory. We make the assumption that there are finitely many nonnumber objects. We propose to use mathematical induction to address this problem: we identify a natural number feature and prove by mathematical induction that for any values of the feature, the goal is achievable. Both the basis and induction steps are verified using first-order theorem provers. We propose a simple method to identify potential features which are the number of objects satisfying a certain formula by generating small models of the action theory and calling a classical planner to achieve the goal. We also propose to regress the goal via different actions and then verify whether the resulting goals are achievable. We implemented the proposed method and experimented with the blocks world domain and a number of other domains from the literature. Experimental results showed that most goals can be verified within a reasonable amount of time.

TCS Journal 2020 Journal Article

Discrete-time formulation, control, solution and verification of pendulum systems with zeroing neural dynamics

  • Yunong Zhang
  • Huanchang Huang
  • Min Yang
  • Jian Li

As a typical kind of nonlinear system, pendulum systems have drawn attention of numerous researchers for a very long time. This paper focuses mainly on dealing with the discrete-time tracking control problem of both the simple pendulum system and inverted-pendulum-on-a-cart (IPOAC) system. Based on zeroing neural dynamics (ZND), controllers of z2 type are designed respectively for the effective tracking control of the above two pendulum systems. Then, with the aim of possible digital hardware implementation, a 4-node discretization (4ND) formula, which is of square precision in terms of truncation error, is employed to discretize the continuous-time pendulum systems with high precision (i. e. , with discretization error being proportional to the cube of the sampling gap). By comparing with Euler-type discretization, simulative results further substantiate the feasibility, accuracy and superiority of the discrete-time control of both the simple pendulum system and IPOAC system with the 4ND formula.

AAAI Conference 2020 Conference Paper

Fast Learning of Temporal Action Proposal via Dense Boundary Generator

  • Chuming Lin
  • Jian Li
  • Yabiao Wang
  • Ying Tai
  • Donghao Luo
  • Zhipeng Cui
  • Chengjie Wang
  • Jilin Li

Generating temporal action proposals remains a very challenging problem, where the main issue lies in predicting precise temporal proposal boundaries and reliable action confidence in long and untrimmed real-world videos. In this paper, we propose an efficient and unified framework to generate temporal action proposals named Dense Boundary Generator (DBG), which draws inspiration from boundary-sensitive methods and implements boundary classification and action completeness regression for densely distributed proposals. In particular, the DBG consists of two modules: Temporal boundary classification (TBC) and Action-aware completeness regression (ACR). The TBC aims to provide two temporal boundary confidence maps by low-level two-stream features, while the ACR is designed to generate an action completeness score map by high-level action-aware features. Moreover, we introduce a dual stream BaseNet (DSB) to encode RGB and optical flow information, which helps to capture discriminative boundary and actionness features. Extensive experiments on popular benchmarks ActivityNet-1. 3 and THUMOS14 demonstrate the superiority of DBG over the state-of-the-art proposal generator (e. g. , MGG and BMN).

TCS Journal 2020 Journal Article

From mathematical equivalence such as Ma equivalence to generalized Zhang equivalency including gradient equivalency

  • Yunong Zhang
  • Min Yang
  • Binbin Qiu
  • Jian Li
  • Mingjie Zhu

The authors carried out time-varying problems solving (TVPS) including robot problems solving in 2001, and began to figure out the reasons for the problems solving via diverse layers. After eight years' thinking, i. e. , in 2009, the authors began to manifest, put forward and carry out the thought of “physical equivalency”. By another eight years' practicing and experimenting, i. e. , in 2017, the authors basically finished establishing the framework of Zhang equivalency. Now, it is the time to establish the complete theory in a brief manner. Therefore, concepts about mathematical equivalence simply termed equivalence are presented firstly including Ma equivalence (especially for robotics), and then concepts about physical equivalency simply termed equivalency are proposed. Meanwhile, concepts about Zhang equivalency as a kind of equivalency are further proposed, and concepts about gradient-dynamics equivalency simply termed gradient equivalency as a kind of equivalency are proposed as well. Furthermore, two specific applications are considered and investigated, which substantiate the efficacy of Zhang equivalency.

NeurIPS Conference 2020 Conference Paper

Improved Algorithms for Convex-Concave Minimax Optimization

  • Yuanhao Wang
  • Jian Li

This paper studies minimax optimization problems $\min_\x \max_\y f(\x, \y)$, where $f(\x, \y)$ is $m_\x$-strongly convex with respect to $\x$, $m_\y$-strongly concave with respect to $\y$ and $(L_\x, L_{\x\y}, L_\y)$-smooth. Zhang et al. \cite{zhang2019lower} provided the following lower bound of the gradient complexity for any first-order method: $\Omega\Bigl(\sqrt{\frac{L_\x}{m_\x}+\frac{L_{\x\y}^2}{m_\x m_\y}+\frac{L_\y}{m_\y}}\ln(1/\epsilon)\Bigr). $ This paper proposes a new algorithm and proved a gradient complexity bound of $\Tilde{O}\Bigl(\sqrt{\frac{L_\x}{m_\x}+\frac{L\cdot L_{\x\y}}{m_\x m_\y}+\frac{L_\y}{m_\y}}\ln\left(1/\epsilon\right)\Bigr), $ where $L=\max\{L_\x, L_{\x\y}, L_\y\}$. This improves over the best known upper bound $\Tilde{O}\left(\sqrt{\nicefrac{L^2}{m_\x m_\y}} \ln^3\left(1/\epsilon\right)\right)$ by Lin et al. \cite{lin2020near}. Our bound achieves linear convergence rate and tighter dependency on condition numbers, especially when $L_{\x\y}\ll L$ (i. e. , the weak interaction regime). Via simple reduction, our new bound also implies improved bounds for strongly convex-concave problems and convex-concave problems. When $f$ is quadratic, we can further improve the bound to $O\Bigl(\sqrt{\frac{L_\x}{m_\x}+\frac{L_{\x\y}^2}{m_\x m_\y}+\frac{L_\y}{m_\y}}\left(\frac{L^2}{m_\x m_\y}\right)^{o(1)}\ln(1/\epsilon)\Bigr)$, which matches the lower bound up to a sub-polynomial factor.

NeurIPS Conference 2020 Conference Paper

Kalman Filtering Attention for User Behavior Modeling in CTR Prediction

  • Hu Liu
  • Jing Lu
  • Xiwei Zhao
  • Sulong Xu
  • Hao Peng
  • Yutong Liu
  • Zehua Zhang
  • Jian Li

Click-through rate (CTR) prediction is one of the fundamental tasks for e-commerce search engines. As search becomes more personalized, it is necessary to capture the user interest from rich behavior data. Existing user behavior modeling algorithms develop different attention mechanisms to emphasize query-relevant behaviors and suppress irrelevant ones. Despite being extensively studied, these attentions still suffer from two limitations. First, conventional attentions mostly limit the attention field only to a single user's behaviors, which is not suitable in e-commerce where users often hunt for new demands that are irrelevant to any historical behaviors. Second, these attentions are usually biased towards frequent behaviors, which is unreasonable since high frequency does not necessarily indicate great importance. To tackle the two limitations, we propose a novel attention mechanism, termed Kalman Filtering Attention (KFAtt), that considers the weighted pooling in attention as a maximum a posteriori (MAP) estimation. By incorporating a priori, KFAtt resorts to global statistics when few user behaviors are relevant. Moreover, a frequency capping mechanism is incorporated to correct the bias towards frequent behaviors. Offline experiments on both benchmark and a 10 billion scale real production dataset, together with an Online A/B test, show that KFAtt outperforms all compared state-of-the-arts. KFAtt has been deployed in the ranking system of JD. com, one of the largest B2C e-commerce websites in China, serving the main traffic of hundreds of millions of active users.

AAAI Conference 2020 Conference Paper

Neuron Interaction Based Representation Composition for Neural Machine Translation

  • Jian Li
  • Xing Wang
  • Baosong Yang
  • Shuming Shi
  • Michael R. Lyu
  • Zhaopeng Tu

Recent NLP studies reveal that substantial linguistic information can be attributed to single neurons, i. e. , individual dimensions of the representation vectors. We hypothesize that modeling strong interactions among neurons helps to better capture complex information by composing the linguistic properties embedded in individual neurons. Starting from this intuition, we propose a novel approach to compose representations learned by different components in neural machine translation (e. g. , multi-layer networks or multihead attention), based on modeling strong interactions among neurons in the representation vectors. Specifically, we leverage bilinear pooling to model pairwise multiplicative interactions among individual neurons, and a low-rank approximation to make the model computationally feasible. We further propose extended bilinear pooling to incorporate first-order representations. Experiments on WMT14 English⇒German and English⇒French translation tasks show that our model consistently improves performances over the SOTA TRANS- FORMER baseline. Further analyses demonstrate that our approach indeed captures more syntactic and semantic information as expected.

NeurIPS Conference 2020 Conference Paper

Online Algorithms for Multi-shop Ski Rental with Machine Learned Advice

  • Shufan Wang
  • Jian Li
  • Shiqiang Wang

We study the problem of augmenting online algorithms with machine learned (ML) advice. In particular, we consider the \emph{multi-shop ski rental} (MSSR) problem, which is a generalization of the classical ski rental problem. In MSSR, each shop has different prices for buying and renting a pair of skis, and a skier has to make decisions on when and where to buy. We obtain both deterministic and randomized online algorithms with provably improved performance when either a single or multiple ML predictions are used to make decisions. These online algorithms have no knowledge about the quality or the prediction error type of the ML prediction. The performance of these online algorithms are robust to the poor performance of the predictors, but improve with better predictions. Extensive experiments using both synthetic and real world data traces verify our theoretical observations and show better performance against algorithms that purely rely on online decision making.

AAAI Conference 2020 Conference Paper

Policy Search by Target Distribution Learning for Continuous Control

  • Chuheng Zhang
  • Yuanqi Li
  • Jian Li

It is known that existing policy gradient methods (such as vanilla policy gradient, PPO, A2C) may suffer from overly large gradients when the current policy is close to deterministic, leading to an unstable training process. We show that such instability can happen even in a very simple environment. To address this issue, we propose a new method, called target distribution learning (TDL), for policy improvement in reinforcement learning. TDL alternates between proposing a target distribution and training the policy network to approach the target distribution. TDL is more effective in constraining the KL divergence between updated policies, and hence leads to more stable policy improvements over iterations. Our experiments show that TDL algorithms perform comparably to (or better than) state-of-the-art algorithms for most continuous control tasks in the MuJoCo environment while being more stable in training.

IJCAI Conference 2019 Conference Paper

AddGraph: Anomaly Detection in Dynamic Graph Using Attention-based Temporal GCN

  • Li Zheng
  • Zhenpeng Li
  • Jian Li
  • Zhao Li
  • Jun Gao

Anomaly detection in dynamic graphs becomes very critical in many different application scenarios, e. g. , recommender systems, while it also raises huge challenges due to the high flexible nature of anomaly and lack of sufficient labelled data. It is better to learn the anomaly patterns by considering all possible features including the structural, content and temporal features, rather than utilizing heuristic rules over the partial features. In this paper, we propose AddGraph, a general end-to-end anomalous edge detection framework using an extended temporal GCN (Graph Convolutional Network) with an attention model, which can capture both long-term patterns and the short-term patterns in dynamic graphs. In order to cope with insufficient explicit labelled data, we employ the negative sampling and margin loss in training of AddGraph in a semi-supervised fashion. We conduct extensive experiments on real-world datasets, and illustrate that AddGraph can outperform the state-of-the-art competitors in anomaly detection significantly.

IJCAI Conference 2019 Conference Paper

Approximate Manifold Regularization: Scalable Algorithm and Generalization Analysis

  • Jian Li
  • Yong Liu
  • Rong Yin
  • Weiping Wang

Graph-based semi-supervised learning is one of the most popular and successful semi-supervised learning approaches. Unfortunately, it suffers from high time and space complexity, at least quadratic with the number of training samples. In this paper, we propose an efficient graph-based semi-supervised algorithm with a sound theoretical guarantee. The proposed method combines Nystrom subsampling and preconditioned conjugate gradient descent, substantially improving computational efficiency and reducing memory requirements. Extensive empirical results reveal that our method achieves the state-of-the-art performance in a short time even with limited computing resources.

AAAI Conference 2019 Conference Paper

Context-Aware Self-Attention Networks

  • Baosong Yang
  • Jian Li
  • Derek F. Wong
  • Lidia S. Chao
  • Xing Wang
  • Zhaopeng Tu

Self-attention model has shown its flexibility in parallel computation and the effectiveness on modeling both long- and short-term dependencies. However, it calculates the dependencies between representations without considering the contextual information, which has proven useful for modeling dependencies among neural representations in various natural language tasks. In this work, we focus on improving self-attention networks through capturing the richness of context. To maintain the simplicity and flexibility of the selfattention networks, we propose to contextualize the transformations of the query and key layers, which are used to calculate the relevance between elements. Specifically, we leverage the internal representations that embed both global and deep contexts, thus avoid relying on external resources. Experimental results on WMT14 English⇒German and WMT17 Chinese⇒English translation tasks demonstrate the effectiveness and universality of the proposed methods. Furthermore, we conducted extensive analyses to quantify how the context vectors participate in the self-attention model.

IJCAI Conference 2019 Conference Paper

Gradient Boosting with Piece-Wise Linear Regression Trees

  • Yu Shi
  • Jian Li
  • Zhize Li

Gradient Boosted Decision Trees (GBDT) is a very successful ensemble learning algorithm widely used across a variety of applications. Recently, several variants of GBDT training algorithms and implementations have been designed and heavily optimized in some very popular open sourced toolkits including XGBoost, LightGBM and CatBoost. In this paper, we show that both the accuracy and efficiency of GBDT can be further enhanced by using more complex base learners. Specifically, we extend gradient boosting to use piecewise linear regression trees (PL Trees), instead of piecewise constant regression trees, as base learners. We show that PL Trees can accelerate convergence of GBDT and improve the accuracy. We also propose some optimization tricks to substantially reduce the training time of PL Trees, with little sacrifice of accuracy. Moreover, we propose several implementation techniques to speedup our algorithm on modern computer architectures with powerful Single Instruction Multiple Data (SIMD) parallelism. The experimental results show that GBDT with PL Trees can provide very competitive testing accuracy with comparable or less training time.

IJCAI Conference 2019 Conference Paper

Multi-Class Learning using Unlabeled Samples: Theory and Algorithm

  • Jian Li
  • Yong Liu
  • Rong Yin
  • Weiping Wang

In this paper, we investigate the generalization performance of multi-class classification, for which we obtain a shaper error bound by using the notion of local Rademacher complexity and additional unlabeled samples, substantially improving the state-of-the-art bounds in existing multi-class learning methods. The statistical learning motivates us to devise an efficient multi-class learning framework with the local Rademacher complexity and Laplacian regularization. Coinciding with the theoretical analysis, experimental results demonstrate that the stated approach achieves better performance.

NeurIPS Conference 2018 Conference Paper

A Simple Proximal Stochastic Gradient Method for Nonsmooth Nonconvex Optimization

  • Zhize Li
  • Jian Li

We analyze stochastic gradient algorithms for optimizing nonconvex, nonsmooth finite-sum problems. In particular, the objective function is given by the summation of a differentiable (possibly nonconvex) component, together with a possibly non-differentiable but convex component. We propose a proximal stochastic gradient algorithm based on variance reduction, called ProxSVRG+. Our main contribution lies in the analysis of ProxSVRG+. It recovers several existing convergence results and improves/generalizes them (in terms of the number of stochastic gradient oracle calls and proximal oracle calls). In particular, ProxSVRG+ generalizes the best results given by the SCSG algorithm, recently proposed by [Lei et al. , NIPS'17] for the smooth nonconvex case. ProxSVRG+ is also more straightforward than SCSG and yields simpler analysis. Moreover, ProxSVRG+ outperforms the deterministic proximal gradient descent (ProxGD) for a wide range of minibatch sizes, which partially solves an open problem proposed in [Reddi et al. , NIPS'16]. Also, ProxSVRG+ uses much less proximal oracle calls than ProxSVRG [Reddi et al. , NIPS'16]. Moreover, for nonconvex functions satisfied Polyak-\L{}ojasiewicz condition, we prove that ProxSVRG+ achieves a global linear convergence rate without restart unlike ProxSVRG. Thus, it can \emph{automatically} switch to the faster linear convergence in some regions as long as the objective function satisfies the PL condition locally in these regions. Finally, we conduct several experiments and the experimental results are consistent with the theoretical results.

YNIMG Journal 2018 Journal Article

Are you thinking what I'm thinking? Synchronization of resting fMRI time-series across subjects

  • Anand A. Joshi
  • Minqi Chong
  • Jian Li
  • Soyoung Choi
  • Richard M. Leahy

We describe BrainSync, an orthogonal transform that allows direct comparison of resting fMRI (rfMRI) time-series across subjects. For this purpose, we exploit the geometry of the rfMRI signal space to propose a novel orthogonal transformation that synchronizes rfMRI time-series across sessions and subjects. When synchronized, rfMRI signals become approximately equal at homologous locations across subjects. The method is based on the observation that rfMRI data exhibit similar connectivity patterns across subjects, as reflected in the pairwise correlations between different brain regions. We show that if the data for two subjects have similar correlation patterns then their time courses can be approximately synchronized by an orthogonal transformation. This transform is unique, invertible, efficient to compute, and preserves the connectivity structure of the original data for all subjects. Analogously to image registration, where we spatially align structural brain images, this temporal synchronization of brain signals across a population, or within-subject across sessions, facilitates cross-sectional and longitudinal studies of rfMRI data. The utility of the BrainSync transform is illustrated through demonstrative simulations and applications including quantification of rfMRI variability across subjects and sessions, cortical functional parcellation across a population, timing recovery in task fMRI data, comparison of task and resting state data, and an application to complex naturalistic stimuli for annotation prediction.

NeurIPS Conference 2018 Conference Paper

BRITS: Bidirectional Recurrent Imputation for Time Series

  • Wei Cao
  • Dong Wang
  • Jian Li
  • Hao Zhou
  • Lei Li
  • Yitan Li

Time series are widely used as signals in many classification/regression tasks. It is ubiquitous that time series contains many missing values. Given multiple correlated time series data, how to fill in missing values and to predict their class labels? Existing imputation methods often impose strong assumptions of the underlying data generating process, such as linear dynamics in the state space. In this paper, we propose BRITS, a novel method based on recurrent neural networks for missing value imputation in time series data. Our proposed method directly learns the missing values in a bidirectional recurrent dynamical system, without any specific assumption. The imputed values are treated as variables of RNN graph and can be effectively updated during the backpropagation. BRITS has three advantages: (a) it can handle multiple correlated missing values in time series; (b) it generalizes to time series with nonlinear dynamics underlying; (c) it provides a data-driven imputation procedure and applies to general settings with missing data. We evaluate our model on three real-world datasets, including an air quality dataset, a health-care data, and a localization data for human activity. Experiments show that our model outperforms the state-of-the-art methods in both imputation and classification/regression accuracies.

IJCAI Conference 2018 Conference Paper

Code Completion with Neural Attention and Pointer Networks

  • Jian Li
  • Yue Wang
  • Michael R. Lyu
  • Irwin King

Intelligent code completion has become an essential research task to accelerate modern software development. To facilitate effective code completion for dynamically-typed programming languages, we apply neural language models by learning from large codebases, and develop a tailored attention mechanism for code completion. However, standard neural language models even with attention mechanism cannot correctly predict the out-of-vocabulary (OoV) words that restrict the code completion performance. In this paper, inspired by the prevalence of locally repeated terms in program source code, and the recently proposed pointer copy mechanism, we propose a pointer mixture network for better predicting OoV words in code completion. Based on the context, the pointer mixture network learns to either generate a within-vocabulary word through an RNN component, or regenerate an OoV word from local context through a pointer component. Experiments on two benchmarked datasets demonstrate the effectiveness of our attention mechanism and pointer mixture network on the code completion task.

NeurIPS Conference 2018 Conference Paper

Multi-Class Learning: From Theory to Algorithm

  • Jian Li
  • Yong Liu
  • Rong Yin
  • Hua Zhang
  • Lizhong Ding
  • Weiping Wang

In this paper, we study the generalization performance of multi-class classification and obtain a shaper data-dependent generalization error bound with fast convergence rate, substantially improving the state-of-art bounds in the existing data-dependent generalization analysis. The theoretical analysis motivates us to devise two effective multi-class kernel learning algorithms with statistical guarantees. Experimental results show that our proposed methods can significantly outperform the existing multi-class classification methods.

TCS Journal 2018 Journal Article

Near-linear time approximation schemes for geometric maximum coverage

  • Kai Jin
  • Jian Li
  • Haitao Wang
  • Bowei Zhang
  • Ningye Zhang

We study approximation algorithms for the following geometric version of the maximum coverage problem: Let P be a set of n weighted points in the plane. Let D represent a planar object, such as a rectangle, or a disk. We want to place m copies of D such that the sum of the weights of the points in P covered by these copies is maximized. For any fixed ε > 0, we present efficient approximation schemes that can find a ( 1 − ε ) -approximation to the optimal solution. In particular, for m = 1 and for the special case where D is a rectangle, our algorithm runs in time O ( n log ⁡ ( 1 ε ) ), improving on the previous result. For m > 1 and the rectangular case, our algorithm runs in O ( n ε log ⁡ ( 1 ε ) + m ε log ⁡ m + m ( 1 ε ) O ( min ⁡ ( m, 1 ε ) ) ) time. For a more general class of shapes (including disks, polygons with O ( 1 ) edges), our algorithm runs in O ( n ( 1 ε ) O ( 1 ) + m ϵ log ⁡ m + m ( 1 ε ) O ( min ⁡ ( m, 1 ε 2 ) ) ) time.

AAAI Conference 2018 Conference Paper

When Will You Arrive? Estimating Travel Time Based on Deep Neural Networks

  • Dong Wang
  • Junbo Zhang
  • Wei Cao
  • Jian Li
  • Yu Zheng

Estimating the travel time of any path (denoted by a sequence of connected road segments) in a city is of great importance to traffic monitoring, route planning, ridesharing, taxi/Uber dispatching, etc. However, it is a very challenging problem, affected by diverse complex factors, including spatial correlations, temporal dependencies, external conditions (e. g. weather, traffic lights). Prior work usually focuses on estimating the travel times of individual road segments or sub-paths and then summing up these times, which leads to an inaccurate estimation because such approaches do not consider road intersections/traffic lights, and local errors may accumulate. To address these issues, we propose an end-to-end Deep learning framework for Travel Time Estimation (called DeepTTE) that estimates the travel time of the whole path directly. More specifically, we present a geo-convolution operation by integrating the geographic information into the classical convolution, capable of capturing spatial correlations. By stacking recurrent unit on the geo-convoluton layer, our DeepTTE can capture the temporal dependencies as well. A multi-task learning component is given on the top of DeepTTE, that learns to estimate the travel time of both the entire path and each local path simultaneously during the training phase. Extensive experiments on two trajectory datasets show our DeepTTE significantly outperforms the state-of-the-art methods.

IJCAI Conference 2017 Conference Paper

Efficient Kernel Selection via Spectral Analysis

  • Jian Li
  • Yong Liu
  • Hailun Lin
  • Yinliang Yue
  • Weiping Wang

Kernel selection is a fundamental problem of kernel methods. Existing measures for kernel selection either provide less theoretical guarantee or have high computational complexity. In this paper, we propose a novel kernel selection criterion based on a newly defined spectral measure of a kernel matrix, with sound theoretical foundation and high computational efficiency. We first show that the spectral measure can be used to derive generalization bounds for some kernel-based algorithms. By minimizing the derived generalization bounds, we propose the kernel selection criterion with spectral measure. Moreover, we demonstrate that the popular minimum graph cut and maximum mean discrepancy are two special cases of the proposed criterion. Experimental results on lots of data sets show that our proposed criterion can not only give the comparable results as the state-of-the-art criterion, but also significantly improve the efficiency.

IJCAI Conference 2017 Conference Paper

Single-Pass PCA of Large High-Dimensional Data

  • Wenjian Yu
  • Yu Gu
  • Jian Li
  • Shenghua Liu
  • Yaohang Li

Principal component analysis (PCA) is a fundamental dimension reduction tool in statistics and machine learning. For large and high-dimensional data, computing the PCA (i. e. , the top singular vectors of the data matrix) becomes a challenging task. In this work, a single-pass randomized algorithm is proposed to compute PCA with only one pass over the data. It is suitable for processing extremely large and high-dimensional data stored in slow memory (hard disk) or the data generated in a streaming fashion. Experiments with synthetic and real data validate the algorithm's accuracy, which has orders of magnitude smaller error than an existing single-pass algorithm. For a set of high-dimensional data stored as a 150 GB file, the algorithm is able to compute the first 50 principal components in just 24 minutes on a typical 24-core computer, with less than 1 GB memory cost.

NeurIPS Conference 2016 Conference Paper

Combinatorial Multi-Armed Bandit with General Reward Functions

  • Wei Chen
  • Wei Hu
  • Fu Li
  • Jian Li
  • Yu Liu
  • Pinyan Lu

In this paper, we study the stochastic combinatorial multi-armed bandit (CMAB) framework that allows a general nonlinear reward function, whose expected value may not depend only on the means of the input random variables but possibly on the entire distributions of these variables. Our framework enables a much larger class of reward functions such as the $\max()$ function and nonlinear utility functions. Existing techniques relying on accurate estimations of the means of random variables, such as the upper confidence bound (UCB) technique, do not work directly on these functions. We propose a new algorithm called stochastically dominant confidence bound (SDCB), which estimates the distributions of underlying random variables and their stochastically dominant confidence bounds. We prove that SDCB can achieve $O(\log T)$ distribution-dependent regret and $\tilde{O}(\sqrt{T})$ distribution-independent regret, where $T$ is the time horizon. We apply our results to the $K$-MAX problem and expected utility maximization problems. In particular, for $K$-MAX, we provide the first polynomial-time approximation scheme (PTAS) for its offline problem, and give the first $\tilde{O}(\sqrt T)$ bound on the $(1-\epsilon)$-approximation regret of its online problem, for any $\epsilon>0$.

YNIMG Journal 2016 Journal Article

High-definition tDCS alters impulsivity in a baseline-dependent manner

  • Bo Shen
  • Yunlu Yin
  • Jiashu Wang
  • Xiaolin Zhou
  • Samuel M. McClure
  • Jian Li

In intertemporal choice (ITC), people discount future rewards in proportion to the time delay until reward receipt. Despite recent non-invasive brain stimulation studies suggesting a general causal link between dorsolateral prefrontal cortex (dlPFC) activity and ITC impulsivity, results regarding the functional specificity of dlPFC are mixed. We used high-definition transcranial direct current stimulation (HD-tDCS) to map changes in causal impulsivity through bi-directional modulation of left and right dlPFC during ITC. Model-free and model-based analyses demonstrated that anodal and cathodal stimulation of left dlPFC, but not right dlPFC, decreased and increased impulsivity, respectively. Critically, an individual differences analysis revealed that modulation of impulsivity was contingent on participants' baseline impulsivity. Overall, our results might reconcile the discrepancies in the existing literature and suggest a baseline-dependent role for left dlPFC during ITC.

TCS Journal 2016 Journal Article

Range queries on uncertain data

  • Jian Li
  • Haitao Wang

Given a set P of n uncertain points on the real line, each represented by its one-dimensional probability density function, we consider the problem of building data structures on P to answer range queries of the following three types for any query interval I: (1) top-1 query: find a point in P that lies in I with the highest probability, (2) top-k query: given any integer k ≤ n as part of the query, return the k points in P that lie in I with the highest probabilities, and (3) threshold query: given any threshold τ as part of the query, return all points of P that lie in I with probabilities at least τ. We present data structures for these range queries with linear or nearly linear space and efficient query time.

TCS Journal 2015 Journal Article

Efficient algorithms for the one-dimensional k-center problem

  • Danny Z. Chen
  • Jian Li
  • Haitao Wang

We consider the problem of finding k centers for n weighted points on a real line. This (weighted) k-center problem was solved in O ( n log ⁡ n ) time previously by using Cole's parametric search and other complicated approaches. In this paper, we present an easier O ( n log ⁡ n ) time algorithm that avoids the parametric search, and in certain special cases our algorithm solves the problem in O ( n ) time. In addition, our techniques involve developing interesting data structures for processing queries that find a lowest point in the common intersection of a certain subset of half-planes. This subproblem is interesting in its own right and our solution for it may find other applications as well.

NeurIPS Conference 2015 Conference Paper

On Top-k Selection in Multi-Armed Bandits and Hidden Bipartite Graphs

  • Wei Cao
  • Jian Li
  • Yufei Tao
  • Zhize Li

This paper discusses how to efficiently choose from $n$ unknowndistributions the $k$ ones whose means are the greatest by a certainmetric, up to a small relative error. We study the topic under twostandard settings---multi-armed bandits and hidden bipartitegraphs---which differ in the nature of the input distributions. In theformer setting, each distribution can be sampled (in the i. i. d. manner) an arbitrary number of times, whereas in the latter, eachdistribution is defined on a population of a finite size $m$ (andhence, is fully revealed after $m$ samples). For both settings, weprove lower bounds on the total number of samples needed, and proposeoptimal algorithms whose sample complexities match those lower bounds.

NeurIPS Conference 2015 Conference Paper

Stochastic Online Greedy Learning with Semi-bandit Feedbacks

  • Tian Lin
  • Jian Li
  • Wei Chen

The greedy algorithm is extensively studied in the field of combinatorial optimization for decades. In this paper, we address the online learning problem when the input to the greedy algorithm is stochastic with unknown parameters that have to be learned over time. We first propose the greedy regret and $\epsilon$-quasi greedy regret as learning metrics comparing with the performance of offline greedy algorithm. We then propose two online greedy learning algorithms with semi-bandit feedbacks, which use multi-armed bandit and pure exploration bandit policies at each level of greedy learning, one for each of the regret metrics respectively. Both algorithms achieve $O(\log T)$ problem-dependent regret bound ($T$ being the time horizon) for a general class of combinatorial structures and reward functions that allow greedy solutions. We further show that the bound is tight in $T$ and other problem instance parameters.

ICRA Conference 2014 Conference Paper

Contact dynamics of massage compliant robotic arm and its coupled stability

  • Yuancan Huang
  • Philippe Souères
  • Jian Li

In this paper, contact dynamics of robot massage is described by the port-Hamiltonian modelling approach. In order to capture accurately the inherent characteristics of the human body in lumped-parameter manners, the conventional linear Kelvin-Voigt models are replaced by the nonlinear Hunt-Crossley models. As an application of the contact dynamics, coupled stability of compliant robotic arm with impedance control is theoretically analyzed from energetic viewpoints. Experiments are done to verify the massage stability. The proposed contact dynamics evidently has great potential on performance improvement of robot massage, which will be our research subject.

IROS Conference 2013 Conference Paper

Design and control of anthropomorphic BIT soft arms for TCM remedial massage

  • Yuancan Huang
  • Jian Li
  • Qiang Huang 0002
  • Changxin Liu

For reproducing the manipulation of TCM remedial massage and meanwhile guaranteeing safety, a 4-DOF anthropomorphic BIT soft arm with integrated elastic joints is developed, and a passivity-based impedance control is used. Due to their series elasticity, the integrated joints may minimize large forces which occur during accidental impacts and, further, may offer more accurate and stable force control and a capacity for energy storage. Then, human expert's fingertip force curve in the process of massage therapy is acquired in vivo by a dedicated measurement device. Three massage techniques, pressing, kneading and plucking, are implemented by the soft arm, respectively, on torso model in vitro and on human body in vivo. Experimental results show that the developed robotic arm can effectively imitates the TCM remedial massage techniques.

RLDM Conference 2013 Conference Abstract

How instructed knowledge shapes aversive learning

  • Lauren Atlas
  • Bradley Doll
  • Nathaniel Daw
  • Jian Li

In humans, expectations reflect prior experience and instructed knowledge. Most models of aver- sive learning make predictions about brain responses as a function of reinforcement alone. Recent studies of reward learning indicate that striatal learning is modulated when participants are instructed about stimulus contingencies. The aim of this study was to test whether instructed knowledge modulates associative fear learning. Participants performed a Pavlovian aversive learning paradigm. Two cues were presented: One (the CS+) was paired with shock on 30 % of trials, whereas the second (the CS-) was never paired with shock. Fol- lowing 20 trials, contingencies reversed. There were three reversals across the session. Participants were assigned to two groups: The Instructed Group was informed about contingencies prior to learning and upon each reversal, whereas the Feedback Group received no information. We analyzed skin conductance responses (SCRs) and brain responses to cues. Fear expression tracked con- tingency reversals (i. e. larger SCRs for current CS+ than CS-), and the Instructed Group showed stronger differential responses. The Instructed Group showed greater activation in right DLPFC, while the Feedback Group showed greater activation in bilateral striatum. We fit a quantitative model with a dynamic learning rate to SCRs to isolate the timecourse of learning in the Feedback Group, focusing on prediction error and associability. We then tested whether Instructions modulated the neural correlates of feedback-driven sig- nals. We observed group differences in bilateral ventral striatum, such that only the Feedback Group showed striatal prediction errors. These results reveal that instructed knowledge influences aversive learning. Instructions enhance fear acqui- sition and expression, and prediction errors are not observed when instructions are veridical. The DLPFC is likely to play a key role in maintaining instructions, which in turn modulate fear expression.

YNIMG Journal 2011 Journal Article

Parallel contributions of distinct human memory systems during probabilistic learning

  • Kathryn C. Dickerson
  • Jian Li
  • Mauricio R. Delgado

Regions within the medial temporal lobe and basal ganglia are thought to subserve distinct memory systems underlying declarative and nondeclarative processes, respectively. One question of interest is how these multiple memory systems interact during learning to contribute to goal directed behavior. While some hypotheses suggest that regions such as the striatum and the hippocampus interact in a competitive manner, alternative views posit that these structures may operate in a parallel manner to facilitate learning. In the current experiment, we probed the functional connectivity between regions in the striatum and hippocampus in the human brain during an event related probabilistic learning task that varied with respect to type of difficulty (easy or hard cues) and type of learning (via feedback or observation). We hypothesized that the hippocampus and striatum would interact in a parallel manner during learning. We identified regions of interest (ROI) in the striatum and hippocampus that showed an effect of cue difficulty during learning and found that such ROIs displayed a similar pattern of blood oxygen level dependent (BOLD) responses, irrespective of learning type, and were functionally correlated as assessed by a Granger causality analysis. Given the connectivity of both structures with dopaminergic midbrain centers, we further applied a reinforcement learning algorithm often used to highlight the role of dopamine in human reward related learning paradigms. Activity in both the striatum and hippocampus positively correlated with a prediction error signal during feedback learning. These results suggest that distinct human memory systems operate in parallel during probabilistic learning, and may act synergistically particularly when a violation of expectation occurs, to jointly contribute to learning and decision making.

NeurIPS Conference 2007 Conference Paper

Parallelizing Support Vector Machines on Distributed Computers

  • Kaihua Zhu
  • Hao Wang
  • Hongjie Bai
  • Jian Li
  • Zhihuan Qiu
  • Hang Cui
  • Edward Chang

Support Vector Machines (SVMs) suffer from a widely recognized scalability problem in both memory use and computational time. To improve scalability, we have developed a parallel SVM algorithm (PSVM), which reduces memory use through performing a row-based, approximate matrix factorization, and which loads only essential data to each machine to perform parallel computation. Let $n$ denote the number of training instances, $p$ the reduced matrix dimension after factorization ($p$ is significantly smaller than $n$), and $m$ the number of machines. PSVM reduces the memory requirement from $\MO$($n^2$) to $\MO$($np/m$), and improves computation time to $\MO$($np^2/m$). Empirical studies on up to $500$ computers shows PSVM to be effective.

v2026.09.13