Arrow Research search

Author name cluster

Siqi Cai

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
1 author row

Possible papers

8

AAAI Conference 2026 Conference Paper

Rep Deep & Machine Learning: Exemplar-Free Continual Video Action Recognition via Slow-Fast Collaborative Learning

  • Xueyi Zhang
  • Chengwei Zhang
  • Zheng Li
  • Xiyu Wang
  • Siqi Cai
  • Mingrui Lao
  • Yanming Guo
  • Huiping Zhuang

In real-world applications, video action recognition models must continuously learn new action categories while retaining previously acquired knowledge. However, most existing approaches rely on storing historical data for replay, which introduces storage burdens and raises data privacy concerns. To address these challenges, we investigate the problem of Exemplar-Free Continual Video Action Recognition (EF-CVAR) and propose a novel framework named Slow-Fast Collaborative Learning (SFCL). SFCL integrates two complementary learning paradigms: a slow branch based on gradient-driven deep learning, which provides strong adaptability to new tasks, and a fast branch based on analytic learning (e.g., Recursive Least Squares), which efficiently preserves old knowledge without requiring access to past samples. To enable effective collaboration between the two branches, we design the Slow-Fast Dynamic Re-parameterization (SFDR) mechanism for adaptive fusion, and the Knowledge Reflection Mechanism (KRM), which mitigates forgetting and task-recency bias via pseudo-feature generation and dual-level knowledge distillation. Extensive experiments on UCF101, HMDB51, and Something-Something V2 demonstrate that SFCL achieves superior performance compared to existing replay-based methods, despite being exemplar-free. Notably, in long-duration continual learning scenarios, SFCL exhibits remarkable robustness, achieving up to a 30.39\% improvement in accuracy over baselines while maintaining a low forgetting rate, highlighting its scalability and effectiveness in real-world video recognition tasks.

NeurIPS Conference 2025 Conference Paper

Listening to the Brain: Multi-Band sEEG Auditory Reconstruction via Dynamic Spatio-Temporal Hypergraphs

  • Xueyi Zhang
  • Ruicong Wang
  • Jialu Sun
  • Siqi Cai
  • Haizhou Li

Speech is a fundamental form of human communication, and speech perception constitutes the initial stage of language comprehension. Although brain-to-speech interface technologies have made significant progress in recent years, most existing studies focus on neural decoding during speech production. Such approaches heavily rely on articulatory motor regions, rendering them unsuitable for individuals with speech motor impairments, such as those with aphasia or locked-in syndrome. To address this limitation, we construct and release NeuroListen, the first publicly available stereo-electroencephalography (sEEG) dataset specifically designed for auditory reconstruction. It contains over 10 hours of neural–speech paired recordings from 5 clinical participants, covering a wide range of semantic categories. Building on this dataset, we propose HyperSpeech, a multi-band neural decoding framework that employs dynamic spatio-temporal hypergraph neural networks to capture high-order dependencies across frequency, spatial, and temporal dimensions. Experimental results demonstrate that HyperSpeech significantly outperforms existing methods across multiple objective speech quality metrics, and achieves superior performance in human subjective evaluations, validating its effectiveness and advancement. This study provides a dedicated dataset and modeling framework for auditory speech decoding, offering foundations for neural language processing and assistive communication systems.

AAAI Conference 2025 Conference Paper

Mjölnir: Breaking the Shield of Perturbation-Protected Gradients via Adaptive Diffusion

  • Xuan Liu
  • Siqi Cai
  • Qihua Zhou
  • Song Guo
  • Ruibin Li
  • Kaiwei Lin

Perturbation-based mechanisms, such as differential privacy, mitigate gradient leakage attacks by introducing noise into the gradients, thereby preventing attackers from reconstructing clients' private data from the leaked gradients. However, can gradient perturbation protection mechanisms truly defend against all gradient leakage attacks? In this paper, we present the first attempt to break the shield of gradient perturbation protection in Federated Learning for the extraction of private information. We focus on common noise distributions, specifically Gaussian and Laplace, and apply our approach to DNN and CNN models. We introduce Mjölnir, a perturbation-resilient gradient leakage attack that is capable of removing perturbations from gradients without requiring additional access to the original model structure or external data. Specifically, we leverage the inherent diffusion properties of gradient perturbation protection to develop a novel diffusion-based gradient denoising model for Mjölnir. By constructing a surrogate client model that captures the structure of perturbed gradients, we obtain crucial gradient data for training the diffusion model. We further utilize the insight that monitoring disturbance levels during the reverse diffusion process can enhance gradient denoising capabilities, allowing Mjölnir to generate gradients that closely approximate the original, unperturbed versions through adaptive sampling steps. Extensive experiments demonstrate that Mjölnir effectively recovers the protected gradients and exposes the Federated Learning process to the threat of gradient leakage, achieving superior performance in gradient denoising and private data recovery.

NeurIPS Conference 2025 Conference Paper

S$^2$M-Former: Spiking Symmetric Mixing Branchformer for Brain Auditory Attention Detection

  • Jiaqi Wang
  • Zhengyu Ma
  • Xiongri Shen
  • Chenlin Zhou
  • Leilei Zhao
  • Han Zhang
  • Yi Zhong
  • Siqi Cai

Auditory attention detection (AAD) aims to decode listeners' focus in complex auditory environments from electroencephalography (EEG) recordings, which is crucial for developing neuro-steered hearing devices. Despite recent advancements, EEG-based AAD remains hindered by the absence of synergistic frameworks that can fully leverage complementary EEG features under energy-efficiency constraints. We propose ***S$^2$M-Former***, a novel ***s***piking ***s***ymmetric ***m***ixing framework to address this limitation through two key innovations: i) Presenting a spike-driven symmetric architecture composed of parallel spatial and frequency branches with mirrored modular design, leveraging biologically plausible token-channel mixers to enhance complementary learning across branches; ii) Introducing lightweight 1D token sequences to replace conventional 3D operations, reducing parameters by 14. 7$\times$. The brain-inspired spiking architecture further reduces power consumption, achieving a 5. 8$\times$ energy reduction compared to recent ANN methods, while also surpassing existing SNN baselines in terms of parameter efficiency and performance. Comprehensive experiments on three AAD benchmarks (KUL, DTU and AV-GC-AAD) across three settings (within-trial, cross-trial and cross-subject) demonstrate that S$^2$M-Former achieves comparable state-of-the-art (SOTA) decoding accuracy, making it a promising low-power, high-performance solution for AAD tasks. Code is available at https: //github. com/JackieWang9811/S2M-Former.

NeurIPS Conference 2024 Conference Paper

Language Without Borders: A Dataset and Benchmark for Code-Switching Lip Reading

  • Xueyi Zhang
  • Chengwei Zhang
  • Mingrui Lao
  • Peng Zhao
  • Jun Tang
  • Yanming Guo
  • Siqi Cai
  • Xianghu Yue

Lip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy environments. With the increasing interpersonal communications in social media owing to globalization, the existing monolingual datasets for lip reading may not be sufficient to meet the exponential proliferation of bilingual and even multilingual users. However, to our best knowledge, research on code-switching is only explored in speech recognition, while the attempts in lip reading are seriously neglected. To bridge this gap, we have collected a bilingual code-switching lip reading benchmark composed of Chinese and English, dubbed CSLR. As the pioneering work, we recruited 62 speakers with proficient foundations in bothspoken Chinese and English to express sentences containing both involved languages. Through rigorous criteria in data selection, CSLR benchmark has accumulated 85, 560 video samples with a resolution of 1080x1920, totaling over 71. 3 hours of high-quality code-switching lip movement data. To systematically evaluate the technical challenges in CSLR, we implement commonly-used lip reading backbones, as well as competitive solutions in code-switching speech for benchmark testing. Experiments show CSLR to be a challenging and under-explored lip reading task. We hope our proposed benchmark will extend the applicability of code-switching lip reading, and further contribute to the communities of cross-lingual communication and collaboration. Our dataset and benchmark are accessible at https: //github. com/cslr-lipreading/CSLR.

EAAI Journal 2023 Journal Article

Automatic contour correction of pectus excavatum using computer-aided diagnosis and convolutional neural network

  • Siqi Cai
  • Yizhi Liao
  • Lixuan Lai
  • Haiyu Zhou
  • Longhan Xie

Pectus excavatum (PE) is a common congenital sternal malformation disease that significantly impacts the physical and psychological well-being of affected individuals. Traditional methods for diagnosing and correcting PE rely heavily on physician expertise, leading to potential uncertainties and errors. While current research has primarily focused on automatic indices extraction and deformity evaluation, there is a lack of emphasis on generating corrective solutions for patients with PE. To address these limitations, we present a novel convolutional neural network (CNN)-based computer-aided diagnosis (CAD) approach for automatically generating recommended corrections for patients with PE. Specifically, our approach involves training a CNN model using sternum contours from normal individuals to predict corrected sternum contours for patients. Through block-wise fine-tuning using transfer learning, we optimize the regression performance for three PE indices. The complete contours of patients are then depicted based on the predicted indices using our CAD method. To validate our approach, we collected a dataset comprising 11, 755 chest CT images from 40 PE patients and 40 healthy individuals. The results suggest there is no significant difference between the predicted contours generated by our model and the actual postoperative contours by skilled surgeons, underscoring the promising efficacy of our model. In summary, our novel approach goes beyond the limitations of existing techniques and offers a significant advancement in the diagnosis and correction of PE.

AAAI Conference 2023 Short Paper

MGIA: Mutual Gradient Inversion Attack in Multi-Modal Federated Learning (Student Abstract)

  • Xuan Liu
  • Siqi Cai
  • Lin Li
  • Rui Zhang
  • Song Guo

Recent studies have demonstrated that local training data in Federated Learning can be recovered from gradients, which are called gradient inversion attacks. These attacks display powerful effects on either computer vision or natural language processing tasks. As it is known that there are certain correlations between multi-modality data, we argue that the threat of such attacks combined with Multi-modal Learning may cause more severe effects. Different modalities may communicate through gradients to provide richer information for the attackers, thus improving the strength and efficiency of the gradient inversion attacks. In this paper, we propose the Mutual Gradient Inversion Attack (MGIA), by utilizing the shared labels between image and text modalities combined with the idea of knowledge distillation. Our experimental results show that MGIA achieves the best quality of both modality data and label recoveries in comparison with other methods. In the meanwhile, MGIA verifies that multi-modality gradient inversion attacks are more likely to disclose private information than the existing single-modality attacks.

JBHI Journal 2020 Journal Article

Real-Time Detection of Compensatory Patterns in Patients With Stroke to Reduce Compensation During Robotic Rehabilitation Therapy

  • Siqi Cai
  • Guofeng Li
  • Enze Su
  • Xuyang Wei
  • Shuangyuan Huang
  • Ke Ma
  • Haiqing Zheng
  • Longhan Xie

Objectives: Compensations are commonly employed by patients with stroke during rehabilitation without therapist supervision, leading to suboptimal recovery outcomes. This study investigated the feasibility of the real-time monitoring of compensation in patients with stroke by using pressure distribution data and machine learning algorithms. Whether trunk compensation can be reduced by combining the online detection of compensation and haptic feedback of a rehabilitation robot was also investigated. Methods: Six patients with stroke did three forms of reaching movements while pressure distribution data were recorded as Dataset1. A support vector machine (SVM) classifier was trained with features extracted from Dataset1. Then, two other patients with stroke performed reaching tasks, and the SVM classifier trained by Dataset1 was employed to classify the compensatory patterns online. Based on the real-time monitoring of compensation, a rehabilitation robot provided an assistive force to patients with stroke to reduce compensations. Results: Good classification performance (F1 score > 0. 95) was obtained in both offline and online compensation analysis using the SVM classifier and pressure distribution data of patients with stroke. Based on the real-time detection of compensatory patterns, the angles of trunk rotation, trunk lean-forward and trunk-scapula elevation decreased by 46. 95%, 32. 35% and 23. 75%, respectively. Conclusion: High classification accuracies verified the feasibility of detecting compensation in patients with stroke based on pressure distribution data. Since the validity and reliability of the online detection of compensation has been verified, this classifier can be incorporated into a rehabilitation robot to reduce trunk compensations in patients with stroke.

v2026.09.13