Arrow Research search

Author name cluster

Xinjian Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
1 author row

Possible papers

6

AAAI Conference 2026 Conference Paper

Spontaneous Yet Predictable: Shapelet-Driven, Channel-Aware Intention Decoding from Multi-Region ECoG

  • Keren Cao
  • Yuhang Tian
  • Kaizhong Zheng
  • Wei Xi
  • Xinjian Li
  • Liangjun Chen

Proactive intention decoding remains a critical yet underexplored challenge in brain–machine interfaces (BMIs), especially under naturalistic, self-initiated behavior. Existing systems rely on reactive decoding of motor cortex signals, resulting in substantial latency. To address this, we leverage the common marmoset’s spontaneous vocalizations and develop a high-resolution, dual-region ECoG recording paradigm targeting the prefrontal and auditory cortices and a neural decoding framework that integrates shapelet-based temporal encoding, position-aware attention, frequency-aware channel masking, contrastive clustering and a minimum error entropy-based robust loss. Our approach achieves 91.9% accuracy up to 200 ms before vocal onset—substantially outperforming 13 competitive baselines. Our model also uncovers a functional decoupling between auditory and prefrontal regions. Furthermore, joint modeling in time and frequency domains reveals novel preparatory neural signatures preceding volitional vocal output. Together, our findings bridge the gap between foundational neuroscience and applied BMI engineering, and establish a generalizable framework for intention decoding from ecologically valid, asynchronous behaviors.

YNIMG Journal 2025 Journal Article

Mesoscale functional connectivity of amygdala to the auditory and prefrontal cortex of macaque monkeys revealed by INS-fMRI

  • Qianbing Li
  • An Ping
  • Yuqi Feng
  • Bin Xu
  • Baorong Zhang
  • Anna Wang Roe
  • Lixia Gao
  • Xinjian Li

Mammals rely heavily on their auditory system to perceive environmental threats, socially communicate, and care for the young. As an extension of the multiple sensory system including the auditory system, the amygdala evaluates the emotional salience of acoustic stimuli, and mediates its impact on sensory, cognitive, and physiological aspects of emotional processing via the lateral amygdala (LA), basal amygdala (BA), and central amygdala (CeA) nuclei of the amygdala in acoustic domain. However, the functional connections of LA, BA, and CeA with the auditory cortex (AC) and the prefrontal cortex (PFC) remain unclear, particularly at the mesoscale level. Here we employed a novel method called INS-fMRI (Infrared Neural Stimulation combined with high-resolution functional magnetic resonance imaging) in Macaque monkeys, this method permits stimulation of multiple sites within single animals in vivo, so that the relative organization of auditory networks can be studied. We found that: (1) Focal INS stimulation of the amygdala elicited robust and reliable responses in both the AC and the PFC; (2) Amygdala stimulation mainly activated ipsilateral AC and PFC; (3) The stimulation of the amygdala mainly activated the secondary AC, and the dorsolateral PFC; (4) The connection between the amygdala and the cortex is mainly mediated by neurons in LA and BA connection area. Our study further revealed the functional connectivity among the amygdala subnucleus, the auditory cortex and the prefrontal cortex, and will shed light on the research for processing biologically meaningful complex sounds.

IJCAI Conference 2023 Conference Paper

Learning to Speak from Text: Zero-Shot Multilingual Text-to-Speech with Unsupervised Text Pretraining

  • Takaaki Saeki
  • Soumi Maiti
  • Xinjian Li
  • Shinji Watanabe
  • Shinnosuke Takamichi
  • Hiroshi Saruwatari

While neural text-to-speech (TTS) has achieved human-like natural synthetic speech, multilingual TTS systems are limited to resource-rich languages due to the need for paired text and studio-quality audio data. This paper proposes a method for zero-shot multilingual TTS using text-only data for the target language. The use of text-only data allows the development of TTS systems for low-resource languages for which only textual resources are available, making TTS accessible to thousands of languages. Inspired by the strong cross-lingual transferability of multilingual language models, our framework first performs masked language model pretraining with multilingual text-only data. Then we train this model with a paired data in a supervised manner, while freezing a language-aware embedding layer. This allows inference even for languages not included in the paired data but present in the text-only data. Evaluation results demonstrate highly intelligible zero-shot TTS with a character error rate of less than 12% for an unseen language.

JBHI Journal 2020 Journal Article

Machine Listening for Heart Status Monitoring: Introducing and Benchmarking HSS—The Heart Sounds Shenzhen Corpus

  • Fengquan Dong
  • Kun Qian
  • Zhao Ren
  • Alice Baird
  • Xinjian Li
  • Zhenyu Dai
  • Bo Dong
  • Florian Metze

Auscultation of the heart is a widely studied technique, which requires precise hearing from practitioners as a means of distinguishing subtle differences in heart-beat rhythm. This technique is popular due to its non-invasive nature, and can be an early diagnosis aid for a range of cardiac conditions. Machine listening approaches can support this process, monitoring continuously and allowing for a representation of both mild and chronic heart conditions. Despite this potential, relevant databases and benchmark studies are scarce. In this paper, we introduce our publicly accessible database, the Heart Sounds Shenzhen Corpus (HSS), which was first released during the recent INTERSPEECH 2018 ComParE Heart Sound sub-challenge. Additionally, we provide a survey of machine learning work in the area of heart sound recognition, as well as a benchmark for HSS utilising standard acoustic features and machine learning models. At best our support vector machine with Log Mel features achieves 49. 7% unweighted average recall on a three category task (normal, mild, moderate/severe).

AAAI Conference 2020 Conference Paper

Towards Zero-Shot Learning for Automatic Phonemic Transcription

  • Xinjian Li
  • Siddharth Dalmia
  • David Mortensen
  • Juncheng Li
  • Alan Black
  • Florian Metze

Automatic phonemic transcription tools are useful for lowresource language documentation. However, due to the lack of training sets, only a tiny fraction of languages have phonemic transcription tools. Fortunately, multilingual acoustic modeling provides a solution given limited audio training data. A more challenging problem is to build phonemic transcribers for languages with zero training data. The difficulty of this task is that phoneme inventories often differ between the training languages and the target language, making it infeasible to recognize unseen phonemes. In this work, we address this problem by adopting the idea of zero-shot learning. Our model is able to recognize unseen phonemes in the target language without any training data. In our model, we decompose phonemes into corresponding articulatory attributes such as vowel and consonant. Instead of predicting phonemes directly, we first predict distributions over articulatory attributes, and then compute phoneme distributions with a customized acoustic model. We evaluate our model by training it using 13 languages and testing it using 7 unseen languages. We find that it achieves 7. 7% better phoneme error rate on average over a standard multilingual model.

NeurIPS Conference 2019 Conference Paper

Adversarial Music: Real world Audio Adversary against Wake-word Detection System

  • Juncheng Li
  • Shuhui Qu
  • Xinjian Li
  • Joseph Szurley
  • J. Zico Kolter
  • Florian Metze

Voice Assistants (VAs) such as Amazon Alexa or Google Assistant rely on wake-word detection to respond to people's commands, which could potentially be vulnerable to audio adversarial examples. In this work, we target our attack on the wake-word detection system. Our goal is to jam the model with some inconspicuous background music to deactivate the VAs while our audio adversary is present. We implemented an emulated wake-word detection system of Amazon Alexa based on recent publications. We validated our models against the real Alexa in terms of wake-word detection accuracy. Then we computed our audio adversaries with consideration of expectation over transform and we implemented our audio adversary with a differentiable synthesizer. Next we verified our audio adversaries digitally on hundreds of samples of utterances collected from the real world. Our experiments show that we can effectively reduce the recognition F1 score of our emulated model from 93. 4% to 11. 0%. Finally, we tested our audio adversary over the air, and verified it works effectively against Alexa, reducing its F1 score from 92. 5% to 11. 0%. To the best of our knowledge, this is the first real-world adversarial attack against a commercial grade VA wake-word detection system. Our demo video is included in the supplementary material.

v2026.09.13