Arrow Research search

Author name cluster

Jia Lu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

4 papers
1 author row

Possible papers

4

JBHI Journal 2026 Journal Article

Decoding Covert Speech from EEG by Functional Areas Spatio-Temporal Transformer

  • Muyun Jiang
  • Wei Zhang
  • Yi Ding
  • Kok Ann Colin Teo
  • LaiGuan Fong
  • Shuailei Zhang
  • Zhiwei Guo
  • Chenyu Liu

Covert speech involves imagining speaking without audible sound or any movements. Decoding covert speech from electroencephalogram (EEG) is challenging due to a limited understanding of neural pronunciation mapping and the low signal-to-noise ratio of the signal. In this study, we developed a large-scale multi-utterance speech EEG dataset from 57 right-handed native English-speaking subjects, each performing covert and overt speech tasks by repeating the same word in five utterances within a ten-second duration. Given the spatio-temporal nature of the neural activation process during speech pronunciation, we developed a Functional Areas Spatio-temporal Transformer (FAST), an effective framework for converting EEG signals into tokens and utilizing transformer architecture for sequence encoding. Our results reveal distinct and interpretable speech neural features by the visualization of FAST-generated activation maps across frontal and temporal brain regions with each word being covertly spoken, providing new insights into the discriminative features of the neural representation of covert speech. This is the first report of such a study, which provides interpretable evidence for speech decoding from EEG. The code for this work has been made public at https://github.com/Jiang-Muyun/FAST

AAAI Conference 2026 Conference Paper

WorldGrow: Generating Infinite 3D World

  • Sikuang Li
  • Chen Yang
  • Jiemin Fang
  • Taoran Yi
  • Jia Lu
  • Jiazhong Cen
  • Lingxi Xie
  • Wei Shen

We tackle the challenge of generating the infinitely extendable 3D world -- large, continuous environments with coherent geometry and realistic appearance. Existing methods face key challenges: 2D-lifting approaches suffer from geometric and appearance inconsistencies across views, 3D implicit representations are hard to scale up, and current 3D foundation models are mostly object-centric, limiting their applicability to scene-level generation. Our key insight is leveraging strong generation priors from pre-trained 3D models for structured scene block generation. To this end, we propose WorldGrow, a hierarchical framework for unbounded 3D scene synthesis. Our method features three core components: (1) a data curation pipeline that extracts high-quality scene blocks for training, making the 3D structured latent representations suitable for scene generation; (2) a 3D block inpainting mechanism that enables context-aware scene extension; and (3) a coarse-to-fine generation strategy that ensures both global layout plausibility and local geometric/textural fidelity. Evaluated on the large-scale 3D-FRONT dataset, WorldGrow achieves SOTA performance in geometry reconstruction, while uniquely supporting infinite scene generation with photorealistic and structurally consistent outputs. These results highlight its capability for constructing large-scale virtual environments and potential for building future world models.

EAAI Journal 2024 Journal Article

Embedding enhancement with foreground feature alignment and primitive knowledge for few-shot learning

  • Xiaoqi Zheng
  • Jia Lu

Few-Shot Learning (FSL) targets a model to quickly discriminate new categories with limited samples. While most methods struggle to effectively utilize the knowledge learned in the base classes, resulting in models that fail to alleviate the domain gap between training and evaluation. Conversely, humans often enable rapid discrimination by two ways: reinforcing impressions through comparing unknown categories with related objects in memory; Searching for the maximum shared information within the intra-class. Motivated by this, in this paper, we propose a novel Transformer-based Embedding Enhancement Network (TEEN) that adaptively leverages the knowledge learned from the base classes, which we refer to as ‘primitive knowledge’, to complete and distinguish the embeddings for novel classes. Additionally, to reduce interference from background features, we introduce the Transformer-based Foreground Feature Alignment (TFFA) to enhance the representation of image foreground. Ultimately, TEEN enhances the embeddings through foreground alignment and primitive knowledge completion, achieving improved performance in few-shot classification. Extensive experiments on inductive few-shot tasks demonstrate the effectiveness of our approach, achieving state-of-the-art results in cross-domain few-shot tasks.

YNIMG Journal 2024 Journal Article

Revealing the spatiotemporal brain dynamics of covert speech compared with overt speech: A simultaneous EEG-fMRI study

  • Wei Zhang
  • Muyun Jiang
  • Kok Ann Colin Teo
  • Raghavan Bhuvanakantham
  • LaiGuan Fong
  • Wei Khang Jeremy Sim
  • Zhiwei Guo
  • Chuan Huat Vince Foo

Covert speech (CS) refers to speaking internally to oneself without producing any sound or movement. CS is involved in multiple cognitive functions and disorders. Reconstructing CS content by brain-computer interface (BCI) is also an emerging technique. However, it is still controversial whether CS is a truncated neural process of overt speech (OS) or involves independent patterns. Here, we performed a word-speaking experiment with simultaneous EEG-fMRI. It involved 32 participants, who generated words both overtly and covertly. By integrating spatial constraints from fMRI into EEG source localization, we precisely estimated the spatiotemporal dynamics of neural activity. During CS, EEG source activity was localized in three regions: the left precentral gyrus, the left supplementary motor area, and the left putamen. Although OS involved more brain regions with stronger activations, CS was characterized by an earlier event-locked activation in the left putamen (peak at 262 ms versus 1170 ms). The left putamen was also identified as the only hub node within the functional connectivity (FC) networks of both OS and CS, while showing weaker FC strength towards speech-related regions in the dominant hemisphere during CS. Path analysis revealed significant multivariate associations, indicating an indirect association between the earlier activation in the left putamen and CS, which was mediated by reduced FC towards speech-related regions. These findings revealed the specific spatiotemporal dynamics of CS, offering insights into CS mechanisms that are potentially relevant for future treatment of self-regulation deficits, speech disorders, and development of BCI speech applications.

v2026.09.13