Arrow Research search

Author name cluster

Linlin Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

7 papers
1 author row

Possible papers

7

JBHI Journal 2026 Journal Article

EEG-Based Emotion Recognition Using Spatial-Temporal Graph-Aware Network with Channel Selection

  • Linlin Li
  • Wanzhong Chen

Electroencephalogram (EEG)-based emotion recognition holds great potential in intelligent human computer interaction and brain-computer interface systems, as the brain generates distinct electrical activity patterns under different emotional states. However, EEG information often contains data from numerous channels, leading to high computational cost and potential redundancy. Existing channel selection methods often rely on uniform rules, lacking frequency-specific adaptability and inter-channel modeling, which can cause information loss and reduced performance during dimensionality reduction. To address this issue, we propose a novel framework that combines discriminative channel selection with hierarchical spatial-temporal modeling to enhance both per formance and efficiency. In preprocessing, wavelet coherence and mutual information are used to adaptively select informative channels across multiple frequency bands. The selected signals are then processed by a Spatial Temporal Graph-aware Network (STG-Net), which models spatial relationships between channels through graph convolution, extracting spatial features from each time frame. Coupled with a temporal modeling module, the network further captures the evolving temporal patterns of emotional states across consecutive frames. Finally, frequency spatial-temporal features are fused for emotion classification. Compared to the state-of-the-art methods, our approach achieves superior performance in both recognition accuracy and model efficiency.

AAAI Conference 2023 Conference Paper

Controllable Image Captioning via Prompting

  • Ning Wang
  • Jiahao Xie
  • Jihao Wu
  • Mingbo Jia
  • Linlin Li

Despite the remarkable progress of image captioning, existing captioners typically lack the controllable capability to generate desired image captions, e.g., describing the image in a rough or detailed manner, in a factual or emotional view, etc. In this paper, we show that a unified model is qualified to perform well in diverse domains and freely switch among multiple styles. Such a controllable capability is achieved by embedding the prompt learning into the image captioning framework. To be specific, we design a set of prompts to fine-tune the pre-trained image captioner. These prompts allow the model to absorb stylized data from different domains for joint training, without performance degradation in each domain. Furthermore, we optimize the prompts with learnable vectors in the continuous word embedding space, avoiding the heuristic prompt engineering and meanwhile exhibiting superior performance. In the inference stage, our model is able to generate desired stylized captions by choosing the corresponding prompts. Extensive experiments verify the controllable capability of the proposed method. Notably, we achieve outstanding performance on two diverse image captioning benchmarks including COCO Karpathy split and TextCaps using a unified model.

AAAI Conference 2023 Conference Paper

Efficient Image Captioning for Edge Devices

  • Ning Wang
  • Jiangrong Xie
  • Hang Luo
  • Qinglin Cheng
  • Jihao Wu
  • Mingbo Jia
  • Linlin Li

Recent years have witnessed the rapid progress of image captioning. However, the demands for large memory storage and heavy computational burden prevent these captioning models from being deployed on mobile devices. The main obstacles lie in the heavyweight visual feature extractors (i.e., object detectors) and complicated cross-modal fusion networks. To this end, we propose LightCap, a lightweight image captioner for resource-limited devices. The core design is built on the recent CLIP model for efficient image captioning. To be specific, on the one hand, we leverage the CLIP model to extract the compact grid features without relying on the time-consuming object detectors. On the other hand, we transfer the image-text retrieval design of CLIP to image captioning scenarios by devising a novel visual concept extractor and a cross-modal modulator. We further optimize the cross-modal fusion model and parallel prediction heads via sequential and ensemble distillations. With the carefully designed architecture, our model merely contains 40M parameters, saving the model size by more than 75% and the FLOPs by more than 98% in comparison with the current state-of-the-art methods. In spite of the low capacity, our model still exhibits state-of-the-art performance on prevalent datasets, e.g., 136.6 CIDEr on COCO Karpathy test split. Testing on the smartphone with only a single CPU, the proposed LightCap exhibits a fast inference speed of 188ms per image, which is ready for practical applications.

JBHI Journal 2022 Journal Article

DCPR-GAN: Dental Crown Prosthesis Restoration Using Two-Stage Generative Adversarial Networks

  • Sukun Tian
  • Miaohui Wang
  • Ning Dai
  • Haifeng Ma
  • Linlin Li
  • Luca Fiorenza
  • Yuchun Sun
  • Yangmin Li

Restoring the correct masticatory function of broken teeth is the basis of dental crown prosthesis rehabilitation. However, it is a challenging task primarily due to the complex and personalized morphology of the occlusal surface. In this article, we address this problem by designing a new two-stage generative adversarial network (GAN) to reconstruct a dental crown surface in the data-driven perspective. Specifically, in the first stage, a conditional GAN (CGAN) is designed to learn the inherent relationship between the defective tooth and the target crown, which can solve the problem of the occlusal relationship restoration. In the second stage, an improved CGAN is further devised by considering an occlusal groove parsing network (GroNet) and an occlusal fingerprint constraint to enforce the generator to enrich the functional characteristics of the occlusal surface. Experimental results demonstrate that the proposed framework significantly outperforms the state-of-the-art deep learning methods in functional occlusal surface reconstruction using a real-world patient database. Moreover, the standard deviation (SD) and root mean square (RMS) between the generated occlusal surface and the target crown calculated by our method are both less than 0. 161 mm. Importantly, the designed dental crown have enough anatomical morphology and higher clinical applicability.

AAAI Conference 2022 Conference Paper

Text Is No More Enough! A Benchmark for Profile-Based Spoken Language Understanding

  • Xiao Xu
  • Libo Qin
  • Kaiji Chen
  • Guoxing Wu
  • Linlin Li
  • Wanxiang Che

Current researches on spoken language understanding (SLU) heavily are limited to a simple setting: the plain text-based SLU that takes the user utterance as input and generates its corresponding semantic frames (e. g. , intent and slots). Unfortunately, such a simple setting may fail to work in complex real-world scenarios when an utterance is semantically ambiguous, which cannot be achieved by the text-based SLU models. In this paper, we first introduce a new and important task, Profile-based Spoken Language Understanding (PROSLU), which requires the model that not only relies on the plain text but also the supporting profile information to predict the correct intents and slots. To this end, we further introduce a large-scale human-annotated Chinese dataset with over 5K utterances and their corresponding supporting profile information (Knowledge Graph (KG), User Profile (UP), Context Awareness (CA)). In addition, we evaluate several state-of-the-art baseline models and explore a multi-level knowledge adapter to effectively incorporate profile information. Experimental results reveal that all existing text-based SLU models fail to work when the utterances are semantically ambiguous and our proposed framework can effectively fuse the supporting information for sentence-level intent detection and token-level slot filling. Finally, we summarize key challenges and provide new points for future directions, which hopes to facilitate the research.

AAAI Conference 2019 Short Paper

Location-Based End-to-End Speech Recognition with Multiple Language Models

  • Zhijie Lin
  • Kaiyang Lin
  • Shiling Chen
  • Linlin Li
  • Zhou Zhao

End-to-End deep learning approaches for Automatic Speech Recognition (ASR) has been a new trend. In those approaches, starting active in many areas, language model can be considered as an important and effective method for semantic error correction. Many existing systems use one language model. In this paper, however, multiple language models (LMs) are applied into decoding. One LM is used for selecting appropriate answers and others, considering both context and grammar, for further decision. Experiment on a general location-based dataset show the effectiveness of our method.

AAAI Conference 2019 Conference Paper

Unsupervised Learning Helps Supervised Neural Word Segmentation

  • Xiaobin Wang
  • Deng Cai
  • Linlin Li
  • Guangwei Xu
  • Hai Zhao
  • Luo Si

By exploiting unlabeled data for further performance improvement for Chinese word segmentation, this work makes the first attempt at exploring adding unsupervised segmentation information into neural supervised segmenter. We survey various effective strategies, including extending the character embedding, augmenting the word score and applying multi-task learning, for leveraging unsupervised information derived from abundant unlabeled data. Experiments on standard data sets show that the explored strategies indeed improve the recall rate of out-of-vocabulary words and thus boost the segmentation accuracy. Moreover, the model enhanced by the proposed methods outperforms state-of-theart models in closed test and shows promising improvement trend when adopting three different strategies with the help of a large unlabeled data set. Our thorough empirical study eventually verifies the proposed approach outperforms the widelyused pre-training approach in terms of effectively making use of freely abundant unlabeled data.

v2026.09.13