Arrow Research search

Author name cluster

Runnan Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

2 papers
1 author row

Possible papers

2

IJCAI Conference 2019 Conference Paper

Towards Discriminative Representation Learning for Speech Emotion Recognition

  • Runnan Li
  • Zhiyong Wu
  • Jia Jia
  • Yaohua Bu
  • Sheng Zhao
  • Helen Meng

In intelligent speech interaction, automatic speech emotion recognition (SER) plays an important role in understanding user intention. While sentimental speech has different speaker characteristics but similar acoustic attributes, one vital challenge in SER is how to learn robust and discriminative representations for emotion inferring. In this paper, inspired by human emotion perception, we propose a novel representation learning component (RLC) for SER system, which is constructed with Multi-head Self-attention and Global Context-aware Attention Long Short-Term Memory Recurrent Neutral Network (GCA-LSTM). With the ability of Multi-head Self-attention mechanism in modeling the element-wise correlative dependencies, RLC can exploit the common patterns of sentimental speech features to enhance emotion-salient information importing in representation learning. By employing GCA-LSTM, RLC can selectively focus on emotion-salient factors with the consideration of entire utterance context, and gradually produce discriminative representation for emotion inferring. Experiments on public emotional benchmark database IEMOCAP and a tremendous realistic interaction database demonstrate the outperformance of the proposed SER framework, with 6. 6% to 26. 7% relative improvement on unweighted accuracy compared to state-of-the-art techniques.

AAAI Conference 2017 Conference Paper

Multi-Task Deep Learning for User Intention Understanding in Speech Interaction Systems

  • Yishuang Ning
  • Jia Jia
  • Zhiyong Wu
  • Runnan Li
  • Yongsheng An
  • Yanfeng Wang
  • Helen Meng

Speech interaction systems have been gaining popularity in recent years. The main purpose of these systems is to generate more satisfactory responses according to users’ speech utterances, in which the most critical problem is to analyze user intention. Researches show that user intention conveyed through speech is not only expressed by content, but also closely related with users’ speaking manners (e. g. with or without acoustic emphasis). How to incorporate these heterogeneous attributes to infer user intention remains an open problem. In this paper, we define Intention Prominence (IP) as the semantic combination of focus by text and emphasis by speech, and propose a multi-task deep learning framework to predict IP. Specifically, we first use long short-term memory (LSTM) which is capable of modeling long short-term contextual dependencies to detect focus and emphasis, and incorporate the tasks for focus and emphasis detection with multi-task learning (MTL) to reinforce the performance of each other. We then employ Bayesian network (BN) to incorporate multimodal features (focus, emphasis, and location reflecting users’ dialect conventions) to predict IP based on feature correlations. Experiments on a data set of 135, 566 utterances collected from real-world Sogou Voice Assistant illustrate that our method can outperform the comparison methods over 6. 9-24. 5% in terms of F1-measure. Moreover, a real practice in the Sogou Voice Assistant indicates that our method can improve the performance on user intention understanding by 7%.

v2026.09.13