Abstract
Deep learning-based speech emotion recognition has been applied for social living assistance, health monitoring, authentication, and other human-to-machine interaction applications. Because of the ubiquitous nature of the applications, computationally efficient and robust speech emotion recognition models are required. The nature of the speech signal requires tracking of time steps, analyzing long-term dependencies and the contexts of the utterances as well as the spatial cues. Recurrent neural networks like long short-term memory and gated recurrent units coupled with attention mechanisms are often used to consider long-term dependencies and context in the speech signal. However, they do not take care of the spatial cues that may exist in the speech signal. Moreover, the operation of most of these systems is sequential which causes slow convergence, and sluggish training. Therefore, we propose a model that employs dilated convolutions layers in combination with hybrid attention mechanisms. The model uses multi-head attention to extract the global context in the feature representations which are fed into the bidirectional long short-term memory configured with self-attention to further handle the context and long-term dependencies. The model uses spectral and voice quality features extracted from the raw speech signals as input. The proposed model achieves comparable performance in terms of F1 score and accuracy. The proposed model's performance is also presented in terms of confusion matrices.
| Original language | English |
|---|---|
| Title of host publication | APCC 2022 - 27th Asia-Pacific Conference on Communications |
| Subtitle of host publication | Creating Innovative Communication Technologies for Post-Pandemic Era |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| Pages | 601-604 |
| Number of pages | 4 |
| ISBN (Electronic) | 9781665499279 |
| DOIs | |
| State | Published - 2022 |
| Event | 27th Asia-Pacific Conference on Communications, APCC 2022 - Jeju Island, Korea, Republic of Duration: 19 Oct 2022 → 21 Oct 2022 |
Publication series
| Name | APCC 2022 - 27th Asia-Pacific Conference on Communications: Creating Innovative Communication Technologies for Post-Pandemic Era |
|---|
Conference
| Conference | 27th Asia-Pacific Conference on Communications, APCC 2022 |
|---|---|
| Country/Territory | Korea, Republic of |
| City | Jeju Island |
| Period | 19/10/22 → 21/10/22 |
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 3 Good Health and Well-being
Keywords
- context-aware emotion recognition
- dilated convolution
- multi-head attention
Fingerprint
Dive into the research topics of 'Speech Emotion Recognition using Context-Aware Dilated Convolution Network'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver