Parallel Network Speech Emotion Recognition Based on Hybrid Attention Mechanism.

Saved in:
Bibliographic Details
Title: Parallel Network Speech Emotion Recognition Based on Hybrid Attention Mechanism.
Authors: Hu, Zhangfang1 huzf@cqupt.edu.cn, Wang, Yulong2 2423245845@qq.com, Tang, Yicheng2 1035937714@qq.com
Source: IAENG International Journal of Computer Science. Jul2026, Vol. 53 Issue 7, p2740-2749. 10p.
Subjects: Emotion recognition, Parallel processing, Long short-term memory
Abstract: In speech emotion recognition tasks, relying on a single feature often limits the representational capacity for capturing emotional information, and using a single network model may lead to insufficient deep feature extraction, both of which can result in low classification accuracy. To address these issues, this paper proposes a parallel network architecture based on a hybrid attention mechanism, utilizing a combination of multiple features as network inputs to enhance emotion classification performance. The model first maps the 81-dimensional fused features into a 128-dimensional embedding space through an embedding layer, and then feeds them into three parallel subnetworks. Each subnetwork consists of an MSDC module, a bidirectional LSTM module, and a hybrid attention mechanism. Experimental results on the RAVDESS dataset show that the proposed method achieves an accuracy of 96.62% and a precision of 96.55% in an 8-class emotion classification task, outperforming other models and demonstrating the effectiveness and superiority of the proposed approach. [ABSTRACT FROM AUTHOR]
Copyright of IAENG International Journal of Computer Science is the property of International Association of Engineers (IAENG) and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:In speech emotion recognition tasks, relying on a single feature often limits the representational capacity for capturing emotional information, and using a single network model may lead to insufficient deep feature extraction, both of which can result in low classification accuracy. To address these issues, this paper proposes a parallel network architecture based on a hybrid attention mechanism, utilizing a combination of multiple features as network inputs to enhance emotion classification performance. The model first maps the 81-dimensional fused features into a 128-dimensional embedding space through an embedding layer, and then feeds them into three parallel subnetworks. Each subnetwork consists of an MSDC module, a bidirectional LSTM module, and a hybrid attention mechanism. Experimental results on the RAVDESS dataset show that the proposed method achieves an accuracy of 96.62% and a precision of 96.55% in an 8-class emotion classification task, outperforming other models and demonstrating the effectiveness and superiority of the proposed approach. [ABSTRACT FROM AUTHOR]
ISSN:1819656X