HemoGAT: Heterogeneous multimodal speech emotion recognition with cross-modal transformer and graph attention network

Loading...
Thumbnail Image

Downloads

Date issued

Journal Title

Journal ISSN

Volume Title

Publisher

Vysoká škola báňská - Technická univerzita Ostrava

Location

Signature

Abstract

Multimodal speech emotion recognition (SER) is a promising field, yet effectively fusing diverse information streams remains challenging. Addressing this requires architectures capable of modeling structural relationships across modalities with fine-grained, feature- level interactions. This paper proposes HemoGAT, a novel heterogeneous multimodal SER architecture that integrates a dual-stream architecture with two core mod- ules: a heterogeneous multimodal graph attention net- work (HM-GAT) and a cross-modal transformer (CMT) to address this. The HM-GAT module captures complex structural and contextual dependencies using a hetero- geneous graph constructed from deep embeddings. The CMT module enables precise cross-modal feature fusion through bidirectional cross-attention. This design effec- tively captures both high-level relationships and immedi- ate cross-modal influences. HemoGAT achieves state-of- the-art (SOTA) performance on the IEMOCAP dataset and highly competitive results on the MELD dataset, demonstrating its superiority over existing methods. Extensive ablation studies were conducted to evaluate HemoGAT. We assessed the impact of the Top-K algo- rithm for heterogeneous graph construction and com- pared unimodal and multimodal fusion strategies. We also examined the contributions of the HM-GAT and CMT modules, analyzed the role of the graph attention network (GAT) in graph learning, and evaluated the effect of GAT layer depth on performance

Description

Delayed publication

Available after

Subject(s)

heterogeneous graph construction, graph attention network, cross-modal transformer, feature fusion, multimodal speech emotion recognition

Citation

Advances in electrical and electronic engineering. 2026, vol. 24, no. 2, pp.144 – 159 : ill.