HemoGAT: Heterogeneous multimodal speech emotion recognition with cross-modal transformer and graph attention network

dc.contributor.authorNguyen, Nhut Min
dc.contributor.authorNguyen, Thanh Trung
dc.contributor.authorNguyen, Tien-Dat
dc.contributor.authorDang, Duc Ngoc Minh
dc.date.accessioned2026-07-17T07:10:11Z
dc.date.available2026-07-17T07:10:11Z
dc.date.issued2026
dc.description.abstractMultimodal speech emotion recognition (SER) is a promising field, yet effectively fusing diverse information streams remains challenging. Addressing this requires architectures capable of modeling structural relationships across modalities with fine-grained, feature- level interactions. This paper proposes HemoGAT, a novel heterogeneous multimodal SER architecture that integrates a dual-stream architecture with two core mod- ules: a heterogeneous multimodal graph attention net- work (HM-GAT) and a cross-modal transformer (CMT) to address this. The HM-GAT module captures complex structural and contextual dependencies using a hetero- geneous graph constructed from deep embeddings. The CMT module enables precise cross-modal feature fusion through bidirectional cross-attention. This design effec- tively captures both high-level relationships and immedi- ate cross-modal influences. HemoGAT achieves state-of- the-art (SOTA) performance on the IEMOCAP dataset and highly competitive results on the MELD dataset, demonstrating its superiority over existing methods. Extensive ablation studies were conducted to evaluate HemoGAT. We assessed the impact of the Top-K algo- rithm for heterogeneous graph construction and com- pared unimodal and multimodal fusion strategies. We also examined the contributions of the HM-GAT and CMT modules, analyzed the role of the graph attention network (GAT) in graph learning, and evaluated the effect of GAT layer depth on performance
dc.identifier.citationAdvances in electrical and electronic engineering. 2026, vol. 24, no. 2, pp.144 – 159 : ill.
dc.identifier.doi10.15598/aeee.v24i2.250415
dc.identifier.issn1336-1376
dc.identifier.issn1804-3119
dc.identifier.urihttp://hdl.handle.net/10084/158807
dc.language.isoen
dc.publisherVysoká škola báňská - Technická univerzita Ostrava
dc.relation.ispartofseriesAdvances in electrical and electronic engineering
dc.relation.urihttps://doi.org/10.15598/aeee.v24i2.250415
dc.rights© Vysoká škola báňská - Technická univerzita Ostrava
dc.rightsAttribution-NoDerivatives 4.0 Internationalen
dc.rights.accessopenAccess
dc.rights.urihttp://creativecommons.org/licenses/by-nd/4.0/
dc.subjectheterogeneous graph construction
dc.subjectgraph attention network
dc.subjectcross-modal transformer
dc.subjectfeature fusion
dc.subjectmultimodal speech emotion recognition
dc.titleHemoGAT: Heterogeneous multimodal speech emotion recognition with cross-modal transformer and graph attention network
dc.typearticle
dc.type.statusPeer-reviewed
dc.type.versionpublishedVersion
local.files.count1
local.files.size8359879
local.has.filesyes

Files

Original bundle

Now showing 1 - 1 out of 1 results
Loading...
Thumbnail Image
Name:
Nguzen_aj..pdf
Size:
7.97 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 out of 1 results
Loading...
Thumbnail Image
Name:
license.txt
Size:
718 B
Format:
Item-specific license agreed upon to submission
Description: