arXiv · 2609.13857
ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation
Abstract
Multimodal emotion recognition in conversation (ERC) requires adapting to the instance-dependent reliability of different evidence sources. Lexical content may be decisive, vocal expression may provide complementary cues, or accurate recognition may require cross-modal interaction; fixed fusion does not explicitly account for this variation. We propose ReH-FUSE, a reliability-aware framework with dialogue-aware text, audio, and cross-modal experts. Its decision-level router first models the relative preference between text and audio and then balances the resulting unimodal mixture against the cross-modal expert. This factorization separates unimodal competition from cross-modal selection. Across three independent runs on IEMOCAP, ReH-FUSE achieves 74.34% weighted F1 and 73.11% macro F1; on MELD, it achieves 68.03% weighted F1. Controlled ablations show that learned routing outperforms uniform expert averaging and benefits from cross-modal interaction.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
Guan-Hua Wen, Hou-Chiang Tseng, Kuan-Yu Chen. 2026-09-12. ReH-FUSE: Reliability-Aware Hierarchical Fusion of Experts for Multimodal Emotion Recognition in Conversation. https://arxiv.org/abs/2609.13857
Cite the original work for its findings. Save a collection to share your selection of sources.