arXiv · 2609.11164
Multimodal Temporal Modeling for Continuous Group Emotion Recognition in Multi-party Dialogues
Abstract
To realize natural behavior in dialogue agents in multi-party dialogue scenarios, it is important to understand group emotion such as valence and arousal as a whole. Most prior work addressed this task at the utterance level or using a coarse-grained time window, which is not sufficient to capture emotional dynamics. In this study, we formulate continuous recognition of the Group Emotion at a one-second resolution. Moreover, we also introduce the Mixed state, which captures the emotional divergence among participants in the group. We constructed a dataset with frame-level soft labels based on the TEIDAN corpus and propose a multimodal temporal framework that integrates audio and video information using a sliding-window context. Experimental results demonstrate that the temporal Transformer outperforms simple baselines and shows stronger temporal agreement with the ground-truth labels than the LLM-based model. The effect of context length is limited, whereas audio-visual input outperforms either unimodal input on the continuous-label metrics. Additionally, our analysis shows larger Group Emotion recognition errors in intervals with high Mixed values, exposing emotional divergence as a key challenge for group emotion recognition.
Explore related subjects
Keep this discovery
Soma Iwata, Koji Inoue, Muyun Wu, Taiga Mori, Divesh Lala, Tatsuya Kawahara. 2026-09-10. Multimodal Temporal Modeling for Continuous Group Emotion Recognition in Multi-party Dialogues. https://doi.org/10.1145/3776591.3837045
Cite the original work for its findings. Save a collection to share your selection of sources.