VAE-Integrated Multiscale Generative Diffusion Modeling for Incomplete Multimodal Emotion Recognition

Wang, Zi Xuan, Yao, Zhu Yi, Shen, Jia Yue, Wang, Pan, Chen, Xue Jiao, Zhang, Xu ORCID: https://orcid.org/0000-0001-6557-6607, Gu, Huang Liang and Zhou, Xiao Kang (2026) VAE-Integrated Multiscale Generative Diffusion Modeling for Incomplete Multimodal Emotion Recognition. IEEE Transactions on Computational Social Systems, 13 (3). pp. 4096-4110. ISSN 2329-924X

[thumbnail of TCSS_manu]
Preview
PDF (TCSS_manu) - Accepted Version
Available under License Creative Commons Attribution.

Download (2MB) | Preview

Abstract

Multimodal emotion recognition (MER) has been widely adopted in affective computing and human–computer interaction. However, real-world multimodal streams frequently suffer from missing or incomplete modalities due to sensor failures, privacy constraints, and heterogeneous acquisition costs, which often causes severe performance degradation for models trained under complete inputs. To address this issue, we propose hierarchical VAE–diffusion for emotion reconstruction (HVDER), a latent-space generative framework for incomplete MER. HVDER first employs modality-specific variational autoencoders (VAEs) to project language, visual, and audio features into a unified low-dimensional latent space with KL-regularized structure. Then, a two-stage coarse-to-fine conditional diffusion module completes missing modality latents by recovering global emotion semantics followed by refining local discriminative details. Finally, the completed and observed modality representations are fused for emotion prediction, optimized end-to-end with a multiobjective loss integrating reconstruction, diffusion matching, cross-modal alignment, and classification supervision. Extensive experiments on CMU-MOSEI and CMU-MOSI validate the effectiveness and robustness of HVDER. Under fixed modality-missing settings, HVDER achieves average ACC2/F1 of 76.8/75.4 on CMU-MOSEI and 73.4/72.9 on CMU-MOSI, outperforming representative baselines across all modality combinations. Under random missing with a high missing rate of MR = 0.7, HVDER maintains ACC2/F1 of 75.4/72.5 on CMU-MOSEI and 68.0/67.2 on CMU-MOSI, indicating a stronger performance lower bound in severely incomplete scenarios. Ablation studies further confirm that latent-space modeling and the coarse-to-fine diffusion design jointly contribute to the main performance gains.

Item Type: Article
Uncontrolled Keywords: emotion ai,fusion reconstruction,multimodal emotion recognition (mer),modelling and simulation,social sciences (miscellaneous),human-computer interaction ,/dk/atira/pure/subjectarea/asjc/2600/2611
Faculty \ School: Faculty of Science > School of Computing Sciences
Related URLs:
Depositing User: LivePure Connector
Date Deposited: 17 Sep 2026 09:42
Last Modified: 17 Sep 2026 14:54
URI: https://ueaeprints.uea.ac.uk/id/eprint/104574
DOI: 10.1109/TCSS.2026.3681252

Downloads

Downloads per month over past year

Actions (login required)

View Item View Item