Multimodal Representations Learning Based on Mutual Information Maximization and Minimization and Identity Embedding for Multimodal Sentiment Analysis

Zheng, Jiawei; Zhang, Sen; Wang, Xiaoping; Zeng, Zhigang

doi:10.48550/arxiv.2201.03969

Cited by 1 publication

(1 citation statement)

References 26 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…MMIM also applies MI maximization between the fusion result and input modalities again, ensuring that the fusion output sufficiently captures modality-invariant clues among the modalities. Like Han et al [96], Zheng et al [97] also introduced a multimodal representation model, MMMIE, grounded in the principles of mutual information maximization, minimization, and identity embedding. This model aims to maximize the mutual information between modalities while minimizing the mutual information between input data and its features, extracting modality-invariant and task-related information.…”

Section: Simple Concatenation Fusionmentioning

confidence: 99%

A Survey of Deep Learning-Based Multimodal Emotion Recognition: Speech, Text, and Face

Lian,

Lu,

et al. 2023

Entropy

View full text Add to dashboard Cite

Multimodal emotion recognition (MER) refers to the identification and understanding of human emotional states by combining different signals, including—but not limited to—text, speech, and face cues. MER plays a crucial role in the human–computer interaction (HCI) domain. With the recent progression of deep learning technologies and the increasing availability of multimodal datasets, the MER domain has witnessed considerable development, resulting in numerous significant research breakthroughs. However, a conspicuous absence of thorough and focused reviews on these deep learning-based MER achievements is observed. This survey aims to bridge this gap by providing a comprehensive overview of the recent advancements in MER based on deep learning. For an orderly exposition, this paper first outlines a meticulous analysis of the current multimodal datasets, emphasizing their advantages and constraints. Subsequently, we thoroughly scrutinize diverse methods for multimodal emotional feature extraction, highlighting the merits and demerits of each method. Moreover, we perform an exhaustive analysis of various MER algorithms, with particular focus on the model-agnostic fusion methods (including early fusion, late fusion, and hybrid fusion) and fusion based on intermediate layers of deep models (encompassing simple concatenation fusion, utterance-level interaction fusion, and fine-grained interaction fusion). We assess the strengths and weaknesses of these fusion strategies, providing guidance to researchers to help them select the most suitable techniques for their studies. In summary, this survey aims to provide a thorough and insightful review of the field of deep learning-based MER. It is intended as a valuable guide to aid researchers in furthering the evolution of this dynamic and impactful field.

show abstract

Section: Simple Concatenation Fusionmentioning

confidence: 99%