2024
DOI: 10.1007/s10462-024-10825-z
|View full text |Cite
|
Sign up to set email alerts
|

A survey on knowledge-enhanced multimodal learning

Maria Lymperaiou,
Giorgos Stamou

Abstract: Multimodal learning has been a field of increasing interest, aiming to combine various modalities in a single joint representation. Especially in the area of visiolinguistic (VL) learning multiple models and techniques have been developed, targeting a variety of tasks that involve images and text. VL models have reached unprecedented performances by extending the idea of Transformers, so that both modalities can learn from each other. Massive pre-training procedures enable VL models to acquire a certain level … Show more

Help me understand this report
View preprint versions

Search citation statements

Order By: Relevance

Paper Sections

Select...

Citation Types

0
0
0

Publication Types

Select...

Relationship

0
0

Authors

Journals

citations
Cited by 0 publications
references
References 296 publications
0
0
0
Order By: Relevance

No citations

Set email alert for when this publication receives citations?