2022
DOI: 10.48550/arxiv.2210.02762
|View full text |Cite
Preprint
|
Sign up to set email alerts
|

Vision Transformer Based Model for Describing a Set of Images as a Story

Zainy M. Malakan,
Ghulam Mubashar Hassan,
Ajmal Mian

Abstract: Visual Story-Telling is the process of forming a multi sentence story from a set of images. Appropriately including visual variation and contextual information captured inside the input images is one of the most challenging aspects of visual storytelling. Consequently, stories developed from a set of images often lack cohesiveness, relevance, and semantic relationship. In this paper, we propose a novel Vision Transformer Based Model for describing a set of images as a story. The proposed method extracts the di… Show more

Help me understand this report
View published versions

Search citation statements

Order By: Relevance

Paper Sections

Select...

Citation Types

0
0
0

Publication Types

Select...

Relationship

0
0

Authors

Journals

citations
Cited by 0 publications
references
References 28 publications
0
0
0
Order By: Relevance

No citations

Set email alert for when this publication receives citations?