Comprehensive Image Captioning via Scene Graph Decomposition

Zhong, Yiwu; Wang, Liwei; Chen, Jianshu; Yu, Dong; Li, Yin

doi:10.48550/arxiv.2007.11731

Cited by 1 publication

(1 citation statement)

References 65 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…‡ Yong Zhang, Baoyuan Wu and Yujiu Yang are the corresponding authors. as image captioning [33,42] and visual question answering [36]. Intuitively, the latter would get greater benefit from more human-like scene graphs.…”

Section: Introductionmentioning

confidence: 99%

Probabilistic Modeling of Semantic Ambiguity for Scene Graph Generation

Yang¹,

Zhang²,

Zhang³

et al. 2021

Preprint

View full text Add to dashboard Cite

To generate "accurate" scene graphs, almost all existing methods predict pairwise relationships in a deterministic manner. However, we argue that visual relationships are often semantically ambiguous. Specifically, inspired by linguistic knowledge, we classify the ambiguity into three types: Synonymy Ambiguity, Hyponymy Ambiguity, and Multi-view Ambiguity. The ambiguity naturally leads to the issue of implicit multi-label, motivating the need for diverse predictions. In this work, we propose a novel plug-andplay Probabilistic Uncertainty Modeling (PUM) module. It models each union region as a Gaussian distribution, whose variance measures the uncertainty of the corresponding visual content. Compared to the conventional deterministic methods, such uncertainty modeling brings stochasticity of feature representation, which naturally enables diverse predictions. As a byproduct, PUM also manages to cover more fine-grained relationships and thus alleviates the issue of bias towards frequent relationships. Extensive experiments on the large-scale Visual Genome benchmark show that combining PUM with newly proposed ResCAGCN can achieve state-of-the-art performances, especially under the mean recall metric. Furthermore, we prove the universal effectiveness of PUM by plugging it into some existing models and provide insightful analysis of its ability to generate diverse yet plausible visual relationships.

show abstract

Section: Introductionmentioning

confidence: 99%