Image caption automatic generation method based on weighted feature

Xi, Su Mei; Cho, Young Im

doi:10.1109/iccas.2013.6703998

Cited by 8 publications

(1 citation statement)

References 6 publications

Supporting

Mentioning

Contrasting

Order By: Relevance

“…A hybrid of the network using CNN and RNN has proven to be very good in solving the problem of generating a textual description of an image [16][17][18][19]. Deep learning applications, which use architectures made of several different networks, are used in industries ranging from automated driving to medical devices.…”

Section: Introductionmentioning

confidence: 99%

Automatic Image Caption Generation Based on Some Machine Learning Algorithms

Predić

Manic

Saračević

et al. 2022

Mathematical Problems in Engineering

View full text Add to dashboard Cite

This paper is dedicated to machine learning, the branches of machine learning, which include the methods for solving this issue, and the practical implementation of the solution to the automatic image description generation. Automatic image caption generation is one of the frequent goals of computer vision. Image description generation models must solve a larger number of complex problems to have this task successfully solved. The objects in the image must be detected and recognized, after which a logical and syntactically correct textual description is generated. For that reason, description generation is a complex problem. It is an extremely important challenge for machine learning algorithms because it represents an impersonation of a complicated human ability to encapsulate huge amounts of highlighted visual pieces of information in descriptive language. The results of the generated descriptions are compared depending on the used pretrained convolutional networks. The BLEU metrics are used to calculate the quality of the image description. Although the solution to the problem of image description automatic generation does provide us with good results, there is yet room for improvement since there are images that are not adequately described.

show abstract

Section: Introductionmentioning

confidence: 99%