What are the objective of image captioning?
The aim of image captioning is to automatically de- scribe an image with one or more natural language sen- tences. This is a problem that integrates computer vision and natural language processing, so its main challenges arise from the need of translating between two distinct, but usually paired, modalities [4].
What is image captioning in deep learning?
Image Captioning is the process of generating a textual description for given images. It has been a very important and fundamental task in the Deep Learning domain. Image captioning has a huge amount of application.
What are the different types of image captioning?
Based on the technique adopted, we classify image captioning approaches into different categories. Representative methods in each category are summarized, and their strengths and limitations are talked about. In this paper, we first discuss methods used in early work which are mainly retrieval and template based.
What does it mean to automatically generate a caption for an image?
Image captioning means automatically generating a caption for an image. As a recently emerged research area, it is attracting more and more attention. To achieve the goal of image captioning, semantic information of images needs to be captured and expressed in natural languages.
Why is image captioning a difficult problem to solve?
Image captioning is a challenging problem owing to the complexity in understanding the image content and di- verse ways of describing it in natural language. Recent advances in deep neural networks have substantially im- proved the performance of this task.
What is the fourth part of image caption generation?
The fourth part introduces the common datasets come up by the image caption and compares the results on different models. Different evaluation methods are discussed. The fifth part summarizes the existing work and proposes the direction and expectations of future work. 2. Feature Extraction Methods