Study and Development of Image Caption Generation using Various Encoders for Different Image Categories

preprint OA: closed CC-BY-4.0
📄 Open PDF View at publisher

Abstract

Abstract Images can act as the source of information or way of communication. Captioning an image by machines in a manner such that it conveys the true meaning is considered the most difficult task. It is indeed a process of deep analysis and most researched area. This paper presents a comparative analysis of different deep learning models such as Inception V3, Resnet50 and VGG16 based on a novel image captioning methodology applied on different picture categories. The hybrid approach is developed to get high BLEU score for each input images. The data set generation, implementation of hybrid approach, and the challenges along with the future work are discussed.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-28T02:00:01.590549+00:00
License: CC-BY-4.0