CATANIC: Automatic generation model of image captions based on multiple attention mechanism

preprint OA: closed
View at publisher

Abstract

As a cross-field of computer vision and natural language processing, image caption generation has been a widely concerned researched topic. However, it is challenging because the automatically generated captioning needs to conform to the image content and language logic. In this paper, we propose a novel image caption generation model based on the DenseNet169 pre-training network and the modified Transformer model, which follows an encoding-decoding architecture. First, we use the DenseNet169 network as the encoder to extract the initial features of the image, and then introduce the “Attention on Attention” module between the encoder and the decoder to correct the initial features of the picture and filter out the attention results that do not match the picture content. At the same time, we introduce the “Attention on Attention” module into the internal decoder of the Transformer model, which strengthens the correlation between the image feature and textual descriptions, and applies the modified Transformer model as a decoder to transform the modified image feature vectors into textual descriptions. Experimental results show that our model is not only able to generate captions that match the content of the pictures as well as the linguistic logic, but also has fewer training parameters and requires less training time cost than a number of different image captioning models with pre-trained convolutional networks as encoders.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00