A region-based image caption generator with refined descriptions

Kinghorn, Philip, Zhang, Li and Shao, Ling (2018) A region-based image caption generator with refined descriptions. Neurocomputing, 272. pp. 416-424. ISSN 0925-2312

[thumbnail of Accepted manuscript]
PDF (Accepted manuscript) - Accepted Version
Available under License Creative Commons Attribution Non-commercial No Derivatives.

Download (983kB) | Preview


Describing the content of an image is a challenging task. To enable detailed description, it requires the detection and recognition of objects, people, relationships and associated attributes. Currently, the majority of the existing research relies on holistic techniques, which may lose details relating to important aspects in a scene. In order to deal with such a challenge, we propose a novel region-based deep learning architecture for image description generation. It employs a regional object detector, recurrent neural network (RNN)-based attribute prediction, and an encoder-decoder language generator embedded with two RNNs to produce refined and detailed descriptions of a given image. Most importantly, the proposed system focuses on a local based approach to further improve upon existing holistic methods, which relates specifically to image regions of people and objects in an image. Evaluated with the IAPR TC-12 dataset, the proposed system shows impressive performance, and outperforms state-of-the-art methods using various evaluation metrics. In particular, the proposed system shows superiority over existing methods when dealing with cross-domain indoor scene images.

Item Type: Article
Uncontrolled Keywords: image description generation,convolutional and recurrent neural networks,description generation
Faculty \ School: Faculty of Science > School of Computing Sciences
Related URLs:
Depositing User: Pure Connector
Date Deposited: 15 Jul 2017 05:06
Last Modified: 22 Oct 2022 02:54
URI: https://ueaeprints.uea.ac.uk/id/eprint/64124
DOI: 10.1016/j.neucom.2017.07.014


Downloads per month over past year

Actions (login required)

View Item View Item