Oboyob: A sequential-semantic Bengali image captioning engine

Deb, Tonmoay; Ali, Mohammad Zariff Ahsham; Bhowmik, Sanchita; Firoze, Adnan; Ahmed, Syed Shahir; Tahmeed, Muhammad Abeer; Rahman, N.S.M. Rezaur; Rahman, Rashedur M.

doi:10.3233/JIFS-179351

Oboyob: A sequential-semantic Bengali image captioning engine

Issue title: Special Section: Collective intelligence in information systems

Guest editors: Ngoc Thanh Nguyen, Edward Szczerbicki, Bogdan Trawiński and Van Du Nguyen

Article type: Research Article

Authors: Deb, Tonmoay | Ali, Mohammad Zariff Ahsham | Bhowmik, Sanchita | Firoze, Adnan | Ahmed, Syed Shahir | Tahmeed, Muhammad Abeer | Rahman, N.S.M. Rezaur | Rahman, Rashedur M.^{; *}

Affiliations: Department of Electrical & Computer Engineering, North South University, Bangladesh

Correspondence: [*] Corresponding author. Rashedur M. Rahman, Department of Electrical & Computer Engineering, North South University, Bangladesh. E-mail: rashedur.rahman@northsouth.edu.

Abstract: Understanding the context with generation of textual description from an input image is an active and challenging research topic in computer vision and natural language processing. However, in the case of Bengali language, the problem is still unexplored. In this paper, we address a standard approach for Bengali image caption generation though subsampling the machine translated dataset. Later, we use several pre-processing techniques with the state-of-the-art CNN-LSTM architecture-based models. The experiment is conducted on standard Flickr-8K dataset, along with several modifications applied to adapt with the Bengali language. The training caption subsampled dataset is computed for both Bengali and English languages for further experiments with 16 distinct models developed in the entire training process. The trained models for both languages are analyzed with respect to several caption evaluation metrics. Further, we establish a baseline performance in Bengali image captioning defining the limitation of current word embedding approaches compared to internal local embedding.

Keywords: Image captioning, CNN, LSTM, natural language processing, computer vision, Bengali image captioning, merge architecture, par-inject architecture, machine translated caption subsampling

DOI: 10.3233/JIFS-179351

Journal: Journal of Intelligent & Fuzzy Systems, vol. 37, no. 6, pp. 7427-7439, 2019

Published: 23 December 2019

Price: EUR 27.50

North America

IOS Press, Inc.
6751 Tepper Drive
Clifton, VA 20124
USA

Tel: +1 703 830 6300
Fax: +1 703 830 2300
sales@iospress.com

For editorial issues, like the status of your submitted paper or proposals, write to editorial@iospress.nl

Europe

IOS Press
Nieuwe Hemweg 6B
1013 BG Amsterdam
The Netherlands

Tel: +31 20 688 3355
Fax: +31 20 687 0091
info@iospress.nl

For editorial issues, permissions, book requests, submissions and proceedings, contact the Amsterdam office info@iospress.nl

Asia

Inspirees International (China Office)
Ciyunsi Beili 207(CapitaLand), Bld 1, 7-901
100025, Beijing
China

Free service line: 400 661 8717
Fax: +86 10 8446 7947
china@iospress.cn

For editorial issues, like the status of your submitted paper or proposals, write to editorial@iospress.nl

如果您在出版方面需要帮助或有任何建, 件至: editorial@iospress.nl

Share this:

North America

Europe

Asia