{"ID":54604,"CreatedAt":"2026-02-27T13:00:40Z","UpdatedAt":"2026-02-27T13:00:40Z","DeletedAt":null,"paper_url":"https://paperswithcode.com/paper/chittron-an-automatic-bangla-image-captioning","arxiv_id":"1809.00339","title":"Chittron: An Automatic Bangla Image Captioning System","abstract":"Automatic image caption generation aims to produce an accurate description of\nan image in natural language automatically. However, Bangla, the fifth most\nwidely spoken language in the world, is lagging considerably in the research\nand development of such domain. Besides, while there are many established data\nsets to related to image annotation in English, no such resource exists for\nBangla yet. Hence, this paper outlines the development of \"Chittron\", an\nautomatic image captioning system in Bangla. Moreover, to address the data set\navailability issue, a collection of 16,000 Bangladeshi contextual images has\nbeen accumulated and manually annotated in Bangla. This data set is then used\nto train a model which integrates a pre-trained VGG16 image embedding model\nwith stacked LSTM layers. The model is trained to predict the caption when the\ninput is an image, one word at a time. The results show that the model has\nsuccessfully been able to learn a working language model and to generate\ncaptions of images quite accurately in many cases. The results are evaluated\nmainly qualitatively. However, BLEU scores are also reported. It is expected\nthat a better result can be obtained with a bigger and more varied data set.","url_abs":"http://arxiv.org/abs/1809.00339v1","url_pdf":"http://arxiv.org/pdf/1809.00339v1.pdf","authors":"[\"Motiur Rahman\", \"Nabeel Mohammed\", \"Nafees Mansoor\", \"Sifat Momen\"]","published":"2018-09-02T00:00:00Z","tasks":"[\"Caption Generation\", \"Image Captioning\", \"Language Modeling\", \"Language Modelling\"]","methods":"[\"Sigmoid Activation\", \"Tanh Activation\", \"LSTM\"]","has_code":false}
