Rizal Setya Perdana, Yoshiteru Ishida
Image captioning task achieves imposing result in generating text description of image by training on the large image and sentence pairs dataset (e.g., MSCOCO). Applied the large publicly available datasets to the specific new task will lead to the problem known as domain shifting due to the different probability distribution. In this research, we propose a Multimodal Instance-based Deep Transfer Learning (MIBTL) for cross-domain image captioning. Generally, transferring knowledge from the source domain into the target domain is pertinent due to the limited number of the dataset from the target domain. The limited number of the dataset also lead to a problem called overfitting. The instance-based strategy intuitively takes into account the influence of data by selecting the most representative data from source domain as the supplement for training on the target domain. We employ deep hash binary code representation of image and text pair to determine the distance between two data points. The experiment conducted by two types of transfer condition includes a slight shift (from MSCOCO to Flickr30k) and significant shift (from MSCOCO to CUB-200 and Oxford-102). The result shows that MIBTL can outperform other baselines to achieve state-of-the-art performance in cross-domain image captioning transfer learning. © 2019 IEEE.
Toyohashi University of Technology, Department of Computer Science and Engineering, Toyohashi, Japan; Department of Information System, Faculty of Computer Science, Brawijaya University, Malang, Indonesia