نوع مقاله : مقاله پژوهشی
عنوان مقاله English
نویسندگان English
In order to solve the problems of poor image description quality, insufficient use of image features, and single-level recurrent neural network in image description generation, this paper proposes an image description generation method based on multi-scale features and computer vision. This algorithm uses a pre-trained object detection network to extract image features in different layers of the convolutional neural network, inputs the image features layer by layer into the multi-attention structure, and connects the multi-attention structure with the multilayer recurrent neural network in turn, and builds a multi-level image description generation network model. Adding residual connections to the multilayer recurrent neural network can improve the network performance and can effectively prevent the network degradation caused by the network deepening. Experimental results show that on the MSCOCO test set, the Bleu-1 and Cider scores of the proposed algorithm can reach 0.804 and 1.167, respectively, which is significantly better than the top-down image description generation algorithm based on a single-attention structure.
کلیدواژهها English