علوم زیست محیطی و دانش جغرافیا

علوم زیست محیطی و دانش جغرافیا

تولید نرم‌افزار توصیف تصویر بر اساس داده‌کاوی و بینایی کامپیوتر

نوع مقاله : مقاله پژوهشی

نویسندگان
1 رئیس اداره فناوری اداره کل آموزش و پرورش استان کهگیلویه و بویراحمد ، یاسوج ، ایران
2 دانشجو،دبیر انجمن علمی، مهندسی نرم افزار، دانشگاه ملی مهارت دختران یاسوج
3 دانشجو، مهندسی نرم افزار، دانشگاه ملی مهارت دختران یاسوج
چکیده
به منظور حل مشکلات کیفیت پایین توصیف تصویر، استفاده ناکافی از ویژگی‌های تصویر و تک سطحی بودن شبکه عصبی بازگشتی در تولید توصیف تصویر، این مقاله یک روش تولید توصیف تصویر مبتنی بر ویژگی‌های چندمقیاسی و بینایی کامپیوتر ارائه می‌دهد. این الگوریتم از شبکه تشخیص هدف از پیش آموزش‌دیده برای استخراج ویژگی‌های تصویر در لایه‌های مختلف شبکه عصبی کانولوشن استفاده می‌کند، ویژگی‌های تصویر را لایه به لایه وارد ساختار چندتوجهی می‌کند، ساختار چندتوجهی را به نوبه خود با شبکه عصبی بازگشتی چندلایه متصل می‌کند و یک مدل شبکه تولید توصیف تصویر چندسطحی می‌سازد. افزودن اتصالات باقیمانده به شبکه‌های عصبی بازگشتی چندلایه می‌تواند عملکرد شبکه را بهبود بخشد و می‌تواند به طور مؤثر از تخریب شبکه ناشی از عمیق شدن شبکه جلوگیری کند. نتایج تجربی نشان می‌دهد که در مجموعه تست mscoco، نمرات bleu-1 و cider الگوریتم پیشنهادی می‌توانند به ترتیب به 0.804 و 1.167 برسند که به طور قابل توجهی بهتر از الگوریتم تولید توصیف تصویر از بالا به پایین مبتنی بر یک ساختار تکتوجهی است.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

Producing image description software based on data mining and computer vision

نویسندگان English

sajad mantegihi 1
asma beheshti por 2
Zahra Ranjbar 3
forogh ahmadi shad 3
1 Head of Technology Department, General Education Department, Kohgiluyeh and Boyer Ahmad Province, Yasuj, Iran
2 Student, Secretary of the Scientific Association, Software Engineering, Yasuj National University of Girls' Skills
3 Student, Software Engineering, Yasuj National Girls' Skills University
چکیده English

In order to solve the problems of poor image description quality, insufficient use of image features, and single-level recurrent neural network in image description generation, this paper proposes an image description generation method based on multi-scale features and computer vision. This algorithm uses a pre-trained object detection network to extract image features in different layers of the convolutional neural network, inputs the image features layer by layer into the multi-attention structure, and connects the multi-attention structure with the multilayer recurrent neural network in turn, and builds a multi-level image description generation network model. Adding residual connections to the multilayer recurrent neural network can improve the network performance and can effectively prevent the network degradation caused by the network deepening. Experimental results show that on the MSCOCO test set, the Bleu-1 and Cider scores of the proposed algorithm can reach 0.804 and 1.167, respectively, which is significantly better than the top-down image description generation algorithm based on a single-attention structure.

کلیدواژه‌ها English

Short-term and long-term memory network
image description
multiple attention mechanism
multi-scale feature fusion
deep neural network