ViLT: Vision-and-Language Transformer for Image Captioning with Unaligned Data - CVPR 2021
The paper 'ViLT: Vision-and-Language Transformer for Image Captioning with Unaligned Data' was published by Liunian Harold Li et al. in 2021. It was presented at the Computer Vision and Pattern Recognition (CVPR) 2021 conference, a prestigious international conference in the field of computer vision.
原文地址: https://www.cveoy.top/t/topic/qxwr 著作权归作者所有。请勿转载和采集!