In order to further explore useful information in entity images, this article additionally uses a multi-modal pre-training model to perform semantic matching on the description images of entities, and removes the description images with poor semantic connections. Then, an OCR model is used to extract possible text information in the images as auxiliary knowledge for the visual encoding obtained only by using pre-trained visual models.

多模态预训练模型提升实体图像信息提取

原文地址: https://www.cveoy.top/t/topic/m9Ep 著作权归作者所有。请勿转载和采集!

免费AI点我,无需注册和登录