多模态预训练模型提升实体图像信息提取
In order to further explore useful information in entity images, this article additionally uses a multi-modal pre-training model to perform semantic matching on the description images of entities, and removes the description images with poor semantic connections. Then, an OCR model is used to extract possible text information in the images as auxiliary knowledge for the visual encoding obtained only by using pre-trained visual models.
原文地址: https://www.cveoy.top/t/topic/m9Ep 著作权归作者所有。请勿转载和采集!