We use the pre-trained multi-modal pre-training model CLIP to encode entity names and entity images separately and calculate their similarity. For entity images, we directly use CLIP for image encoding, while for entity names, we modify them to 'A photo of entity name' and input them into CLIP for text encoding, and then delete images with similarity below the set threshold.

使用CLIP模型进行实体名称和图像相似度匹配

原文地址: https://www.cveoy.top/t/topic/m9EQ 著作权归作者所有。请勿转载和采集!

免费AI点我,无需注册和登录