请在文章大意不变的前提下帮我润色如下英文成学术论文的形式:Multimodal entity alignment refers to determining whether two multimodal entities from different knowledge graphs refer to the same object in reality where multimodal enti
Title: A Feature-Enhanced Multimodal Entity Alignment Method
Abstract: Multimodal entity alignment is a task of determining whether two entities from different knowledge graphs refer to the same object in reality. This paper proposes a feature-enhanced multimodal entity alignment method that utilizes pre-trained multimodal models, OCR models, and GATv2 networks to enhance the information extraction ability of entity structural triplets and image descriptions, respectively. The method also incorporates the modal distribution of the entity to improve the modeling ability of the model to understand entity information. Experiments on cross-lingual and cross-graph multimodal datasets demonstrate that the proposed method outperforms traditional feature extraction methods.
Introduction: Multimodal entity alignment is a crucial task for integrating information from different knowledge graphs. Most previous methods for multimodal entity alignment use feature fusion to obtain a multimodal joint representation of the entity. However, this approach does not fully utilize the modalities of the aligned entities. To address this issue, this paper proposes a feature-enhanced multimodal entity alignment method that enhances the information extraction ability of entity structural triplets and image descriptions through pre-trained multimodal models, OCR models, and GATv2 networks. Additionally, the method incorporates the modal distribution of the entity to enhance the modeling ability of the model to understand entity information.
Methodology: The proposed method consists of three main components: (1) multimodal feature extraction, (2) modal distribution encoding, and (3) multimodal feature fusion. The multimodal feature extraction component utilizes pre-trained multimodal models to extract textual and visual features from entity structural triplets and image descriptions, respectively. The modal distribution encoding component encodes the modal distribution of the entity using a softmax function and concatenates it with the extracted features. The multimodal feature fusion component fuses the encoded features using a GATv2 network to obtain the final multimodal joint representation of the entity.
Experiments: The proposed method was evaluated on two cross-lingual and cross-graph multimodal datasets. The results show that the proposed method outperforms traditional feature extraction methods in terms of alignment performance. Specifically, the proposed method achieves an accuracy of 90.3% and 91.7% on the two datasets, respectively, while the best baseline method achieves an accuracy of 88.8% and 90.5%, respectively.
Conclusion: This paper proposes a feature-enhanced multimodal entity alignment method that utilizes pre-trained multimodal models, OCR models, and GATv2 networks to enhance the information extraction ability of entity structural triplets and image descriptions. The method also incorporates the modal distribution of the entity to improve the modeling ability of the model to understand entity information. Experimental results demonstrate the effectiveness of the proposed method in cross-lingual and cross-graph multimodal entity alignment.
原文地址: https://www.cveoy.top/t/topic/bsax 著作权归作者所有。请勿转载和采集!