This paper proposes the CoT block, which is a Transformer-style feature extraction block that not only has the advantages of self-attention mechanism but also can connect the contextual information of nearby convolution, enhancing the visual expression ability.

This paper proposes a multi-scale feature fusion module, where features at different scales contain and express different feature information. By fusing the multi-layer features of the decoding part, it can simultaneously enhance the expression of spatial geometric feature information and semantic feature information, enabling the network to accurately segment the target.

CoT Block: A Transformer-Style Feature Extraction Block for Enhanced Visual Representation

原文地址: https://www.cveoy.top/t/topic/pepv 著作权归作者所有。请勿转载和采集!

免费AI点我,无需注册和登录