This study employs CoTNet-50, an optimized network derived from Contextual Transformer Networks (CoTNet-50) built upon ResNet-50, as the CNN backbone. The network architecture and parameters of CoTNet-50 are displayed in Table 1. CoTNet-50 leverages a Transformer-style architecture, enabling it to efficiently extract both global and adjacent context information within the image. This approach enhances the learning of self-attention in a resource-efficient manner, thereby boosting the expressive capability of the output features. For a more in-depth understanding of CoTNet-50, please refer to [40] for a comprehensive explanation.

CoTNet-50: A Transformer-Based CNN Backbone for Enhanced Image Representation

原文地址: https://www.cveoy.top/t/topic/qxxJ 著作权归作者所有。请勿转载和采集!

免费AI点我,无需注册和登录