[Submitted on 23 Nov 2022 (v1), last revised 3 Mar 2023 (this version, v2)] · arXiv.org

View PDF HTML (experimental)

Abstract:Extensive research works demonstrate that the attention mechanism in convolutional neural networks (CNNs) effectively improves accuracy. Nevertheless, few works design attention mechanisms using large receptive fields. In this work, we propose a novel attention method named Rega-net to increase CNN accuracy by enlarging the receptive field. Inspired by the mechanism of the human retina, we design convolutional kernels to resemble the non-uniformly distributed structure of the human retina. Then, we sample variable-resolution values in the Gabor function distribution and fill these values in retina-like kernels. This distribution allows essential features to be more visible in the center position of the receptive field. We further design an attention module including these retina-like kernels. Experiments demonstrate that our Rega-Net achieves 79.96% top-1 accuracy on ImageNet-1K classification and 43.1% mAP on COCO2017 object detection. The mAP of the Rega-Net increased by up to 3.5% compared to baseline networks.
Subjects: Computer Vision and Pattern Recognition (cs.CV)
Cite as: arXiv:2211.12698 [cs.CV]
  (or arXiv:2211.12698v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2211.12698

arXiv-issued DOI via DataCite

Related DOI: https://doi.org/10.1109/LGRS.2023.3270186

DOI(s) linking to related resources

Submission history

From: Chun Bao [view email]
[v1] Wed, 23 Nov 2022 04:24:21 UTC (3,547 KB)
[v2] Fri, 3 Mar 2023 07:24:23 UTC (2,996 KB)

Read the original on arxiv.org ↗