2026年 05期

结合注意力互斥部分表征的枸杞细粒度图像分类方法

Fine-grained vision classification for boxthorn combining attention-based mutually exclusive partial representation


摘要(Abstract):

针对现有细粒度图像分类方法由特征提取冗余导致的流程复杂和关键区域定位不准确的问题,提出一种结合注意力互斥部分表征的不同品种枸杞果实细粒度图像分类方法,构建多模态卷积加性自注意力视觉变换器(MMCAS-ViT)模型,将重参数化轻量级视觉转换器(RCViT-T)作为骨干网络,融合通道注意力令牌混合器(CATM)、指数移动平均(EMA)和部件空间混合器(PSM)模块,实现空间域和通道域的多重信息交互,准确定位图像关键区域,分别在自建枸杞数据集、鸟类数据集CUB-200-2011、车类数据集Stanford Cars和狗类数据集Stanford Dogs中对比验证MMCAS-ViT模型的分类准确率。结果表明,MMCAS-ViT模型通过多模态注意力机制和令牌混合策略的协同作用,在保持模型轻量化的同时显著提升了细粒度图像分类性能,枸杞、 CUB-200-2011、 Stanford Cars和Stanford Dogs 4个数据集的分类准确率分别为89.2%、 89.5%、 92.2%和93.6%,消融实验进一步证实不同模块在细粒度图像分类中的有效性。

关键词(KeyWords):细粒度图像分类;注意力机制;互斥表征;枸杞

基金项目(Foundation):宁夏回族自治区自然科学基金项目(2025AAC030196)

作者(Author):张芳艳,胡春生,巫鹏举,安巍,李英

DOI:10.13349/j.cnki.jdxbn.20260701.002

参考文献(References):

[1]Wah C,Branson S,Welinder P,et al.The Caltech-UCSD Birds-200-2011 Dataset:Technical Report CNS-TR-2011-001[R].Pasadena:California Institute of Technology,2011:2.

[2]Van Hom G,Branson S,Farrell R,et al.Building a bird recognition app and large scale dataset with citizen scientists:the fine print in fine-grained dataset collection[C]//2015 IEEE Conference on Computer Vision and Pattern Recognition(CVPR),June06-12,2015,Boston,MA,USA.New York:IEEE,2015:595.

[3]Russakovsky O,Deng Jia,Su Hao,et al.ImageNet large scale visual recognition challenge[J].International Journal of Computer Vision,2015,115:211.

[4]Krizhevsky A,Sutskever I,Hinton G E.ImageNet classification with deep convolutional neural networks[J].Communications of the ACM,2017,60(6):84.

[5]Simonyan K,Zisserman A.Very deep convolutional networks for large-scale image recognition[PP/OL].arXiv(2014-09-04)[2025-05-10].https://doi.org/10.48550/arXiv.1409.1556.

[6]He Kaiming,Zhang Xiangyu,Ren Shaoqing,et al.Deep residual learning for image recognition[C]//2016 IEEE Conference on Computer Vision and Pattern Recognition(CVPR),June 27-30,2016,Las Vegas,NV,USA.New York:IEEE,2016:770.

[7]Ge Weifeng,Lin Xiangru,Yu Yizhong.Weakly supervised complementary parts models for fine-grained image classification from the bottom up[C]//2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR),June 15-20,2019,Long Beach,CA,USA.New York:IEEE,2019:3034.

[8]Liu Chuanbin,Xie Hongtao,Zha Zhengjun,et al.Filtration and distillation:enhancing region attention for fine-grained visual categorization[J].Proceedings of the AAAI Conference on Artificial Intelligence,2020,34(7):11555.

[9]Ding Yao,Zhou Yanzhao,Zhu Yi,et al.Selective sparse sampling for fine-grained image recognition[C]//2019 IEEE/CVF International Conference on Computer Vision(ICCV),October 27-November 02,2019,Seoul,Korea(South).New York:IEEE,2019:6599.

[10]Chen Wei,Liu Tieyan,Lan Yanyan,et al.Ranking measures and loss functions in learning to rank[C]//Bengio Y,Schuurmans D,Lafferty J D,et al.NIPS'09:Proceedings of the 23rd International Conference on Neural Information Processing Systems.Red Hook:Curran Associates Inc.,2009:315.

[11]Dosovitskiy A,Beyer L,Kolesnikov A,et al.An image is worth16×16 words:transformers for image recognition at scale[PP/OL].arXiv(2020-10-22)[2025-05-06].https://doi.org/10.48550/arXiv.2010.11929.

[12]Carion N,Massa F,Synnaeve G,et al.End-to-end object detection with transformers[C]//Vedaldi A,Bischof H,Brox T,et al.Computer Vision-ECCV 2020:Lecture Notes in Computer Science,Vol 12346.Cham:Springer,2020:213-229.

[13]Zheng Sixiao,Lu Jiachen,Zhao Hengshuang,et al.Rethinking semantic segmentation from a sequence-to-sequence perspective with transformers[C]//2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR),June 20-25,2021,Nashville,TN,USA.New York:IEEE,2021:6881-6890.

[14]Xie Enze,Wang Wenjia,Wang Wenhai,et al.Segmenting transparent object segmentation in the wild with Transformer[PP/OL].arXiv(2021-01-21)[2025-05-13].https://doi.org/10.48550/arXiv.2101.08461.

[15]Chen Jieneng,Lu Yongyi,Yu Qihang,et al.TransUNet:transformers make strong encoders for medical image segmentation[PP/OL].arXiv(2021-02-08)[2025-05-17].https://doi.org/10.48550/arXiv.2102.04306.

[16]Zhang Tianfang,Li Lei,Zhou Yang,et al.CAS-ViT:convolutional additive self-attention vision transformers for efficient mobile applications[PP/OL].arXiv(2024-08-07)[2025-05-13].https://doi.org/10.48550/arXiv.2408.03703.

[17]Wang Chuanming,Fu Huiyuan,Ma Huadong.Learning mutually exclusive part representations for fine-grained image classification[J].IEEE Transactions on Multimedia,2023,26:3113.

[18]Glorot X,Bordes A,Bengio Y.Deep sparse rectifier neural networks[C]//Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics,April 11-13,2011,Fort Lauderdale,FL,USA.[S.l.]:JMLR,2011:315-323.

[19]Wang Yaming,Morariu V I,Davis L S.Learning a discriminative filter bank within a CNN for fine-grained recognition[C]//2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition,June 18-23,2018,Salt Lake City,UT,USA.New York:IEEE,2018:4148.

[20]Yang Xuhui,Wang Yaowei,Chen Ke,et al.Fine-grained objectclassification via self-supervised pose alignment[C]//2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR),June 18-24,2022,New Orleans,LA,USA.New York:IEEE,2022:7389.

[21]Du Ruoyi,Chang Dongliang,Bhunia A K,et al.Fine-grained visual classification via progressive multi-granularity training of jigsaw patches[C]//Vedaldi A,Bischof H,Brox T,et al.Computer Vision-ECCV 2020:Lecture Notes in Computer Science,Vol 12365.Cham:Springer,2020:153.

[22]Liu Xinyu,Peng Houwen,Zheng Ningxin,et al.EfficientViT:memory efficient vision transformer with cascaded group attention[C]//2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR),June 17-24,2023,Vancouver,BC,Canada.New York:IEEE,2023:14420.

[23]Pan Junting,Bulat A,Tan Fuwen,et al.EdgeViTs:competing light-weight cnns on mobile devices with vision transformers[C]//AVIDAN S,BROSTOW G,CISS??M,et al.Computer VisionECCV 2022:Lecture Notes in Computer Science,Vol 13671.Cham:Springer,2022:294.

[24]Shaker A,Maaz M,Rasheed H,et al.SwiftFormer:efficient additive attention for Transformer-based real-time mobile vision applications[PP/OL].arXiv(2023-03-27)[2025-05-09].https://doi.org/10.48550/arXiv.2303.15446.

[25]Krause J,Stark M,Deng Jia,et al.3D object representations for fine-grained categorization[C]//2013 IEEE International Conference on Computer Vision Workshops,December 02-08,2013,Sydney,Australia.New York:IEEE,2013:554.

[26]Khosla A,Jayadevaprakash N,Yao Bangpeng,et al.Novel dataset for fine-grained image categorization:Stanford Dogs[C]//First Workshop on Fine-grained Visual Categorization(FGVC),2011 IEEE Conference on Computer Vision and Pattern Recognition(CVPR),June 21-23,2011,Colorado Springs,CO,USA.New York:IEEE,2011:1.

[27]Hu Tao,Xu Jizheng,Huang Cong,et al.Weakly supervised bilinear attention network for fine-grained visual classification[PP/OL].arXiv(2019-02-1 8)[2025-05-12].https://doi.org/10.48550/arXiv.1808.02152.