用户登录
期刊信息
  • 主管单位:
  • 中国科学技术协会
  • 主办单位:
  • 中国仪器仪表学会、上海光学仪器研究所、中国光学学会工程光学专业委员会
  • 主  编:
  • 庄松林
  • 地  址:
  • 上海市军工路516号上海理工大学《光学仪器》编辑部
  • 邮政编码:
  • 200093
  • 联系电话:
  • 021-55270110
  • 电子邮件:
  • gxyq@usst.edu.cn
  • 国际标准刊号:
  • 1005-5630
  • 国内统一刊号:
  • 31-1504/TH
  • 邮发代号:
  • 单  价:
  • 15.00
  • 定  价:
  • 90.00
AllMix用于标签噪声学习的图像分类方法
Image classification approach using AllMix for label noise learning
投稿时间:2024-02-27  
DOI:10.3969/j.issn.1005-5630.202402270026
中文关键词:  标签噪声学习  图像分类  半监督学习  对比学习
英文关键词:label noise learning  image classification  semi-supervised learning  contrastive learning
基金项目:基础科研条件与重大科学仪器设备研发计划(2022YFF0706003)
作者单位E-mail
邱烨敏 上海理工大学 光电信息与计算机工程学院,上海 200093  
张荣福 上海理工大学 光电信息与计算机工程学院,上海 200093 zrf@usst.edu.cn 
何晨 上海理工大学 光电信息与计算机工程学院,上海 200093  
杨紫叶 上海理工大学 光电信息与计算机工程学院,上海 200093  
高顾昱 上海理工大学 光电信息与计算机工程学院,上海 200093  
摘要点击次数: 1986
全文下载次数: 1512
中文摘要:
      人工收集和标注的数据集不可避免会含有标签噪声,其对图像分类模型的泛化能力存在负面影响。因此,设计针对含有标签噪声数据集的鲁棒性分类算法成为研究热点。自监督学习预训练耗时且在样本选择后仍包含较多噪声样本是现有方法的主要问题。提出了AllMix模型,该模型节省了预训练的时间,在DivideMix模型的基础上,使用AllMatch训练策略替换原有的MixMatch训练策略。AllMatch训练策略使用焦损和广义交叉熵损失来优化含有标签噪声样本的损失计算,同时引入了高置信度样本半监督学习模块和对比学习模块充分学习无标签噪声样本。实验结果表明,AllMix模型比现有的无训练的标签噪声分类算法的性能,在CIFAR10数据集上,针对50%、80%和90%的对称噪声分别高出了0.7%、0.7%和5.0%,针对含有80%和90%对称噪声的CIFAR100数据集分别提高了2.8%和10.1%。
英文摘要:
      Datasets collected and annotated manually are inevitably contaminated with label noise, which negatively affects the generalization ability of image classification models. Therefore, designing robust classification algorithms for datasets with label noise has become a hot research topic. The main issue with existing methods is that self-supervised learning pre-training is time-consuming and still includes a large number of noisy samples after sample selection. This paper introduces the AllMix model, which reduces the time required for pre-training. Based on the DivideMix model, the AllMatch training strategy replaces the original MixMatch training strategy. The AllMatch training strategy uses focal loss and generalized cross-entropy loss to optimize the loss calculation for labeled samples. Additionally, it introduces a high-confidence sample semi-supervised learning module and a contrastive learning module to fully learn from unlabeled samples. Experimental results show that on the CIFAR10 dataset, the existing pre-trained label noise classification algorithms are 0.7%, 0.7%, and 5.0% higher in performance than those without pre-training for 50%, 80%, and 90% symmetric noise ratios, respectively. On the CIFAR100 dataset with 80% and 90% symmetric noise ratios, the model performance is 2.8% and 10.1% higher, respectively.
HTML   查看全文  查看/发表评论  下载PDF阅读器
关闭