|
| AllMix用于标签噪声学习的图像分类方法 |
| Image classification approach using AllMix for label noise learning |
| 投稿时间:2024-02-27 |
| DOI:10.3969/j.issn.1005-5630.202402270026 |
| 中文关键词: 标签噪声学习 图像分类 半监督学习 对比学习 |
| 英文关键词:label noise learning image classification semi-supervised learning contrastive learning |
| 基金项目:基础科研条件与重大科学仪器设备研发计划(2022YFF0706003) |
|
| 摘要点击次数: 1986 |
| 全文下载次数: 1512 |
| 中文摘要: |
| 人工收集和标注的数据集不可避免会含有标签噪声,其对图像分类模型的泛化能力存在负面影响。因此,设计针对含有标签噪声数据集的鲁棒性分类算法成为研究热点。自监督学习预训练耗时且在样本选择后仍包含较多噪声样本是现有方法的主要问题。提出了AllMix模型,该模型节省了预训练的时间,在DivideMix模型的基础上,使用AllMatch训练策略替换原有的MixMatch训练策略。AllMatch训练策略使用焦损和广义交叉熵损失来优化含有标签噪声样本的损失计算,同时引入了高置信度样本半监督学习模块和对比学习模块充分学习无标签噪声样本。实验结果表明,AllMix模型比现有的无训练的标签噪声分类算法的性能,在CIFAR10数据集上,针对50%、80%和90%的对称噪声分别高出了0.7%、0.7%和5.0%,针对含有80%和90%对称噪声的CIFAR100数据集分别提高了2.8%和10.1%。 |
| 英文摘要: |
| Datasets collected and annotated manually are inevitably contaminated with label noise, which negatively affects the generalization ability of image classification models. Therefore, designing robust classification algorithms for datasets with label noise has become a hot research topic. The main issue with existing methods is that self-supervised learning pre-training is time-consuming and still includes a large number of noisy samples after sample selection. This paper introduces the AllMix model, which reduces the time required for pre-training. Based on the DivideMix model, the AllMatch training strategy replaces the original MixMatch training strategy. The AllMatch training strategy uses focal loss and generalized cross-entropy loss to optimize the loss calculation for labeled samples. Additionally, it introduces a high-confidence sample semi-supervised learning module and a contrastive learning module to fully learn from unlabeled samples. Experimental results show that on the CIFAR10 dataset, the existing pre-trained label noise classification algorithms are 0.7%, 0.7%, and 5.0% higher in performance than those without pre-training for 50%, 80%, and 90% symmetric noise ratios, respectively. On the CIFAR100 dataset with 80% and 90% symmetric noise ratios, the model performance is 2.8% and 10.1% higher, respectively. |
| HTML 查看全文 查看/发表评论 下载PDF阅读器 |
| 关闭 |