Learning from Heterogeneous Label Quality: Strategies for Robust Model Training

Main Article Content

  Wu Shukun
  Hafiz Muhammad Kurniawan
  Adolf Asih Suprianto

Abstract

Background: Training data collected from crowdsourcing platforms, automated labeling systems, and multiple annotators often contain labels with unequal levels of reliability. Conventional machine-learning methods generally treat all labels as equally trustworthy, which can reduce model performance when label errors vary across samples and annotation sources.
Aims: This study proposes a quality-aware training method that explicitly models the heterogeneity of label reliability. It examines whether sample-level quality weighting can improve predictive performance and robustness under synthetic and real-world label-noise conditions.
Methods: The proposed method combines annotation or neighborhood agreement, loss-based confidence, and source-level reliability to estimate a continuous quality score for each training sample. The scores are incorporated into a weighted cross-entropy objective and periodically updated during training. Experiments were conducted on CIFAR-10, CIFAR-100, CIFAR-10N, and CIFAR-100N using ResNet-18. The method was compared with cross-entropy, label smoothing, generalized cross-entropy, Co-teaching, and DivideMix.
Result: Under high heterogeneous noise, the proposed method achieved accuracies of 83.21% on CIFAR-10 and 54.62% on CIFAR-100, compared with 73.48% and 45.92% for conventional cross-entropy. It also achieved the highest accuracy under the evaluated real-world noisy-label conditions. Ablation and sensitivity analyses showed that the combined quality indicators were complementary and that the method remained effective under moderate quality-estimation errors.
Conclusion: Explicitly modeling label reliability improves robustness when training data contain unequal annotation quality. The proposed method provides a practical and computationally efficient strategy for learning from crowdsourced, weakly supervised, and multi-source datasets.

Article Details

How to Cite
Shukun, W., Kurniawan, H. M., & Suprianto, A. A. (2026). Learning from Heterogeneous Label Quality: Strategies for Robust Model Training. International Journal of Advances in Artificial Intelligence and Machine Learning, 3(2), 188–207. https://doi.org/10.58723/ijaaiml.v3i2.799
Section
Articles

References

Bernhardt, M., Castro, D. C., Tanno, R., Schwaighofer, A., Tezcan, K. C., Monteiro, M., Bannur, S., Lungren, M. P., Nori, A., Glocker, B., Alvarez-Valle, J., & Oktay, O. (2022). Active label cleaning for improved dataset quality under resource constraints. Nature Communications, 13(1). https://doi.org/10.1038/s41467-022-28818-3

Braun, D. (2024). I beg to differ: how disagreement is handled in the annotation of legal machine learning data sets. Artificial Intelligence and Law, 32(3), 839–862. https://doi.org/10.1007/s10506-023-09369-4

Brigato, L., Barz, B., Iocchi, L., & Denzler, J. (2022). Image Classification With Small Datasets: Overview and Benchmark. IEEE Access, 10, 49233–49250. https://doi.org/10.1109/ACCESS.2022.3172939

Capstick, A., Palermo, F., Cui, T., & Barnaghi, P. (2025). Training neural networks on data sources with unknown reliability. Information Fusion, 123, 103239. https://doi.org/10.1016/J.INFFUS.2025.103239

Chase, R. J., Harrison, D. R., Burke, A., Lackmann, G. M., & McGovern, A. (2022). A Machine Learning Tutorial for Operational Meteorology. Part I: Traditional Machine Learning. Weather and Forecasting, 37(8), 1509–1529. https://doi.org/10.1175/WAF-D-22-0070.1

Chen, Y., Xu, K., Zhou, P., Ban, X., & He, D. (2022). Improved cross entropy loss for noisy labels in vision leaf disease classification. IET Image Processing, 16(6), 1511–1519. https://doi.org/10.1049/ipr2.12402

de Vos, B. D., Jansen, G. E., & Išgum, I. (2023). Stochastic co-teaching for training neural networks with unknown levels of label noise. Scientific Reports 2023 13:1, 13(1), 16875-. https://doi.org/10.1038/s41598-023-43864-7

Gao, Y., Fu, J., Wang, Y., & Guo, Y. (2024). Typicality- and instance-dependent label noise-combating : a novel framework for simulating and combating real-world noisy labels for endoscopic polyp classification. Visual Computing for Industry, Biomedicine, and Art, 7. https://doi.org/10.1186/s42492-024-00162-x

Gil-González, J., Daza-Santacoloma, G., Cárdenas-Peña, D., Orozco-Gutiérrez, A., & Álvarez-Meza, A. (2025). Generalized cross-entropy for learning from crowds based on correlated chained Gaussian processes. Results in Engineering, 25(January), 103863. https://doi.org/10.1016/j.rineng.2024.103863

González-Santoyo, C., Renza, D., & Moya-Albor, E. (2025). Identifying and Mitigating Label Noise in Deep Learning for Image Classification. Technologies 2025, Vol. 13, Page 132, 13(4), 132. https://doi.org/10.3390/TECHNOLOGIES13040132

Hafiz, A. M., Bhat, R. A., & Hassaballah, M. (2022). Image classification using convolutional neural network tree ensembles. Multimedia Tools and Applications 2022 82:5, 82(5), 6867–6884. https://doi.org/10.1007/S11042-022-13604-6

Haliburton, L., Leusmann, J., Welsch, R., Ghebremedhin, S., Isaakidis, P., Schmidt, A., & Mayer, S. (2025). Uncovering labeler bias in machine learning annotation tasks. AI and Ethics, 5(3), 2515–2528. https://doi.org/10.1007/s43681-024-00572-w

Han, B., Sun, Y. X., Zhang, Y. L., Zhang, L., Hu, H., Li, L., Zhou, J., Ye, G., & He, H. (2024). Collaborative Refining for Learning from Inaccurate Labels. Advances in Neural Information Processing Systems, 37, 92745–92768. https://doi.org/10.52202/079017-2945

Herde, M., Huseljic, D., Sick, B., & Calma, A. (2021). A Survey on Cost Types, Interaction Schemes, and Annotator Performance Models in Selection Algorithms for Active Learning in Classification. IEEE Access, 9, 166970–166989. https://doi.org/10.1109/ACCESS.2021.3135514

Higashimoto, R., Yoshida, S., & Muneyasu, M. (2024). CRAS: Curriculum Regularization and Adaptive Semi-Supervised Learning with Noisy Labels. Applied Sciences (Switzerland), 14(3). https://doi.org/10.3390/app14031208

Klie, J. C., de Castilho, R. E., & Gurevych, I. (2024). Analyzing Dataset Annotation Quality Management in the Wild. Computational Linguistics, 50(3), 817–866. https://doi.org/10.1162/COLI_A_00516

Klie, J. C., Webber, B., & Gurevych, I. (2023). Annotation Error Detection: Analyzing the Past and Present for a More Coherent Future. Computational Linguistics, 49(1), 157–198. https://doi.org/10.1162/coli_a_00464

Kukreja, V., Abraham, A., Kalaiselvi, K., Thilak, K. D., Hariharan, S., & Chen, S. Y. (2023). Machine Learning for Data Fusion: A Fuzzy AHP Approach for Open Issues. Computers, Materials and Continua, 77(3), 2899–2914. https://doi.org/10.32604/cmc.2023.045136

Lazaros, K., Vrahatis, A. G., & Kotsiantis, S. (2026). Human-in-the-Loop Artificial Intelligence: A Systematic Review of Concepts, Methods, and Applications. Entropy, 28(4), 1–50. https://doi.org/10.3390/e28040377

Liu, Z., & Letchmunan, S. (2024). Representing uncertainty and imprecision in machine learning: A survey on belief functions. Journal of King Saud University - Computer and Information Sciences, 36(1). https://doi.org/10.1016/j.jksuci.2023.101904

Liu, Z., Qian, P., Wang, X., Zhuang, Y., Qiu, L., & Wang, X. (2023). Combining Graph Neural Networks with Expert Knowledge for Smart Contract Vulnerability Detection. IEEE Transactions on Knowledge and Data Engineering, 35(2), 1296–1310. https://doi.org/10.1109/TKDE.2021.3095196

Mohammed, S., Budach, L., Feuerpfeil, M., Ihde, N., Nathansen, A., Noack, N., Patzlaff, H., Naumann, F., & Harmouch, H. (2025). The effects of data quality on machine learning performance on tabular data. Information Systems, 132(March), 102549. https://doi.org/10.1016/j.is.2025.102549

Northcutt, C. G., Athalye, A., & Mueller, J. (2021). Pervasive Label Errors in Test Sets Destabilize Machine Learning Benchmarks. Advances in Neural Information Processing Systems. https://arxiv.org/pdf/2103.14749

Northcutt, C. G., Jiang, L., & Chuang, I. L. (2021). Confident Learning: Estimating Uncertainty in Dataset Labels. Journal of Artificial Intelligence Research, 70, 1373–1411. https://doi.org/10.1613/JAIR.1.12125

Özdemir, Z., Keles, H. Y., & Tanriöver, Ö. Ö. (2026). CLoE: Curriculum Learning on Endoscopic Images for Robust MES Classification. IEEE Access, 14, 30441–30454. https://doi.org/10.1109/ACCESS.2026.3666962

Pradhan, V. K., Schaekermann, M., & Lease, M. (2022). In Search of Ambiguity: A Three-Stage Workflow Design to Clarify Annotation Guidelines for Crowd Workers. Frontiers in Artificial Intelligence, 5(May), 1–16. https://doi.org/10.3389/frai.2022.828187

Priestley, M., O’Donnell, F., & Simperl, E. (2023). A Survey of Data Quality Requirements That Matter in ML Development Pipelines. Journal of Data and Information Quality, 15(2). https://doi.org/10.1145/3592616

Rashidi, H. H., Tran, N., Albahra, S., & Dang, L. T. (2021). Machine learning in health care and laboratory medicine: General overview of supervised learning and Auto-ML. International Journal of Laboratory Hematology, 43(S1), 15–22. https://doi.org/10.1111/ijlh.13537

Ratner, A., Bach, S. H., Ehrenberg, H., Fries, J., Wu, S., & Ré, C. (2019). Snorkel: rapid training data creation with weak supervision. The VLDB Journal 2019 29:2, 29(2), 709–730. https://doi.org/10.1007/S00778-019-00552-1

Shetty, S. H., Shetty, S., Singh, C., & Rao, A. (2022). Supervised machine learning: Algorithms and applications. Fundamentals and Methods of Machine and Deep Learning: Algorithms, Tools, and Applications, 1–16. https://doi.org/10.1002/9781119821908.CH1;CTYPE:STRING:BOOK

Wang, Q., Han, B., Liu, T., Niu, G., Yang, J., & Gong, C. (2021). Tackling Instance-Dependent Label Noise via a Universal Probabilistic Model. 35th AAAI Conference on Artificial Intelligence, AAAI 2021, 11B, 10183–10191. https://doi.org/10.1609/aaai.v35i11.17221

Wang, Y., Han, X., Li, C., Luo, L., Yin, Q., Zhang, J., Peng, G., Shi, D., & He, M. (2024). Impact of Gold-Standard Label Errors on Evaluating Performance of Deep Learning Models in Diabetic Retinopathy Screening: Nationwide Real-World Validation Study. Journal of Medical Internet Research, 26, 1–12. https://doi.org/10.2196/52506

Xu, C., Song, P., Tang, S., Guo, D., & Yang, X. (2025). Alleviating Confirmation Bias in Learning with Noisy Labels via Two-Network Collaboration. ACM Transactions on Intelligent Systems and Technology, 16(4). https://doi.org/10.1145/3723009

Ye, C., Duan, H., Zhang, H., Zhang, H., Wang, H., & Dai, G. (2023a). Multi-Source Data Repairing: A Comprehensive Survey. Mathematics, 11(10), 1–22. https://doi.org/10.3390/math11102314

Ye, C., Duan, H., Zhang, H., Zhang, H., Wang, H., & Dai, G. (2023b). Multi-Source Data Repairing: A Comprehensive Survey. Mathematics, 11(10), 1–22. https://doi.org/10.3390/math11102314

Zhang, T., Yi, C., Song, Y., Wei, B., & Chen, Z. (2026). Hyperbolic DivideMix: Leveraging Poincaré Ball Embeddings for Robust Human Activity Recognition under Label Noise. IEEE Sensors Journal. https://doi.org/10.1109/JSEN.2026.3680607

Zhang, Z., & Sabuncu, M. R. (2018). Generalized Cross Entropy Loss for Training Deep Neural Networks with Noisy Labels. Advances in Neural Information Processing Systems, 2018-December, 8778–8788. https://arxiv.org/pdf/1805.07836

Zhao, X., Huang, P., & Shu, X. (2022). Wavelet-Attention CNN for image classification. Multimedia Systems 2022 28:3, 28(3), 915–924. https://doi.org/10.1007/S00530-022-00889-8

Zhong, Y., Du, B., & Xu, C. (2021). Learning to reweight examples in multi-label classification. Neural Networks, 142, 428–436. https://doi.org/10.1016/J.NEUNET.2021.03.022