Self-Supervised Learning Strategies for Reducing Label Dependency in Large-Scale Data Science Applications

Main Article Content

  Kit Ling Chan
  Mohd Zaki Zakaria
  Eka Puji Agustini

Abstract

Background: The rapid expansion of large-scale data ecosystems has intensified the challenge of acquiring high-quality labeled datasets for supervised learning models. Although raw data availability continues to grow exponentially, manual annotation remains costly, time-consuming, and susceptible to inconsistency, creating a structural bottleneck in scalable machine learning deployment. Existing self-supervised learning (SSL) approaches have demonstrated promising results in domain-specific applications, particularly in computer vision. However, many lack a unified, modality-agnostic framework and systematic evaluation under realistic constraints such as limited-label availability and noisy annotations.
Aims: This study aims to develop a unified self-supervised learning framework that reduces label dependency across heterogeneous data modalities. The scope of the research includes image and tabular datasets, focusing on evaluating model robustness under low-label conditions and label noise scenarios.
Methods: The proposed framework integrates contrastive representation learning with reconstruction-based objectives in a two-stage training strategy. First, large-scale unlabeled pre-training is conducted to learn robust feature representations. This is followed by supervised fine-tuning under controlled label fractions (1%–20%) and synthetic label corruption levels of up to 30%. Performance is evaluated against fully supervised baselines using accuracy and data efficiency metrics.
Result: Experimental findings demonstrate that the proposed SSL framework achieves up to 15% higher accuracy than fully supervised baselines under 1% labeled data conditions. Moreover, it maintains approximately 85–90% of clean-label performance even with 30% label corruption. The data efficiency ratio improves by up to 2.3× in low-label regimes, indicating substantial gains in learning effectiveness under constrained supervision.
Conclusion: The results confirm that self-supervised pre-training significantly enhances representation quality, robustness, and scalability while reducing annotation requirements. This research establishes SSL as a foundational paradigm for developing data-efficient and resilient AI systems in large-scale analytics environments.

Article Details

How to Cite
Ling Chan, K., Zakaria, M. Z., & Agustini, E. P. (2026). Self-Supervised Learning Strategies for Reducing Label Dependency in Large-Scale Data Science Applications. International Journal of Advances in Artificial Intelligence and Machine Learning, 3(2), 113–125. https://doi.org/10.58723/ijaaiml.v3i2.802
Section
Articles

References

Adegboye, M. A., Karnik, A., & Fung, W.-K. (2021). Numerical study of pipeline leak detection for gas-liquid stratified flow. Journal of Natural Gas Science and Engineering, 94. https://doi.org/10.1016/j.jngse.2021.104054

Ali, A. A., & Hammad, M. T. (2025). A Co-Evolutionary Genetic Algorithm Approach to Optimizing Deep Learning for Brain Tumor Classification. IEEE Access, 13, 21229–21248. https://doi.org/10.1109/ACCESS.2025.3535844

Alomar, K., Aysel, H. I., & Cai, X. (2023). Data Augmentation in Classification and Segmentation : A Survey and New Strategies. Journal of Imaging, 9(2), 46. https://doi.org/10.3390/jimaging9020046

Bernhardt, M., Castro, D. C., Tanno, R., Schwaighofer, A., Tezcan, K. C., Monteiro, M., Bannur, S., Lungren, M. P., Nori, A., Glocker, B., Alvarez-valle, J., & Oktay, O. (2022). Active label cleaning for improved dataset quality under resource constraints. Nature Communications, 13(1). https://doi.org/10.1038/s41467-022-28818-3

Chikhaoui, K., & Alfarraj, M. (2024). Self-supervised learning for efficient seismic facies classification. Geophysics. https://doi.org/10.1190/geo2023-0508.1

Dong, N., Kampffmeyer, M., Su, H., & Xing, E. (2024). An exploratory study of self-supervised pre-training on partially supervised multi-label classification on chest X-ray images. Applied Soft Computing, 163, 111855. https://doi.org/10.1016/j.asoc.2024.111855

Ericsson, L., Gouk, H., Loy, C. C., & Hospedales, T. M. (2022). Self-Supervised Representation Learning: Introduction, advances, and challenges. IEEE Signal Processing Magazine, 39(3), 42–62. https://doi.org/10.1109/MSP.2021.3134634

Gui, J., Chen, T., Zhang, J., Cao, Q., Sun, Z., & Luo, H. (2024). A Survey on Self-Supervised Learning: Algorithms, Applications, and Future Trends. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(12), 9052–9071. https://doi.org/10.1109/TPAMI.2024.3415112

Jalal, M. E., & Elmaghraby, A. (2025). Toward Deep Semi-Supervised Continual Learning : A Unified Survey for Scalable and Adaptive AI. IEEE Access, 13. https://doi.org/10.1109/ACCESS.2025.3556569

Lattuada, M., Gianniti, E., Ardagna, D., & Zhang, L. (2022). Performance prediction of deep learning applications training in GPU as a service systems. Cluster Computing, 25, 1279–1302. https://doi.org/10.1007/s10586-021-03428-8

Li, S., Wang, Y., Hanson, E., Chang, A., Ki, Y. S., & Li, H. (2024). NDRec: A Near-Data Processing System for Training Large-Scale Recommendation Models. IEEE Transactions on Computers, 73(5), 1248–1261. https://doi.org/10.1109/TC.2024.3365939

Liu, H., Zhan, Y., Xia, H., Mao, Q., & Tan, Y. (2022). Self-supervised transformer-based pre-training method using latent semantic masking auto-encoder for pest and disease classification. Computers and Electronics in Agriculture, 203. https://doi.org/10.1016/j.compag.2022.107448

Lugo, L., & Vielzeuf, V. (2024). Efficiency-oriented approaches for self-supervised speech representation learning. International Journal of Speech Technology, 27, 765–779. https://doi.org/10.1007/s10772-024-10121-9

Maleki, F., & Spatz, A. (2023). Generalizability of Machine Learning Models : Quantitative Evaluation of Three Methodological Pitfalls. Radiology: Artificial Intelligence, 5(1). https://doi.org/10.1148/ryai.220028open_in_new

Marks, M., Knott, M., Kondapaneni, N., Cole, E., Perez-cruz, F., & Perona, P. (2025). A Closer Look at Benchmarking Self-supervised Pre-training with Image Classification. International Journal of Computer Vision, 133(8), 5013–5025. https://doi.org/10.1007/s11263-025-02402-w

Meyers, B. E., & Boyd, S. P. (2023). Signal Decomposition Using Masked Proximal Operators. Foundations and Trends in Signal Processing, 17(1), 1–78. https://doi.org/10.1561/2000000122

Miseta, T., Fodor, A., & Vathy-fogarassy, Á. (2024). Neurocomputing Surpassing early stopping : A novel correlation-based stopping criterion for neural networks. Neurocomputing, 567, 127028. https://doi.org/10.1016/j.neucom.2023.127028

Mithas, S., & Silveira, A. D. O. (2022). How will artificial intelligence and Industry 4.0 emerging technologies transform operations management? Production and Operations Management, 31(12). https://doi.org/10.1111/poms.13864

Mohammadi, S., & Livani, M. A. (2026). A two-stage self-supervised learning framework for breast cancer detection with multi-scale vision transformers. Information Sciences, 735. https://doi.org/10.1016/j.ins.2025.123061

Qiu, S., Zaheer, Q., Shah, M. A. H., & Shah, S. F. H. (2025). LiDAR-Simulated Multimodal and Self-Supervised Contrastive Digital Twin Approach for Probabilistic Point Cloud Generation of Rail Fasteners. Journal of Computing in Civil Engineering, 39(2). https://doi.org/10.1061/jccee5.cpeng-6137

Rani, V., Kumar, M., Gupta, A., Sachdeva, M., Mittal, A., & Kumar, K. (2024). Self-supervised learning for medical image analysis: a comprehensive review. Evolving Systems, 15, 1607–1633. https://doi.org/10.1007/s12530-024-09581-w

Rani, V., Nabi, S. T., Kumar, M., Mittal, A., & Kumar, K. (2023). Self-supervised Learning: A Succinct Review. Archives of Computational Methods in Engineering, 30, 2761–2775. https://doi.org/10.1007/s11831-023-09884-2

Rastogi, D., Johri, P., Kadry, S., Kim, S., Kumar, L., Baghela, V. S., & Khan, A. A. (2025). XcepFusion for brain tumor detection using a hybrid transfer learning framework with layer pruning and freezing. Scientific Reports, 16, 1–30. https://doi.org/10.1038/s41598-025-33970-z

Schwabe, D., Becker, K., Seyferth, M., Klaß, A., & Schaeffter, T. (2024). The METRIC-framework for assessing data quality for trustworthy AI in medicine : a systematic review. Npj Digital Medicine. https://doi.org/10.1038/s41746-024-01196-4

Sha, X., Si, X., Zhu, Y., A, S. W., & Zhao, Y. (2025). Automatic three-dimensional reconstruction of transparent objects with multiple optimization strategies under limited constraints. Image and Vision Computing, 160. https://doi.org/10.1016/j.imavis.2025.105580

Tan, T., Shull, P. B., Hicks, J. L., Uhlrich, S. D., & Chaudhari, A. S. (2024). Self-Supervised Learning Improves Accuracy and Data Efficiency for IMU-Based Ground Reaction Force Estimation. IEEE Transactions on Biomedical Engineering, 71(7), 2095–2104. https://doi.org/10.1109/TBME.2024.3361888

Wang, D., Wang, X., Wang, L., Li, M., & Da, Q. (2023). A Real-world Dataset and Benchmark For Foundation Model Adaptation in Medical Image Classification. Scientific Data, 10(1), 1–9. https://doi.org/10.1038/s41597-023-02460-0

Wen, C., Li, X., Huang, H., Liu, Y.-S., & Fang, Y. (2023). 3D Shape Contrastive Representation Learning With Adversarial Examples. IEEE Transactions on Multimedia, 27, 679–692. https://doi.org/10.1109/TMM.2023.3265177

Wu, C., Pfrommer, J., Zhou, M., & Beyerer, J. (2024). Self-Supervised Generative-Contrastive Learning of Multi-Modal Euclidean Input for 3D Shape Latent Representations : A Dynamic Switching Approach. IEEE Transactions on Multimedia, 26, 8432–8441. https://doi.org/10.1109/TMM.2023.3338079

Xu, L., Xie, H., Member, S., Qin, S. J., Tao, X., Member, S., Wang, F. L., & Member, S. (2026). Parameter-Efficient Fine-Tuning Methods for Pretrained Language Models : A Critical Review and Assessment. IEEE Transactions on Pattern Analysis and Machine Intelligence, 48(6), 6107–6126. https://doi.org/10.1109/TPAMI.2026.3657354

Xue, Y., Yang, R., Chen, X., Liu, W., Wang, Z., & Liu, X. (2024). A Review on Transferability Estimation in Deep Transfer Learning. IEEE Transactions on Artificial Intelligence, 5(12), 5894–5914. https://doi.org/10.1109/TAI.2024.3445892

Yang, X., Liu, F., & Lin, G. (2023). Effective End-to-End Vision Language Pretraining With Semantic Visual Loss. IEEE Transactions on Multimedia, 25, 8408–8417. https://doi.org/10.1109/TMM.2023.3237166

Yao, W., Wei, W., Du, W., Xu, D., Wang, W., & Chih, W. (2025). A survey on self ‑ supervised learning for non ‑ sequential tabular data. Machine Learning, 114(1), 1–19. https://doi.org/10.1007/s10994-024-06674-0

Yu, Z., Xia, Z., Xu, D., Zhang, Z., Zhang, L., Zhang, P., & Wu, L. (2026). Structure-aware generalization for heterogeneous histopathology via prototype-based multiple instance learning. Npj Digital Medicine, 9, 1–13. https://doi.org/10.1038/s41746-025-02289-4

Zhang, J., Lin, L., Yang, S., & Liu, J. (2026). Self-supervised skeleton-based action representation learning: A Benchmark and Beyond. International Journal of Computer Vision, 134. https://doi.org/10.1007/s11263-025-02644-8

Zhang, P.-F., Huang, Z., Xu, X.-S., & Bai, G. (2024). Effective and Robust Adversarial Training Against Data and Label Corruptions. IEEE Transactions on Multimedia, 26, 9477–9488. https://doi.org/10.1109/TMM.2024.3394677

Zhang, T., Cui, W., Qin, Y., Zhu, H., Liu, S., & Jiang, F. (2025). Mask-Informed Spatial-Spectral Attention Unfolding Network for Snapshot Compressive Reconstruction. IEEE Journal of Selected Topics in Signal Processing, 19(8), 1850–1860. https://doi.org/10.1109/JSTSP.2025.3628534

Zong, Y., Aodha, O. Mac, & Hospedales, T. M. (2025). Self-Supervised Multimodal Learning: A Survey. IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(7), 5299–5318. https://doi.org/10.1109/TPAMI.2024.3429301