Interpretable Surrogate Models for Explaining Black-Box Predictions in Data Science

Main Article Content

  Jaesik Jeong
  Lusiana Efrizoni
  Susi Erlinda

Abstract

Background: Black-box machine-learning models often achieve strong predictive performance but provide limited transparency. Interpretable surrogate models approximate these models using simpler structures, although their explanations may not always be sufficiently faithful or stable.
Aims: This study evaluates global and local surrogate models for explaining black-box predictions and examines the trade-offs among fidelity, stability, complexity, sparsity, and computational efficiency.
Methods: Experiments were conducted on four public binary-classification datasets: Adult Income, Bank Marketing, Breast Cancer Wisconsin Diagnostic, and German Credit. Random Forest, XGBoost, and Multilayer Perceptron served as black-box classifiers. Their predictions were explained using global logistic regression, global decision tree, rule-based, local linear, and local decision tree surrogates. Each experiment was repeated using five random seeds. Explanation quality was assessed using fidelity, probability error, stability, directional consistency, structural complexity, sparsity, and computation time.
Result: The 40-rule global surrogate achieved the highest global fidelity of 0.918, followed by the depth-six decision tree at 0.915. The 15-feature local linear surrogate achieved the highest local fidelity of 0.927, while the depth-four local tree reached 0.916. Greater complexity improved fidelity but reduced the stability and simplicity of explanations. Global surrogates were generally more stable, whereas local surrogates provided stronger neighborhood-level approximation.
Conclusion: Interpretable surrogate models provide useful but approximate explanations of black-box predictions. Medium-complexity configurations offer a more practical balance among fidelity, stability, and interpretability. Surrogate selection should therefore consider the purpose of explanation, dataset characteristics, model architecture, and explanation reliability.

Article Details

How to Cite
Jeong, J., Efrizoni, L., & Erlinda, S. (2026). Interpretable Surrogate Models for Explaining Black-Box Predictions in Data Science. International Journal of Advances in Artificial Intelligence and Machine Learning, 3(3), 208–226. https://doi.org/10.58723/ijaaiml.v3i3.800
Section
Articles

References

Ali, S., Abuhmed, T., El-Sappagh, S., Muhammad, K., Alonso-Moral, J. M., Confalonieri, R., Guidotti, R., Del Ser, J., Díaz-Rodríguez, N., & Herrera, F. (2023). Explainable Artificial Intelligence (XAI): What we know and what is left to attain Trustworthy Artificial Intelligence. Information Fusion, 99(January), 101805. https://doi.org/10.1016/j.inffus.2023.101805

Amparore, E., Perotti, A., & Bajardi, P. (2021). To trust or not to trust an explanation: using LEAF to evaluate local linear XAI methods. PeerJ Computer Science, 7, 1–26. https://doi.org/10.7717/peerj-cs.479

Babichev, S., & Yarema, O. (2026). Comparative Evaluation of Machine-Learning Classifiers for High-Dimensional Gene Expression Data. IEEE Access. https://doi.org/10.1109/ACCESS.2026.3662150

Bodria, F., Giannotti, F., Guidotti, R., Naretto, F., Pedreschi, D., & Rinzivillo, S. (2023). Benchmarking and survey of explanation methods for black box models. In Data Mining and Knowledge Discovery (Vol. 37, Number 5). Springer US. https://doi.org/10.1007/s10618-023-00933-9

Carmichael, Z., & Scheirer, W. J. (2023). Unfooling Perturbation-Based Post Hoc Explainers. Proceedings of the 37th AAAI Conference on Artificial Intelligence, AAAI 2023, 37, 6925–6934. https://doi.org/10.1609/aaai.v37i6.25847

Chatterjee, A., Riegler, M. A., Johnson, M. S., Das, J., Pahari, N., Ramachandra, R., Ghosh, B., Saha, A., & Bajpai, R. (2024). Exploring online public survey lifestyle datasets with statistical analysis, machine learning and semantic ontology. Scientific Reports 2024 14:1, 14(1), 24190-. https://doi.org/10.1038/s41598-024-74539-6

Dasgupta, S., Frost, N., & Moshkovitz, M. (2022). Framework for Evaluating Faithfulness of Local Explanations. Proceedings of Machine Learning Research, 162, 4794–4815. https://arxiv.org/pdf/2202.00734

Fridkin, S., & Bendersky, M. (2026). Interpretable Machine Learning: A Comprehensive Review of Foundations, Methods, and the Path Forward. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 16(1), e70075. https://doi.org/10.1002/WIDM.70075

Fryer, D. V., Strumke, I., & Nguyen, H. (2021). Model Independent Feature Attributions: Shapley Values That Uncover Non-linear Dependencies. PeerJ Computer Science, 7, 1–23. https://doi.org/10.7717/PEERJ-CS.582

Gaudel, R., Galárraga, L., Delaunay, J., Rozé, L., & Bhargava, V. (2022). s-LIME: Reconciling Locality and Fidelity in Linear Explanations. Lecture Notes in Computer Science (Including Subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 13205 LNCS, 102–114. https://doi.org/10.1007/978-3-031-01333-1_9/SAVE-RESEARCH

Goode, K., & Hofmann, H. (2021). Visual diagnostics of an explainer model: Tools for the assessment of LIME explanations. Statistical Analysis and Data Mining, 14(2), 185–200. https://doi.org/10.1002/sam.11500

Guidotti, R., Monreale, A., Ruggieri, S., Naretto, F., Turini, F., Pedreschi, D., & Giannotti, F. (2024). Stable and actionable explanations of black-box models through factual and counterfactual rules. Data Mining and Knowledge Discovery, 38(5), 2825–2862. https://doi.org/10.1007/s10618-022-00878-5

Guidotti, R., Monreale, A., Ruggieri, S., Turini, F., Giannotti, F., & Pedreschi, D. (2019). A survey of methods for explaining black box models. ACM Computing Surveys, 51(5). https://doi.org/10.1145/3236009;SERIALTOPIC:TOPIC:ACM-PUBTYPE

Hassan, S. U., Abdulkadir, S. J., Zahid, M. S. M., & Al-Selwi, S. M. (2025). Local interpretable model-agnostic explanation approach for medical imaging analysis: A systematic literature review. Computers in Biology and Medicine, 185(April 2024), 109569. https://doi.org/10.1016/j.compbiomed.2024.109569

Huang, A. A., & Huang, S. Y. (2023). Increasing transparency in machine learning through bootstrap simulation and shapely additive explanations. PLoS ONE, 18(2 February), 1–15. https://doi.org/10.1371/journal.pone.0281922

Khasawneh, T., & Azzeh, M. (2024). Polishing the black box: flexible model-based partitioning surrogate models for interpretable machine learning model. International Journal of Data Science and Analytics 2024 20:4, 20(4), 3663–3691. https://doi.org/10.1007/S41060-024-00687-7

Lakkaraju, H., Arsov, N., & Bastani, O. (2020). Robust and Stable Black Box Explanations. 37th International Conference on Machine Learning, ICML 2020, PartF168147-8, 5584–5594. https://arxiv.org/pdf/2011.06169

Li, M., Sun, H., Huang, Y., & Chen, H. (2024). Shapley value: from cooperative game to explainable artificial intelligence. Autonomous Intelligent Systems, 4(1). https://doi.org/10.1007/s43684-023-00060-8

Lin, T., Lau, R. Y. K., & Hu, S. (2026). A critical review of state-of-the-art explainable artificial intelligence (XAI) methods and their business applications. Artificial Intelligence Review, 59(7). https://doi.org/10.1007/s10462-026-11565-y

Lyu, Q., Apidianaki, M., & Callison-Burch, C. (2024). Towards Faithful Model Explanation in NLP: A Survey. Computational Linguistics, 50(2), 657–723. https://doi.org/10.1162/coli_a_00511

Marcinkevičs, R., & Vogt, J. E. (2023). Interpretable and explainable machine learning: A methods-centric overview with concrete examples. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery, 13(3), 1–32. https://doi.org/10.1002/widm.1493

Martínez-Ramírez, J. M., Cueto-Ureña, C., Ramírez-Expósito, M. J., & Martínez-Martos, J. M. (2025). Machine Learning Models Utilizing Oxidative Stress Biomarkers for Breast Cancer Prediction: Efficacy and Limitations in Sentinel Lymph Node Metastasis Detection. Biomedicines 2025, Vol. 13 Page 3107, 13(12), 3107. https://doi.org/10.3390/BIOMEDICINES13123107

Min, C., Liao, G., Wen, G., Li, Y., & Guo, X. (2023). Ensemble Interpretation: A Unified Method for Interpretable Machine Learning. https://arxiv.org/pdf/2312.06255

Nambiar, A., Harikrishnaa, S., & Sharanprasath, S. (2023). Model-agnostic explainable artificial intelligence tools for severity prediction and symptom analysis on Indian COVID-19 data. Frontiers in Artificial Intelligence, 6. https://doi.org/10.3389/frai.2023.1272506

Petch, J., Di, S., & Nelson, W. (2022). Opening the Black Box: The Promise and Limitations of Explainable Machine Learning in Cardiology. Canadian Journal of Cardiology, 38(2), 204–213. https://doi.org/10.1016/J.CJCA.2021.09.004

Qiu, L., Yang, Y., Cao, C. C., Zheng, Y., Ngai, H., Hsiao, J., & Chen, L. (2022). Generating Perturbation-based Explanations with Robustness to Out-of-Distribution Data. Proceedings of the ACM Web Conference 2022, 3594–3605. https://doi.org/10.1145/3485447.3512254

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). “Why should i trust you?” Explaining the predictions of any classifier. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 13-17-August-2016, 1135–1144. https://doi.org/10.1145/2939672.2939778

Ronco, M., & Camps-Valls, G. (2023). Role of locality, fidelity and symmetry regularization in learning explainable representations. Neurocomputing, 562(October), 126884. https://doi.org/10.1016/j.neucom.2023.126884

Rudin, C. (2019). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead. Nature Machine Intelligence 2019 1:5, 1(5), 206–215. https://doi.org/10.1038/s42256-019-0048-x

Scholbeck, C. A., Casalicchio, G., Molnar, C., Bischl, B., & Heumann, C. (2024). Marginal effects for non-linear prediction functions. Data Mining and Knowledge Discovery, 38(5), 2997–3042. https://doi.org/10.1007/s10618-023-00993-x

Shaikh, T. A., Rasool, T., Verma, P., & Mir, W. A. (2024). A fundamental overview of ensemble deep learning models and applications: systematic literature and state of the art. Annals of Operations Research 2024, 1–77. https://doi.org/10.1007/S10479-024-06444-0

Slack, D., Hilgard, S., Jia, E., Singh, S., & Lakkaraju, H. (2020). Fooling LIME and SHAP: Adversarial attacks on post hoc explanation methods. AIES 2020 - Proceedings of the AAAI/ACM Conference on AI, Ethics, and Society, 180–186. https://doi.org/10.1145/3375627.3375830

Tian, Y., Xu, D., Tong, E., Sun, R., Chen, K., Li, Y., Baker, T., Niu, W., & Liu, J. (2024). Toward Learning Model-Agnostic Explanations for Deep Learning-Based Signal Modulation Classifiers. IEEE Transactions on Reliability, 73(3), 1529–1543. https://doi.org/10.1109/TR.2024.3367780

Viana, F. A. C., Gogu, C., & Goel, T. (2021). Surrogate modeling: tricks that endured the test of time and some recent developments. Structural and Multidisciplinary Optimization 2021 64:5, 64(5), 2881–2908. https://doi.org/10.1007/S00158-021-03001-2

Wu, K., Hu, X., Liu, L., Zhu, X., Huang, J., & Pedrycz, W. (2026). Local Surrogate Models With Residual Fuzzy Rules for Model-Agnostic Explanations. IEEE Transactions on Cybernetics. https://doi.org/10.1109/TCYB.2025.3556464

Yan, F., Wen, S., Nepal, S., Paris, C., & Xiang, Y. (2022). Explainable machine learning in cybersecurity: A survey. International Journal of Intelligent Systems, 37(12), 12305–12334. https://doi.org/10.1002/int.23088

Yuen, K. F., Wang, X., Kyriazos, T., & Poga, M. (2024). Application of Machine Learning Models in Social Sciences: Managing Nonlinear Relationships. Encyclopedia 2024, Vol. 4, Pages 1790-1805, 4(4), 1790–1805. https://doi.org/10.3390/ENCYCLOPEDIA4040118

Zafar, M. R., & Khan, N. (2021). Deterministic Local Interpretable Model-Agnostic Explanations for Stable Explainability. Machine Learning and Knowledge Extraction, 3(3), 525–541. https://doi.org/10.3390/make3030027

Zhu, X., Wang, D., Pedrycz, W., & Li, Z. (2023). Fuzzy Rule-Based Local Surrogate Models for Black-Box Model Explanation. IEEE Transactions on Fuzzy Systems, 31(6), 2056–2064. https://doi.org/10.1109/TFUZZ.2022.3218426