An Explainable AutoML Pipeline for Multi-Task Tabular Data Using Optuna and SHAP
view PDF
view PDF

How to Cite

S., Abarna, Rakshitha A., Shanmathi K., and Yogesh G. 2026. “An Explainable AutoML Pipeline for Multi-Task Tabular Data Using Optuna and SHAP”. Journal of ISMAC 8 (3): 292-307. https://doi.org/10.36548/jismac.2026.3.007.

Keywords

SHapley Additive exPlanations (SHAP)
Hyperparameter Optimization
Automated Machine Learning (AutoML)
Multi-Task Tabular Learning
Explainable Artificial Intelligence (XAI)

Abstract

Developing interpretable models on structured tabular data is difficult because there is a sequence dependency among data pre-processing, model learning, hyperparameter tuning, and result interpretation. The existing AutoML systems focus more on predictive accuracy rather than having an integrated approach towards explainability for different types of learning tasks. In this research, we introduce an Explainable AutoML Pipeline for performing automated data pre-processing, multi-model learning, hyperparameter optimization using Optuna, model selection using a leaderboard, and finally interpreting the results using SHAP. The framework automatically performs preprocessing on the input data, optimizes various candidate models, and determines the most suitable learning algorithm without any human intervention while also offering explanations about the decision-making process on a per-feature level. Model validation is then done by applying it to the California Housing and Titanic datasets. In particular, the optimized LightGBM model scored 0.8470 on the regression R2 metric, while the optimized XGBoost classifier managed to reach 81.56% accuracy, an F1-score of 0.7481, and an AUC of 0.8117. The SHAP analysis successfully detected the most influential predictive features, which increases the model interpretability

References

  1. Khan, Muhammad Salman, Tianbo Peng, Hanzlah Akhlaq, and Muhammad Adeel Khan. "Comparative Analysis of Automated Machine Learning for Hyperparameter Optimization and Explainable Artificial Intelligence Models." IEEE Access 2025, vol. 13: 84966-84991.
  2. Bifarin, Olatomiwa O., and Facundo M. Fernández. "Automated Machine Learning and Explainable AI (AutoML-XAI) for Metabolomics: Improving Cancer Diagnostics." Journal of the American Society for Mass Spectrometry 2024, vol. 35, no. 6: 1089-1100.
  3. Strudwick, James, Laura-Jayne Gardiner, Kate Denning-James, Niina Haiminen, Ashley Evans, Jennifer Kelly, Matthew Madgwick et al. "AutoXAI4Omics: An Automated Explainable AI Tool for Omics and Tabular Data." Briefings in Bioinformatics 2025, vol. 26, no. 1: bbae593.
  4. Cañete-Sifuentes, Leonardo, Victor Robles, Ernestina Menasalvas, and Raul Monroy. "Comparing Automated Machine Learning Against an Off-the-Shelf Pattern-based Classifier in a Class Imbalance Problem: Predicting University Dropout." IEEE Access 2023, vol. 11: 139147-139156.
  5. Kirubavathi, G., Dhanyatha Shree CA, and Nithish Kumar. "Enhanced DNS Cybersecurity: Optimization of XGBoost using Automated Hyperparameter Tuning via Optuna." In 2024 4th International Conference on Mobile Networks and Wireless Communications (ICMNWC), IEEE, 2024: 1-7.
  6. Gijsbers, Pieter, Marcos LP Bueno, Stefan Coors, Erin LeDell, Sébastien Poirier, Janek Thomas, Bernd Bischl, and Joaquin Vanschoren. "Amlb: An Automl Benchmark." Journal of Machine Learning Research 2024, vol. 25, no. 101: 1-65.
  7. Grinsztajn, Léo, Edouard Oyallon, and Gaël Varoquaux. "Why do tree-based models still outperform deep learning on typical tabular data." Advances in Neural Information Processing Systems 2022, vol. 35: 507-520.
  8. Roshinta, Trisna Ari, and Szűcs Gábor. "A Comparative Study of Lime and Shap for Enhancing Trustworthiness and Efficiency in Explainable AI Systems." In 2024 IEEE International Conference on Computing (ICOCO), IEEE, 2024: 134-139.
  9. Rao, Sannidhi, Shikha Mehta, Shreya Kulkarni, Harshal Dalvi, Neha Katre, and Meera Narvekar. "A Study of LIME and SHAP Model Explainers for Autonomous Disease Predictions." In 2022 IEEE Bombay Section Signature Conference (IBSSC), IEEE, 2022: 1-6.
  10. Gezici, Bahar, and Ayça Kolukisa Tarhan. "Explainable AI for Software Defect Prediction with Gradient Boosting Classifier." In 2022 7th International Conference on Computer Science and Engineering (UBMK), IEEE, 2022: 1-6.
  11. Abraham, Asha, Habeeb Shaik Mohideen, and R. Kayalvizhi. "A Tabular Variational Auto Encoderbased Hybrid Model for Imbalanced Data Classification with Feature Selection." IEEE Access 2023, vol. 11: 122760-122771.
  12. Jaiswal, Sushma, and Priyanka Gupta. "Ensemble approach: XGBoost, CATBoost, and LightGBM for Diabetes Mellitus Risk Prediction." In 2022 Second International Conference on Computer Science, Engineering and Applications (ICCSEA), IEEE, 2022: 1-6.
  13. Scikit-learn Developers, "California Housing Dataset," Scikit-learn Documentation. Available: https://scikit-learn.org/stable/modules/generated/sklearn.datasets.fetch_california_housing.html
  14. Kaggle, "Titanic: Machine Learning from Disaster," Dataset. Available: https://www.kaggle.com/competitions/titanic