Shah, N. B., & Ganatra, A. P. (2026). Explainable ViT and SCNN-LSTM Framework for Audio-Assisted Image Captioning. Journal of Innovative Image Processing, 8(3), 978-999. https://doi.org/10.36548/jiip.2026.3.012