Shah, N.B. and Ganatra, A.P. (2026) “Explainable ViT and SCNN-LSTM Framework for Audio-Assisted Image Captioning”, Journal of Innovative Image Processing, 8(3), pp. 978–999. doi:10.36548/jiip.2026.3.012.