[1]
Shah, N.B. and Ganatra, A.P. 2026. Explainable ViT and SCNN-LSTM Framework for Audio-Assisted Image Captioning. Journal of Innovative Image Processing. 8, 3 (Jul. 2026), 978–999. DOI:https://doi.org/10.36548/jiip.2026.3.012.