[1]
N. B. Shah and A. P. Ganatra, “Explainable ViT and SCNN-LSTM Framework for Audio-Assisted Image Captioning”, J. Innov. Image Process., vol. 8, no. 3, pp. 978–999, Jul. 2026, doi: 10.36548/jiip.2026.3.012.