SHAH, Nidhi B.; GANATRA, Amit P. Explainable ViT and SCNN-LSTM Framework for Audio-Assisted Image Captioning. Journal of Innovative Image Processing, [S. l.], v. 8, n. 3, p. 978–999, 2026. DOI: 10.36548/jiip.2026.3.012. Disponível em: https://irojournals.com/iroiip/article/view/2324. Acesso em: 20 sep. 2026.