Explainable ViT and SCNN-LSTM Framework for Audio-Assisted Image Captioning
view PDF
view PDF

How to Cite

Shah, Nidhi B., and Amit P. Ganatra. 2026. “Explainable ViT and SCNN-LSTM Framework for Audio-Assisted Image Captioning”. Journal of Innovative Image Processing 8 (3): 978-99. https://doi.org/10.36548/jiip.2026.3.012.

Keywords

Separable CNN-LSTM
Explainable AI (XAI)
Image Captioning
Vision Transformer (ViT)
Audio-Assisted Visual Understanding

Abstract

Assistive technologies are crucial for image captioning, a process that involves generating textual captions of visual scenes. Many of the current methods, however, lack semantic understanding, are computationally complex, and are not interpretable. Based on these challenges, in this paper, we present an Explainable Vision Transformer and Depthwise Separable CNN-LSTM (ViT-SCNN-LSTM) architecture to achieve accurate and explainable image caption generation. The framework is based on a pre-trained Vision Transformer (ViT-B/16) to extract rich visual features and a lightweight Depthwise Separable SCNN-LSTM decoder to generate descriptive captions without high computational complexity. Grad-CAM is integrated to get visual explanations by identifying image areas that affect the generation of captions. The audio-assisted module also provides an option to turn captions into speech, making it easier for visually impaired users to access the generated captions. This model was developed using Python and tested on the Flickr8k image corpus (8,000 images, 40,000 captions). To evaluate the generalisation capability, cross-dataset experiments were also performed with Flickr30k and MSCOCO 2017. Experimental results achieved ROUGE-1, ROUGE-L, ROUGE-S, BLEU-1, BLEU-2, BLEU-3, and BLEU-4 scores of 0.87, 0.86, 0.85, 0.8848, 0.8750, 0.8649, and 0.8547, respectively. The results show the proposed approach offers an efficient, interpretable, and user-friendly solution for automatic image understanding and caption generation.

References

  1. Aoun, Muhammad, Tehseen Mazhar, Tariq Shahzad, Wajahat Waheed, and Habib Hamam. "Double‐Attention Transformer for Cross‐Modal Image Captioning: Enhancing Visual–Linguistic Alignment on Low‐Resource Datasets." Applied Computational Intelligence and Soft Computing 2026, no. 1 (2026): 5733967.
  2. Zhang, Shengcai, Junkai Fu, Junxiang Xue, and Dezhi An. "Generative Image Steganography Based on Mapping-Guided Stable Diffusion with Enhanced Robustness." Journal of King Saud University Computer and Information Sciences (2026).
  3. Li, Li, Yingzhe Peng, Xu Yang, Ruoxi Chen, Xu Haiyang, Yan Ming, and Huang Fei. "L-Clipscore: A Lightweight Embedding-Based Captioning Metric for Evaluating and Training." IEEE Transactions on Multimedia (2026).
  4. Martin, Antoinette Deborah, and Inkyu Moon. "Privacy-Preserving Image Captioning Using Virtual Photon-Limited Imaging and Federated Learning." Results in Optics (2026): 100970.
  5. Chauhan, Harshil Narendrabhai, and Chintan Thacker. "A Comprehensive Survey on Automatic Image Captioning-Deep Learning Techniques, Datasets and Evaluation Parameters." International Journal of Electrical and Computer Engineering (IJECE) 15, no. 3 (2025): 3257-3266.
  6. Khan, Abdullah, and Jaswinder Singh. "A Novel Image Captioning Technique Using Deep Learning Methodology." ICCK Transactions on Machine Intelligence 1, no. 2 (2025): 52-68.
  7. Kaur, Mehzabeen, and Harpreet Kaur. "An Efficient CNN-LSTM Based Framework For Improved Image Captioning." Procedia Computer Science 258 (2025): 3601-3607.
  8. Al Badarneh, Israa, Bassam H. Hammo, and Omar Al-Kadi. "An Ensemble Model With Attention Based Mechanism for Image Captioning." Computers and Electrical Engineering 123 (2025): 110077.
  9. Sathyanarayana, K. B., and Dinesh Naik. "An Ensemble of Vision-Language Transformer-Based Captioning Model with Rotatory Positional Embeddings." IEEE Access, vol. 13, 2025, 59841–59865.
  10. Asiri, Mashael M., Kholoud Alghamdi, Fahad Alzahrani, and Mahir Mohammed Sharif. "An Innovative Multi-Head Attention Mechanism-Driven Recurrent Neural Network Model with Feature Representation Fusion for Enhanced Image Captioning to Assist Individuals with Visual Impairments." Scientific Reports 15, no. 1 (2025): 35845.
  11. Chawla, Monika, and Rashmi Agrawal. "Analysis of Image Captioning Approaches from a Deep Learning Perspective." Journal of Mobile Multimedia 21, no. 3-4 (2025): 363-378.
  12. Patel, Madhvi, Pranay Deepak Reddy Vaka, Dhirendra Pratap Singh, Jaytrilok Choudhary, and Surendra Solanki. "Enhanced Image Captioning with Advanced Context-Aware Object Relational Model." Discover Computing 28, no. 1 (2025): 337.
  13. Alkhaldi, Tareq M., Mashael M. Asiri, Fahad Alzahrani, and Mahir Mohammed Sharif. "Fusion of Deep Transfer Learning Models with Gannet Optimisation Algorithm for an Advanced Image Captioning System for Visual Disabilities." Scientific Reports 15, no. 1 (2025): 40446.
  14. Hoseini, Farnaz, and Anaram Yaghoobi Notash. "Image Captioning Using Bidirectional LSTM Neural Network." Discover Artificial Intelligence 5, no. 1 (2025): 80.
  15. Huang, Jia-Hong, Hongyi Zhu, Yixian Shen, Stevan Rudinac, and Evangelos Kanoulas. "Image2text2image: A Novel Framework for Label-Free Evaluation of Image-to-Text Generation with Text-To-Image Diffusion Models." In International Conference on Multimedia Modeling, Singapore: Springer Nature Singapore, 2025, 413-427.
  16. Celona, Luigi, Simone Bianco, Marco Donzella, and Paolo Napoletano. "Improving Image Captioning Descriptiveness by Ranking and LLM-Based Fusion." Neural Computing and Applications 37, no. 32 (2025): 27279-27299.
  17. Haque, Anwar Ul, Sayeed Ghani, and Muhammad Saeed. "Knowledge-Driven Image Captioning." IEEE Access 13 (2025): 211352-211369.
  18. Afnan, Yasir, Kifayat Ullah, Bilal Ur Rehman, Inam Ul Hassan, Maria Zulfiqar, Zawish Asif, Wasim Habib, Muhammad Amir, and Muhammad Arshad. "A Hybrid Image Captioning Framework with EfficientNetB0 and Transformer Networks." Spectrum of Engineering Sciences (2025): 888-898.
  19. Dr. Shashidhar Kini K, Mahesh Timmanna Hegde, Image to Audio Conversion Using Machine Learning, In: Future Trends in Internet of Things, Internet of Everything and its Applications V5B19, IIP Series, Volume 5, 2025, 88-92.
  20. Tyagi, Shourya, Olukayode Ayodele Oki, Vineet Verma, Swati Gupta, Meenu Vijarania, Joseph Bamidele Awotunde, and Abdulrauph Olanrewaju Babatunde. "Novel Advance Image Caption Generation Utilizing Vision Transformer and Generative Adversarial Networks." Computers 13, no. 12 (2024): 305.
  21. Sasibhooshan, Reshmi, Suresh Kumaraswamy, and Santhoshkumar Sasidharan. "Image Caption Generation Using Visual Attention Prediction and Contextual Spatial Relation Extraction." Journal of Big Data 10, no. 1 (2023): 18.
  22. Agrawal, Vaishnavi, Shariva Dhekane, Neha Tuniya, and Vibha Vyas. "Image Caption Generator Using Attention Mechanism." In 2021 12th International Conference on Computing Communication and Networking Technologies (ICCCNT), IEEE, 2021, 1-6.
  23. Hodosh, Micah, Peter Young, and Julia Hockenmaier. "Framing Image Description as a Ranking Task: Data, Models and Evaluation Metrics." Journal of Artificial Intelligence Research 47 (2013): 853-899.
  24. Young, Peter, Alice Lai, Micah Hodosh, and Julia Hockenmaier. "From Image Descriptions to Visual Denotations: New Similarity Metrics for Semantic Inference Over Event Descriptions." Transactions of the association for computational linguistics 2 (2014): 67-78.
  25. Lin, Tsung-Yi, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. "Microsoft Coco: Common Objects in Context." In European conference on computer vision, Cham: Springer International Publishing, 2014, 740-755.