⚠ Official Notice: www.ijisrt.com is the official website of the International Journal of Innovative Science and Research Technology (IJISRT) Journal for research paper submission and publication. Please beware of fake or duplicate websites using the IJISRT name.



PictoVoice — Pictures That Talk: An AI Based Image Captioning and Speech Generation System


Authors : Deepthi R. S.; Bindushree U.; Amrutha Raghupathy; Shivani N.

Volume/Issue : Volume 11 - 2026, Issue 9 - September


Google Scholar : https://tinyurl.com/meh4n4f

DOI : https://doi.org/10.38124/ijisrt/26sep311

Note : A published paper may take 4-5 working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and ResearchGate.


Abstract : PictoVoice is an image-to-speech system designed to assist users in understanding visual information through spoken descriptions. The system accepts an image as input and processes it to extract relevant visual features, which are then used to generate a meaningful textual caption. The generated caption is subsequently converted into speech, enabling users to understand the contents of an image through audio output. The system integrates image processing, visual feature extraction, image captioning, and speech generation into a unified pipeline. Multilingual support for languages is currently being implemented to improve accessibility for users from diverse linguistic backgrounds.

Keywords : Image Captioning, Image Processing, Visual Feature Extraction, Speech Generation, Text-to-Speech, Accessibility.

References :

  1. S. Shanthi, P. Gowthami, T. S. Harshitha, T. Kavya, and K. Sangeetha, “Deep Learning based Audio Description of Visual Content by Enhanc-ing Accessibility for the Visually Impaired,” in Proc. IEEE Int. Conf. Sustainable Communication Networks and Application, 2023. [Online]. Available: https://ieeexplore.ieee.org/document/10370414/
  2. “Computer Vision and Voice Assisted Image Captioning Framework for Visually Impaired Individuals using Deep Learning Approach,” IEEE, 2023. [Online]. Available: https://ieeexplore.ieee.org/document/ 10353449/
  3. “An Efficient Image to Speech Generation System for Visually Impaired Individuals Using Deep Learning Techniques,” IEEE, 2024. [Online]. Available: https://ieeexplore.ieee.org/document/10717315/
  4. “Machine Learning based approach to Image Description for the Visu-ally Impaired,” IEEE, 2021. [Online]. Available: https://ieeexplore.ieee. org/document/9544867/
  5. “Comparative Analysis between InceptionResnetV2 and InceptionV3 for Attention based Image Captioning,” IEEE, 2021. [Online]. Available: https://ieeexplore.ieee.org/document/9587514/
  6. “Deep Learning Based Voice Assistant for the Visually Impaired,” IEEE, 2021. [Online]. Available: https://ieeexplore.ieee.org/document/ 9532661/
  7. “Deep Fusion: A CNN-LSTM Image Caption Generator for En-hanced Visual Understanding,” IEEE, 2023. [Online]. Available: https://ieeexplore.ieee.org/document/10351389/
  8. “Deep Learning for Automated Image Captioning: A CNN and Trans-former Model Analysis,” IEEE, 2025. [Online]. Available: https:// ieeexplore.ieee.org/document/10914740/
  9. “Transforming Disability Into Ability: An Explainable Vision-to-Voice Image Captioning Framework Using Transformer Models and Edge Computing,” IEEE Journal, 2025. [Online]. Available: https://ieeexplore. ieee.org/document/11195080/
  10. “Image Captioning for the Visually Impaired and Blind: A Recipe for Low-Resource Languages,” IEEE, 2023. [Online]. Available: https://ieeexplore.ieee.org/document/10340575/

PictoVoice is an image-to-speech system designed to assist users in understanding visual information through spoken descriptions. The system accepts an image as input and processes it to extract relevant visual features, which are then used to generate a meaningful textual caption. The generated caption is subsequently converted into speech, enabling users to understand the contents of an image through audio output. The system integrates image processing, visual feature extraction, image captioning, and speech generation into a unified pipeline. Multilingual support for languages is currently being implemented to improve accessibility for users from diverse linguistic backgrounds.

Keywords : Image Captioning, Image Processing, Visual Feature Extraction, Speech Generation, Text-to-Speech, Accessibility.

Paper Submission Last Date
30 - September - 2026

SUBMIT YOUR PAPER CALL FOR PAPERS
Video Explanation for Published paper

Never miss an update from Papermashup

Get notified about the latest tutorials and downloads.

Subscribe by Email

Get alerts directly into your inbox after each post and stay updated.
Subscribe
OR

Subscribe by RSS

Add our RSS to your feedreader to get regular updates from us.
Subscribe