Authors :
Deepthi R. S.; Bindushree U.; Amrutha Raghupathy; Shivani N.
Volume/Issue :
Volume 11 - 2026, Issue 9 - September
Google Scholar :
https://tinyurl.com/meh4n4f
DOI :
https://doi.org/10.38124/ijisrt/26sep311
Note : A published paper may take 4-5
working days from the publication date to appear in PlumX Metrics, Semantic Scholar, and
ResearchGate.
Abstract :
PictoVoice is an image-to-speech system designed to assist users in understanding visual information through
spoken descriptions. The system accepts an image as input and processes it to extract relevant visual features, which are
then used to generate a meaningful textual caption. The generated caption is subsequently converted into speech, enabling
users to understand the contents of an image through audio output. The system integrates image processing, visual feature
extraction, image captioning, and speech generation into a unified pipeline. Multilingual support for languages is currently
being implemented to improve accessibility for users from diverse linguistic backgrounds.
Keywords :
Image Captioning, Image Processing, Visual Feature Extraction, Speech Generation, Text-to-Speech, Accessibility.
References :
- S. Shanthi, P. Gowthami, T. S. Harshitha, T. Kavya, and K. Sangeetha, “Deep Learning based Audio Description of Visual Content by Enhanc-ing Accessibility for the Visually Impaired,” in Proc. IEEE Int. Conf. Sustainable Communication Networks and Application, 2023. [Online]. Available: https://ieeexplore.ieee.org/document/10370414/
- “Computer Vision and Voice Assisted Image Captioning Framework for Visually Impaired Individuals using Deep Learning Approach,” IEEE, 2023. [Online]. Available: https://ieeexplore.ieee.org/document/ 10353449/
- “An Efficient Image to Speech Generation System for Visually Impaired Individuals Using Deep Learning Techniques,” IEEE, 2024. [Online]. Available: https://ieeexplore.ieee.org/document/10717315/
- “Machine Learning based approach to Image Description for the Visu-ally Impaired,” IEEE, 2021. [Online]. Available: https://ieeexplore.ieee. org/document/9544867/
- “Comparative Analysis between InceptionResnetV2 and InceptionV3 for Attention based Image Captioning,” IEEE, 2021. [Online]. Available: https://ieeexplore.ieee.org/document/9587514/
- “Deep Learning Based Voice Assistant for the Visually Impaired,” IEEE, 2021. [Online]. Available: https://ieeexplore.ieee.org/document/ 9532661/
- “Deep Fusion: A CNN-LSTM Image Caption Generator for En-hanced Visual Understanding,” IEEE, 2023. [Online]. Available: https://ieeexplore.ieee.org/document/10351389/
- “Deep Learning for Automated Image Captioning: A CNN and Trans-former Model Analysis,” IEEE, 2025. [Online]. Available: https:// ieeexplore.ieee.org/document/10914740/
- “Transforming Disability Into Ability: An Explainable Vision-to-Voice Image Captioning Framework Using Transformer Models and Edge Computing,” IEEE Journal, 2025. [Online]. Available: https://ieeexplore. ieee.org/document/11195080/
- “Image Captioning for the Visually Impaired and Blind: A Recipe for Low-Resource Languages,” IEEE, 2023. [Online]. Available: https://ieeexplore.ieee.org/document/10340575/
PictoVoice is an image-to-speech system designed to assist users in understanding visual information through
spoken descriptions. The system accepts an image as input and processes it to extract relevant visual features, which are
then used to generate a meaningful textual caption. The generated caption is subsequently converted into speech, enabling
users to understand the contents of an image through audio output. The system integrates image processing, visual feature
extraction, image captioning, and speech generation into a unified pipeline. Multilingual support for languages is currently
being implemented to improve accessibility for users from diverse linguistic backgrounds.
Keywords :
Image Captioning, Image Processing, Visual Feature Extraction, Speech Generation, Text-to-Speech, Accessibility.