TY - GEN
T1 - EchoVision
T2 - 8th International Conference on Information Science and Systems, ICISS 2025
AU - Bhatnagar, Parth
AU - Sodhi, Manjit
AU - Gururaj, H. L.
AU - Kumar, Kusum
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2026/6/4
Y1 - 2026/6/4
N2 - Assistive technology has advanced significantly with the integration of artificial intelligence (AI), enabling visually impaired individuals to interact with their surroundings more effectively. EchoVision is a deep learning-based system that provides real-Time image captioning and text-To-speech (TTS) synthesis to help users comprehend visual scenes. The system leverages transformer-based models for improved caption generation, outperforming traditional CNN-LSTM architectures in fluency, accuracy, and contextual relevance. Additionally, a speech synthesis module based on Tacotron 2 ensures clear and natural audio output.The model's performance was evaluated using standard image captioning metrics such as BLEU, CIDEr, METEOR, and ROUGE-L, demonstrating superior results over conventional approaches. A user study with visually impaired individuals confirmed the system's effectiveness in diverse environments, highlighting its potential for real-world applications. Furthermore, EchoVision is optimized for edge-device deployment, making it practical for low-power mobile applications. Despite its advantages, the system faces challenges such as handling complex scenes, rare object recognition, and speech pronunciation of uncommon words. Future improvements will focus on scene graph integration, few-shot learning, and phoneme-based speech enhancements. Ethical considerations, including bias in AI-generated captions and user privacy, are also discussed to ensure responsible deployment.Overall, EchoVision represents a significant step toward AI-powered accessibility solutions, enhancing situational awareness and independence for visually impaired individuals. The project contributes to the growing field of vision-language models in assistive technology, paving the way for more inclusive AI applications.
AB - Assistive technology has advanced significantly with the integration of artificial intelligence (AI), enabling visually impaired individuals to interact with their surroundings more effectively. EchoVision is a deep learning-based system that provides real-Time image captioning and text-To-speech (TTS) synthesis to help users comprehend visual scenes. The system leverages transformer-based models for improved caption generation, outperforming traditional CNN-LSTM architectures in fluency, accuracy, and contextual relevance. Additionally, a speech synthesis module based on Tacotron 2 ensures clear and natural audio output.The model's performance was evaluated using standard image captioning metrics such as BLEU, CIDEr, METEOR, and ROUGE-L, demonstrating superior results over conventional approaches. A user study with visually impaired individuals confirmed the system's effectiveness in diverse environments, highlighting its potential for real-world applications. Furthermore, EchoVision is optimized for edge-device deployment, making it practical for low-power mobile applications. Despite its advantages, the system faces challenges such as handling complex scenes, rare object recognition, and speech pronunciation of uncommon words. Future improvements will focus on scene graph integration, few-shot learning, and phoneme-based speech enhancements. Ethical considerations, including bias in AI-generated captions and user privacy, are also discussed to ensure responsible deployment.Overall, EchoVision represents a significant step toward AI-powered accessibility solutions, enhancing situational awareness and independence for visually impaired individuals. The project contributes to the growing field of vision-language models in assistive technology, paving the way for more inclusive AI applications.
UR - https://www.scopus.com/pages/publications/105042409730
UR - https://www.scopus.com/pages/publications/105042409730#tab=citedBy
U2 - 10.1145/3803722.3803724
DO - 10.1145/3803722.3803724
M3 - Conference contribution
AN - SCOPUS:105042409730
T3 - ICISS 2025 - Proceedings of the 8th International Conference on Information Science and Systems
SP - 8
EP - 15
BT - ICISS 2025 - Proceedings of the 8th International Conference on Information Science and Systems
PB - Association for Computing Machinery, Inc
Y2 - 14 August 2025 through 16 August 2025
ER -