Skip to main navigation Skip to search Skip to main content

EchoVision: A Deep Learning-Based Assistive Captioning System for the Visually Impaired

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

Assistive technology has advanced significantly with the integration of artificial intelligence (AI), enabling visually impaired individuals to interact with their surroundings more effectively. EchoVision is a deep learning-based system that provides real-Time image captioning and text-To-speech (TTS) synthesis to help users comprehend visual scenes. The system leverages transformer-based models for improved caption generation, outperforming traditional CNN-LSTM architectures in fluency, accuracy, and contextual relevance. Additionally, a speech synthesis module based on Tacotron 2 ensures clear and natural audio output.The model's performance was evaluated using standard image captioning metrics such as BLEU, CIDEr, METEOR, and ROUGE-L, demonstrating superior results over conventional approaches. A user study with visually impaired individuals confirmed the system's effectiveness in diverse environments, highlighting its potential for real-world applications. Furthermore, EchoVision is optimized for edge-device deployment, making it practical for low-power mobile applications. Despite its advantages, the system faces challenges such as handling complex scenes, rare object recognition, and speech pronunciation of uncommon words. Future improvements will focus on scene graph integration, few-shot learning, and phoneme-based speech enhancements. Ethical considerations, including bias in AI-generated captions and user privacy, are also discussed to ensure responsible deployment.Overall, EchoVision represents a significant step toward AI-powered accessibility solutions, enhancing situational awareness and independence for visually impaired individuals. The project contributes to the growing field of vision-language models in assistive technology, paving the way for more inclusive AI applications.

Original languageEnglish
Title of host publicationICISS 2025 - Proceedings of the 8th International Conference on Information Science and Systems
PublisherAssociation for Computing Machinery, Inc
Pages8-15
Number of pages8
ISBN (Electronic)9798400713200
DOIs
Publication statusPublished - 04-06-2026
Event8th International Conference on Information Science and Systems, ICISS 2025 - Oxford, United Kingdom
Duration: 14-08-202516-08-2025

Publication series

NameICISS 2025 - Proceedings of the 8th International Conference on Information Science and Systems

Conference

Conference8th International Conference on Information Science and Systems, ICISS 2025
Country/TerritoryUnited Kingdom
CityOxford
Period14-08-2516-08-25

All Science Journal Classification (ASJC) codes

  • Computer Networks and Communications
  • Computer Science Applications
  • Information Systems
  • Signal Processing
  • Control and Systems Engineering

Fingerprint

Dive into the research topics of 'EchoVision: A Deep Learning-Based Assistive Captioning System for the Visually Impaired'. Together they form a unique fingerprint.

Cite this