TY - GEN
T1 - Sign Language Recognition Using Temporalspatial Feature Fusion Via a ResNet50-LSTM Hybrid Approach
AU - Modali, Venkata Aditi
AU - Vernekar, Nidhi
AU - Shivaprasad, Sakshi
AU - Chadaga, Krishnaraj
AU - James, Jimcymol
AU - Mahadeva, Rajesh
N1 - Publisher Copyright:
© 2025 IEEE.
PY - 2025
Y1 - 2025
N2 - Sign language recognition (SLR) plays a critical role in overcoming communication barriers for deaf and hard-of-hearing individuals. This study presents a vision-based SLR system that integrates spatial feature extraction using a pretrained ResNet50 convolutional neural network and temporal sequence modeling via a Long Short-Term Memory (LSTM) network. The system is evaluated on the LSA64 dataset, achieving an F1-score of 95.96% and an overall accuracy of approximately 97%. Comparative analysis with recent state-of-the-art methods (2021-2024) demonstrates that the proposed model offers superior performance while maintaining computational efficiency. The paper's contributions are the combination of the use of transfer learning for spatial information, mathematical description of the architecture in detail, and comprehensive comparison of the results to state-of-the-art methods. The results confirm the effectiveness of temporal-spatial feature fusion for achieving robust and realtime sign language recognition.
AB - Sign language recognition (SLR) plays a critical role in overcoming communication barriers for deaf and hard-of-hearing individuals. This study presents a vision-based SLR system that integrates spatial feature extraction using a pretrained ResNet50 convolutional neural network and temporal sequence modeling via a Long Short-Term Memory (LSTM) network. The system is evaluated on the LSA64 dataset, achieving an F1-score of 95.96% and an overall accuracy of approximately 97%. Comparative analysis with recent state-of-the-art methods (2021-2024) demonstrates that the proposed model offers superior performance while maintaining computational efficiency. The paper's contributions are the combination of the use of transfer learning for spatial information, mathematical description of the architecture in detail, and comprehensive comparison of the results to state-of-the-art methods. The results confirm the effectiveness of temporal-spatial feature fusion for achieving robust and realtime sign language recognition.
UR - https://www.scopus.com/pages/publications/105020836226
UR - https://www.scopus.com/pages/publications/105020836226#tab=citedBy
U2 - 10.1109/NMITCON65824.2025.11188061
DO - 10.1109/NMITCON65824.2025.11188061
M3 - Conference contribution
AN - SCOPUS:105020836226
T3 - 3rd IEEE International Conference on Networks, Multimedia and Information Technology, NMITCON 2025
BT - 3rd IEEE International Conference on Networks, Multimedia and Information Technology, NMITCON 2025
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 3rd IEEE International Conference on Networks, Multimedia and Information Technology, NMITCON 2025
Y2 - 1 August 2025 through 2 August 2025
ER -