Leveraging Sparse Motion Learning And Hybrid Feature Fusion For Accurate Sign Language Recognition

Authors

  • Kiran C. Kulkarni
  • Manoj A. Wakchaure

DOI:

https://doi.org/10.51483/IJAIML.6.8s.2026.25-34

Keywords:

Sign Language Recognition; Sparse Motion Sequence Extraction; Deep Learning; Spatio-Temporal Features; Siamese Network; Gesture Classification; Hybrid Feature Layer.

Abstract

The proposed Sparse Motion Sequence Extraction Network (SMSE-Net) provides a sophisticated framework for fast and accurate hand sign language identification. Unlike typical deep learning models, which analyse dense frame sequences, SMSE-Net employs a sparse motion extraction approach to identify and prioritise essential gesture transitions while removing unnecessary frames. The architecture uses a Sparse Layer to extract spatial features and a Hybrid Layer to learn temporal motion, both of which are merged with a Siamese similarity model for exact classification. When compared to traditional models such as CNN, LSTM, and CNN+LSTM, experimental assessments show that SMSE-Net outperforms them all in terms of accuracy, precision, recall, and F1-score. The suggested technique has an average accuracy of 95.4% and a recall of 99%, demonstrating its robustness, generalisability, and applicability for real-time sign recognition applications.

Downloads

Published

2026-08-01

How to Cite

Kulkarni, K. C., & Wakchaure, M. A. (2026). Leveraging Sparse Motion Learning And Hybrid Feature Fusion For Accurate Sign Language Recognition. International Journal of Artificial Intelligence and Machine Learning, 6(8s), 25–34. https://doi.org/10.51483/IJAIML.6.8s.2026.25-34