Leveraging Sparse Motion Learning And Hybrid Feature Fusion For Accurate Sign Language Recognition
DOI:
https://doi.org/10.51483/IJAIML.6.8s.2026.25-34Keywords:
Sign Language Recognition; Sparse Motion Sequence Extraction; Deep Learning; Spatio-Temporal Features; Siamese Network; Gesture Classification; Hybrid Feature Layer.Abstract
The proposed Sparse Motion Sequence Extraction Network (SMSE-Net) provides a sophisticated framework for fast and accurate hand sign language identification. Unlike typical deep learning models, which analyse dense frame sequences, SMSE-Net employs a sparse motion extraction approach to identify and prioritise essential gesture transitions while removing unnecessary frames. The architecture uses a Sparse Layer to extract spatial features and a Hybrid Layer to learn temporal motion, both of which are merged with a Siamese similarity model for exact classification. When compared to traditional models such as CNN, LSTM, and CNN+LSTM, experimental assessments show that SMSE-Net outperforms them all in terms of accuracy, precision, recall, and F1-score. The suggested technique has an average accuracy of 95.4% and a recall of 99%, demonstrating its robustness, generalisability, and applicability for real-time sign recognition applications.





