An Optimized Deep Learning Model For Automatic Detection Of Stuttering Disfluencies In Speech Signals

Authors

  • Mrs. Rajeswary Nair
  • K.S Kannan

DOI:

https://doi.org/10.51483/IJAIML.6.6s.2026.1104-1114

Keywords:

Internet of Vehicles (IoV), Traffic Flow, Road, Improved Dolphin Swarm Dynamic Recurrent Neural Network (IDS-DRNN)

Abstract

Stuttering disfluencies, marked by disruptions such as repetitions, prolongations, and speech blocks, significantly impact an individual’s communication fluency and social interaction. However, current research predominantly focuses on limited feature sets and conventional deep learning models, which may not fully capture the complex temporal variations inherent in stuttered speech. To address this gap, an optimized deep learning framework is developed for effectively identifying stuttering disfluencies with improved accuracy and adaptability. The UCLASS dataset, which consists of a variety of stuttered and fluent speech samples, was utilized to ensure diversity in speech patterns. Data augmentation techniques, including time stretching and pitch shifting, were applied to balance the dataset and enrich its variability. Further, noise reduction was conducted using spectral gating to enhance signal clarity. For feature extraction, Gammatone Frequency Cepstral Coefficients (GFCC) were employed to capture robust auditory-inspired features that better represent the perceptual characteristics of speech. The proposed Adaptive Sheep Flock-driven Gate customized Long Short-Term Memory Network (ASF-GLSTM-NET) integrates adaptive gate mechanisms inspired by sheep flock optimization, enabling the model to dynamically focus on critical temporal speech patterns while reducing irrelevant noise. This architecture efficiently learns both local and long-range dependencies for precise classification of stuttering and non-stuttering events. Implemented in Python, the findings show that the ASF-GLSTM-NET approach performs better than multimodal baseline architectures, achieving superior results, with accuracy, F1-score, recall, and precision ranging from 95% to 99%. These findings support the clinical relevance of the proposed system for future diagnostic and assistive applications.

Downloads

Published

2026-06-24

How to Cite

Nair, M. R., & Kannan, K. (2026). An Optimized Deep Learning Model For Automatic Detection Of Stuttering Disfluencies In Speech Signals. International Journal of Artificial Intelligence and Machine Learning, 6(6s), 1104–1114. https://doi.org/10.51483/IJAIML.6.6s.2026.1104-1114