Hierarchical Convolutional Spectral Feature Learning for Multi-Class Disfluency Classification in the SEP-28k Stuttering Speech Corpus
- DOI
- 10.2991/978-94-6239-799-6_3How to use a DOI?
- Keywords
- Audio feature analysis; Convolutional Neural Network (CNN); Deep Neural Network; Mel-Frequency Cepstral Coefficients; Multiple-class categorization; SEP-28k dataset; Stuttering
- Abstract
This paper presents a deep learning framework that is used to perform automatic multi-classification of stuttering on the Sep28k dataset. This approach extracts Mel-Frequency Cepstral Coefficients (MFCCs), which help represent the spectral information of the stuttered speech audio signal, whereas a convolutional neural network (CNN) is employed for multi-class classification. The audio data are partitioned into fixed 5-s segments and converted into 40-dimensional MFCC representations to ensure uniform inputs. The proposed features are then encoded, and supervised multi-class classifiers are trained with an 80:20 train–test split. The proposed model is composed of sequential 2D convolutional layers, which are supported by max-pooling and batch normalization. The Dropout Layer within a fully connected Layer to improve generalization and minimize overfitting. The Model undergoes 30 epochs of training using Adam and multi-class log loss. This experiment achieved a test accuracy of 87.9% with good precision, recall, and F1-scores. From the confusion matrix, it is demonstrated that most stuttering sub-types are correctly classified, with only small errors between similar-sounding classes. The result shows that CNN using MFCC features is proved effective in detecting the pattern in stuttering dysfluencies. Although the model is planned as a CNN–BiLSTM system, it currently uses only the CNN layer. Although BiLSTM in the future may enhance time-based learning and it will definitely improve ac-curacy.
- Copyright
- © 2026 The Author(s)
- Open Access
- Open Access This chapter is licensed under the terms of the Creative Commons Attribution-NonCommercial 4.0 International License (http://creativecommons.org/licenses/by-nc/4.0/), which permits any noncommercial use, sharing, adaptation, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons license and indicate if changes were made.
Cite this article
TY - CONF AU - Afrin Alam AU - Amritanjali Amritanjali PY - 2026 DA - 2026/09/30 TI - Hierarchical Convolutional Spectral Feature Learning for Multi-Class Disfluency Classification in the SEP-28k Stuttering Speech Corpus BT - Proceedings of the International Conference on Emerging Trends and Technologies in Applications of Computer Science (ICETTACS 2026) PB - Atlantis Press SP - 29 EP - 42 SN - 1951-6851 UR - https://doi.org/10.2991/978-94-6239-799-6_3 DO - 10.2991/978-94-6239-799-6_3 ID - Alam2026 ER -