Nico Arya Divano, Yuita Arum Sari, Sigit Adinugroho
Dysarthria is a neurological speech disorder commonly observed in patients with multiple sclerosis, Parkinson’s disease, and stroke. Traditional diagnostic procedures rely on sebjective and time-consuming clinical evaluations, which motivates the need for automated and reliable assessment methods. however, conventional Convolutional Neural Networks (CNNs) face a fundamental limitation when processing speech recordings of varying durations, often requiring truncation or padding that leads to loss of temporal information and reduced classification performance. This study proposes a robust dysarthria detection framework based on mel-spectrogram features for binary classification, Dysarthria and Control, using the TORGO dataset. The audio signals are converted into mel-spectrogram representations, then fed into a Multiscale CNN architecture. The Multiscale CNN employs parallel convolutional branches with different receptive fields to enrich multiresolution feature extraction, while a Temporal Pyramid Pooling (TPP) layer transforms variable-length temporal features into fixed-dimensional representations, enabling the model to effectively handle recordings of diverse durations. Experimental results demonstrate that the proposed Multiscale CNN-TPP model effectively overcomes the variable-length speech problem and significantly outperforms a standard CNN baseline. The model achieves an accuracy of 97.56% and an F1-score of 97.54%, indicating enhanced spatio-temporal feature extraction and improved generalization. These findings confirm the effectiveness of integrating multiscale and temporal pooling mechanisms and provide a promising foundation for developing reliable automated tools to support early dysarthria detection. © 2025 IEEE.
Faculty of Computer Science, Brawijaya University, Malang, Indonesia