Advances in speech emotion recognition using a CNN and gender-based segmentation framework with feature selection techniques
PeerJ Computer Science, cilt.12, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 12
- Basım Tarihi: 2026
- Doi Numarası: 10.7717/peerj-cs.4016
- Dergi Adı: PeerJ Computer Science
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Applied Science & Technology Source, Compendex, Directory of Open Access Journals, Technology Collection (ProQuest)
- Anahtar Kelimeler: CNN, Feature selection, Fisher score, SER, Speech emotion recognition
- Orta Doğu Teknik Üniversitesi Adresli: Evet
Özet
Speech Emotion Recognition (SER) plays a critical role in affective computing and human-computer interaction. This study introduces a novel and fully automated SER framework that integrates Convolutional Neural Networks (CNNs) with advanced feature selection techniques to enhance classification accuracy while maintaining computational efficiency. Leveraging the librosa toolkit, five spectral feature types—including Mel-frequency Cepstral Coefficients (MFCCs), chroma, mel-spectrogram, spectral contrast, and tonnetz—were extracted, resulting in a 193-dimensional feature space. The Fisher score algorithm, combined with Recursive Feature Elimination (RFE), was used to identify the most discriminative features. Additionally, a gender-specific preprocessing pipeline and recognition have been integrated to mitigate bias and improve model robustness. The proposed model was evaluated on three benchmark datasets: RAVDESS, EMO-DB, and EMOVO, covering eight distinct emotional states. The results demonstrate strong generalization across languages and speaker identities, validated through both conventional and leave-one-out cross-validation settings. These findings underscore the effectiveness of combining deep learning with feature selection for scalable, accurate, and interpretable speech emotion recognition.