Shahrood University of Technology Journal of AI and Data Mining 2322-5211 10 4 2022 11 01 Speech Emotion Recognition using Enriched Spectrogram and Deep Convolutional Neural Network Transfer Learning 539 547 2576 10.22044/jadm.2022.12241.2372 EN B. Z. Mansouri Electrical and Computer Engineering Department, Ferdows branch, Islamic Azad University, Ferdows, Iran. H.R. Ghaffary Electrical and Computer Engineering Department, Ferdows branch, Islamic Azad University, Ferdows, Iran. A. Harimi Electrical and Computer Engineering Department, Ferdows branch, Islamic Azad University, Ferdows, Iran. Electrical and Computer Engineering Department, Shahrood branch, Islamic Azad University, Shahrood, Iran. Journal Article 2022 08 28 Speech emotion recognition (SER) is a challenging field of research that has attracted attention during the last two decades. Feature extraction has been reported as the most challenging issue in SER systems. Deep neural networks could partially solve this problem in some other applications. In order to address this problem, we proposed a novel enriched spectrogram calculated based on the fusion of wide-band and narrow-band spectrograms. The proposed spectrogram benefited from both high temporal and spectral resolution. Then we applied the resultant spectrogram images to the pre-trained deep convolutional neural network, ResNet152. Instead of the last layer of ResNet152, we added five additional layers to adopt the model to the present task. All the experiments performed on the popular EmoDB dataset are based on leaving one speaker out of a technique that guarantees the speaker's independency from the model. The model gains an accuracy rate of 88.97% which shows the efficiency of the proposed approach in contrast to other state-of-the-art methods.

https://jad.shahroodut.ac.ir/article_2576_9ab41e081a2486b986bdbc5e658ee3f4.pdf