Skip to content
Case Study~6 min readupdated

Speech Emotion Recognition

A speech emotion recognition system trained on four datasets. Ended up at 93.27% accuracy. Here's what actually happened.

By , AI Engineer

final accuracy

93.27%

cross-dataset, held-out

datasets

4

CREMA-D · RAVDESS · TESS · SAVEE

inference latency

<200ms

CPU, live mic input

01

Summary

A model that listens to speech and classifies the emotion in it, trained across four public datasets instead of one, with a live pipeline that classifies microphone input in under 200 ms on a CPU. Aimed at customer-service and healthcare use cases.

02

The problem

Most SER systems are trained on one dataset and fall apart on anything else. The goal was to build something that generalises — not just RAVDESS, which every tutorial uses, but across accents, recording conditions, and emotion labels that don't always agree with each other.

03

My role

A personal project, built end to end: dataset preparation and label harmonisation, model design, training and evaluation, and the real-time inference pipeline with its Django REST backend and React frontend.

04

Architecture

Audio is turned into mel spectrograms with Librosa. A CNN extracts local patterns from each spectrogram and an LSTM models how those patterns change over time (CLSTM). For live use, PyAudio reads from the microphone, the same feature extraction runs on each chunk, and a Django REST backend serves predictions to a React frontend.

05

What I tried first (that didn't work)

Started with a plain LSTM on mel spectrograms. Decent on RAVDESS alone (around 78%), fell to 61% when I mixed in CREMA-D. The model was memorising speaker identity, not emotion. Classic.

The model was memorising speaker identity, not emotion. Classic overfitting to a small, homogeneous dataset.

06

What worked

Stacking CNN + LSTM (CLSTM). The CNN extracts local patterns from the spectrogram, the LSTM captures how those patterns evolve over time. Trained on CREMA-D, RAVDESS, TESS, and SAVEE together after normalising the emotion label schema across all four.

07

Results

Final: 93.27% on the held-out test set. Real-time inference pipeline reads from a microphone using PyAudio and Librosa, classifies in under 200ms on CPU.

08

Limitations

  • All four datasets are acted speech. Spontaneous speech in a real call centre is harder, and accuracy there would be lower until it's tested on real recordings.
  • The emotion labels were mapped to one schema by hand, so some categories are approximations.
09

What I'd do differently

The label harmonisation was done by hand. Took a weekend. A proper ontology mapping or using a pre-trained audio encoder (like wav2vec2) would have been faster and probably hit higher accuracy on out-of-distribution audio.

Need a model turned into a product?

I can take a trained model through to a real-time service your application can call.