torchaudio mfcc set too high n_mels

Viewed 65

I am currently working on pytorch for speech recognition model.

When I used torchaudio.transforms.MFCC(sample_rate=16000, n_mfcc=40) for data preprocessing, the warning saying n_mels(128) is set too high or n_freqs(201) too low came out. Of course it's just a warning but I was bit worried. Also, when I used torchaudio.transforms.MFCC(sample_rate=8000, n_mfcc=40) it worked fine with no warnings.

(1) what could possibly the cause of this warning with higher sample rate?

(2) what does this warning exactly telling me? will this affect my model performance?

I am new to stackoverflow and pytorch coding so sorry if I made any mistakes with my questions.

1 Answers

Your audio has a certain frequency content. The high frequency content could be limited by the sampling rate, but also there could be low pass / high cut effects that put the limit even lower.

When you upsample your audio from 8k to 16k, it is possible for the samples to represent higher frequency content - but process does not create any such content - so the top part of the frequency spectrum will then be empty.

The default for torchaudio is to use samplerate/2 as the maximum filterbank frequency. So when you increase the samplerate the filterbank highest frequency goes up, but there is nodata in those bins. This triggers the warning your see.

To get rid of the warning, either:

  1. reduce the samplerate, so that samplerate/2 is reasonable for your audio content
  2. set fmax explicitly to be around the highest frequency content in your audio

Empty filterbank rows may affect model performance. The 0 energy can skew standardization. Whether this is noticeable depends on many factors such as the model, training process, difficulty of the problem, required performance et.c. But the safest approach is not avoid it.

Related